FEAT Let techniques disable judge feedback and use it for GOAT - #2890
Open
dev3 (shashank03-dev) wants to merge 2 commits into
Open
dev3 (shashank03-dev) wants to merge 2 commits into
dev3 (shashank03-dev) wants to merge 2 commits into
Conversation
GOAT's attacker never sees judge output (paper section 3.3), but the goat technique inherited RedTeamingAttack's default of appending the scorer rationale to every attacker turn, and a technique had no way to override the scenario's scoring config. - AttackTechniqueFactory gains use_score_as_feedback, applied to a copy of the scenario's scoring config in create(). - The goat technique turns feedback off; judging and early stopping are kept. - AttackIdentifier records use_score_as_feedback when feedback is off, so scenario resume cannot reuse results across feedback modes. Enabled feedback is omitted, keeping existing hashes stable. Adds the column and migration. Closes microsoft#2889
When the scenario's scoring config is not forwarded (for example a plain AttackScoringConfig for TAP under WARN or SKIP), the override was never applied and the attack built its default config with feedback on. Apply the override to the attack_scoring_config baked into attack_kwargs instead, and raise when there is none rather than silently running with the default.
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Closes #2889.
Summary
GOAT's attacker never sees judge output (paper, section 3.3), but the
goattechnique from #2763 inheritedRedTeamingAttack's default of appending the judge's rationale to every attacker turn. This was deferred in the #2763 review because a technique had no way to set it.This PR adds a small generic factory option, turns feedback off for GOAT, and makes the flag part of the attack's eval identity so scenario resume cannot mix results from the two modes. Per-turn judging and early stopping are unchanged.
Changes
Attack identity
AttackIdentifiergainsuse_score_as_feedback, included in the eval hash.AttackStrategy._create_identifier()records it only when feedback is off, so every existing component and eval hash stays the same.AttackIdentifiers.use_score_as_feedbackcolumn, with a migration generated bybuild_scripts.memory_migrations(revision34a18645c7e9on top ofaca1eba410d9).Factory option
AttackTechniqueFactory(use_score_as_feedback=...).None(default) leaves the scenario's config untouched. When set,create()applies it to a shallow copy of the scenario's scoring config, so the scenario's scorers and config subtype are kept and the shared config object is never modified.attack_scoring_configraises at factory construction instead of being silently ignored.AttackScoringConfigfor TAP underWARN/SKIP), the override is applied to theattack_scoring_configbaked intoattack_kwargs; if there is none,create()raises rather than letting the attack build its default config with feedback on.with_adversarial_system_prompt_prefix()copies carry it.GOAT
goatsetsuse_score_as_feedback=False.goat.yamlno longer lists judge feedback as a difference from the paper.goat_follow_up_prompt.yamlis updated to match.Docs
doc/code/scenarios/0_attack_techniques.py(and the paired notebook, markdown cell only) lists the new option.Notes for review
dataclasses.replace(which the issue mentioned):TAPAttackScoringConfighas its own__init__, and copying keeps its type and threshold without re-running the constructor. Covered by a test.RedTeamingAttackused bygenerate_simulated_conversation_asyncalready sets feedback off, so its identity now records that too. That attack is not persisted as an attack result, and the simulated-conversation seed is keyed on its own configuration, so resume is unaffected.AttackScoringConfig(use_score_as_feedback=False)toRedTeamingAttack,CrescendoAttackor TAP in their own code gets a new eval hash. A scenario started before this change and resumed after it re-runs those attacks instead of reusing results recorded without the flag in their identity. That is the safe direction.True), it is a one-line change, but every existing attack hash with a scoring config would change.Tests
aca1eba410d9upgrades cleanly and passescheck_schema_migrations.attack_scoring_config; skipped scenario config applies to the baked config or raises (including the realTreeOfAttacksWithPruningAttackwith a plain config); identifier; prefixed copies.RedTeamingAttackbuilt through the GOAT factory scores every turn with feedback off, leaves the scenario config unchanged, and carries the flag in its identity.The new identity and factory tests fail on
main; the pass-through and hash-stability tests pass on both.Verification
pre-commiton changed files: all hooks pass, includingCheck Memory MigrationsandEnforce Alembic Revision Immutability.ty check pyrit tests/unit: no new diagnostics (identical output tomain).diff-coveragainstupstream/main: 100% (35/35 changed lines).