You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
Is your feature request related to a problem? Please describe.
In GOAT (paper, section 3.3), the attacker never sees the judge's output. Each turn it only sees the target's latest reply and its own earlier reasoning.
The goat technique added in #2763 is different here. It uses RedTeamingAttack, where use_score_as_feedback defaults to True. So every turn, the judge's rationale ("the response refused because...") is added to what the attacker sees. That gives the GOAT attacker a hint the paper's attacker never gets, so results from our GOAT are not a fair comparison with the paper.
This came up in the #2763 review. The request was to turn feedback off for GOAT while keeping per-turn judging and early stopping, and not to switch to last-turn-only scoring. It was left as a follow-up, and goat.yaml (difference #2, "Judge feedback") documents it as one.
Two things stop us from doing it today:
A technique cannot set the flag. Scenarios pass their own AttackScoringConfig to AttackTechniqueFactory.create(), and it replaces whatever the technique wanted. Putting attack_scoring_config in attack_kwargs does not help, because create() overwrites it.
The flag is not part of the attack's identity. Two attacks that differ only in use_score_as_feedback get the same component hash and the same AtomicAttack.technique_eval_hash. AttackStrategy._create_identifier() reads only the objective scorer from the scoring config. Scenario resume uses technique_eval_hash to decide which objectives are already done. So if we only fixed (1), resuming a GOAT scenario that started before this change would silently reuse results produced with feedback on. That is the same problem Add GOAT attack technique #2763 fixed for adversarial_prompt_template.
Today nothing in scenarios, the CLI or the GUI sets this flag, so (2) cannot happen yet. It starts to matter as soon as (1) lets GOAT turn feedback off, so both parts belong in one change.
Describe the solution you'd like
A small, generic change. No new executor and no GOAT-only code path.
Factory option
Add use_score_as_feedback: bool | None = None to AttackTechniqueFactory.__init__.
None (the default): nothing changes, the scenario's config is used as is.
True / False: in create(), the factory makes a copy of the scenario's scoring config with only this flag changed (dataclasses.replace). The scenario's scorer, auxiliary scorers and config subtype (for example TAPAttackScoringConfig) are all kept, and the scenario's own object is not modified.
Include the option in the factory identifier when it is set.
Set use_score_as_feedback=False on the goat factory in pyrit/setup/initializers/techniques/extra.py, and update goat.yaml difference Fix link to assets #2 to say feedback is now off.
Attack identity
Add a use_score_as_feedback field to AttackIdentifier, included in the eval hash next to adversarial_prompt_template, and fill it in _create_identifier()only when feedback is off. Because True is the default, every existing stored hash stays exactly the same, so saved runs of other techniques still resume. GOAT gets a new hash, so old GOAT runs are correctly not reused. If the field needs a stored column like adversarial_prompt_template did in Add GOAT attack technique #2763, I will add it with a generated Alembic migration.
Tests (the identity and factory tests fail on main today)
component hash and technique_eval_hash differ between feedback on and off, for RedTeamingAttack, CrescendoAttack and TreeOfAttacksWithPruningAttack
hashes for default (feedback on) attacks are unchanged
the identity survives the memory round trip
the created attack has the flag the factory asked for, keeps the scenario's scorer and config type, and does not change the scenario's config object
None leaves current behavior exactly as it is
GOAT's attacker input no longer contains the judge rationale, while judging and early stopping still happen every turn
Docs: update the factory option docs and the GOAT notes.
Describe alternatives you've considered, if relevant
Create a GOAT-specific attack class. Not needed. The Add GOAT attack technique #2763 review preferred reusing RedTeamingAttack, and a new class for one flag would duplicate a lot of code.
Score only the last turn. Rejected in the Add GOAT attack technique #2763 review. It would miss successes on earlier turns and remove early stopping.
Accept a full attack_scoring_config on the factory. Too broad. It would compete with the scenario's scorer, which the scenario should keep controlling. A single flag is the smallest change that solves the problem.
Always record the flag in the identity, even when True. Simpler, but it changes the hash of every existing attack with a scoring config, so older saved runs would stop resuming. Happy to switch if you prefer that.
Additional context
Only GOAT changes behavior. Every other technique keeps None and behaves exactly as before.
The same option can later be used by any technique whose paper has an attacker that never sees the judge.
Is your feature request related to a problem? Please describe.
In GOAT (paper, section 3.3), the attacker never sees the judge's output. Each turn it only sees the target's latest reply and its own earlier reasoning.
The
goattechnique added in #2763 is different here. It usesRedTeamingAttack, whereuse_score_as_feedbackdefaults toTrue. So every turn, the judge's rationale ("the response refused because...") is added to what the attacker sees. That gives the GOAT attacker a hint the paper's attacker never gets, so results from our GOAT are not a fair comparison with the paper.This came up in the #2763 review. The request was to turn feedback off for GOAT while keeping per-turn judging and early stopping, and not to switch to last-turn-only scoring. It was left as a follow-up, and
goat.yaml(difference #2, "Judge feedback") documents it as one.Two things stop us from doing it today:
AttackScoringConfigtoAttackTechniqueFactory.create(), and it replaces whatever the technique wanted. Puttingattack_scoring_configinattack_kwargsdoes not help, becausecreate()overwrites it.use_score_as_feedbackget the same component hash and the sameAtomicAttack.technique_eval_hash.AttackStrategy._create_identifier()reads only the objective scorer from the scoring config. Scenario resume usestechnique_eval_hashto decide which objectives are already done. So if we only fixed (1), resuming a GOAT scenario that started before this change would silently reuse results produced with feedback on. That is the same problem Add GOAT attack technique #2763 fixed foradversarial_prompt_template.Today nothing in scenarios, the CLI or the GUI sets this flag, so (2) cannot happen yet. It starts to matter as soon as (1) lets GOAT turn feedback off, so both parts belong in one change.
Describe the solution you'd like
A small, generic change. No new executor and no GOAT-only code path.
Factory option
use_score_as_feedback: bool | None = NonetoAttackTechniqueFactory.__init__.None(the default): nothing changes, the scenario's config is used as is.True/False: increate(), the factory makes a copy of the scenario's scoring config with only this flag changed (dataclasses.replace). The scenario's scorer, auxiliary scorers and config subtype (for exampleTAPAttackScoringConfig) are all kept, and the scenario's own object is not modified.use_score_as_feedback=Falseon thegoatfactory inpyrit/setup/initializers/techniques/extra.py, and updategoat.yamldifference Fix link to assets #2 to say feedback is now off.Attack identity
use_score_as_feedbackfield toAttackIdentifier, included in the eval hash next toadversarial_prompt_template, and fill it in_create_identifier()only when feedback is off. BecauseTrueis the default, every existing stored hash stays exactly the same, so saved runs of other techniques still resume. GOAT gets a new hash, so old GOAT runs are correctly not reused. If the field needs a stored column likeadversarial_prompt_templatedid in Add GOAT attack technique #2763, I will add it with a generated Alembic migration.Tests (the identity and factory tests fail on
maintoday)technique_eval_hashdiffer between feedback on and off, forRedTeamingAttack,CrescendoAttackandTreeOfAttacksWithPruningAttackNoneleaves current behavior exactly as it isDocs: update the factory option docs and the GOAT notes.
Describe alternatives you've considered, if relevant
RedTeamingAttack, and a new class for one flag would duplicate a lot of code.attack_scoring_configon the factory. Too broad. It would compete with the scenario's scorer, which the scenario should keep controlling. A single flag is the smallest change that solves the problem.True. Simpler, but it changes the hash of every existing attack with a scoring config, so older saved runs would stop resuming. Happy to switch if you prefer that.Additional context
Noneand behaves exactly as before.I would like to take this.