Skip to content

FEAT Let attack techniques turn off judge feedback, and use it for GOAT to match the paper #2889

Description

@shashank03-dev

Is your feature request related to a problem? Please describe.

In GOAT (paper, section 3.3), the attacker never sees the judge's output. Each turn it only sees the target's latest reply and its own earlier reasoning.

The goat technique added in #2763 is different here. It uses RedTeamingAttack, where use_score_as_feedback defaults to True. So every turn, the judge's rationale ("the response refused because...") is added to what the attacker sees. That gives the GOAT attacker a hint the paper's attacker never gets, so results from our GOAT are not a fair comparison with the paper.

This came up in the #2763 review. The request was to turn feedback off for GOAT while keeping per-turn judging and early stopping, and not to switch to last-turn-only scoring. It was left as a follow-up, and goat.yaml (difference #2, "Judge feedback") documents it as one.

Two things stop us from doing it today:

  1. A technique cannot set the flag. Scenarios pass their own AttackScoringConfig to AttackTechniqueFactory.create(), and it replaces whatever the technique wanted. Putting attack_scoring_config in attack_kwargs does not help, because create() overwrites it.
  2. The flag is not part of the attack's identity. Two attacks that differ only in use_score_as_feedback get the same component hash and the same AtomicAttack.technique_eval_hash. AttackStrategy._create_identifier() reads only the objective scorer from the scoring config. Scenario resume uses technique_eval_hash to decide which objectives are already done. So if we only fixed (1), resuming a GOAT scenario that started before this change would silently reuse results produced with feedback on. That is the same problem Add GOAT attack technique #2763 fixed for adversarial_prompt_template.

Today nothing in scenarios, the CLI or the GUI sets this flag, so (2) cannot happen yet. It starts to matter as soon as (1) lets GOAT turn feedback off, so both parts belong in one change.

Describe the solution you'd like

A small, generic change. No new executor and no GOAT-only code path.

Factory option

  1. Add use_score_as_feedback: bool | None = None to AttackTechniqueFactory.__init__.
    • None (the default): nothing changes, the scenario's config is used as is.
    • True / False: in create(), the factory makes a copy of the scenario's scoring config with only this flag changed (dataclasses.replace). The scenario's scorer, auxiliary scorers and config subtype (for example TAPAttackScoringConfig) are all kept, and the scenario's own object is not modified.
  2. Include the option in the factory identifier when it is set.
  3. Set use_score_as_feedback=False on the goat factory in pyrit/setup/initializers/techniques/extra.py, and update goat.yaml difference Fix link to assets #2 to say feedback is now off.

Attack identity

  1. Add a use_score_as_feedback field to AttackIdentifier, included in the eval hash next to adversarial_prompt_template, and fill it in _create_identifier() only when feedback is off. Because True is the default, every existing stored hash stays exactly the same, so saved runs of other techniques still resume. GOAT gets a new hash, so old GOAT runs are correctly not reused. If the field needs a stored column like adversarial_prompt_template did in Add GOAT attack technique #2763, I will add it with a generated Alembic migration.

Tests (the identity and factory tests fail on main today)

  • component hash and technique_eval_hash differ between feedback on and off, for RedTeamingAttack, CrescendoAttack and TreeOfAttacksWithPruningAttack
  • hashes for default (feedback on) attacks are unchanged
  • the identity survives the memory round trip
  • the created attack has the flag the factory asked for, keeps the scenario's scorer and config type, and does not change the scenario's config object
  • None leaves current behavior exactly as it is
  • GOAT's attacker input no longer contains the judge rationale, while judging and early stopping still happen every turn

Docs: update the factory option docs and the GOAT notes.

Describe alternatives you've considered, if relevant

  • Create a GOAT-specific attack class. Not needed. The Add GOAT attack technique #2763 review preferred reusing RedTeamingAttack, and a new class for one flag would duplicate a lot of code.
  • Score only the last turn. Rejected in the Add GOAT attack technique #2763 review. It would miss successes on earlier turns and remove early stopping.
  • Accept a full attack_scoring_config on the factory. Too broad. It would compete with the scenario's scorer, which the scenario should keep controlling. A single flag is the smallest change that solves the problem.
  • Always record the flag in the identity, even when True. Simpler, but it changes the hash of every existing attack with a scoring config, so older saved runs would stop resuming. Happy to switch if you prefer that.

Additional context

  • Only GOAT changes behavior. Every other technique keeps None and behaves exactly as before.
  • The same option can later be used by any technique whose paper has an attacker that never sees the judge.

I would like to take this.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions