Skip to content

FEAT Let techniques disable judge feedback and use it for GOAT - #2890

Open
dev3 (shashank03-dev) wants to merge 2 commits into
microsoft:mainfrom
shashank03-dev:feat/goat-disable-judge-feedback
Open

dev3 (shashank03-dev) wants to merge 2 commits into
microsoft:mainfrom
shashank03-dev:feat/goat-disable-judge-feedback

Conversation

@shashank03-dev

@shashank03-dev dev3 (shashank03-dev) commented Sep 27, 2026 •

Copy link
Copy Markdown
Contributor

Closes #2889.

Summary

GOAT's attacker never sees judge output (paper, section 3.3), but the goat technique from #2763 inherited RedTeamingAttack's default of appending the judge's rationale to every attacker turn. This was deferred in the #2763 review because a technique had no way to set it.

This PR adds a small generic factory option, turns feedback off for GOAT, and makes the flag part of the attack's eval identity so scenario resume cannot mix results from the two modes. Per-turn judging and early stopping are unchanged.

Changes

Attack identity

  • AttackIdentifier gains use_score_as_feedback, included in the eval hash. AttackStrategy._create_identifier() records it only when feedback is off, so every existing component and eval hash stays the same.
  • New nullable AttackIdentifiers.use_score_as_feedback column, with a migration generated by build_scripts.memory_migrations (revision 34a18645c7e9 on top of aca1eba410d9).

Factory option

  • AttackTechniqueFactory(use_score_as_feedback=...). None (default) leaves the scenario's config untouched. When set, create() applies it to a shallow copy of the scenario's scoring config, so the scenario's scorers and config subtype are kept and the shared config object is never modified.
  • Setting it for an attack that does not accept attack_scoring_config raises at factory construction instead of being silently ignored.
  • When the scenario's config is not forwarded (e.g. a plain AttackScoringConfig for TAP under WARN/SKIP), the override is applied to the attack_scoring_config baked into attack_kwargs; if there is none, create() raises rather than letting the attack build its default config with feedback on.
  • The option is part of the factory identifier, and with_adversarial_system_prompt_prefix() copies carry it.

GOAT

  • goat sets use_score_as_feedback=False.
  • goat.yaml no longer lists judge feedback as a difference from the paper. goat_follow_up_prompt.yaml is updated to match.

Docs

  • doc/code/scenarios/0_attack_techniques.py (and the paired notebook, markdown cell only) lists the new option.

Notes for review

  • Shallow copy instead of dataclasses.replace (which the issue mentioned): TAPAttackScoringConfig has its own __init__, and copying keeps its type and threshold without re-running the constructor. Covered by a test.
  • Hash impact: only attacks with feedback off get a new hash. GOAT's hash changes, so GOAT runs started before this change are not reused on resume, which is intended. The internal RedTeamingAttack used by generate_simulated_conversation_async already sets feedback off, so its identity now records that too. That attack is not persisted as an attack result, and the simulated-conversation seed is keyed on its own configuration, so resume is unaffected.
  • Upgrade effect for existing users: nothing in scenarios, the CLI or the GUI sets this flag, but a user who passed AttackScoringConfig(use_score_as_feedback=False) to RedTeamingAttack, CrescendoAttack or TAP in their own code gets a new eval hash. A scenario started before this change and resumed after it re-runs those attacks instead of reusing results recorded without the flag in their identity. That is the safe direction.
  • If you would rather always record the flag (including True), it is a one-line change, but every existing attack hash with a scoring config would change.

Tests

  • Identity: feedback on and off produce different component and eval hashes; default and no-config attacks keep their hashes; persisted identifiers round-trip through memory with the flag.
  • Migration: a database at aca1eba410d9 upgrades cleanly and passes check_schema_migrations.
  • Factory: pass-through when unset; override applied to a copy while keeping the scenario's scorers; TAP subtype and threshold kept; validation error for attacks without attack_scoring_config; skipped scenario config applies to the baked config or raises (including the real TreeOfAttacksWithPruningAttack with a plain config); identifier; prefixed copies.
  • GOAT end to end: a real RedTeamingAttack built through the GOAT factory scores every turn with feedback off, leaves the scenario config unchanged, and carries the flag in its identity.

The new identity and factory tests fail on main; the pass-through and hash-stability tests pass on both.

Verification

  • pre-commit on changed files: all hooks pass, including Check Memory Migrations and Enforce Alembic Revision Immutability.
  • ty check pyrit tests/unit: no new diagnostics (identical output to main).
  • Unit tests: 20,638 passed, 146 skipped; total coverage 92.56%.
  • diff-cover against upstream/main: 100% (35/35 changed lines).

GOAT's attacker never sees judge output (paper section 3.3), but the goat
technique inherited RedTeamingAttack's default of appending the scorer
rationale to every attacker turn, and a technique had no way to override the
scenario's scoring config.

- AttackTechniqueFactory gains use_score_as_feedback, applied to a copy of
  the scenario's scoring config in create().
- The goat technique turns feedback off; judging and early stopping are kept.
- AttackIdentifier records use_score_as_feedback when feedback is off, so
  scenario resume cannot reuse results across feedback modes. Enabled
  feedback is omitted, keeping existing hashes stable. Adds the column and
  migration.

Closes microsoft#2889
Comment thread pyrit/scenario/core/attack_technique_factory.py
When the scenario's scoring config is not forwarded (for example a plain
AttackScoringConfig for TAP under WARN or SKIP), the override was never
applied and the attack built its default config with feedback on.

Apply the override to the attack_scoring_config baked into attack_kwargs
instead, and raise when there is none rather than silently running with
the default.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

FEAT Let attack techniques turn off judge feedback, and use it for GOAT to match the paper

2 participants