Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
3 changes: 3 additions & 0 deletions doc/code/scenarios/0_attack_techniques.ipynb
Original file line number Diff line number Diff line change
Expand Up @@ -34,6 +34,9 @@
"- a **`AttackTechniqueSeedGroup`** (`seed_technique`) of general-technique seeds, which can carry a\n",
" **system prompt**, a **prepended_conversation**, a **simulated_conversation**\n",
" (`SeedSimulatedConversation`), and a **next_message**;\n",
"- a **score-feedback override** (`use_score_as_feedback`): the scenario still supplies the scorer,\n",
" but the technique can decide whether its attacker sees the scorer's rationale each turn (the\n",
" `goat` technique turns this off to match its paper);\n",
"- the selection metadata that lets a scenario pick it: its `name` and `technique_tags`.\n",
"\n",
"The objective is *not* part of the technique — it stays separate and is supplied by the dataset at\n",
Expand Down
3 changes: 3 additions & 0 deletions doc/code/scenarios/0_attack_techniques.py
Original file line number Diff line number Diff line change
Expand Up @@ -38,6 +38,9 @@
# - a **`AttackTechniqueSeedGroup`** (`seed_technique`) of general-technique seeds, which can carry a
# **system prompt**, a **prepended_conversation**, a **simulated_conversation**
# (`SeedSimulatedConversation`), and a **next_message**;
# - a **score-feedback override** (`use_score_as_feedback`): the scenario still supplies the scorer,
# but the technique can decide whether its attacker sees the scorer's rationale each turn (the
# `goat` technique turns this off to match its paper);
# - the selection metadata that lets a scenario pick it: its `name` and `technique_tags`.
#
# The objective is *not* part of the technique — it stays separate and is supplied by the dataset at
Expand Down
16 changes: 7 additions & 9 deletions pyrit/datasets/executors/red_teaming/goat.yaml
Original file line number Diff line number Diff line change
Expand Up @@ -15,21 +15,19 @@ description: |
objective is met. This technique keeps RedTeamingAttack's per-turn judging and early stopping
rather than reproducing GOAT's fixed-turn, judge-free loop, since that is existing,
already-tested PyRIT behavior and changing it would mean a new executor.
2. Judge feedback. RedTeamingAttack's AttackScoringConfig.use_score_as_feedback defaults to
True, so the objective judge's rationale is appended to the feedback the attacker sees each
turn. GOAT's own Chain-of-Attack-Thought loop (paper section 3.3) never sees judge output --
only the raw target response and its own prior reasoning. This technique keeps the
RedTeamingAttack default (judge feedback included) rather than adding a new per-technique
scoring-config override to the factory, which is a separate, more invasive change than a
prompt/schema swap; disabling it is a reasonable follow-up if a closer match is wanted.
3. Per-turn follow-up prompt. GOAT's follow-up prompt (paper Figure A.3) re-supplies the
2. Per-turn follow-up prompt. GOAT's follow-up prompt (paper Figure A.3) re-supplies the
attacker's own previous prompt (P) as an explicit field, alongside the goal and the target's
latest response. RedTeamingAttack's per-turn adversarial template only receives feedback_text
(the target's latest response, optionally with judge feedback) and objective -- the
(built from the target's latest response, without judge feedback) and objective -- the
attacker's own previous prompt is not currently exposed to per-turn templates. This
technique's follow-up prompt (goat_follow_up_prompt.yaml) omits that field rather than
approximate it, since the attacker's full prior reasoning already stays in the adversarial
chat's own conversation history regardless.

Judge feedback matches the paper: GOAT's Chain-of-Attack-Thought loop (paper section 3.3) never
sees judge output -- only the raw target response and its own prior reasoning. The technique
sets use_score_as_feedback=False, so the objective judge still scores every turn (keeping early
stopping) but its rationale is not appended to what the attacker sees.
groups:
- AI Red Team
source: AI Red Team
Expand Down
Original file line number Diff line number Diff line change
Expand Up @@ -5,8 +5,8 @@ description: |
generate follow-up adversarial reply given a target LLM response and prior conversation
prompt"), with one deliberate omission: Figure A.3 also re-supplies the attacker's own previous
prompt (P) as an explicit field. RedTeamingAttack's per-turn adversarial template only receives
feedback_text (built from the Defender's latest response, optionally with judge feedback) and
objective -- the attacker's own previous prompt is not currently exposed to per-turn templates,
feedback_text (built from the Defender's latest response; the technique disables judge feedback)
and objective -- the attacker's own previous prompt is not currently exposed to per-turn templates,
so that field is left out here rather than approximated. It stays available to the attacker
regardless, since the adversarial chat's own conversation history already includes everything
it said before.
Expand Down
7 changes: 7 additions & 0 deletions pyrit/executor/attack/core/attack_strategy.py
Original file line number Diff line number Diff line change
Expand Up @@ -689,6 +689,12 @@ def _create_identifier(
scoring_config.objective_scorer.get_identifier()
)

# Disabled feedback changes what the adversarial chat sees, so it must change the eval hash.
# Enabled feedback is the default and is omitted to keep existing hashes stable.
use_score_as_feedback: bool | None = None
if scoring_config is not None and not scoring_config.use_score_as_feedback:
use_score_as_feedback = False

# Add adversarial chat target and its effective prompts if present. The adversarial
# target becomes a child (filtered to model params by the eval rule), while the
# effective system/seed prompts land on the attack-strategy node so they are included
Expand Down Expand Up @@ -738,6 +744,7 @@ def _create_identifier(
adversarial_system_prompt=adversarial_system_prompt,
adversarial_seed_prompt=adversarial_seed_prompt,
adversarial_prompt_template=adversarial_prompt_template,
use_score_as_feedback=use_score_as_feedback,
)

@staticmethod
Expand Down
Original file line number Diff line number Diff line change
@@ -0,0 +1,35 @@
# Copyright (c) Microsoft Corporation.
# Licensed under the MIT license.

"""
add use_score_as_feedback to attack identifiers.

Revision ID: 34a18645c7e9
Revises: aca1eba410d9
Create Date: 2026-09-28 00:27:55.099128
"""

from collections.abc import Sequence

import sqlalchemy as sa
from alembic import op

# revision identifiers, used by Alembic.
revision: str = "34a18645c7e9"
down_revision: str | None = "aca1eba410d9"
branch_labels: str | Sequence[str] | None = None
depends_on: str | Sequence[str] | None = None


def upgrade() -> None:
"""Apply this schema upgrade."""
# ### commands auto generated by Alembic - please adjust! ###
op.add_column("AttackIdentifiers", sa.Column("use_score_as_feedback", sa.Boolean(), nullable=True))
# ### end Alembic commands ###


def downgrade() -> None:
"""Revert this schema upgrade."""
# ### commands auto generated by Alembic - please adjust! ###
op.drop_column("AttackIdentifiers", "use_score_as_feedback")
# ### end Alembic commands ###
1 change: 1 addition & 0 deletions pyrit/memory/memory_models.py
Original file line number Diff line number Diff line change
Expand Up @@ -841,6 +841,7 @@ class AttackIdentifierEntry(ComponentIdentifierEntry[AttackIdentifier]):
adversarial_system_prompt: Mapped[str | None] = mapped_column(Unicode, nullable=True)
adversarial_seed_prompt: Mapped[str | None] = mapped_column(Unicode, nullable=True)
adversarial_prompt_template: Mapped[str | None] = mapped_column(Unicode, nullable=True)
use_score_as_feedback: Mapped[bool | None] = mapped_column(Boolean, nullable=True)
objective_target_hash: Mapped[str | None] = mapped_column(
String(64), ForeignKey(f"{TargetIdentifierEntry.__tablename__}.hash"), nullable=True
)
Expand Down
2 changes: 2 additions & 0 deletions pyrit/models/identifiers/attack_identifier.py
Original file line number Diff line number Diff line change
Expand Up @@ -41,6 +41,8 @@ class AttackIdentifier(ComponentIdentifier):
adversarial_seed_prompt: Annotated[str | None, Evaluate.Include()] = None
#: Effective per-turn adversarial prompt template text, if the strategy uses one.
adversarial_prompt_template: Annotated[str | None, Evaluate.Include()] = None
#: ``False`` when the adversarial chat does not see scorer rationales; omitted when enabled (the default).
use_score_as_feedback: Annotated[bool | None, Evaluate.Include()] = None
#: The objective target the attack drives.
objective_target: Annotated[TargetIdentifier | None, Evaluate.Include(only_params=frozenset({"temperature"}))] = (
None
Expand Down
83 changes: 76 additions & 7 deletions pyrit/scenario/core/attack_technique_factory.py
Original file line number Diff line number Diff line change
Expand Up @@ -96,6 +96,7 @@ def __init__(
uses_adversarial: bool | None = None,
supports_additional_request_converters: bool = False,
scorer_override_policy: ScorerOverridePolicy = ScorerOverridePolicy.WARN,
use_score_as_feedback: bool | None = None,
) -> None:
"""
Initialize the factory with a technique-specific configuration.
Expand Down Expand Up @@ -145,14 +146,24 @@ class constructor signature and seed-technique shape.
scorer_override_policy: What to do when a scenario's scorer is
incompatible with the attack's ``attack_scoring_config`` type
annotation. Defaults to WARN.
use_score_as_feedback: Optional technique-level override for
``AttackScoringConfig.use_score_as_feedback``. When set, ``create()``
applies it to a copy of the scenario's scoring config, keeping the
scenario's scorers. Use ``False`` for techniques whose attacker must
not see scorer rationales while every turn is still scored. ``None``
(the default) leaves the scenario's config unchanged. When the
scenario's config is not forwarded (see ``scorer_override_policy``),
the override is applied to the ``attack_scoring_config`` in
``attack_kwargs`` instead, and ``create()`` raises if there is none.

Raises:
TypeError: If any kwarg name is not a valid constructor parameter,
or if the attack class constructor uses ``**kwargs``.
ValueError: If ``objective_target`` or
``attack_adversarial_config`` is included in ``attack_kwargs``,
or if ``uses_adversarial=False`` while an adversarial chat or
prompt is wired.
if ``uses_adversarial=False`` while an adversarial chat or
prompt is wired, or if ``use_score_as_feedback`` is set for an
attack that does not accept ``attack_scoring_config``.
"""
self._name = name
self._attack_class = attack_class
Expand All @@ -172,12 +183,14 @@ class constructor signature and seed-technique shape.
self._seed_technique = seed_technique
self._supports_additional_request_converters = supports_additional_request_converters
self._scorer_override_policy = scorer_override_policy
self._use_score_as_feedback = use_score_as_feedback

self._uses_adversarial = uses_adversarial if uses_adversarial is not None else self._derive_uses_adversarial()

self._validate_kwargs()
self._validate_converter_composition()
self._validate_adversarial_flags()
self._validate_score_feedback_override()

@classmethod
def with_simulated_conversation(
Expand Down Expand Up @@ -393,6 +406,20 @@ def _validate_converter_composition(self) -> None:
f"but {self._attack_class.__name__} does not accept 'attack_converter_config'."
)

def _validate_score_feedback_override(self) -> None:
"""
Validate that a feedback override can reach the attack's scoring config.

Raises:
ValueError: If ``use_score_as_feedback`` is set but the attack class
does not accept ``attack_scoring_config``.
"""
if self._use_score_as_feedback is not None and "attack_scoring_config" not in self._get_accepted_params():
raise ValueError(
f"Factory '{self._name}': use_score_as_feedback requires {self._attack_class.__name__} "
f"to accept 'attack_scoring_config'."
)

def _validate_kwargs(self) -> None:
"""
Validate that all kwargs are valid parameters for the attack class constructor.
Expand Down Expand Up @@ -695,8 +722,9 @@ class constructor accepts ``attack_converter_config``.

Raises:
ValueError: If a create-time adversarial chat is supplied while the
factory already baked one, or if ``scorer_override_policy`` is RAISE
and the scenario scorer is incompatible with the attack's type annotation.
factory already baked one, if ``scorer_override_policy`` is RAISE
and the scenario scorer is incompatible with the attack's type annotation,
or if ``use_score_as_feedback`` is set but no scoring config reaches the attack.
"""
create_time_target: PromptTarget | None = adversarial_chat

Expand Down Expand Up @@ -725,7 +753,21 @@ class constructor accepts ``attack_converter_config``.
attack_scoring_config=attack_scoring_config,
accepted_params=accepted_params,
):
Comment thread
shashank03-dev marked this conversation as resolved.
kwargs["attack_scoring_config"] = attack_scoring_config
kwargs["attack_scoring_config"] = self._apply_score_feedback_override(
attack_scoring_config=attack_scoring_config
)
elif self._use_score_as_feedback is not None:
# The scenario's config was skipped, so the override must reach the config the attack
# will actually use. Without a baked config the attack builds its own default, which
# cannot honor the override, so reject instead of silently running with the default.
baked_config = kwargs.get("attack_scoring_config")
if baked_config is None:
raise ValueError(
f"Factory '{self._name}': use_score_as_feedback={self._use_score_as_feedback} cannot be "
f"applied because the {type(attack_scoring_config).__name__} was not forwarded to "
f"{self._attack_class.__name__} and no attack_scoring_config is set in attack_kwargs."
)
kwargs["attack_scoring_config"] = self._apply_score_feedback_override(attack_scoring_config=baked_config)
if "attack_adversarial_config" in accepted_params and (
create_time_target is not None
or adversarial_system_prompt is not None
Expand All @@ -751,6 +793,30 @@ class constructor accepts ``attack_converter_config``.
attack = self._attack_class(**kwargs)
return AttackTechnique(attack=attack, seed_technique=self._seed_technique)

def _apply_score_feedback_override(self, *, attack_scoring_config: AttackScoringConfig) -> AttackScoringConfig:
"""
Apply this technique's ``use_score_as_feedback`` override to a scoring config.

A shallow copy keeps the config's scorers and subtype (e.g. TAP's) without
re-running its constructor, and leaves the original config unchanged.

Args:
attack_scoring_config: The scenario's config, or the baked config when the
scenario's config is not forwarded.

Returns:
AttackScoringConfig: The caller's config when no override applies, otherwise
a copy with the override applied.
"""
if (
self._use_score_as_feedback is None
or attack_scoring_config.use_score_as_feedback == self._use_score_as_feedback
):
return attack_scoring_config
overridden = copy.copy(attack_scoring_config)
overridden.use_score_as_feedback = self._use_score_as_feedback
return overridden

def _compose_converter_config(
self,
*,
Expand Down Expand Up @@ -1036,8 +1102,9 @@ def _build_identifier(self) -> ComponentIdentifier:
Build the behavioral identity for this factory.

Includes the factory name, attack class, kwargs, adversarial chat, the
adversarial system-prompt prefix, and the adversarial-flag booleans so
factories with different configurations produce different hashes. When a
adversarial system-prompt prefix, the score-feedback override, and the
adversarial-flag booleans so factories with different configurations
produce different hashes. When a
seed technique is present, its seeds are added as ``children["technique_seeds"]``.

Returns:
Expand All @@ -1063,6 +1130,8 @@ def _build_identifier(self) -> ComponentIdentifier:
params["adversarial_prompt_template"] = self._serialize_value(self._adversarial_prompt_template)
if self._adversarial_system_prompt_prefix is not None:
params["adversarial_system_prompt_prefix"] = self._adversarial_system_prompt_prefix
if self._use_score_as_feedback is not None:
params["use_score_as_feedback"] = self._use_score_as_feedback

children: dict[str, Any] = {}
if self._seed_technique is not None:
Expand Down
2 changes: 2 additions & 0 deletions pyrit/setup/initializers/techniques/extra.py
Original file line number Diff line number Diff line change
Expand Up @@ -98,6 +98,8 @@ def get_technique_factories() -> list[AttackTechniqueFactory]:
adversarial_prompt_template=SeedPrompt.from_yaml_file(
EXECUTOR_RED_TEAM_PATH / "goat_follow_up_prompt.yaml"
),
# GOAT's attacker never sees judge output (paper section 3.3); every turn is still scored.
use_score_as_feedback=False,
),
AttackTechniqueFactory(
name="split_payload",
Expand Down
Loading