Skip to content

FEAT: Add A2ATarget for Agent-to-Agent protocol communication - #2771

Open
Radoslaw Brus (rbrus) wants to merge 6 commits into
microsoft:mainfrom
rbrus:feat/a2a-prompt-target
Open

Radoslaw Brus (rbrus) wants to merge 6 commits into
microsoft:mainfrom
rbrus:feat/a2a-prompt-target

Conversation

@rbrus

@rbrus Radoslaw Brus (rbrus) commented Sep 22, 2026 •

Copy link
Copy Markdown

Closes #2824

Summary

This PR adds A2ATarget, so PyRIT can red-team agents exposed over the Agent-to-Agent (A2A) protocol, such as Microsoft Foundry Agent Service incoming A2A endpoints and Google ADK agents.

Key Capabilities

  1. Protocol negotiation: speaks A2A 0.3 (message/send) by default. If the agent returns JSON-RPC -32601 (method not found), it falls back to 0.2 (tasks/send) and pins that dialect.
  2. Multi-turn context continuity: each PyRIT conversation maps to one A2A context (contextId in 0.3, sessionId in 0.2), so the agent keeps its own server-side state across turns. A task waiting for input (input-required / auth-required) is continued with its task ID; completed tasks are not reused. reset_conversation_async forgets a conversation's context.
  3. Task lifecycle: requests set configuration.blocking, and tasks still submitted or working are polled with tasks/get up to task_timeout_seconds.
  4. Rate limits: HTTP 429, and 429s the agent relays from its model as JSON-RPC errors, raise RateLimitException and are retried with pyrit_target_retry.
  5. Errors: JSON-RPC errors come back as PyRIT error responses (response_type="error"), marked blocked when they match CONTENT_FILTER_MARKERS and unknown otherwise.
  6. Authentication: Bearer tokens (auth_token) and API keys (api_key with a configurable api_key_header).

System prompts and prepended conversations are handled by PyRIT's existing normalization pipeline; the target declares multi-turn support but not system prompts or editable history.

Testing

Unit tests in tests/unit/prompt_target/target/test_a2a_target.py cover 0.3 success, Message results, task polling and timeout, context continuity, input-required continuation, conversation isolation, reset, 0.2 fallback and session continuity, error and content-filter responses, HTTP and relayed rate-limit retries, and auth headers.

  • pytest tests/unit/prompt_target: 1360 passed, 31 skipped
  • ruff check, ruff format, ty check on the changed files: clean

Live test against a Microsoft Foundry Agent Service prompt agent (gpt-5-nano) with incoming A2A enabled (A2A 0.3, JSON-RPC, Entra bearer auth):

Check Result
Single turn Reply returned, dialect pinned to 0.3
Server-side continuity Turn 1 gave an employee ID; turn 2 sent only the new message with the stored contextId, and the agent recalled the ID
Conversation isolation A different conversation asking the same question got no ID
Content-filtered prompt Returned as an error response (see note)
30 concurrent requests Relayed 429s were retried; 27 of 30 completed with retry waits capped at 30s for the test

Note: when Azure content filtering blocks a prompt, Foundry's A2A endpoint returns a generic JSON-RPC -32603 "Received 400 from a service request" with no content-filter detail, so those responses are classified as unknown rather than blocked.

Introduces A2ATarget, enabling PyRIT to red-team multi-agent systems and services communicating over the Agent-to-Agent (A2A) protocol.

Key features:
- Dual-spec protocol negotiation: supports A2A v0.3 (message/send) with automatic fallback to v0.2 (tasks/send) on JSON-RPC -32601.
- Multi-turn conversation task continuity: tracks and preserves active server task IDs per conversation session.
- Robust response and refusal handling: captures JSON-RPC error refusals as conversational evidence rather than raising unhandled exceptions.
- Agent card discovery: provides get_agent_card_async() to inspect declared agent capabilities at standard discovery paths.
- Auth support: handles Bearer token and custom API-Key headers.
- Comprehensive unit test suite covering v0.3, v0.2 fallback, task continuity, refusals, and agent card discovery.
@rbrus

Copy link
Copy Markdown
Author

@microsoft-github-policy-service agree

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks for adding this! We are interested in including A2A support in PyRIT. The overall direction looks good, and the target follows the main PyRIT target contract.

The main change I would like before merging is to use the official A2A Python SDK (a2a-sdk), rather than maintain our own protocol implementation. It is published on PyPI and supports A2A v1.0 with v0.3 compatibility. Using its typed messages, agent-card discovery, and client transport would reduce the version-specific code we need to maintain. A2ATarget can remain a thin adapter that owns PyRIT message mapping, identifiers/capabilities, upstream context tracking, and error mapping. An optional a2a extra would fit our existing dependency pattern.

A few gaps to address during that change:

  • Safe polling retries: A 429 from tasks/get currently retries the whole send method, which submits the prompt again with a new message ID. Retry the poll against the existing task instead of creating duplicate agent work.
  • Task result handling: Failed or canceled tasks currently become ordinary text responses with no error flag. An input-required task with partial artifacts can also return those artifacts instead of the question in its status message. Handle these states explicitly, and do not use str(result) as an agent answer when there is no supported text.
  • Timeouts: The blocking send uses HTTPX's five-second default timeout. The 120-second task timeout only starts after that response arrives, so it does not protect a slow initial request.
  • Protocol versions: The v02 branch does not match the published v0.2.5 format, which already uses message/send and kind. The 0.3+ claim also needs narrowing because v1.0 has breaking changes. This is a strong reason to let the SDK handle protocol compatibility.
  • Identity and capability limits: Include non-secret protocol/routing settings that affect behavior in the target identifier. If stored history exists but the upstream context is missing, fail explicitly or provide a supported restoration path instead of continuing with only the latest message. Capability overrides should not advertise features the adapter cannot send.
  • Integration coverage: We should have a real HTTP integration test against a deterministic SDK-based agent, not only mocked responses. Cover multi-turn context, input-required continuation, failed tasks, and a polling 429 that does not resubmit the prompt. An environment-gated test against an owned Foundry agent would also be useful for real authentication and service interoperability.

Requesting changes mainly for the SDK direction and integration coverage. This is a useful addition, and I would like to get it into PyRIT with the protocol details handled by the maintained SDK.

@rbrus

Copy link
Copy Markdown
Author

Thanks for the thorough and constructive review Richard Lundeen (@richlundeen)! I completely agree with the direction. Adopting the official a2a-sdk as an optional dependency (pyrit[a2a]) is a much cleaner architecture for PyRIT long-term than maintaining raw protocol serialization and dialect fallbacks.

I will update the PR with the following changes:

  1. a2a-sdk integration:

    • Add a2a optional extra to pyproject.toml (and sync into all).
    • Refactor A2ATarget to use a2a-sdk's typed models, client, and agent card discovery, keeping A2ATarget focused on PyRIT message adaptation, capability limits, and session/context mapping.
    • Follow the standard lazy-import pattern with a descriptive ImportError (pip install pyrit[a2a]).
  2. Safe polling & retry handling:

    • Scope 429 retries during tasks/get polling to the polling loop itself, ensuring a polling rate limit does not re-trigger _send_prompt_to_target_async or resubmit the prompt with a new message ID.
  3. Explicit task lifecycle states:

    • Properly map failed and canceled task states to response_type="error".
    • For input-required, prioritize the question from status.message over partial artifacts, and eliminate fallback to raw str(result).
  4. Timeouts:

    • Configure HTTP client / request timeouts to respect task_timeout_seconds (or an explicit request timeout) so blocking sends are properly protected against slow initial responses.
  5. Target identity & context validation:

    • Include relevant non-secret routing/protocol settings in the target identifier.
    • Explicitly fail if multi-turn history exists but upstream context is missing, rather than silently dropping prior turns. Ensure capability declarations match adapter capabilities.
  6. Integration tests:

    • Add deterministic HTTP integration test coverage against a local SDK-based agent server covering multi-turn context continuity, input-required continuation, task failure states, and polling 429 without prompt resubmission.
    • Add an environment-gated integration test for Microsoft Foundry A2A endpoints (@pytest.mark.skipif(...)).

I'll push the updates shortly. Thanks again!

@rbrus

Copy link
Copy Markdown
Author

Also for additional context on motivation and broader ecosystem integration: I built agent-redteam-benchmark and actively develop agent-probe, using them to evaluate agent targets across diverse scenarios. I am also preparing a second red-teaming agent with more advanced orchestration strategies that will build upon this A2A target integration.

@rbrus

Copy link
Copy Markdown
Author

The updates are pushed in commit 0f76903:

  • Official a2a-sdk adoption: Added a2a optional dependency (a2a-sdk>=1.1.0) to pyproject.toml (synced into all), with lazy import error messaging (pip install pyrit[a2a]). Replaced raw JSON-RPC dialect implementation with a2a-sdk client transport and typed models.
  • Safe polling retries: Scoped 429 handling during task polling to retry get_task against the existing task ID without restarting _send_prompt_to_target_async or resubmitting the prompt.
  • Task lifecycle handling: Mapped failed, canceled, and rejected task states to response_type="error". For input-required tasks, the agent question is extracted from status.message (preventing partial artifacts from overwriting the question) without raw str(result) fallbacks.
  • Timeouts: Configured httpx.Timeout using request_timeout_seconds (defaulting to task_timeout_seconds), protecting initial blocking sends from default HTTPX timeouts.
  • Identity & capability limits: Added non-secret configuration (endpoint, protocol_version, task_timeout_seconds) to ComponentIdentifier. Explicitly raise ValueError if multi-turn history exists without an upstream context ID (with set_conversation_context available for manual restoration). Validated custom configuration capabilities to reject modalities other than text or system prompts.
  • Testing:
    • 20 unit tests covering typed message results, task completion, input-required questions and continuation, failed/canceled error mappings, content filter handling, polling 429 without resubmission, timeout handling, and capability validation.
    • 4 real HTTP integration tests against a local deterministic A2A server (multi-turn memory recall, input-required flow, failed task, polling 429 with verified single prompt submission) + environment-gated test for Microsoft Foundry.
    • Full prompt target unit suite: 1364 passed, 31 skipped; ruff check, ruff format, and ty check clean.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

FEAT: A2ATarget for Agent-to-Agent (A2A) protocol endpoints

2 participants