FEAT: Add A2ATarget for Agent-to-Agent protocol communication - #2771
Radoslaw Brus (rbrus) wants to merge 6 commits into
Conversation
Introduces A2ATarget, enabling PyRIT to red-team multi-agent systems and services communicating over the Agent-to-Agent (A2A) protocol. Key features: - Dual-spec protocol negotiation: supports A2A v0.3 (message/send) with automatic fallback to v0.2 (tasks/send) on JSON-RPC -32601. - Multi-turn conversation task continuity: tracks and preserves active server task IDs per conversation session. - Robust response and refusal handling: captures JSON-RPC error refusals as conversational evidence rather than raising unhandled exceptions. - Agent card discovery: provides get_agent_card_async() to inspect declared agent capabilities at standard discovery paths. - Auth support: handles Bearer token and custom API-Key headers. - Comprehensive unit test suite covering v0.3, v0.2 fallback, task continuity, refusals, and agent card discovery.
|
@microsoft-github-policy-service agree |
5293bb3 to
439d91f
Compare
Richard Lundeen (richlundeen)
left a comment
There was a problem hiding this comment.
Thanks for adding this! We are interested in including A2A support in PyRIT. The overall direction looks good, and the target follows the main PyRIT target contract.
The main change I would like before merging is to use the official A2A Python SDK (a2a-sdk), rather than maintain our own protocol implementation. It is published on PyPI and supports A2A v1.0 with v0.3 compatibility. Using its typed messages, agent-card discovery, and client transport would reduce the version-specific code we need to maintain. A2ATarget can remain a thin adapter that owns PyRIT message mapping, identifiers/capabilities, upstream context tracking, and error mapping. An optional a2a extra would fit our existing dependency pattern.
A few gaps to address during that change:
- Safe polling retries: A 429 from
tasks/getcurrently retries the whole send method, which submits the prompt again with a new message ID. Retry the poll against the existing task instead of creating duplicate agent work. - Task result handling: Failed or canceled tasks currently become ordinary text responses with no error flag. An
input-requiredtask with partial artifacts can also return those artifacts instead of the question in its status message. Handle these states explicitly, and do not usestr(result)as an agent answer when there is no supported text. - Timeouts: The blocking send uses HTTPX's five-second default timeout. The 120-second task timeout only starts after that response arrives, so it does not protect a slow initial request.
- Protocol versions: The
v02branch does not match the published v0.2.5 format, which already usesmessage/sendandkind. The0.3+claim also needs narrowing because v1.0 has breaking changes. This is a strong reason to let the SDK handle protocol compatibility. - Identity and capability limits: Include non-secret protocol/routing settings that affect behavior in the target identifier. If stored history exists but the upstream context is missing, fail explicitly or provide a supported restoration path instead of continuing with only the latest message. Capability overrides should not advertise features the adapter cannot send.
- Integration coverage: We should have a real HTTP integration test against a deterministic SDK-based agent, not only mocked responses. Cover multi-turn context, input-required continuation, failed tasks, and a polling 429 that does not resubmit the prompt. An environment-gated test against an owned Foundry agent would also be useful for real authentication and service interoperability.
Requesting changes mainly for the SDK direction and integration coverage. This is a useful addition, and I would like to get it into PyRIT with the protocol details handled by the maintained SDK.
|
Thanks for the thorough and constructive review Richard Lundeen (@richlundeen)! I completely agree with the direction. Adopting the official I will update the PR with the following changes:
I'll push the updates shortly. Thanks again! |
|
Also for additional context on motivation and broader ecosystem integration: I built agent-redteam-benchmark and actively develop agent-probe, using them to evaluate agent targets across diverse scenarios. I am also preparing a second red-teaming agent with more advanced orchestration strategies that will build upon this A2A target integration. |
|
The updates are pushed in commit 0f76903:
|
Closes #2824
Summary
This PR adds
A2ATarget, so PyRIT can red-team agents exposed over the Agent-to-Agent (A2A) protocol, such as Microsoft Foundry Agent Service incoming A2A endpoints and Google ADK agents.Key Capabilities
message/send) by default. If the agent returns JSON-RPC-32601(method not found), it falls back to 0.2 (tasks/send) and pins that dialect.contextIdin 0.3,sessionIdin 0.2), so the agent keeps its own server-side state across turns. A task waiting for input (input-required/auth-required) is continued with its task ID; completed tasks are not reused.reset_conversation_asyncforgets a conversation's context.configuration.blocking, and tasks stillsubmittedorworkingare polled withtasks/getup totask_timeout_seconds.429, and429s the agent relays from its model as JSON-RPC errors, raiseRateLimitExceptionand are retried withpyrit_target_retry.response_type="error"), markedblockedwhen they matchCONTENT_FILTER_MARKERSandunknownotherwise.auth_token) and API keys (api_keywith a configurableapi_key_header).System prompts and prepended conversations are handled by PyRIT's existing normalization pipeline; the target declares multi-turn support but not system prompts or editable history.
Testing
Unit tests in
tests/unit/prompt_target/target/test_a2a_target.pycover 0.3 success,Messageresults, task polling and timeout, context continuity, input-required continuation, conversation isolation, reset, 0.2 fallback and session continuity, error and content-filter responses, HTTP and relayed rate-limit retries, and auth headers.pytest tests/unit/prompt_target: 1360 passed, 31 skippedruff check,ruff format,ty checkon the changed files: cleanLive test against a Microsoft Foundry Agent Service prompt agent (gpt-5-nano) with incoming A2A enabled (A2A 0.3, JSON-RPC, Entra bearer auth):
contextId, and the agent recalled the ID429s were retried; 27 of 30 completed with retry waits capped at 30s for the testNote: when Azure content filtering blocks a prompt, Foundry's A2A endpoint returns a generic JSON-RPC
-32603 "Received 400 from a service request"with no content-filter detail, so those responses are classified asunknownrather thanblocked.