Preserve structured output stop reasons - #1850
sylvesterkaczmarek wants to merge 2 commits into
Conversation
|
@sylvesterkaczmarek the streaming path has the same problem, and this PR doesn't touch it. The stream accumulator parses at Offline repro (stub transport, the reply is cut at repro# Offline: messages.stream(output_format=Model) where the reply is cut off at
# max_tokens. Does the caller get stop_reason="max_tokens", or a raw pydantic
# ValidationError before stop_reason is known? Stub transport, synthetic SSE.
import json
import anthropic
try:
import httpx2 as httpx # anthropic>=1.0 ships on httpx2
except ImportError:
import httpx
import pydantic
from pydantic import BaseModel
class Rate(BaseModel):
lane: str
price: float
def sse(events):
out = []
for ev in events:
out.append(f"event: {ev['type']}\ndata: {json.dumps(ev)}\n\n")
return "".join(out).encode()
EVENTS = [
{"type": "message_start", "message": {"id": "msg_1", "type": "message", "role": "assistant",
"model": "claude-test", "content": [], "stop_reason": None, "stop_sequence": None,
"usage": {"input_tokens": 10, "output_tokens": 1}}},
{"type": "content_block_start", "index": 0, "content_block": {"type": "text", "text": ""}},
{"type": "content_block_delta", "index": 0, "delta": {"type": "text_delta", "text": '{"lane": "A-B", "pri'}},
{"type": "content_block_stop", "index": 0},
{"type": "message_delta", "delta": {"stop_reason": "max_tokens", "stop_sequence": None},
"usage": {"output_tokens": 8}},
{"type": "message_stop"},
]
def handler(request: httpx.Request) -> httpx.Response:
return httpx.Response(200, headers={"content-type": "text/event-stream"}, content=sse(EVENTS))
client = anthropic.Anthropic(api_key="sk-test", http_client=httpx.Client(transport=httpx.MockTransport(handler)), max_retries=0)
print("anthropic", anthropic.__version__)
try:
with client.messages.stream(model="claude-test", max_tokens=8, output_format=Rate,
messages=[{"role": "user", "content": "x"}]) as stream:
msg = stream.get_final_message()
print("stop_reason:", msg.stop_reason, "parsed_output:", msg.parsed_output)
except pydantic.ValidationError as e:
print("pydantic.ValidationError raised mid-stream, stop_reason never reached:", e.errors()[0]["type"])Same on 0.87.0. Could this PR move that parse to |
|
@HardMax71 Confirmed. The streaming path had the same event-ordering issue: content_block_stop runs before message_delta, so structured-output validation could fail before stop_reason was available. Updated in signed commit 32d3d1f:
Focused validation: 66 streaming/parsing tests passed, including public-stream regressions for max_tokens. |
Summary
Preserve structured-output responses when generation terminates with
stop_reason="refusal"or"max_tokens"instead of replacing the response with a schema-validation exception.messages.parse()and the beta equivalent currently attempt to validate every text block againstoutput_formatunconditionally. Refusal text and max-token-truncated JSON are not guaranteed to satisfy the requested output schema, so those terminal responses can fail inside Pydantic/JSON parsing before the caller can inspect the model's actualstop_reason.For example, a refusal containing ordinary refusal text is currently fed to
TypeAdapter.validate_json(), and a response truncated at{"value":is treated as malformed structured output. In both cases the SDK loses the more important information: why generation stopped.Fix
Skip structured-output validation only for:
refusalmax_tokensThe text remains available unchanged,
parsed_outputisNone, and the originalstop_reasonis preserved.Normal completed responses continue through the existing validation path, so malformed structured output on an ordinary completed turn still raises as before.
The behavior is applied symmetrically to GA and beta parsed messages.
Regression coverage
Adds tests verifying:
parsed_output=Nonefor GA and beta;max_tokensreturns withparsed_output=None;