Streaming and long turns: the 200 is not the finish line
Anthropic API
This page covers tools outside your selection. You can still read it. Find matching guides
Streaming trades one big wait for many small events, and moves failure to after the response starts. The event flow and the mid-stream error are the two things to get right.
A long non-streaming request is a single bet: wait, and either a full response arrives or the connection dies with nothing to show. Streaming changes the shape. "When creating a Message, you can set "stream": true to incrementally stream the response using server-sent events (SSE)". With the different shape come the things a naive integration gets wrong: the event flow, and where failure now lives.
This is the streaming companion to error handling, which covers the retryable -versus-not sorting; here the focus is the stream itself.
The event flow, enough to consume it
A stream is not a text firehose; it is a structured sequence. The spine: "message_start : contains a Message object with empty content", then a series of content blocks — each a content_block_start, some content_block_delta events, a content_block_stop — then one or more message_delta events and a final message_stop. You assemble the answer by accumulating deltas onto the block at each index.
A couple of flow facts save real debugging time. First, "Event streams may also include any number of ping events" — keep-alives with no content, and a consumer that does not expect them treats a normal stream as malformed. Second, on usage accounting: "The token counts shown in the usage field of the message_delta event are cumulative", so you read the running total, not a per-event increment to sum yourself.
For structured output, deltas carry chunked partial JSON; accumulate the string and parse once the content_block_stop for that block arrives, rather than trying to parse each fragment.
Failure moved, and that is the whole point
The trap streaming introduces is the one error handling flags from the other side: success is reported early, and the failure can come later. "The API may occasionally send errors in the event stream": an error event, arriving after the initial 200, on its own schedule. A handler that checks the HTTP status and moves on has already called a half-delivered response a success.
The documented recovery is plain and correct: "Save all content that was successfully received before the error occurred", then decide whether to retry or surface what you have. Which means your consumer must hold the partial response as it accumulates — not only for display, but so that a mid-stream error leaves you with something rather than nothing.
Do not hand-roll what the SDK assembles
The single most useful instruction on the page is to lean on the SDK's "built-in message accumulation and error handling capabilities". The SDKs turn the event sequence into an accumulated message and handle the error events, which is exactly the fiddly, easy-to-get-subtly-wrong part. Reassembling deltas by hand is a choice to own that code; make it deliberately, and only if you cannot use an SDK.
Streaming also earns its keep on the long-turn problem from error handling: networks drop idle connections, so a large max_tokens without streaming can fail in ways that look like API errors and are not. A stream produces events throughout, which keeps the connection alive and turns "did it hang?" into visible progress. For anything genuinely long-running, the Batches API is the other answer.
What goes wrong
Trusting the 200. The status said success; the error event came three seconds later, inside the stream. Status codes are necessary and not sufficient for a stream.
Choking on ping events. A keep-alive with no content read as a protocol violation, because the parser assumed every event carries text.
Summing cumulative usage. Treating each message_delta's token count as an increment double-counts wildly. It is already the total.
Discarding the partial on error. No accumulation buffer, so a mid-stream failure throws away everything received and turns a recoverable blip into a blank screen.
Hand-assembling deltas for no reason. Reimplementing the SDK's accumulator, then owning its edge cases — chunked JSON, block indices, fallback content blocks — in production.
How to check it worked
Force the ugly case, do not wait for it. Interrupt a stream mid-flight (kill the connection, or use a test hook that emits an error event after a few deltas) and confirm two things: your consumer still holds the partial content, and it reports a failure rather than a success. A streaming client that has only ever been watched completing cleanly is untested exactly where streaming differs from a plain call.
Sources
- Streaming Messages — Claude Platform Docs Tier 1 2026-09-04
Something wrong with this page?
Say what you expected and what you got. That is usually the shortest route to a correction, and it goes on the public issue tracker so the fix is visible.