Streaming and what it costs you
Streaming is the most boring feature with the highest production impact. It changes nothing about what you pay and everything about whether your app feels broken.
It also hides one of the nastiest bugs in this whole course, and the bug does not throw. It just quietly returns the wrong string.
The event sequence
A streamed response is server-sent events. The envelope is stable enough to memorise in one sitting.
message_start -> the shell, with input usage
content_block_start -> index 0 begins
content_block_delta -> ...tokens...
content_block_stop -> index 0 done
(repeat per block)
message_delta -> final stop_reason and output usage
message_stop -> doneping events can show up anywhere and carry no meaning. An error event can arrive mid-stream, which is the part people forget: a stream that started successfully can still fail halfway, so a 200 on the initial response is not a guarantee of a complete answer.
Reassemble by index, never by concatenation
Picture three people dictating into one phone line, each announcing a seat number before they speak. If you write down every word in the order it arrives and ignore the seat numbers, you end up with one paragraph containing three people’s sentences shuffled together. Precisely: content blocks are indexed, several can appear in a single response, and the only thing distinguishing their deltas is the index carried on content_block_start.
So a model that writes a sentence and then calls two tools produces three block sequences interleaved in one stream.
If your client does buffer += delta.text you will be fine right up until the day a tool call appears, and then you will silently glue a tool’s JSON input onto the end of your prose. Key by index from content_block_start. Parse a tool’s arguments only once its content_block_stop has arrived, because the JSON arrives in fragments and is not valid until the last one.
Why you stream even when nobody is watching
The obvious reason is perceived latency. The less obvious one is that long generations are better served incrementally, and a non-streaming request has to hold a single connection open for the entire generation with nothing coming back down it. Timeouts, proxies and load balancers all hate that.
So the rule of thumb is: if the response could be long, stream it, even for a background job that has no human reading the output. You throw the deltas away and keep the final message. The cost is identical.
What streaming does not change
It does not change the price. It does not change the model’s behaviour, the sampling, the tool loop, or the stop reasons. It does not give you the ability to cancel mid-generation and pay less for what you already received.
It is a delivery mechanism sitting on top of exactly the same request body, with "stream": true added. If a question offers you streaming as an answer to a cost problem, the answer it wants is Batches or caching.
Try it yourself
The shape of a stream
The event names are easy to half-remember. What is worth holding is the envelope, because the envelope tells you where to hang your handlers.
What is the overall structure of a streamed Messages API response, and where do the actual tokens arrive?
Reveal answer
A stream opens with message_start, which carries the message shell including usage so far. Then, for each content block, a content_block_start, a run of content_block_delta events, and a content_block_stop. The tokens arrive in the deltas. It closes with message_delta, which carries the final stop_reason and the output token usage, then message_stop. ping events may appear anywhere and mean nothing. Because content blocks are indexed, a response containing text plus two tool calls streams as three separate block sequences that you reassemble by index.
Does streaming change your bill
You switch a high-volume endpoint from non-streaming to streaming. What happens to the cost per request?
Show answer
Correct answer: D — It is unchanged, because you pay for input and output tokens either way
Streaming is a transport choice. You pay for the same input and output tokens whichever way they arrive, and there is no per-event surcharge. The 50 percent figure is the tempting distractor because it is a real number, but it belongs to the Batches API, which is the opposite trade: you give up interactivity entirely in exchange for half price.
Predict what this accumulator produces
A working streaming client, written when the integration only ever returned prose.
let buffer = "";
for await (const event of stream) {
if (event.type === "content_block_delta") {
buffer += event.delta.text ?? event.delta.partial_json ?? "";
}
}
return buffer;The response now contains one text block followed by two tool_use blocks.
Show answer
Correct answer: A — The prose with both tools' JSON arguments glued onto the end of it
All content blocks stream down one connection, interleaved and distinguished only by their index. A single flat accumulator has no way to tell them apart, so it concatenates prose and partial tool JSON into one string that is neither valid prose nor valid JSON. The tempting answer is a thrown exception, because a corrupted result feels like it ought to fail loudly, and the actual behaviour is worse precisely because it does not: you get a plausible-looking string and discover the problem downstream. Key deltas by the index from content_block_start.
Where final usage lands in a stream
Your billing dashboard reads output token counts off the stream. Which event must it listen for?
Show answer
Correct answer: C — message_delta, near the end of the stream
message_start carries usage as it stands at the start, so input tokens are known but output tokens are not yet generated. The running deltas carry text, not counts. message_delta near the end is where the final stop_reason and output token usage arrive, which is why a client that hangs up early gets a bill it never saw. content_block_stop carries no usage at all.
What a 200 does and does not promise
A distinction people flatten, usually while writing the retry logic. It decides where your error handling has to live.
If the initial streaming response returns HTTP 200, what have you actually been guaranteed about the completeness of the answer?
Reveal answer
Almost nothing beyond the fact that the stream opened. An error event can arrive mid-stream, after any number of successful deltas, so a stream that started cleanly can still fail partway through. This means error handling cannot live only around the initial request; it has to live inside the event loop as well. The practical test for a complete response is that you saw message_delta with a final stop_reason and then message_stop, not that the HTTP status was 200.
Confirm you reassemble by index
Look at your own streaming client, or sketch one, and check how it handles a response with two tool calls in it.
Deltas are accumulated into a map or array keyed by the block index from content_block_start, never appended to a single flat string. A response containing one text block and two tool_use blocks reconstructs into three distinct blocks, and the tool inputs are parsed only after their content_block_stop arrives.