Whetstone.
The Messages API and Its MechanicsStreaming and what it costs you
Module 1, Lesson 322 min

Streaming and what it costs you

Streaming is the most boring feature with the highest production impact. It changes nothing about what you pay and everything about whether your app feels broken.

It also hides one of the nastiest bugs in this whole course, and the bug does not throw. It just quietly returns the wrong string.

The event sequence

A streamed response is server-sent events. The envelope is stable enough to memorise in one sitting.

The order, every time
message_start          -> the shell, with input usage
  content_block_start  -> index 0 begins
  content_block_delta  -> ...tokens...
  content_block_stop   -> index 0 done
  (repeat per block)
message_delta          -> final stop_reason and output usage
message_stop           -> done

ping events can show up anywhere and carry no meaning. An error event can arrive mid-stream, which is the part people forget: a stream that started successfully can still fail halfway, so a 200 on the initial response is not a guarantee of a complete answer.

Reassemble by index, never by concatenation

Picture three people dictating into one phone line, each announcing a seat number before they speak. If you write down every word in the order it arrives and ignore the seat numbers, you end up with one paragraph containing three people’s sentences shuffled together. Precisely: content blocks are indexed, several can appear in a single response, and the only thing distinguishing their deltas is the index carried on content_block_start.

So a model that writes a sentence and then calls two tools produces three block sequences interleaved in one stream.

If your client does buffer += delta.text you will be fine right up until the day a tool call appears, and then you will silently glue a tool’s JSON input onto the end of your prose. Key by index from content_block_start. Parse a tool’s arguments only once its content_block_stop has arrived, because the JSON arrives in fragments and is not valid until the last one.

Why you stream even when nobody is watching

The obvious reason is perceived latency. The less obvious one is that long generations are better served incrementally, and a non-streaming request has to hold a single connection open for the entire generation with nothing coming back down it. Timeouts, proxies and load balancers all hate that.

So the rule of thumb is: if the response could be long, stream it, even for a background job that has no human reading the output. You throw the deltas away and keep the final message. The cost is identical.

What streaming does not change

It does not change the price. It does not change the model’s behaviour, the sampling, the tool loop, or the stop reasons. It does not give you the ability to cancel mid-generation and pay less for what you already received.

It is a delivery mechanism sitting on top of exactly the same request body, with "stream": true added. If a question offers you streaming as an answer to a cost problem, the answer it wants is Batches or caching.

Practice

Try it yourself

Recall

The shape of a stream

The event names are easy to half-remember. What is worth holding is the envelope, because the envelope tells you where to hang your handlers.

What is the overall structure of a streamed Messages API response, and where do the actual tokens arrive?

Reveal answer

A stream opens with message_start, which carries the message shell including usage so far. Then, for each content block, a content_block_start, a run of content_block_delta events, and a content_block_stop. The tokens arrive in the deltas. It closes with message_delta, which carries the final stop_reason and the output token usage, then message_stop. ping events may appear anywhere and mean nothing. Because content blocks are indexed, a response containing text plus two tool calls streams as three separate block sequences that you reassemble by index.

Quiz

Does streaming change your bill

You switch a high-volume endpoint from non-streaming to streaming. What happens to the cost per request?

  1. AIt drops, because you can stop early once you have enough
  2. BIt rises slightly, because of the per-event overhead
  3. CIt drops by 50 percent, the same discount as Batches
  4. DIt is unchanged, because you pay for input and output tokens either way
Show answer

Correct answer: D — It is unchanged, because you pay for input and output tokens either way

Streaming is a transport choice. You pay for the same input and output tokens whichever way they arrive, and there is no per-event surcharge. The 50 percent figure is the tempting distractor because it is a real number, but it belongs to the Batches API, which is the opposite trade: you give up interactivity entirely in exchange for half price.

Quiz

Predict what this accumulator produces

A working streaming client, written when the integration only ever returned prose.

The accumulator
let buffer = "";
for await (const event of stream) {
  if (event.type === "content_block_delta") {
    buffer += event.delta.text ?? event.delta.partial_json ?? "";
  }
}
return buffer;

The response now contains one text block followed by two tool_use blocks.

  1. AThe prose with both tools' JSON arguments glued onto the end of it
  2. BOnly the prose, because tool deltas are delivered on a separate stream
  3. CAn exception, because the two block types cannot share a stream
  4. DThree separate strings, because the SDK splits the buffer per block
Show answer

Correct answer: A — The prose with both tools' JSON arguments glued onto the end of it

All content blocks stream down one connection, interleaved and distinguished only by their index. A single flat accumulator has no way to tell them apart, so it concatenates prose and partial tool JSON into one string that is neither valid prose nor valid JSON. The tempting answer is a thrown exception, because a corrupted result feels like it ought to fail loudly, and the actual behaviour is worse precisely because it does not: you get a plausible-looking string and discover the problem downstream. Key deltas by the index from content_block_start.

Quiz

Where final usage lands in a stream

Your billing dashboard reads output token counts off the stream. Which event must it listen for?

  1. Amessage_start, which carries the full usage object
  2. BEach content_block_delta, summed
  3. Cmessage_delta, near the end of the stream
  4. Dcontent_block_stop for the final block
Show answer

Correct answer: C — message_delta, near the end of the stream

message_start carries usage as it stands at the start, so input tokens are known but output tokens are not yet generated. The running deltas carry text, not counts. message_delta near the end is where the final stop_reason and output token usage arrive, which is why a client that hangs up early gets a bill it never saw. content_block_stop carries no usage at all.

Recall

What a 200 does and does not promise

A distinction people flatten, usually while writing the retry logic. It decides where your error handling has to live.

If the initial streaming response returns HTTP 200, what have you actually been guaranteed about the completeness of the answer?

Reveal answer

Almost nothing beyond the fact that the stream opened. An error event can arrive mid-stream, after any number of successful deltas, so a stream that started cleanly can still fail partway through. This means error handling cannot live only around the initial request; it has to live inside the event loop as well. The practical test for a complete response is that you saw message_delta with a final stop_reason and then message_stop, not that the HTTP status was 200.

Check

Confirm you reassemble by index

Look at your own streaming client, or sketch one, and check how it handles a response with two tool calls in it.

You should see

Deltas are accumulated into a map or array keyed by the block index from content_block_start, never appended to a single flat string. A response containing one text block and two tool_use blocks reconstructs into three distinct blocks, and the tool inputs are parsed only after their content_block_stop arrives.

Sign in to track your progress →