Whetstone.
The Messages API and Its MechanicsConversation state, system prompts, and the prefill rule
Module 1, Lesson 226 min

Conversation state, system prompts, and the prefill rule

The API has no memory. None. Every call is the first call as far as the server is concerned, and the entire conversation is something you carry in and out yourself.

If that sounds familiar, it should. It is the same shape as any agent that wakes up with no idea what it did yesterday: continuity is not something the substrate provides, it is a file somebody has to carry in. Here that file is the messages array.

Stateless is a design decision, not an omission

There is no session, no conversation_id, no thread that lives on Anthropic’s side. To continue a conversation you resend it:

Turn three carries turns one and two
{
  "model": "claude-sonnet-5",
  "max_tokens": 1024,
  "system": "You are a terse code reviewer.",
  "messages": [
    { "role": "user", "content": "Review this diff." },
    { "role": "assistant", "content": "Two problems. First..." },
    { "role": "user", "content": "Fix the second one." }
  ]
}

Two consequences fall straight out of this, and both are examinable.

Input tokens grow with the conversation. Turn twenty pays for turns one through nineteen all over again. That is not waste caused by a bug, it is how the protocol works, and it is precisely the bill that prompt caching exists to cut.

Your app owns persistence. If you want conversations to survive a process restart, that is your database, your schema, your problem. The Agent SDK and Managed Agents give you harnesses that do this for you, which is one of the real reasons to pick them over a hand-rolled loop.

The system prompt is configuration, not conversation

Because system is a top-level parameter it sits outside the alternation entirely. It does not count as a turn, it does not need a matching reply, and it is byte-identical on every request in a well-built app.

That last property is why it is the natural front of your cached prefix. Hold that thought until module four.

system also accepts an array of content blocks rather than a bare string, which is how you attach cache control to part of it while leaving the rest uncached.

Prefill: alive in the middle, dead at the end

Assistant prefill was the trick where you end your messages array with a partial assistant message and let the model continue it, forcing a format:

This pattern now returns 400 on Claude 4.6 and later
"messages": [
  { "role": "user", "content": "List three risks as JSON." },
  { "role": "assistant", "content": "[" }
]

The replacement for last-turn prefill is structured outputs, which is a real feature with a real schema rather than a hack that leans on the model’s completion behaviour. That is module two, lesson three.

Why this changed at all

Prefill worked by exploiting the fact that the model just continues text. Newer models do more between turns, including thinking, and injecting a half-finished assistant turn in front of that is ambiguous at best. Rather than let it silently degrade, the API rejects it loudly.

Loud rejection is a gift. A 400 you can read beats a subtly worse answer you cannot detect, and that framing shows up more than once in this course.

Practice

Try it yourself

Recall

The prefill rule, stated precisely

This one is worth more marks than its size suggests, because it is a rule with a boundary and the boundary is the whole question.

On Claude 4.6 and later, what exactly happens if a request ends with an assistant message, and is assistant prefill therefore dead?

Reveal answer

A request whose FINAL message has role assistant returns a 400 on Claude 4.6 and later. That specific pattern, last-turn prefill, is what was removed. Assistant messages earlier in the conversation are completely normal and still required, because that is how you replay a multi-turn history and how you feed back a tool_use turn. So prefill as a steering trick on the last turn is gone; assistant messages as conversation history are untouched.

Quiz

What statelessness actually costs you

The Messages API keeps no server-side conversation state. Which consequence follows directly from that?

  1. AEvery turn re-sends the whole prior conversation, so input tokens grow with conversation length
  2. BYou must supply a conversation_id on every request after the first to link the turns
  3. CLong conversations are impossible without a server-side database holding the history
  4. DThe model cannot refer back to anything said earlier unless you summarise it yourself
Show answer

Correct answer: A — Every turn re-sends the whole prior conversation, so input tokens grow with conversation length

Stateless means the request body IS the conversation. You resend everything every turn, so input token count grows roughly linearly with history length, which is exactly why prompt caching exists and why it pays off most on long chats. There is no conversation_id parameter to supply, and the model refers back to earlier turns perfectly well, because you handed it those turns in the array.

Quiz

Which of these two arrays runs

Both of these are sent to claude-opus-5. One of them returns 400.

Array A
"messages": [
  { "role": "user", "content": "List three risks as JSON." },
  { "role": "assistant", "content": "[" }
]
Array B
"messages": [
  { "role": "user", "content": "Review this diff." },
  { "role": "assistant", "content": "Two problems. First..." },
  { "role": "user", "content": "Fix the second one." }
]
  1. AA runs, B fails, because B has too many turns
  2. BB runs, A fails, because A ends with an assistant message
  3. CBoth fail, because assistant messages are no longer accepted
  4. DBoth run, because prefill is only discouraged and not blocked
Show answer

Correct answer: B — B runs, A fails, because A ends with an assistant message

Only the final position matters. Array A ends on an assistant message, which is last-turn prefill, and that returns 400 on Claude 4.6 and later. Array B contains an assistant message too, but it sits in the middle as ordinary conversation history, which is not merely allowed but required for any multi-turn call. The tempting answer is that both fail, because the rule gets remembered as a blanket ban on assistant messages rather than as a rule about the last element.

Quiz

Why the system prompt is the natural cache prefix

You are adding prompt caching to a chatbot. Which part of the request is the obvious first thing to cache, and why?

  1. AThe most recent user message, because it is the one being processed right now
  2. BThe assistant's previous reply, because it is already generated and will not change
  3. CNothing, because cache_control applies to output tokens and this request generates few
  4. DThe system prompt and any static tool definitions, because they are byte-identical on every request
Show answer

Correct answer: D — The system prompt and any static tool definitions, because they are byte-identical on every request

A cache hit needs an exact prefix match, so the only things worth caching are the parts that do not change between calls. System prompt and tool definitions sit at the front of every request and are identical every time, which makes them the textbook prefix. The newest user message is the tempting wrong answer, because it is the part you are thinking about, but it is different on every single request and can therefore never hit. Caching also applies to input, never output.

Quiz

Finding the prefill rule in two minutes

A teammate insists assistant prefill still works because it worked on their last project. You have the Claude documentation open and two minutes to settle it. Where do you look first?

  1. AThe Messages API reference entry for the messages parameter
  2. BThe prompt engineering guide's section on controlling output format
  3. CThe model-specific migration and deprecation notes for the version they are on
  4. DThe pricing page, to check whether prefill has a separate rate
Show answer

Correct answer: C — The model-specific migration and deprecation notes for the version they are on

This is a behaviour that changed at a specific model version, so the authoritative place is the material that tracks version-to-version changes, where a removal is stated along with the version boundary it starts at. The Messages API reference is the tempting answer because the parameter in question is messages, but a reference page describes the current contract rather than telling you when and why it changed, so it will not settle an argument about which version broke it. Prompt engineering guidance may still describe prefill as a technique without flagging the version limit, which is exactly how a stale belief survives.

Check

Audit your own turn loop

On paper, trace what your integration sends on turn five of a conversation.

You should see

Turn five sends the system parameter plus all eight prior messages (four user, four assistant) plus the new user message, in one array, in order, with roles strictly alternating around any tool_result pairs. No message has role system. The final element has role user, never assistant.

Sign in to track your progress →