Conversation state, system prompts, and the prefill rule
The API has no memory. None. Every call is the first call as far as the server is concerned, and the entire conversation is something you carry in and out yourself.
If that sounds familiar, it should. It is the same shape as any agent that wakes up with no idea what it did yesterday: continuity is not something the substrate provides, it is a file somebody has to carry in. Here that file is the messages array.
Stateless is a design decision, not an omission
There is no session, no conversation_id, no thread that lives on Anthropic’s side. To continue a conversation you resend it:
{
"model": "claude-sonnet-5",
"max_tokens": 1024,
"system": "You are a terse code reviewer.",
"messages": [
{ "role": "user", "content": "Review this diff." },
{ "role": "assistant", "content": "Two problems. First..." },
{ "role": "user", "content": "Fix the second one." }
]
}Two consequences fall straight out of this, and both are examinable.
Input tokens grow with the conversation. Turn twenty pays for turns one through nineteen all over again. That is not waste caused by a bug, it is how the protocol works, and it is precisely the bill that prompt caching exists to cut.
Your app owns persistence. If you want conversations to survive a process restart, that is your database, your schema, your problem. The Agent SDK and Managed Agents give you harnesses that do this for you, which is one of the real reasons to pick them over a hand-rolled loop.
The system prompt is configuration, not conversation
Because system is a top-level parameter it sits outside the alternation entirely. It does not count as a turn, it does not need a matching reply, and it is byte-identical on every request in a well-built app.
That last property is why it is the natural front of your cached prefix. Hold that thought until module four.
system also accepts an array of content blocks rather than a bare string, which is how you attach cache control to part of it while leaving the rest uncached.
Prefill: alive in the middle, dead at the end
Assistant prefill was the trick where you end your messages array with a partial assistant message and let the model continue it, forcing a format:
"messages": [
{ "role": "user", "content": "List three risks as JSON." },
{ "role": "assistant", "content": "[" }
]The replacement for last-turn prefill is structured outputs, which is a real feature with a real schema rather than a hack that leans on the model’s completion behaviour. That is module two, lesson three.
Why this changed at all
Prefill worked by exploiting the fact that the model just continues text. Newer models do more between turns, including thinking, and injecting a half-finished assistant turn in front of that is ambiguous at best. Rather than let it silently degrade, the API rejects it loudly.
Loud rejection is a gift. A 400 you can read beats a subtly worse answer you cannot detect, and that framing shows up more than once in this course.
Try it yourself
The prefill rule, stated precisely
This one is worth more marks than its size suggests, because it is a rule with a boundary and the boundary is the whole question.
On Claude 4.6 and later, what exactly happens if a request ends with an assistant message, and is assistant prefill therefore dead?
Reveal answer
A request whose FINAL message has role assistant returns a 400 on Claude 4.6 and later. That specific pattern, last-turn prefill, is what was removed. Assistant messages earlier in the conversation are completely normal and still required, because that is how you replay a multi-turn history and how you feed back a tool_use turn. So prefill as a steering trick on the last turn is gone; assistant messages as conversation history are untouched.
What statelessness actually costs you
The Messages API keeps no server-side conversation state. Which consequence follows directly from that?
Show answer
Correct answer: A — Every turn re-sends the whole prior conversation, so input tokens grow with conversation length
Stateless means the request body IS the conversation. You resend everything every turn, so input token count grows roughly linearly with history length, which is exactly why prompt caching exists and why it pays off most on long chats. There is no conversation_id parameter to supply, and the model refers back to earlier turns perfectly well, because you handed it those turns in the array.
Which of these two arrays runs
Both of these are sent to claude-opus-5. One of them returns 400.
"messages": [
{ "role": "user", "content": "List three risks as JSON." },
{ "role": "assistant", "content": "[" }
]"messages": [
{ "role": "user", "content": "Review this diff." },
{ "role": "assistant", "content": "Two problems. First..." },
{ "role": "user", "content": "Fix the second one." }
]Show answer
Correct answer: B — B runs, A fails, because A ends with an assistant message
Only the final position matters. Array A ends on an assistant message, which is last-turn prefill, and that returns 400 on Claude 4.6 and later. Array B contains an assistant message too, but it sits in the middle as ordinary conversation history, which is not merely allowed but required for any multi-turn call. The tempting answer is that both fail, because the rule gets remembered as a blanket ban on assistant messages rather than as a rule about the last element.
Why the system prompt is the natural cache prefix
You are adding prompt caching to a chatbot. Which part of the request is the obvious first thing to cache, and why?
Show answer
Correct answer: D — The system prompt and any static tool definitions, because they are byte-identical on every request
A cache hit needs an exact prefix match, so the only things worth caching are the parts that do not change between calls. System prompt and tool definitions sit at the front of every request and are identical every time, which makes them the textbook prefix. The newest user message is the tempting wrong answer, because it is the part you are thinking about, but it is different on every single request and can therefore never hit. Caching also applies to input, never output.
Finding the prefill rule in two minutes
A teammate insists assistant prefill still works because it worked on their last project. You have the Claude documentation open and two minutes to settle it. Where do you look first?
Show answer
Correct answer: C — The model-specific migration and deprecation notes for the version they are on
This is a behaviour that changed at a specific model version, so the authoritative place is the material that tracks version-to-version changes, where a removal is stated along with the version boundary it starts at. The Messages API reference is the tempting answer because the parameter in question is messages, but a reference page describes the current contract rather than telling you when and why it changed, so it will not settle an argument about which version broke it. Prompt engineering guidance may still describe prefill as a technique without flagging the version limit, which is exactly how a stale belief survives.
Audit your own turn loop
On paper, trace what your integration sends on turn five of a conversation.
Turn five sends the system parameter plus all eight prior messages (four user, four assistant) plus the new user message, in one array, in order, with roles strictly alternating around any tool_result pairs. No message has role system. The final element has role user, never assistant.