The Messages API request shape
A third of this exam hangs off one endpoint. One. Learn its shape properly in the next twenty minutes and the biggest domain stops being scary, because most of what looks like a hard question turns out to be a shape question wearing a costume.
And the fastest way to lose those marks is muscle memory. Almost everyone arrives here having written a thousand OpenAI-shaped requests, and two of the habits transfer wrong.
The four keys that matter
A Messages API request is a JSON body. Three fields are required and the fourth is the one people forget is not a message.
{
"model": "claude-sonnet-5",
"max_tokens": 1024,
"messages": [
{ "role": "user", "content": "Summarise this in one line." }
]
}model, max_tokens, messages. That is the whole requirement. max_tokens has no default, which trips people who came from other providers where you can omit it and get whatever the model feels like.
And then system:
{
"model": "claude-sonnet-5",
"max_tokens": 1024,
"system": "You are a terse code reviewer.",
"messages": [{ "role": "user", "content": "Review this diff." }]
}Think of system as a dial on the outside of the box rather than an item inside it. It configures the call; it is not a turn in the conversation. Precisely: system is a top-level request parameter, a sibling of model and messages, and it never appears inside the messages array.
Content blocks are the real unit
content can be a plain string, which is sugar. Underneath, content is always an array of typed blocks, and once you do anything interesting you write the array yourself.
{
"role": "user",
"content": [{ "type": "text", "text": "Summarise this in one line." }]
}The block types you will meet: text, image, document, tool_use, tool_result, and thinking. An assistant response is also a list of blocks, which is why a single reply can contain a paragraph of text and two tool calls at the same time.
This is the second habit that transfers wrong. If your parser assumes content[0].text exists, it will work perfectly right up until the first time the model reaches for a tool, and then it will hand undefined to something downstream that was never designed to receive it. Select blocks by their type, never by their index.
stop_reason is your control flow
The response tells you why generation ended, and your code branches on it. There are seven documented values. Four carry almost all the weight day to day:
end_turn means Claude finished naturally. Ship it.
max_tokens means you cut it off. The text is truncated, not wrong. Raise the budget or ask for less.
stop_sequence means one of your stop_sequences fired.
tool_use means the response contains at least one tool_use block and Claude is waiting on you. This is not an error, it is the middle of a loop, and treating it as terminal is the classic bug in a hand-rolled integration.
The other three are rarer but real, and a switch that does not handle them falls through to a default branch that was written for something else:
pause_turn means a server-tool loop hit its iteration limit, ten by default. Append the assistant turn and send again; the server picks up where it stopped. Do not add a “continue” message, and do not treat it as failure.
refusal means a safety classifier declined. It arrives as a normal HTTP 200, not an error, with a stop_details object naming the policy category. The practical consequence is the one that bites: code that reads content[0] unconditionally breaks here, because on a pre-output refusal content is empty.
model_context_window_exceeded means the context window filled, which is a different failure from max_tokens and wants a different fix: compact or split the conversation rather than raising a budget.
Usage comes back alongside it: input tokens, output tokens, and, once caching is in play, separate counts for cache creation and cache reads. Those fields are where every cost lesson later in this course lands, so notice now that the API tells you what you spent on every single call. You never have to guess.
Try it yourself
Where the system prompt lives
This card is about a cross-provider habit, not about difficulty. The trap is that the wrong answer is the one your fingers already know.
In the Messages API, where does the system prompt go, and why is putting it in the messages array wrong?
Reveal answer
The system prompt is a top-level system parameter on the request body, a sibling of model, messages and max_tokens. It is not a message, and it can never be messages[0]. This differs from the OpenAI chat-completions convention where system is the first element of the messages list, which is exactly why it gets tested. One refinement worth carrying, because the blunt "there is no system role" version is now out of date. Newer models (Fable 5, Mythos 5, Opus 4.8, Opus 5, but not Sonnet 5) do accept a system role inside messages for a mid-conversation operator instruction, with no beta header. It must follow a user turn and it still cannot come first, so the answer to "where does the system prompt go" is unchanged.
Which field is not optional
You are reviewing a colleague's minimal request body. Which of these can you not leave out?
Show answer
Correct answer: A — max_tokens
max_tokens is required on every Messages API request: it is the hard cap on the response, and there is no default for you to fall back on. system, temperature and stop_sequences are all genuinely optional. temperature is the tempting answer because people assume sampling must be configured, but it has a default, and on Opus 4.7 and later setting it to a non-default value is actively rejected.
Spot the request that will not run
One of these bodies is rejected before the model generates a single token. Read them properly rather than skimming.
{
"model": "claude-sonnet-5",
"max_tokens": 1024,
"system": "You are a terse code reviewer.",
"messages": [
{ "role": "user", "content": "Review this diff." }
]
}{
"model": "claude-sonnet-5",
"max_tokens": 1024,
"messages": [
{ "role": "system", "content": "You are a terse code reviewer." },
{ "role": "user", "content": "Review this diff." }
]
}Show answer
Correct answer: C — B is invalid, because a system message can never be the first entry in messages
Request A is the correct shape: system is a top-level parameter sitting beside model and max_tokens. Request B fails on position. Be careful with the reason here, because the once-standard phrasing (there is no system role at all) has gone stale: newer models do accept a system role inside messages for a mid-conversation instruction, but it must follow a user turn and can never be messages[0], which is exactly where B puts it. B is also on claude-sonnet-5, which does not support mid-conversation system messages at all, so it fails twice over. The tempting answer is that both work, because B is precisely the shape that is valid on OpenAI's chat-completions endpoint and it reads as perfectly sensible JSON.
Predict what this parser does
This helper has shipped happily for weeks against a text-only integration. Today you added your first tool.
function getReply(response: Message): string {
return response.content[0].text;
}Claude now returns an assistant message whose content is a tool_use block first, then a text block.
Show answer
Correct answer: B — It returns undefined or throws, because content[0] is a tool_use block with no text field
Content is an ordered array of typed blocks and nothing guarantees a text block sits at index 0. When the first block is a tool_use, reading .text off it gives you undefined, which then propagates into whatever consumes the return value. The tempting answer is that the SDK normalises the order for you, because that would be a friendly thing for an SDK to do, and it is exactly the assumption that makes this bug survive code review. Find the block you want by its type, never by its position.
Reading stop_reason
Your request came back with stop_reason set to tool_use. What does your code do next?
Show answer
Correct answer: C — Execute the requested tool, then send a new request with a user message carrying the tool_result blocks
tool_use means Claude has stopped mid-task to ask you to run something. You execute it and continue the conversation by appending the assistant message you just received plus a new user message containing the matching tool_result blocks. Increasing max_tokens is the distractor built from confusing this with the max_tokens stop_reason, which genuinely means you truncated the response and may want a bigger budget.
Write a minimal request from memory
In a scratch file, write the smallest valid Messages API request body you can, with no optional fields at all.
Exactly four keys, model with a full dated or current model id, max_tokens, and a messages array containing one object with role user and a string or content-block content. No system key, no temperature, and nothing with a system role inside messages.