Adaptive thinking and effort
The way you control how hard Claude thinks changed completely, and the old way does not fail gracefully. It fails with a 400.
Extended thinking is over
The old mechanism was explicit. You set a thinking object with type: "enabled" and a budget_tokens number, and you were effectively telling the model how many tokens it was allowed to spend reasoning before answering.
Its status now has three distinct states and you need all three:
Before 4.6: the working mechanism.
On Claude 4.6: deprecated. Still functions. Do not write new code against it.
On Claude 4.7 and later: rejected with 400. The request does not run.
Adaptive thinking
The replacement inverts who decides. A fixed budget_tokens was a prepaid meter: you decided in advance how much reasoning the job was worth, before you had seen the job. Adaptive thinking hands the meter to the model, so it decides how much thinking a given input warrants. Easy question, little thinking. Hard question, more.
This is strictly better in the general case, because a prepaid budget was always wrong in one of two directions: too small on the hard inputs where it mattered, or wasteful on the easy ones that dominate your volume. And your volume is nearly always dominated by easy ones.
One exception you must remember: claude-haiku-4-5-20251001 has no adaptive thinking. It is the only current-generation model without it. If a design needs deliberate multi-step reasoning and also wants the cheapest model, those two requirements are in conflict and the question is testing whether you notice.
effort is the dial you still have
You are not left with no control. output_config.effort is the current lever, and it defaults to high.
The default being high has a direct cost consequence, and it is the kind that hides well because nothing looks wrong. Thorough deliberation is what you get unless you ask otherwise, so a high-volume classification endpoint that has never touched effort is paying for reasoning that a one-word label does not need, on every single request, for as long as nobody looks.
That is a genuine, easy cost lever and it belongs in the same mental drawer as caching and Batches. Turn it down where the task is simple; leave it alone where the task is hard.
Note also that effort and format live in the same output_config object, which is a tidy way to remember both. One object, two knobs: how hard it thinks, and what shape comes out.
Try it yourself
Extended thinking, version by version
This has three states across three model generations and the exam will pick the one you did not memorise.
What is the status of the old extended thinking parameters (thinking.type enabled plus budget_tokens) on Claude 4.6 versus 4.7 and later, and what replaced them?
The effort default
You send a request to claude-opus-5 with no effort set. What are you getting?
Same body, two models, two outcomes
This exact body is sent twice, once to claude-opus-4-6 and once to claude-opus-4-7. Nothing else changes.
{
"max_tokens": 4096,
"thinking": { "type": "enabled", "budget_tokens": 8000 },
"messages": [
{ "role": "user", "content": "Work through this proof step by step." }
]
}
Porting a thinking budget forward
A service sets thinking with type: enabled and budget_tokens: 8000. You repoint it from claude-opus-4-6 to claude-opus-4-7. What breaks?
The cheapest change on a high-volume endpoint
A classification endpoint runs two million requests a month, returns a single label each time, and has never set effort. Which change most directly reduces its bill without changing the model?
Decide effort deliberately once
For each distinct workload in a design of yours, write down an effort level and a reason.
Every workload has an explicit effort value with a justification tied to task complexity and volume, rather than inheriting the high default by accident, and any workload running on claude-haiku-4-5-20251001 is noted as having no adaptive thinking at all.