Adaptive thinking and effort
The way you control how hard Claude thinks changed completely, and the old way does not fail gracefully. It fails with a 400.
Extended thinking is over
The old mechanism was explicit. You set a thinking object with type: "enabled" and a budget_tokens number, and you were effectively telling the model how many tokens it was allowed to spend reasoning before answering.
Its status now has three distinct states and you need all three:
Before 4.6: the working mechanism.
On Claude 4.6: deprecated. Still functions. Do not write new code against it.
On Claude 4.7 and later: rejected with 400. The request does not run.
Adaptive thinking
The replacement inverts who decides. A fixed budget_tokens was a prepaid meter: you decided in advance how much reasoning the job was worth, before you had seen the job. Adaptive thinking hands the meter to the model, so it decides how much thinking a given input warrants. Easy question, little thinking. Hard question, more.
This is strictly better in the general case, because a prepaid budget was always wrong in one of two directions: too small on the hard inputs where it mattered, or wasteful on the easy ones that dominate your volume. And your volume is nearly always dominated by easy ones.
One exception you must remember: claude-haiku-4-5-20251001 has no adaptive thinking. It is the only current-generation model without it. If a design needs deliberate multi-step reasoning and also wants the cheapest model, those two requirements are in conflict and the question is testing whether you notice.
effort is the dial you still have
You are not left with no control. output_config.effort is the current lever, and it defaults to high.
{
"model": "claude-opus-5",
"max_tokens": 4096,
"output_config": { "effort": "high" },
"messages": [{ "role": "user", "content": "..." }]
}The default being high has a direct cost consequence, and it is the kind that hides well because nothing looks wrong. Thorough deliberation is what you get unless you ask otherwise, so a high-volume classification endpoint that has never touched effort is paying for reasoning that a one-word label does not need, on every single request, for as long as nobody looks.
That is a genuine, easy cost lever and it belongs in the same mental drawer as caching and Batches. Turn it down where the task is simple; leave it alone where the task is hard.
Note also that effort and format live in the same output_config object, which is a tidy way to remember both. One object, two knobs: how hard it thinks, and what shape comes out.
Try it yourself
Extended thinking, version by version
This has three states across three model generations and the exam will pick the one you did not memorise.
What is the status of the old extended thinking parameters (thinking.type enabled plus budget_tokens) on Claude 4.6 versus 4.7 and later, and what replaced them?
Reveal answer
On Claude 4.6 the extended thinking parameters are DEPRECATED, so they still function but you should not write new code against them. On Claude 4.7 and later they are REJECTED WITH 400. Adaptive thinking replaced them: the model decides how much to think rather than you handing it a token budget. The control you have now is output_config.effort, which defaults to high.
The effort default
You send a request to claude-opus-5 with no effort set. What are you getting?
Show answer
Correct answer: D — effort of high, the thorough default
output_config.effort defaults to high. That is worth internalising because it means the expensive-and-thorough behaviour is what you get by default, so if you are running a high-volume, low-complexity workload and have never touched effort, you are paying for deliberation you did not ask for. A medium default would be the sensible guess, which is exactly why it is the distractor.
Same body, two models, two outcomes
This exact body is sent twice, once to claude-opus-4-6 and once to claude-opus-4-7. Nothing else changes.
{
"max_tokens": 4096,
"thinking": { "type": "enabled", "budget_tokens": 8000 },
"messages": [
{ "role": "user", "content": "Work through this proof step by step." }
]
}Show answer
Correct answer: A — 4.6 succeeds because the parameters are deprecated there, and 4.7 returns 400 because they are rejected there
Deprecated and rejected are two different states and 4.6 and 4.7 sit on opposite sides of the line. On 4.6 the extended thinking parameters still function, which is precisely why services built there keep working and their owners never notice; on 4.7 and later the same body returns 400. The tempting answer is that both succeed, because deprecation usually means a warning rather than a wall, and that expectation is what makes this a nasty migration bug: nothing in your code changed, only the model id did.
Porting a thinking budget forward
A service sets thinking with type: enabled and budget_tokens: 8000. You repoint it from claude-opus-4-6 to claude-opus-4-7. What breaks?
Show answer
Correct answer: B — The request returns 400, because those parameters are rejected on 4.7 and later
On 4.6 those parameters were deprecated, which is why the service worked before the repoint. On 4.7 and later they are rejected with 400. The silent-cap and silent-disable options are tempting because deprecated features usually fade rather than fail, and this is the second place in this course where Anthropic chose a hard rejection over a quiet degradation.
The cheapest change on a high-volume endpoint
A classification endpoint runs two million requests a month, returns a single label each time, and has never set effort. Which change most directly reduces its bill without changing the model?
Show answer
Correct answer: C — Turn effort down, since the endpoint is inheriting the high default on a task that needs no deliberation
The default is high, so a workload that never touched effort is buying thorough deliberation on every one of two million trivially easy requests. Turning it down is a one-line change that targets exactly the waste. Setting a thinking budget is the tempting answer because it describes the right intention, but that mechanism is deprecated on 4.6 and returns 400 on 4.7 and later, so it is the previous generation's version of this idea. Streaming does not change the bill at all, and raising max_tokens raises a ceiling rather than reducing what is generated.
Decide effort deliberately once
For each distinct workload in a design of yours, write down an effort level and a reason.
Every workload has an explicit effort value with a justification tied to task complexity and volume, rather than inheriting the high default by accident, and any workload running on claude-haiku-4-5-20251001 is noted as having no adaptive thinking at all.