Whetstone.
Choosing a ModelAdaptive thinking and effort
Module 3, Lesson 324 min

Adaptive thinking and effort

The way you control how hard Claude thinks changed completely, and the old way does not fail gracefully. It fails with a 400.

Extended thinking is over

The old mechanism was explicit. You set a thinking object with type: "enabled" and a budget_tokens number, and you were effectively telling the model how many tokens it was allowed to spend reasoning before answering.

Its status now has three distinct states and you need all three:

Before 4.6: the working mechanism.

On Claude 4.6: deprecated. Still functions. Do not write new code against it.

On Claude 4.7 and later: rejected with 400. The request does not run.

Adaptive thinking

The replacement inverts who decides. A fixed budget_tokens was a prepaid meter: you decided in advance how much reasoning the job was worth, before you had seen the job. Adaptive thinking hands the meter to the model, so it decides how much thinking a given input warrants. Easy question, little thinking. Hard question, more.

This is strictly better in the general case, because a prepaid budget was always wrong in one of two directions: too small on the hard inputs where it mattered, or wasteful on the easy ones that dominate your volume. And your volume is nearly always dominated by easy ones.

One exception you must remember: claude-haiku-4-5-20251001 has no adaptive thinking. It is the only current-generation model without it. If a design needs deliberate multi-step reasoning and also wants the cheapest model, those two requirements are in conflict and the question is testing whether you notice.

effort is the dial you still have

You are not left with no control. output_config.effort is the current lever, and it defaults to high.

effort lives on output_config, alongside format
{
  "model": "claude-opus-5",
  "max_tokens": 4096,
  "output_config": { "effort": "high" },
  "messages": [{ "role": "user", "content": "..." }]
}

The default being high has a direct cost consequence, and it is the kind that hides well because nothing looks wrong. Thorough deliberation is what you get unless you ask otherwise, so a high-volume classification endpoint that has never touched effort is paying for reasoning that a one-word label does not need, on every single request, for as long as nobody looks.

That is a genuine, easy cost lever and it belongs in the same mental drawer as caching and Batches. Turn it down where the task is simple; leave it alone where the task is hard.

Note also that effort and format live in the same output_config object, which is a tidy way to remember both. One object, two knobs: how hard it thinks, and what shape comes out.

Practice

Try it yourself

Recall

Extended thinking, version by version

This has three states across three model generations and the exam will pick the one you did not memorise.

What is the status of the old extended thinking parameters (thinking.type enabled plus budget_tokens) on Claude 4.6 versus 4.7 and later, and what replaced them?

Reveal answer

On Claude 4.6 the extended thinking parameters are DEPRECATED, so they still function but you should not write new code against them. On Claude 4.7 and later they are REJECTED WITH 400. Adaptive thinking replaced them: the model decides how much to think rather than you handing it a token budget. The control you have now is output_config.effort, which defaults to high.

Quiz

The effort default

You send a request to claude-opus-5 with no effort set. What are you getting?

  1. Ano effort setting, so the model does not deliberate at all
  2. Beffort of low, the conservative default
  3. Ceffort of medium, the balanced default
  4. Deffort of high, the thorough default
Show answer

Correct answer: D — effort of high, the thorough default

output_config.effort defaults to high. That is worth internalising because it means the expensive-and-thorough behaviour is what you get by default, so if you are running a high-volume, low-complexity workload and have never touched effort, you are paying for deliberation you did not ask for. A medium default would be the sensible guess, which is exactly why it is the distractor.

Quiz

Same body, two models, two outcomes

This exact body is sent twice, once to claude-opus-4-6 and once to claude-opus-4-7. Nothing else changes.

The body, unchanged between the two calls
{
  "max_tokens": 4096,
  "thinking": { "type": "enabled", "budget_tokens": 8000 },
  "messages": [
    { "role": "user", "content": "Work through this proof step by step." }
  ]
}
  1. A4.6 succeeds because the parameters are deprecated there, and 4.7 returns 400 because they are rejected there
  2. BBoth succeed, because a deprecated parameter keeps functioning until the model it belongs to is retired
  3. CBoth return 400, because the extended thinking parameters were removed outright when 4.6 shipped
  4. D4.6 returns 400 and 4.7 succeeds, because adaptive thinking restored budget_tokens when it arrived at 4.7
Show answer

Correct answer: A — 4.6 succeeds because the parameters are deprecated there, and 4.7 returns 400 because they are rejected there

Deprecated and rejected are two different states and 4.6 and 4.7 sit on opposite sides of the line. On 4.6 the extended thinking parameters still function, which is precisely why services built there keep working and their owners never notice; on 4.7 and later the same body returns 400. The tempting answer is that both succeed, because deprecation usually means a warning rather than a wall, and that expectation is what makes this a nasty migration bug: nothing in your code changed, only the model id did.

Quiz

Porting a thinking budget forward

A service sets thinking with type: enabled and budget_tokens: 8000. You repoint it from claude-opus-4-6 to claude-opus-4-7. What breaks?

  1. ANothing, because thinking.type and budget_tokens are forward compatible
  2. BThe request returns 400, because those parameters are rejected on 4.7 and later
  3. CThe budget is silently capped at the model's maximum and the request succeeds
  4. DThinking is disabled and budget_tokens ignored, but the request still succeeds
Show answer

Correct answer: B — The request returns 400, because those parameters are rejected on 4.7 and later

On 4.6 those parameters were deprecated, which is why the service worked before the repoint. On 4.7 and later they are rejected with 400. The silent-cap and silent-disable options are tempting because deprecated features usually fade rather than fail, and this is the second place in this course where Anthropic chose a hard rejection over a quiet degradation.

Quiz

The cheapest change on a high-volume endpoint

A classification endpoint runs two million requests a month, returns a single label each time, and has never set effort. Which change most directly reduces its bill without changing the model?

  1. AEnable streaming, so the response can be cut short as soon as the label token appears
  2. BSet thinking.budget_tokens, so deliberation is capped rather than left unbounded on each request
  3. CTurn effort down, since the endpoint is inheriting the high default on a task that needs no deliberation
  4. DRaise max_tokens, so a label truncated mid-word never has to be regenerated on a second call
Show answer

Correct answer: C — Turn effort down, since the endpoint is inheriting the high default on a task that needs no deliberation

The default is high, so a workload that never touched effort is buying thorough deliberation on every one of two million trivially easy requests. Turning it down is a one-line change that targets exactly the waste. Setting a thinking budget is the tempting answer because it describes the right intention, but that mechanism is deprecated on 4.6 and returns 400 on 4.7 and later, so it is the previous generation's version of this idea. Streaming does not change the bill at all, and raising max_tokens raises a ceiling rather than reducing what is generated.

Check

Decide effort deliberately once

For each distinct workload in a design of yours, write down an effort level and a reason.

You should see

Every workload has an explicit effort value with a justification tied to task complexity and volume, rather than inheriting the high default by accident, and any workload running on claude-haiku-4-5-20251001 is noted as having no adaptive thinking at all.

Sign in to track your progress →