Whetstone.
Cost and Token ManagementPricing and the multiplier stack
Module 4, Lesson 124 min

Pricing and the multiplier stack

Cost and Token Management is only 2.8 percent of the exam, which is about one or two questions. They will be arithmetic questions with exact numbers in them, so they are free marks if you learned the table and lost marks if you did not. Twenty minutes now, permanently.

The per-million-token table

Input first, output second, dollars per million tokens.

Fable 5 and Mythos 5: 10 and 50. The premium tier. Mythos is invite-only.

Opus 5, and the whole Opus 4.x line: 5 and 25. Opus 5 costs the same as its predecessors, which is a nice detail: upgrading within Opus is free at the price level.

Sonnet 5: 2 and 10 through 2026-08-31, then 3 and 15.

Sonnet 4.6 and Sonnet 4.5: 3 and 15.

Haiku 4.5: 1 and 5.

Notice the pattern under all of it: output is five times input, everywhere. Fable 10 to 50, Opus 5 to 25, Sonnet 3 to 15, Haiku 1 to 5. That regularity is easier to remember than ten separate numbers, and it tells you that verbosity is a cost problem five times bigger than context is.

The multipliers

Now the layer on top. These are the numbers that turn a price into a bill.

Batches: minus 50 percent. The single biggest lever available, and the only cost of it is that the work stops being interactive.

Cache write: 1.25 times for a 5-minute TTL, 2 times for a 1-hour TTL. You pay a premium to put something in the cache.

Cache read: 0.1 times. You pay a tenth to read it back. This is where the money is.

Long context: no premium. Deliberately in the list as a non-multiplier, because people assume one exists.

inference_geo: "us": 1.1 times. Data residency costs 10 percent.

Managed Agents: plus 0.08 dollars per session-hour, on top of tokens. Note the unit. This is not a token multiplier, it is a flat rate against wall-clock session time, so a slow session costs more than a fast one even at identical token counts.

Web search: 10 dollars per 1000 searches. Web fetch: free.

Token multipliers versus flat charges

The whole trick to these questions is sorting the list into two piles before you touch any arithmetic. Some of these numbers scale with how many tokens you moved. Others scale with something else entirely, and multiplying them by your token cost gives you a confidently wrong number.

The practical stacking order for a big offline job: pick the cheapest model that meets the requirement, cache the stable prefix, batch the whole thing, and turn effort down if the task is simple. Four levers, all independent, and a workload that uses all four is not in the same universe as one that uses none.

Practice

Try it yourself

Recall

The price table, cold

Straight recall. These numbers are the most lookup-able thing on the exam, which is exactly why getting them wrong is unforgivable.

State the per-million-token input and output prices for Fable, Mythos, Opus 5, Opus 4.x, Sonnet 5, Sonnet 4.6 and 4.5, and Haiku 4.5.

Reveal answer

Fable 5 and Mythos 5: 10 dollars in, 50 out. Opus 5 and the Opus 4.x line: 5 in, 25 out. Sonnet 5: 2 in and 10 out through 2026-08-31, then 3 in and 15 out. Sonnet 4.6 and Sonnet 4.5: 3 in, 15 out. Haiku 4.5: 1 in, 5 out. Across the whole lineup output costs five times input.

Quiz

The Batches discount

Which multiplier does the Batches API apply?

  1. AMinus 50 percent
  2. BMinus 25 percent
  3. C0.1 times, the same as a cache read
  4. DMinus 90 percent on input only
Show answer

Correct answer: A — Minus 50 percent

Batches is minus 50 percent, on input and output alike. The 0.1 times figure is real but belongs to cache reads, and mixing the two up is the most common cost-question error because both are large discounts you reach for in the same conversation. They also compose differently, since a cache read is per-token and batching is per-job.

Quiz

The data residency multiplier

You set inference_geo to us for a compliance requirement. What does it do to your bill?

  1. ANothing, data residency is free
  2. BIt multiplies token costs by 1.25
  3. CIt multiplies token costs by 1.1
  4. DIt adds 0.08 dollars per session-hour
Show answer

Correct answer: C — It multiplies token costs by 1.1

inference_geo: "us" is a 1.1 times multiplier, so a 10 percent surcharge on tokens. 1.25 is the tempting wrong answer because it is a real multiplier in this same table, but it belongs to a 5-minute-TTL cache write. The 0.08 dollars per session-hour is Managed Agents, which is a per-session charge and not a token multiplier at all.

Quiz

Add up one month of a Managed Agents workload

A month of usage, on claude-sonnet-5 at its current promotional rate, before 2026-08-31.

The month
input tokens      2,000,000
output tokens       500,000
session-hours            20
no caching, no batching, no data residency

Sonnet 5 is 2 dollars per million in and 10 dollars per million out. Managed Agents adds 0.08 dollars per session-hour.

  1. A9.00 dollars, since session-hours are included in the token price
  2. B10.60 dollars
  3. C9.08 dollars
  4. D21.20 dollars, since the session-hour charge multiplies the token cost
Show answer

Correct answer: B — 10.60 dollars

Two independent terms. Tokens are 2 times 2 plus 0.5 times 10, which is 4 plus 5, so 9 dollars. Session-hours are 20 times 0.08, so 1.60 dollars. Total 10.60. The two wrong arithmetic answers are the two ways of mishandling the second term, either by treating a flat per-hour charge as a multiplier on tokens or by fumbling the decimal and adding 8 cents once rather than per hour. Sorting each figure into multiplier or flat charge before you calculate anything is most of the work in these questions.

Quiz

Confirming a rate that has a date on it

You are sizing a budget that runs into September and you need to know what Sonnet 5 costs after the promotional period ends. You have the documentation and the console open. Where do you get the answer?

  1. AThe console's usage dashboard, which shows your current spend per model
  2. BThe models overview, since prices are listed alongside each model's limits
  3. CThe Messages API reference, since billing follows the request parameters
  4. DThe pricing page, which carries the per-million rates and the date the promotional rate ends
Show answer

Correct answer: D — The pricing page, which carries the per-million rates and the date the promotional rate ends

The rate and its expiry date are both published on the pricing page, and the expiry is the part you actually need, since a forward-looking budget depends on a number that has not applied yet. The usage dashboard is the genuinely tempting answer, because it is authoritative about money and it is the instinct of anyone who has ever been surprised by a bill, but it reports what you have already spent at the rate that was in force. It cannot tell you about a rate change that has not happened. A dashboard is a record, not a forecast.

Check

Build a one-line cost model

Estimate the monthly cost of one workload you can describe.

You should see

A single expression of the form (input tokens times input price plus output tokens times output price) times volume, with every applicable multiplier named and applied, and you can state which multipliers are token multipliers and which are flat per-unit charges.

Sign in to track your progress →