Whetstone.
Cost and Token ManagementPrompt caching and the minimum that is not monotonic
Module 4, Lesson 226 min

Prompt caching and the minimum that is not monotonic

Caching is the biggest cost lever that does not change your architecture. It is also the one feature in this course that fails silently, which makes its details worth more attention than its share of the exam suggests.

Every other trap in this course is a smoke alarm going off. This one is the smoke alarm with the battery quietly removed: nothing sounds, nothing looks different, and you find out from the bill.

The economics in one paragraph

Mark a prefix of your request as cacheable. The first request pays a write premium. Every subsequent request that matches that prefix exactly pays a read price of 0.1 times normal input.

Writes cost 1.25 times for a 5-minute TTL and 2 times for a 1-hour TTL.

The break-even falls straight out of that arithmetic, and it is worth doing once so you never have to guess:

5-minute TTL. Two uses costs 1.25 plus 0.1 equals 1.35, against 2.0 uncached. Ahead from the second use.

1-hour TTL. Two uses costs 2.0 plus 0.1 equals 2.1, against 2.0 uncached, so still behind. Three uses costs 2.2 against 3.0. Ahead from the third use.

The pricing docs state the same result counted the other way, as cache reads rather than total uses: one read at the 5-minute TTL, two at the 1-hour. Two uses is one write plus one read, and three uses is one write plus two reads, so the numbers agree. Watch which unit a question is using, because “two” and “three” become “one” and “two” the moment someone counts reads instead of uses, and the gap between those two framings is the easiest place to lose a mark you actually understood.

So the cheaper TTL wins on short bursts and the longer one wins on anything reused across a session. Neither is a default.

What is actually cacheable

Prefix matching, exact bytes, from the front. Which means the only worthwhile candidates are the parts that never change: the system prompt, tool definitions, and any long static document you attach to every call.

The newest user message can never hit, because it is different every time. In a long conversation you can also cache the accumulated history, since turn twenty’s prefix contains turn nineteen’s prefix, and that is where multi-turn chat gets dramatically cheaper.

The minimum, and why it is the best question on this topic

The consequence that makes this a superb exam item: a 3000-token prefix caches on Opus 5 and does not cache on Opus 4.6. Same prompt, same cache_control, same code, different model id, and one of them saves you 90 percent while the other saves you nothing.

It fails silently, and that is the real lesson

Every other trap in this course is a 400. This one is not.

Send a below-minimum prefix and the request succeeds. No error, no warning field, no signal in the response body that says “you asked for caching and did not get it”. Your bill just stays at full price while your architecture diagram says caching is on.

The general form of this is worth carrying out of the course entirely: ask of any check what single fault would make it quieter rather than louder. A feature whose failure mode is silence needs an assertion pointed at it, because nothing else will ever tell you.

How it composes with everything else

Caching is a multiplier on input tokens, so it stacks with Batches, which is a multiplier on the whole job. A batched job over a cached prefix gets both. It does not interact with the output side at all, so a workload that is expensive because it generates a lot of text is not a caching problem, it is an effort and prompt-design problem.

And one platform note that belongs here rather than in module five: automatic prompt caching is available everywhere except Bedrock. If your deployment target is Bedrock, that is a line in your cost model, not a footnote.

Practice

Try it yourself

Recall

The four cache minimums

There is no rule to derive here and that is the point of the card. Four numbers, attached to model groups, and the sequence is the trap rather than the content.

List the prompt cache minimum token counts and which models each applies to, and say why the sequence is worth noticing.

Reveal answer

512 tokens for Opus 5, Fable and Mythos. 1024 for Opus 4.8, Sonnet 5 and Sonnet 4.6. 2048 for Opus 4.7. 4096 for Opus 4.6, Opus 4.5 and Haiku 4.5. It is worth noticing because the sequence is NOT monotonic with model age or capability: the minimum gets smaller on newer models, so a 3000-token prefix caches on Opus 5 and does not cache at all on Opus 4.6. You cannot reason it out, you have to know the table.

Quiz

What happens below the minimum

You set cache_control on a 3000-token prefix and send it to claude-opus-4-6, whose minimum is 4096. What do you observe?

  1. AA 400 telling you the prefix is too short to cache
  2. BA warning field in the response usage object
  3. CThe prefix is padded to the minimum and cached
  4. DNo error at all, and no cache savings, silently
Show answer

Correct answer: D — No error at all, and no cache savings, silently

Nothing goes wrong loudly. The request succeeds, the response looks normal, and you simply never get a cache hit, so your bill quietly stays at full price. This is the opposite failure direction to every other trap in this course, and it is far more dangerous because there is no 400 to notice. The only detection is reading the cache read token counts in the usage object and seeing zeroes.

Quiz

When the 1-hour TTL pays for itself

A cache write at the 1-hour TTL costs 2 times, and a read costs 0.1 times. How many total uses of that prefix before caching is cheaper than not caching?

  1. A2 uses
  2. B3 uses
  3. C10 uses
  4. D21 uses
Show answer

Correct answer: B — 3 uses

At 2 uses you pay 2.0 plus 0.1 equals 2.1 against an uncached 2.0, so you are still behind. At 3 uses you pay 2.0 plus 0.2 equals 2.2 against an uncached 3.0 and you are clearly ahead. Two uses is the tempting answer because it is the correct break-even for the 5-minute TTL at 1.25 times write, which shows why the two TTLs are not interchangeable on short-lived prefixes.

Quiz

Same prefix, two models, one saving

An identical stable prefix, measured at 3000 tokens, is marked for caching in two services.

Service A
{ "model": "claude-opus-5", "system": [ { "type": "text", "text": "...3000 tokens...", "cache_control": { "type": "ephemeral" } } ] }
Service B
{ "model": "claude-opus-4-6", "system": [ { "type": "text", "text": "...3000 tokens...", "cache_control": { "type": "ephemeral" } } ] }
  1. ABoth cache, since 3000 tokens is a substantial prefix on any current model
  2. BNeither caches, since the minimum is a flat 4096 tokens across the lineup
  3. CA caches and B does not, because the minimum is 512 on Opus 5 and 4096 on Opus 4.6
  4. DB caches and A does not, because older models were given the lower minimums
Show answer

Correct answer: C — A caches and B does not, because the minimum is 512 on Opus 5 and 4096 on Opus 4.6

The minimum is a per-model number and it is smaller on the newer model, so the same prefix clears the bar on Opus 5 at 512 and falls well short on Opus 4.6 at 4096. The genuinely tempting answer is that older models have the lower minimum, since capability limits usually loosen as models get newer and it feels like the newer model should be the demanding one. The sequence runs the other way, and there is no reasoning that gets you there. Note what makes this dangerous in practice rather than on paper: Service B does not fail, it just pays full price forever.

Quiz

Spot why this cache never hits

Written by someone who understood the pricing perfectly and the prefix rule not at all.

Marked for caching
{
  "model": "claude-sonnet-5",
  "system": "You are a policy assistant.",
  "messages": [
    { "role": "user", "content": "...30k of policy documents..." },
    { "role": "assistant", "content": "Understood." },
    { "role": "user", "content": [
        { "type": "text", "text": "What is the notice period?", "cache_control": { "type": "ephemeral" } }
    ]}
  ]
}
  1. AThe cache_control sits on the newest user message, which changes on every request and can never match a stored prefix
  2. Bcache_control may only be set on the system parameter and tool definitions, never inside messages
  3. CThe system parameter must be an array of content blocks before any cache_control takes effect
  4. DThe 30k document exceeds the per-block cache size limit, so the request falls back to uncached
Show answer

Correct answer: A — The cache_control sits on the newest user message, which changes on every request and can never match a stored prefix

Caching works on an exact prefix match from the front of the request, so the only useful thing to mark is content that is byte-identical every time. Marking the newest question guarantees a miss, since that is the one part that changes on every call. The 30k of policy documents sitting above it is the thing that was worth caching and it went unmarked. The size answer is the tempting one because 30k feels like it might exceed something, and the actual size rule runs the other way: prefixes fail for being too small, never too large.

Check

Check your own prefix against your own model

Take a real system prompt plus tool definitions from something you have built and check it against the cache minimum for the model it runs on.

You should see

You have a token count for the stable prefix, the exact minimum for that model id, and a yes or no. If it is a no, you have written down which of the two fixes applies, moving to a model with a lower minimum or growing the stable prefix past the bar.

Sign in to track your progress →