Whetstone.
Context engineeringWindows, budgets and the cache that fails silently
Module 2, Lesson 225 min

Windows, budgets and the cache that fails silently

Two separate things get muddled under “context costs money”. Untangle them now, because the exam treats them as different sub-skills and so should you.

Window sizes, and the premium that does not exist

The current models, claude-opus-5, claude-sonnet-5 and claude-fable-5, carry a 1M token context window and 128k output. claude-haiku-4-5-20251001 is 200k context and 64k output.

And the part people get wrong: there is no long-context pricing premium. Filling a 1M window does not move you onto a different rate card. A token near the end of a huge prompt costs the same as a token near the start of a small one.

The three multipliers

Prompt caching prices in multiples of the base input rate.

  • Cache write, 5 minute TTL: 1.25x
  • Cache write, 1 hour TTL: 2x
  • Cache read: 0.1x

A cache read costs a tenth of an ordinary input token. That is the whole reason anyone bothers.

Run the arithmetic yourself rather than trusting a rule of thumb, and be careful setting up the comparison, because there is an off-by-one waiting in it.

The trap is comparing the wrong two things. A cached run of one write plus N reads is N plus 1 requests, so it has to be measured against N plus 1 uncached calls, not against N. Write it the wrong way and you will conclude caching needs one more read than it really does.

Set up properly, with the 5 minute TTL at 1.25x write and 0.1x reads:

1.25 + 0.1N < 1 + N, which solves to N > 0.28. One read is enough. Concretely: write plus one read costs 1.35 against 2.0 for two uncached calls.

With the 1 hour TTL at 2x write: 2 + 0.1N < 1 + N, which solves to N > 1.11. Two reads. That is 2.2 against 3.0.

Those are the same numbers the pricing docs give when they say caching pays off after one cache read at the 5 minute TTL and two at the 1 hour TTL, which is a useful check on your own algebra. Stated as total uses of the prefix rather than reads, it is two and three, which is how the Foundations course puts it. Same fact, and the only thing that changes is whether you are counting the write.

That is a strikingly low bar, and it is why caching is close to free money for any prompt with a stable prefix. It is also why the failure mode in the next section is so expensive: the wins are large enough that you assume you are getting them.

The minimum that fails silently

Here is the one that will actually cost you money in production.

Every model has a minimum cacheable prompt length, and a prompt below it silently does not cache. No error. No warning. Nothing.

It behaves like a minimum order value that the shop never mentions. You hand over a basket that is a pound short, the transaction completes, you get a receipt, and the discount you thought you had arranged simply is not on it. Nothing failed. You just paid list price and were never told.

The values are not monotonic across the lineup, so you cannot reason your way to them:

Minimum Models
512 tokens Opus 5, Fable, Mythos
1024 tokens Opus 4.8, Sonnet 5, Sonnet 4.6
2048 tokens Opus 4.7
4096 tokens Opus 4.6, Opus 4.5, Haiku 4.5

Read that table twice. Opus 5 has the lowest minimum in the list. Haiku 4.5 has the highest. Any intuition of the form “bigger model, bigger minimum” or “newer model, lower minimum” gets contradicted somewhere in that table. It is a lookup, not a derivation.

Now stack that against the silent failure. You mark a 900-token system prompt as cacheable, you ship it against Haiku 4.5, and it has never cached once. Every request is full price. Your latency looks normal. Your logs look normal. The only symptom is a bill that is roughly ten times what you modelled, and you will blame traffic before you blame this.

The habit that saves you

Two checks, both cheap, and both worth running against your own production traffic this week rather than only in an exam.

One: look up the minimum for the exact model you are calling, every time you change models. Switching from Opus 5 to Haiku 4.5 raises your minimum by a factor of eight, and nothing in your code will complain.

Two: verify caching from the usage numbers, not from the fact that you set the marker. The request tells you what was written and what was read. A cache marker is an intention. The usage figures are the measurement, and they are the only thing that can tell you the intention failed.

Practice

Try it yourself

Quiz

A prompt below the minimum

You mark a 700-token system prompt as cacheable and run it against Haiku 4.5, whose minimum cacheable length is 4096 tokens. What happens?

  1. AThe API returns a 400 telling you the prompt is too short to cache
  2. BIt caches anyway, padded up to the model's minimum
  3. CThe request fails and you must remove the cache marker to proceed
  4. DThe request succeeds normally and simply does not cache, with no error
Show answer

Correct answer: D — The request succeeds normally and simply does not cache, with no error

A prompt below the model's minimum cacheable length silently does not cache. There is no error, no warning, and the request behaves normally in every visible way except the bill. Option 1 is the tempting one because a 400 is what a well-behaved API would do, and it is what you would design. The absence of an error is exactly what makes this expensive: you can ship a caching strategy that has never once cached.

Quiz

Reading the minimum table

You are choosing a model and you need a low minimum cacheable length because your static prefix is short. Which statement is correct?

  1. AOpus 5 sits at 512 tokens and Sonnet 5 at 1024, so the minimum cannot be inferred from tier or recency
  2. BAll current models share a 1024-token minimum, so the model choice does not matter here
  3. CLarger models always have larger minimums, because their attention blocks are bigger
  4. DThe minimum scales with the context window, so 1M-window models have the highest
Show answer

Correct answer: A — Opus 5 sits at 512 tokens and Sonnet 5 at 1024, so the minimum cannot be inferred from tier or recency

The minimums are not monotonic across the lineup: Opus 5, Fable and Mythos are at 512; Opus 4.8, Sonnet 5 and Sonnet 4.6 at 1024; Opus 4.7 at 2048; Opus 4.6, Opus 4.5 and Haiku 4.5 at 4096. Option 3 is the tempting one because a bigger-model-bigger-minimum rule feels mechanically plausible, and it is flatly contradicted by Opus 5 sitting at the lowest value and Haiku 4.5 at the highest. Look it up per model, never infer it.

Quiz

Nothing is cached and nothing is wrong

Caching is configured, the static prefix is about 3k tokens and has not changed in a week, traffic is steady at several requests a minute, and no request has ever returned an error. Every response reports the same thing.

What the response reported, every request, for a week
input tokens             3120
cache creation tokens       0
cache read tokens           0
errors                   none

What is the most likely explanation?

  1. ACache entries are expiring before the next request arrives, so every call has to rewrite
  2. BThe API rejected the cache marker and returned a warning you are not logging
  3. CCache reads are not reflected in the usage figures, so these numbers cannot tell you anything
  4. DThe static prefix is below this model's minimum cacheable length, so it silently never cached
Show answer

Correct answer: D — The static prefix is below this model's minimum cacheable length, so it silently never cached

The giveaway is the zero in the cache creation row. If the TTL were expiring between calls you would see a write on every request and no reads, not zero of both, and expiry is option 1: the tempting answer, because expiry is the caching failure most people have actually met and it explains the missing reads. Zero writes means nothing was ever eligible to be written. A 3k prefix against a model with a 4096-token minimum, Haiku 4.5 for instance, never qualifies, and below-minimum is a silent no-op rather than an error, which is why the errors row is also zero.

Recall

The three multipliers

Three numbers, and the third one is the reason the other two are worth paying. Everything in the break-even arithmetic falls out of these, so this is the card to get automatic.

What are the cache write and cache read multipliers?

Reveal answer

Cache write is 1.25x base for the 5 minute TTL and 2x base for the 1 hour TTL. Cache read is 0.1x base. So a read costs a tenth of an ordinary input token and a write costs a quarter more, or double, depending on which TTL you chose.

Do

Compute your own break-even

Pen and paper, two minutes, no network required. Derive it rather than memorising it, because a derived number survives a change to the multipliers and a memorised one does not.

  • Take the 5 minute TTL. Writing the cache costs 1.25 units where an uncached call costs 1.
  • Each subsequent read costs 0.1 units instead of 1.
  • Count the requests on both sides before you write anything: one write plus N reads is N plus 1 requests, so the honest comparison is against N plus 1 uncached calls, not N. Getting this wrong is the classic off-by-one and it makes caching look worse than it is.
  • Write the inequality on that basis and solve for N.
  • Repeat for the 1 hour TTL at 2x write.
Done whenYou got one read for the 5 minute TTL and two for the 1 hour TTL, and you can say why the comparison is against N plus 1 uncached calls rather than N.
Sign in to track your progress →