Whetstone.
Prompting that actually worksModel-specific prompting, and inherited instructions that now hurt
Module 1, Lesson 323 min

Model-specific prompting, and inherited instructions that now hurt

The consolidated reference opens with four model-specific pages: Fable 5, Sonnet 5, Opus 5 and Opus 4.8. They come before the general principles, and the ordering is the point.

General principles barely move between models. Be clear, use examples, structure with tags: that advice was true two years ago and is true now. The model-specific advice is where a prompt that worked last quarter quietly gets worse.

Inherited instructions do not decay into no-ops

This is the mental model to carry into the exam and into your codebase.

You would expect an instruction written for an old model to become harmless on a new one. The model gets better, the crutch becomes unnecessary, the line sits there doing nothing. That intuition is wrong, and the current guidance is basically a list of places it is wrong.

Think of a retry wrapper written around a flaky service. The service gets fixed. Your wrapper does not become a no-op, because it was never conditional on the service being flaky; it just keeps retrying, and now it is amplifying load on something healthy. Prompt instructions age the same way. They are unconditional text, and the condition that justified them lived only in your head.

Dial back the tool language

Concrete instance one. Everyone’s tool descriptions are full of this:

The thing to stop writing
CRITICAL: You MUST call this tool before responding.
IMPORTANT: ALWAYS use search_docs for any factual question.

That style existed because older models under-called tools and needed shouting at. Current models, 4.5 and 4.6 onward, overtrigger on this language. They call the tool when they should not, which costs you latency, tokens and correctness.

The fix is unglamorous: describe when the tool is appropriate, in plain declarative language, and let the model decide. “Use this to look up current pricing. Prices are not in your training data.” That is enough.

Opus 5 already checks its own work

Concrete instance two, and it is the sharpest one. On Opus 5, remove self-verification instructions. The model self-verifies natively, and the inherited instruction causes over-verification: it burns output on checking work it had already checked.

So the migration step nobody performs is: open the prompt, find every “double-check your answer”, “verify each step”, “re-read the input and confirm” block, and delete it when the target is Opus 5.

What this means as a habit

Before you point a prompt at a different model, read that model’s page. Not the general principles, which you already know, but the model page, which exists precisely because it contains the things that differ.

Worth doing on your own systems whether or not an exam is involved, because the symptom is not an error. It is a prompt that got slightly worse on a day nothing was deployed, which is the hardest class of regression to attribute after the fact.

This is inference, not published fact, so treat it as a study heuristic rather than a quotable line: the two instances above are the ones the docs call out explicitly, and both are removals rather than additions. If you are looking for a compression of the whole section, “migrating to a newer model is mostly deleting things” is a decent one.

Practice

Try it yourself

Quiz

The CRITICAL problem

Your tool descriptions are full of phrases like CRITICAL: You MUST call this tool before answering. You upgrade to a current model. What happens?

  1. ANothing, the emphasis is inert and models ignore formatting like that
  2. BThe model overtriggers on the tools, calling them when it should not
  3. CThe request 400s because emphasis markers are no longer permitted in tool descriptions
  4. DTool calling silently disables for that request
Show answer

Correct answer: B — The model overtriggers on the tools, calling them when it should not

Current models (4.5 and 4.6 onward) respond strongly to forceful tool language and overtrigger on it, calling tools in situations that do not warrant it. Option 1 is the tempting wrong answer because it was closer to true on older models, where you genuinely did need to shout to get reliable tool calls. The instruction did not become harmless, it became harmful, which is the worst kind of change to inherit.

Quiz

Opus 5 and self-verification

You are porting a prompt onto Opus 5. It contains a block instructing the model to double-check its work and verify each step before answering. What should you do?

  1. ARemove it; Opus 5 self-verifies natively and the inherited instruction causes over-verification
  2. BKeep it; explicit verification instructions are safe to leave in on any model
  3. CStrengthen it, since a more capable model can verify each step more thoroughly
  4. DMove it into the system prompt, where it will apply more consistently per turn
Show answer

Correct answer: A — Remove it; Opus 5 self-verifies natively and the inherited instruction causes over-verification

Opus 5 self-verifies natively, so an inherited verification instruction stacks on top of behaviour the model already has and produces over-verification. Option 2 is the tempting one because it feels conservative and safe, and for years it was. This is the general shape of the whole model-specific section: instructions written for an older model do not decay into no-ops, they compound with native behaviour.

Quiz

Fix this inherited prompt

Both of these lines were correct when they were written. The target has since moved to Opus 5 and nobody revisited them.

Inherited, now pointed at Opus 5
Tool description:
  CRITICAL: You MUST call search_docs before answering any factual question.

System prompt, final paragraph:
  Before you answer, double-check your work and verify each step.

Which edit matches current guidance?

  1. AKeep both; they were correct when they were written and are harmless on a more capable model
  2. BDelete the tool description entirely, and keep the verification block exactly as it stands
  3. CRewrite the tool description in plain language saying when the tool applies, and delete the verification block
  4. DStrengthen both, since a more capable model can act on more forceful instructions reliably
Show answer

Correct answer: C — Rewrite the tool description in plain language saying when the tool applies, and delete the verification block

These are the two removals the guidance calls out by name, and note that only one of them is a deletion. The tool still needs a description, so that line gets rewritten rather than dropped: say when the tool is appropriate and let the model decide. The verification block goes entirely. Option 1 is the tempting one because harmless-when-obsolete is how we expect instructions to age, and it is the exact intuition this lesson exists to break.

Recall

Why model pages come first

The ordering on the consolidated page is an editorial judgement about risk, and it is worth being able to reconstruct the judgement rather than just the order.

Why does the consolidated prompting reference put four model-specific pages ahead of the general principles?

Reveal answer

Because the advice that changes between models is the advice most likely to make a working prompt worse when you switch models, and general principles are broadly stable across models. Checking the page for the specific model you are calling is the step that catches inherited instructions which have flipped from helpful to harmful.

Check

Audit your own inheritance

Five minutes with one of your own repositories open. Offline, no account, no deploy. Search your prompts and tool descriptions for CRITICAL, MUST, ALWAYS, double-check and verify each step.

You should see

You have found at least one instruction written for an older model that is still live in a prompt or a tool description, and you can say which of the two categories it falls into: now inert, or now actively harmful. If your honest conclusion is that everything you found is inert, write down the reason you believe that rather than assuming it, because the whole point of this lesson is that inert is the wrong default.

Sign in to track your progress →