Whetstone.
Designing a Claude ApplicationFrom requirements to architecture
Module 2, Lesson 424 min

From requirements to architecture

Requirements arrive as a paragraph from someone who does not know what a token is. Turning that into an architecture is a real sub-skill with real marks attached, and the good news is that it decomposes far better than it looks like it should. Three questions do most of the work.

Three questions, in this order

How big is the input, in tokens? Not pages, not megabytes. Tokens, because that is the unit both the context limit and the bill are denominated in. This answer picks your model id, since context window is a model property you cannot configure your way around. A 400-page contract does not fit in a 200k window no matter how clever the prompt is.

Is a human waiting? If yes: synchronous, streamed, and Batches is impossible. If no: Batches takes 50 percent off and you have just halved a line item by asking one question. The trap is work that sounds interactive because it appears in a UI but is actually generated overnight.

What happens when it is wrong? This is the requirement everybody forgets to write down and it is the one that determines whether you need citations, approval gates, strict tool use, or a human in the loop at all. “It writes a draft the user edits” and “it emails the customer” are wildly different systems built from the same API.

The order is load-bearing rather than stylistic. Question one fixes the model, and the model determines which features are even available. Question two fixes the transport, and the transport determines the interaction design. Answer them the other way round and you will design an interaction for a workload that turns out to be a nightly batch.

Non-functional requirements are where the design lives

The functional part of most Claude integrations is a paragraph. Everything expensive is non-functional: throughput, cost ceiling, latency budget, data residency, auditability, and what the failure mode looks like.

Map requirements to features, not to vibes

A short lookup table that covers most of what you will be asked:

“Must show where the answer came from” maps to citations. Not to a source field you invent.

“Must be valid JSON we can insert into a database” maps to structured outputs on a 4.5-or-later model. Not to a stern prompt and a retry.

“Must approve before acting” maps to owning the tool loop. No harness that hides the loop can satisfy it.

“The same 30k of policy documents go into every request” maps to prompt caching, and immediately to checking your model’s cache minimum.

“We have sixty tools” maps to tool search, because sixty tool definitions is a context problem before it is anything else.

“It runs overnight over 50,000 records” maps to Batches, and to noticing that Batches is not available on any third-party platform, which may in turn decide where you deploy.

The systems life cycle bit nobody enjoys

Two lifecycle facts belong in every design document. Model ids retire on published dates, so a hardcoded dated id is a dependency with an expiry, and claude-opus-4-1-20250805 retires on 2026-08-05. And beta headers are versioned by date, so a feature you enabled today is pinned to a contract that will eventually be superseded.

Write both into the design as maintenance commitments rather than discovering them when a request starts failing. A system that cannot survive its own dependencies expiring was never finished.

Practice

Try it yourself

Recall

The latency question decides the transport

One question about the user, answered honestly, eliminates most of the design space.

What question about latency do you ask first, and what does each answer rule in or out?

Reveal answer

Ask whether a human is waiting on this specific response. If yes, it is synchronous, you stream it, and Batches is off the table no matter how much cheaper it is. If no, and the work can tolerate being processed later, Batches applies and takes 50 percent off. The trap is a requirement that sounds interactive but is not, such as a nightly enrichment job someone described as a dashboard feature: that is batch work wearing a UI costume.

Recall

The triage, and why the order is fixed

The order is not arbitrary and that is the part worth holding. Each answer narrows what the next question is even choosing between.

What are the three triage questions, in order, and what does each one decide?

Reveal answer

One, how big is the input in tokens. This decides the model id, because the context window is a property of the model and not something you can configure around. Two, is a human waiting. This decides the transport: synchronous and streamed, or Batches at 50 percent off. Three, what happens when it is wrong. This decides the safety features: citations, approval gates, strict tool use, human in the loop. The order matters because the model choice constrains which features are even available to you, and the transport choice constrains the interaction design, so answering them out of order means redoing work.

Quiz

Mapping a requirement to a feature

The requirement reads: users must be able to see which part of the uploaded contract each answer came from. Which feature is the requirement describing?

  1. AStructured outputs, so the answer carries a source field
  2. BPrompt caching, so the contract stays in context cheaply
  3. CCitations, so each answer points at the span it came from
  4. DThe Files API, so the contract can be referenced by id
Show answer

Correct answer: C — Citations, so each answer points at the span it came from

The requirement is attribution to spans in a supplied document, which is exactly citations. Structured outputs with a source field is the tempting answer because it produces something that looks like provenance, but the model fills that field itself and can be wrong about it, so it provides the appearance of attribution without the substance. Caching and the Files API are both plausible parts of the same system and neither addresses the stated requirement.

Quiz

Which requirement changes the model choice

Four requirements land on your desk. Which one most directly forces a specific model id rather than a feature?

  1. AEvery destructive action needs human approval before it runs
  2. BResponses must be valid JSON against our published schema
  3. CThe system must process 400-page documents in a single request
  4. DThe tool catalogue will grow to roughly 60 tool definitions
Show answer

Correct answer: C — The system must process 400-page documents in a single request

Document size maps straight onto the context window, and the context window is a property of the model id, not something you can configure. A 400-page document pushes you toward the 1M-context models rather than a 200k one. Approval gates are a build-shape decision, JSON is structured outputs, and a 60-tool catalogue is what tool search exists for. All three of those are solvable without changing model.

Quiz

Sixty tools and a context problem

A design has grown to about 60 tool definitions, and every request now carries all of them. What is the correct response?

  1. AMove to a 1M-context model, since the definitions now need the room
  2. BSplit the design into 60 separate agents, one per tool, and route between them
  3. CCache the tool definitions, since they are byte-identical every request
  4. DUse tool search, which exists so a large catalogue does not have to sit in context
Show answer

Correct answer: D — Use tool search, which exists so a large catalogue does not have to sit in context

Tool search exists precisely for this: a large catalogue consumes context before the conversation has started, and the fix is not to carry it all. Caching is the genuinely tempting answer, because tool definitions really are stable and really are cacheable, so it would cut the cost of the problem. It would not cut the context consumption at all, though, and context is what is actually running out. Cheaper is not the same as smaller. A bigger context window is the same mistake with more money attached.

Do

Run the three-question triage

Take one paragraph of vague requirements, real or invented, and triage it in a scratch file.

  • Write the paragraph as someone non-technical would say it.
  • Answer the size question, in tokens rather than pages or characters.
  • Answer the latency question with is a human waiting, yes or no.
  • Answer the trust question with what happens when the model is wrong here.
  • List the model id, the transport, and the two features that fall out of those three answers.
Done whenThe output names one model id, either sync-with-streaming or Batches, and at most three features, with each choice traceable back to one of the three answers.
Sign in to track your progress →