Whetstone.
Putting It All TogetherThe sales assistant, read as an architecture
Module 5, Lesson 215 min

The sales assistant, read as an architecture

The official course finishes by building a sales assistant. This lesson takes the same brief and derives the configuration from everything above it, which is the thing an exam can actually ask you.

To be explicit about sourcing: the official project’s interior is not public, so what follows is a derivation from a realistic brief of that shape, not a reproduction of their code.

The brief

An assistant that reads long contracts and call notes, queries a CRM through MCP tools, runs unattended for hours across a working day, and can send email on your behalf.

Four sentences, and every one of them lands on a different module.

Deriving it

Reads long contracts. That is module 3’s failure mode by definition. Offloading handles it: a contract past 20,000 tokens becomes a path plus a head-and-tail preview, and the default filesystem tools are how the agent reaches back in for the clause it wants. Nothing to configure, but a great deal to understand, because an engineer who does not know offloading exists will see truncated context and reach for a bigger model.

CRM tools over MCP. Two-part configuration from module 1: an mcp_servers entry and a matching mcp_toolset, both required. Their descriptions are permanent context cost, so eight CRM tools is eight descriptions on every turn of a multi-hour run.

Runs for hours unattended. A checkpointer, non-negotiable, for two independent reasons. A long run must survive a crash, and the send gate below needs somewhere to pause.

Can send email. The only irreversible action in the brief, so it is the whole of interrupt_on.

The decision that matters most

The assistant never needs to run code. So it gets no execution-capable backend, execute stays registered and powerless, and no ShellToolMiddleware is registered at all.

That single line is worth more than every other safety decision in the design combined, and it is worth being precise about why. A gate is oversight on a power you handed over. An absent capability is not oversight, it is absence. Module 2 established that sandboxes do not protect against context injection, so anything an agent can do, a sufficiently persuasive input may get it to do. Nothing persuades an agent into a capability it was never granted.

The strongest sandbox is the one you did not need.

Memory against skills

The split falls straight out of module 3’s rule: every task, or some tasks.

Memory: tone of voice, the approval policy, the rule that a discount is never quoted without a named approver. Every interaction touches these, and the interaction where the approval rule matters most is the one where the agent did not think it was doing pricing.

Skills: the renewal negotiation playbook, the discount ladder, the escalation runbook. Long, detailed, and relevant on a minority of conversations. Cheap at level one, loaded when invoked.

Permissions, last because it is the one people botch

Scope the filesystem to the document paths it reads. Specific denies first, then the allow, then an explicit catch-all deny.

The version without that last line looks scoped and is not. One allow rule and nothing else means every path outside it matches no rule, and no matching rule means allowed. A permission list that fails open quietly cancels the careful work done everywhere else in the design, which is precisely what makes it the best question in this domain.

Practice

Try it yourself

Quiz

Where the tone rule goes

The assistant must never quote a discount without a named approver, on every single interaction. Which channel carries that?

  1. AA skill, so the rule is available whenever pricing comes up in the conversation
  2. BThe subagent brief, so any delegated pricing work inherits it
  3. CAlways-loaded memory, because it applies unconditionally to every interaction
  4. DA tool description on the pricing tool, which the model always reads
Show answer

Correct answer: C — Always-loaded memory, because it applies unconditionally to every interaction

Every interaction means always-loaded, which is memory. The first option is the seductive one because the rule is topical and skills are the topical mechanism, so it feels well-organised. But a skill only loads when something invokes it, and the interaction where this rule matters most is precisely the one where the agent did not think it was doing pricing. A rule you might not load is a rule you might break.

Quiz

What gets a gate

The assistant can search the CRM, read contracts, draft replies and send email. Which of these belongs in interrupt_on?

  1. ASending email, because it is the one action that cannot be undone
  2. BAll four, since oversight ought to be applied uniformly
  3. CReading contracts, because they contain sensitive data
  4. DNone, because the permission list already covers filesystem access
Show answer

Correct answer: A — Sending email, because it is the one action that cannot be undone

Reversibility is the axis, and sending is the only irreversible action in the list. Gating all four is the tempting instinct and it is actively harmful: an approval prompt that fires constantly becomes a reflex click, so uniform oversight buys latency and spends the attention that made the gate worth having. Permissions cover paths, not outbound actions, so the last option confuses two different levers.

Quiz

The execution decision

The assistant has no need to run code. What is the correct configuration, and what does it buy?

  1. AGrant a Docker-policy shell anyway, in case a future task turns out to need it
  2. BGrant a host shell but forbid its use in the system prompt
  3. CGrant no execution-capable backend, so execute is registered but powerless
  4. DGrant execution and gate every shell call behind interrupt_on
Show answer

Correct answer: C — Grant no execution-capable backend, so execute is registered but powerless

The capability is granted by the backend, so an agent with no execution-capable backend cannot execute regardless of what it is asked to do, and that is the strongest possible position. The last option is the near-miss worth understanding: a gate is a good answer when you need the capability, but it is oversight on a power you chose to hand over. Not granting it is not oversight, it is absence, and absence cannot be talked around by an injection.

Recall

Derive the whole configuration

This is the exam's actual question, asked once about one system rather than eight times about eight features.

For an unattended assistant that reads long documents, uses CRM tools over MCP, runs for hours and can send email, name the configuration decision you would make for each of the six checklist items and give the reason.

Reveal answer

Backend with no execution capability, because the assistant never needs to run code and an ungranted capability cannot be misused. Shell policy therefore moot, and no ShellToolMiddleware registered at all. Permissions scoped to the document paths it reads, specific denies first, ending in a catch-all deny, because the list fails open otherwise. interrupt_on containing the send action only, because that is the sole irreversible operation and a gate that fires constantly stops being read. Context thresholds left at their defaults unless measurement says otherwise, with a checkpointer present because a multi-hour run must survive a crash and because the send gate needs somewhere to pause. Memory carrying the always-applicable rules such as tone and approval policy, skills carrying the occasional playbooks such as renewal negotiation.

Quiz

Spot the flaw

A colleague's design for the same assistant. One decision undermines several of the others.

Proposed configuration
backend:      host, execution enabled
permissions:  allow /crm-exports/**
interrupt_on: send_email
memory:       tone.md, approval-policy.md
skills:       renewal-negotiation, discount-ladder
  1. AThe skills should be memory, since pricing comes up on most calls
  2. BThe permission list has no catch-all deny, so paths outside /crm-exports are allowed
  3. Cinterrupt_on should include read_contract, since contracts contain sensitive data
  4. Dmemory should be one combined file, since two are loaded on every run
Show answer

Correct answer: B — The permission list has no catch-all deny, so paths outside /crm-exports are allowed

One allow rule and nothing else means the list fails open everywhere it does not match, so the agent can reach any path outside /crm-exports while looking scoped. It reads as a restriction and is the opposite. The first option is a real judgement call worth arguing about but it costs tokens rather than safety, and the point of the question is that a fail-open list quietly cancels the careful work done in the other four lines.

Check

Rebuild it for your own domain

Swap the domain for one you know and redo the derivation from scratch.

You should see

You produced six decisions with six reasons, at least one of your reasons cites a documented default by name, and you can point at the single decision that most reduces your blast radius.

Sign in to track your progress →