Whetstone.
Security, guardrails and Claude CodeGuardrails, secrets and key management
Module 4, Lesson 223 min

Guardrails, secrets and key management

Two sub-skills in one lesson: Guardrails and Safe Deployment (2.3%) and Identity, Secrets and Key Management (1.6%). They are small, they are dry, so let us be quick and exact.

Secrets: two hard rules

Never put a secret in a prompt. Never put a secret in a sandbox.

Both of those are stated plainly in the guidance and both get violated constantly, for sympathetic reasons.

Why not a prompt. Anything in the context is reachable by the model, and anything reachable by the model can come back out in the output. Extraction is not exotic: it is the entire point of the prompt-leak literature. Worse, prompt content flows into your logs, your traces, your evaluation datasets, and any transcript you keep for debugging. One credential in a system prompt becomes a credential in a hundred places you were not thinking about, some of them lower-trust than the original.

That fan-out is the part to hold on to. It is the same shape as a secret committed to a repository: the moment it is in there, remediation is not “delete the line”, it is rotation, because you have lost track of every copy. A system prompt is a string that gets logged, sampled, replayed into evals and pasted into bug reports.

Why not a sandbox. This one is almost funny when you say it out loud. The sandbox exists to contain code you do not trust. Putting a credential inside it hands the credential to the thing you were containing. If sandboxed code needs to reach a paid API, the key lives outside and the sandboxed side calls a broker that holds it.

Guardrails: encouraged versus enforced

The distinction the exam is testing is simple and people blur it.

A prompt instruction is encouragement. “Refuse requests for medical advice” genuinely changes behaviour and genuinely helps. It is also inside the context, competing with everything else in the context, including anything an attacker put there.

A guardrail is enforcement. It sits outside the model, inspects the input or the output, and can block the action regardless of what the model decided. Classifying an input before it reaches the model, validating an output before it reaches the user, gating a tool call behind a policy check: those are boundaries because the model does not get a vote.

Use both. Just do not confuse which one is load-bearing. When a question offers you a well-written system-prompt rule and an external check, the external check is the guardrail.

Safe deployment is about what you can do afterwards

The instinct is to make deployment safe by reviewing harder before launch. That does not work here, and the reason is specific: a natural-language interface has an unenumerable input space. You cannot list the failure modes in advance, so completeness of pre-launch review is the wrong target.

What actually helps is posture after launch.

Ship narrow. Limited audience, limited scope, limited tool access. Expand once you have real traffic telling you what people actually do with it, which is never what you predicted.

Monitor behaviour, not errors. Your logs will be full of successful requests. The failures you care about here are 200s that did the wrong thing. An error rate dashboard is blind to the entire category.

Be able to turn it off. The single most useful property of an LLM deployment is a fast rollback, because your response time to a discovered failure is the thing you actually control.

Practice

Try it yourself

Quiz

The API key in the sandbox

Your agent needs to call a paid API from inside a sandboxed execution environment. Where does the key live?

  1. AOutside the sandbox, behind a service the sandboxed code calls without ever seeing the credential
  2. BIn an environment variable inside the sandbox, since the sandbox is already isolated
  3. CIn the system prompt, so the model can insert it when it writes the API call
  4. DIn a read-only file mounted into the sandbox, so the code cannot modify or move it
Show answer

Correct answer: A — Outside the sandbox, behind a service the sandboxed code calls without ever seeing the credential

Never put a secret in a sandbox or in a prompt. Options 2 and 4 are the tempting ones because isolation feels like it should be enough, and mounting a file read-only sounds like a hardening measure. The sandbox exists to contain code you do not trust; putting a credential inside it hands that credential to the exact thing you were containing. Keep the key on the trusted side and let the untrusted side make requests through a broker that holds it.

Quiz

Where guardrails live

Which describes a guardrail you can actually rely on?

  1. AA line in the system prompt instructing the model to refuse harmful requests outright
  2. BA strongly worded reminder repeated at the end of every user turn to reinforce the rule
  3. CA check outside the model that inspects input or output and can block the action whatever the model decided
  4. DSetting effort to high so the model reasons more carefully about safety before acting
Show answer

Correct answer: C — A check outside the model that inspects input or output and can block the action whatever the model decided

A guardrail you can rely on sits outside the model and can veto the action. Option 1 is the tempting one because prompt-level instructions are real and they do help; they are just not a boundary, since anything reaching the model can argue with them. The distinction the exam wants is between behaviour you have encouraged and behaviour you have enforced.

Quiz

Two ways to reach a paid API

Sandboxed code needs to call a pricing API. Two implementations, both of which work.

A and B both return prices
// A
const systemPrompt = `You can call the pricing API directly.
Use this key in the Authorization header: ${process.env.PRICING_API_KEY}`

// B
const systemPrompt = `Call POST /broker/pricing to look up a price.
The broker attaches credentials. You never see them.`

Which one ships, and what exactly is wrong with the other?

  1. AA, provided the system prompt is never written to your logs, traces or any evaluation dataset
  2. BBoth are equivalent, because the model has to reach the pricing API either way it is written up
  3. CB; A puts the credential into the context, where the model can reach it and it flows on into logs, traces and stored transcripts
  4. DB, but only because A hardcodes the Authorization header name instead of letting the SDK attach it
Show answer

Correct answer: C — B; A puts the credential into the context, where the model can reach it and it flows on into logs, traces and stored transcripts

A violates the never-in-a-prompt rule twice over: the model can reach anything in its context and can therefore emit it, and the prompt itself flows into every log, trace and evaluation dataset downstream. Option 1 is the tempting one because not logging system prompts is a real hardening step that a careful team would already have, and it sounds sufficient. It is not a boundary, it only closes the second of the two exposures and leaves extraction through the output completely open. B is the broker pattern: the untrusted side gets to make requests, never to hold the thing that authorises them.

Recall

Two places a secret never goes

Two rules, and each one has a reason that is worth more than the rule. If you can only recall the prohibitions, you will not spot the variants that are the same mistake in a different shape.

Name two places a secret must never be put, and why.

Reveal answer

Never in a sandbox, because the sandbox exists to contain untrusted execution and putting a credential inside it gives the untrusted side the credential. And never in a prompt, because prompt content is reachable by the model and therefore extractable through the output, and it also flows into logs, traces and any transcript you keep.

Recall

Safe deployment posture

The counterintuitive part is where the safety comes from. Pre-launch review is the instinct and it is not the answer here, for a specific structural reason about the input space.

What does safe deployment of an LLM feature look like in practice?

Reveal answer

Ship narrow and expand: limited scope, limited audience, limited tool access, with monitoring on what the system actually does rather than on whether it errored. The load-bearing property is being able to turn it off or roll it back quickly, because you cannot enumerate the failure modes of a natural-language interface before launch, so the response time to a discovered failure matters more than the completeness of your pre-launch review.

Sign in to track your progress →