Whetstone.
The Execution EnvironmentSandboxes, the local shell, and running code
Module 2, Lesson 219 min

Sandboxes, the local shell, and running code

Two facts in this lesson are worth close to verbatim memorisation, and one default is worth being slightly angry about.

The policy ladder

Separately from backends, ShellToolMiddleware takes an execution policy. There are three.

ShellToolMiddleware execution policies
HostExecutionPolicy          # DEFAULT. Full host access.
DockerExecutionPolicy        # container-isolated
CodexSandboxExecutionPolicy  # sandboxed

That is a good exam question and a better production incident. It is also perfectly defensible as a design choice: this is a library for building capable agents, and it defaults toward capability. The danger is not that the default is wrong, it is that it is the one default a reviewer will not think to question.

One note for you specifically: ShellToolMiddleware is absent from TypeScript, alongside FilesystemFileSearchMiddleware. The camelCase translation rule tells you what it would be called if it existed. It does not.

The two security lines

“Sandboxes isolate code execution from your host system, but they don’t protect against context injection.”

The boundary is about execution, not about input. Whatever reaches the agent’s context still reaches it, sandbox or no sandbox. Injection is an input problem and execution isolation is orthogonal to it.

“Never put secrets inside a sandbox.”

This one trips people because “sandbox” pattern-matches to “safe place”. It is not a vault. The isolation faces outward: it protects your host from the code. It does nothing to protect the sandbox’s contents from the agent, which has execute and read_file inside it and can be talked into using them by exactly the injection the sandbox does not defend against.

Put the two together and you have the real threat model in one line: the sandbox contains the blast, it does not filter the fuse.

Running code, and the honest gap

The official course gives code execution its own lesson. What is solidly sourceable is the chain you now have: the backend grants execution, the execute tool is how the agent reaches it, and the shell policy decides where the command actually runs.

Beyond that chain, the docs are the unresolved ground from lesson 1, so this course teaches the chain and declines to invent an API around it. If a question turns on a specific interpreter class or protocol version, that is a reference lookup and the exam is semi open book for exactly this reason.

The composite risk worth carrying out of this module

An agent that ingests untrusted content and can also execute code is a remote code execution path with extra steps. Neither half is alarming alone. A web fetcher is ordinary. A shell is ordinary.

The sandbox is what is supposed to make the pair survivable, and it only does that if you changed the policy, because the default put the shell on your laptop.

Practice

Try it yourself

Quiz

The default policy

ShellToolMiddleware takes an execution policy. Which one is the default, and what does choosing it by omission actually mean?

  1. ADockerExecutionPolicy, container-isolated by default
  2. BCodexSandboxExecutionPolicy, sandboxed by default
  3. CThere is no default; a policy must be supplied
  4. DHostExecutionPolicy, full access to the host
Show answer

Correct answer: D — HostExecutionPolicy, full access to the host

The default is HostExecutionPolicy, which is full access to your actual machine. Docker is the tempting answer because a decade of tooling has trained everyone to expect safe-by-default containerisation, and that trained expectation is precisely what makes this dangerous. Leaving the policy alone is a decision, and the decision is full host access.

Quiz

Spot the risk

This configuration is reviewed and approved. One line is a production incident waiting to happen.

Under review
agent = create_deep_agent(
    model=model,
    tools=[fetch_url],
    middleware=[ShellToolMiddleware()],
    system_prompt="Summarise the pages the user gives you.",
)
  1. Afetch_url is unsandboxed, so the agent can reach arbitrary URLs
  2. BThe system prompt is too short to constrain the agent's behaviour
  3. Ccreate_deep_agent is missing a checkpointer, so an interrupted run cannot be resumed later
  4. DShellToolMiddleware is constructed with no policy, so shell commands run on the host
Show answer

Correct answer: D — ShellToolMiddleware is constructed with no policy, so shell commands run on the host

No policy means HostExecutionPolicy, so this agent can run shell commands on the machine that hosts it, while also fetching arbitrary attacker-controlled web pages into its own context. The first option is a real observation and the wrong answer: fetching untrusted content is only a catastrophe because of the shell sitting next to it. The two lines are individually defensible and jointly a remote code execution path.

Quiz

What a sandbox does not defend

The docs state this almost as a warning label, and it is the sentence most worth carrying verbatim.

  1. ASandboxes do not protect against context injection
  2. BSandboxes do not protect against filesystem writes
  3. CSandboxes do not protect against network egress
  4. DSandboxes do not protect against long-running processes
Show answer

Correct answer: A — Sandboxes do not protect against context injection

Sandboxes isolate code execution from your host system, but they do not protect against context injection. The other three are all things a sandbox boundary can plausibly address through configuration, which is what makes them weak distractors and this one the real point. Injection is an input problem, and no amount of execution isolation touches what arrives in the agent's context.

Recall

Which way the boundary faces

The second quotable security line, and the reasoning that makes it obvious rather than arbitrary.

Why should you never put secrets inside a sandbox?

Reveal answer

Because the isolation faces outward. It protects your host from the code, not the code's contents from the agent. An agent with execute and read_file inside the sandbox can read anything in there, and since sandboxes explicitly do not protect against context injection, an agent can be talked into reading and exfiltrating exactly that. A secret placed inside gains nothing from a boundary pointed the other way. The sandbox contains the blast; it does not filter the fuse.

Quiz

The TypeScript gap, again

You are asked to reproduce this lesson's shell policy setup in TypeScript. What is the honest answer?

  1. AUse shellToolMiddleware, the camelCase equivalent of the Python class
  2. BShellToolMiddleware is absent from TypeScript, so there is no direct equivalent
  3. CUse createMiddleware to wrap the Python implementation yourself
  4. DPolicies live in backend configuration in TypeScript rather than in middleware at all
Show answer

Correct answer: B — ShellToolMiddleware is absent from TypeScript, so there is no direct equivalent

ShellToolMiddleware is one of the two middlewares genuinely absent from TypeScript, alongside FilesystemFileSearchMiddleware. The first option is the trap built by this course's own casing rule: camelCase is the correct translation pattern for built-ins that exist, which makes the translated name feel automatically valid. A naming convention tells you what something would be called, not that it is there.

Check

State the threat model in two lines

In a scratch file, write what a sandbox protects and what it does not, without hedging either sentence.

You should see

Your first line is about execution isolation protecting the host, your second says injection is untouched by it, and neither sentence uses the word secure on its own.

Sign in to track your progress →