Sandboxes, the local shell, and running code
Two facts in this lesson are worth close to verbatim memorisation, and one default is worth being slightly angry about.
The policy ladder
Separately from backends, ShellToolMiddleware takes an execution policy. There are three.
HostExecutionPolicy # DEFAULT. Full host access.
DockerExecutionPolicy # container-isolated
CodexSandboxExecutionPolicy # sandboxedThat is a good exam question and a better production incident. It is also perfectly defensible as a design choice: this is a library for building capable agents, and it defaults toward capability. The danger is not that the default is wrong, it is that it is the one default a reviewer will not think to question.
One note for you specifically: ShellToolMiddleware is absent from TypeScript, alongside FilesystemFileSearchMiddleware. The camelCase translation rule tells you what it would be called if it existed. It does not.
The two security lines
“Sandboxes isolate code execution from your host system, but they don’t protect against context injection.”
The boundary is about execution, not about input. Whatever reaches the agent’s context still reaches it, sandbox or no sandbox. Injection is an input problem and execution isolation is orthogonal to it.
“Never put secrets inside a sandbox.”
This one trips people because “sandbox” pattern-matches to “safe place”. It is not a vault. The isolation faces outward: it protects your host from the code. It does nothing to protect the sandbox’s contents from the agent, which has execute and read_file inside it and can be talked into using them by exactly the injection the sandbox does not defend against.
Put the two together and you have the real threat model in one line: the sandbox contains the blast, it does not filter the fuse.
Running code, and the honest gap
The official course gives code execution its own lesson. What is solidly sourceable is the chain you now have: the backend grants execution, the execute tool is how the agent reaches it, and the shell policy decides where the command actually runs.
Beyond that chain, the docs are the unresolved ground from lesson 1, so this course teaches the chain and declines to invent an API around it. If a question turns on a specific interpreter class or protocol version, that is a reference lookup and the exam is semi open book for exactly this reason.
The composite risk worth carrying out of this module
An agent that ingests untrusted content and can also execute code is a remote code execution path with extra steps. Neither half is alarming alone. A web fetcher is ordinary. A shell is ordinary.
The sandbox is what is supposed to make the pair survivable, and it only does that if you changed the policy, because the default put the shell on your laptop.
Try it yourself
The default policy
ShellToolMiddleware takes an execution policy. Which one is the default, and what does choosing it by omission actually mean?
Show answer
Correct answer: D — HostExecutionPolicy, full access to the host
The default is HostExecutionPolicy, which is full access to your actual machine. Docker is the tempting answer because a decade of tooling has trained everyone to expect safe-by-default containerisation, and that trained expectation is precisely what makes this dangerous. Leaving the policy alone is a decision, and the decision is full host access.
Spot the risk
This configuration is reviewed and approved. One line is a production incident waiting to happen.
agent = create_deep_agent(
model=model,
tools=[fetch_url],
middleware=[ShellToolMiddleware()],
system_prompt="Summarise the pages the user gives you.",
)Show answer
Correct answer: D — ShellToolMiddleware is constructed with no policy, so shell commands run on the host
No policy means HostExecutionPolicy, so this agent can run shell commands on the machine that hosts it, while also fetching arbitrary attacker-controlled web pages into its own context. The first option is a real observation and the wrong answer: fetching untrusted content is only a catastrophe because of the shell sitting next to it. The two lines are individually defensible and jointly a remote code execution path.
What a sandbox does not defend
The docs state this almost as a warning label, and it is the sentence most worth carrying verbatim.
Show answer
Correct answer: A — Sandboxes do not protect against context injection
Sandboxes isolate code execution from your host system, but they do not protect against context injection. The other three are all things a sandbox boundary can plausibly address through configuration, which is what makes them weak distractors and this one the real point. Injection is an input problem, and no amount of execution isolation touches what arrives in the agent's context.
Which way the boundary faces
The second quotable security line, and the reasoning that makes it obvious rather than arbitrary.
Why should you never put secrets inside a sandbox?
Reveal answer
Because the isolation faces outward. It protects your host from the code, not the code's contents from the agent. An agent with execute and read_file inside the sandbox can read anything in there, and since sandboxes explicitly do not protect against context injection, an agent can be talked into reading and exfiltrating exactly that. A secret placed inside gains nothing from a boundary pointed the other way. The sandbox contains the blast; it does not filter the fuse.
The TypeScript gap, again
You are asked to reproduce this lesson's shell policy setup in TypeScript. What is the honest answer?
Show answer
Correct answer: B — ShellToolMiddleware is absent from TypeScript, so there is no direct equivalent
ShellToolMiddleware is one of the two middlewares genuinely absent from TypeScript, alongside FilesystemFileSearchMiddleware. The first option is the trap built by this course's own casing rule: camelCase is the correct translation pattern for built-ins that exist, which makes the translated name feel automatically valid. A naming convention tells you what something would be called, not that it is there.
State the threat model in two lines
In a scratch file, write what a sandbox protects and what it does not, without hedging either sentence.
Your first line is about execution isolation protecting the host, your second says injection is untouched by it, and neither sentence uses the word secure on its own.