Whetstone.
Building a Deep AgentHuman in the loop, and what a pause really is
Module 1, Lesson 718 min

Human in the loop, and what a pause really is

A gate is not a pause button. It is a saved position that somebody comes back to, and almost every mistake in this area comes from picturing it as the former.

The parameter

interrupt_on is one of the six deep agent additions. It names which tools stop and wait for human sign-off before they execute.

Underneath it is HumanInTheLoopMiddleware, a built-in that exists in both Python and TypeScript and imports cleanly into a plain create_agent call. Same pattern as everything else in this course: the deep agent parameter is a shorter spelling of a middleware you could have wired yourself.

The todoist bridge is the same machine

You have already built this. The needs-review gate in the todoist bridge does exactly this shape: something is proposed, it stops, it waits in a state you can come back to, and your answer resumes it rather than restarting it.

That system taught you the two things that actually matter in production and neither is in the docs.

A gate nobody reads is worse than no gate. If everything is gated, approval becomes a reflex click and you have added latency to buy nothing. The value of a gate is entirely in the attention it receives, and attention is a fixed budget you are spending.

The gate has to reach a human where they already are. A pause waiting in a system nobody opens is not oversight, it is a stalled run with extra steps.

Neither of those is examinable. Both are why the feature is worth having.

What to gate

The useful axis is reversibility, not danger. A tool that writes a file you can delete is not the same risk as one that sends an email, spends money, or force-pushes.

So the practical list is short: anything that reaches outside the agent’s own environment, anything that costs money, anything a human would need to apologise for. execute is a strong candidate whenever the backend is the host machine, which module 2 will show you is the default.

Gates against permissions

These two get confused because both read as safety. They are answering different questions.

  • Permissions encode what is allowed, statically, evaluated in order, and they fail open where you forgot to write a rule. Module 2, lesson 3.
  • interrupt_on encodes who decides, at the moment of the call, with a human in the path.

A static rule cannot express “it depends”. A gate cannot cover a path you did not think about. Real systems use both, and knowing which one you are reaching for is the difference between defence in depth and two half-measures that each assume the other one is doing the work.

Practice

Try it yourself

Recall

The parameter and its prerequisite

Two facts that belong together, because the second one is the part people leave out and it is the part that breaks in production.

What does interrupt_on configure, and what must also be present for it to work across a real delay?

Reveal answer

interrupt_on names which tools pause for human sign-off before executing. It is one of the six deep agent additions. It requires a checkpointer, because the pause is a persisted state rather than a blocked process: the agent stops, the graph state is written down, and a later invocation resumes from it. Without persistence there is nothing for the human's answer to come back to.

Quiz

What is underneath it

interrupt_on is a deep agent parameter. What actually implements the behaviour?

  1. AHumanInTheLoopMiddleware, which is importable into a plain create_agent
  2. BA dedicated pause node injected into the graph by the deep agent constructor
  3. CThe checkpointer, which halts execution when it detects a gated tool
  4. DPatchToolCalls middleware, which rewrites the gated call into a no-op
Show answer

Correct answer: A — HumanInTheLoopMiddleware, which is importable into a plain create_agent

HumanInTheLoopMiddleware is a built-in that exists in both languages and can be imported into a plain create_agent call. interrupt_on is the friendlier spelling, not a new capability. The checkpointer option is the tempting one because persistence is genuinely required, but the checkpointer stores state rather than deciding when to stop, and confusing a prerequisite with a mechanism is a reliable way to get a layer-ownership question wrong.

Quiz

The human answers tomorrow

A gated tool pauses at 6pm. The process is restarted overnight. At 9am the human approves. What has to have been true for the approval to land?

  1. AThe original process must still be running, since an interrupt holds the call open until answered
  2. BThe state was persisted by a checkpointer and the resuming call uses the same thread_id
  3. CThe approval must arrive inside the middleware's configured timeout window
  4. DNothing special; approvals are replayed from the stored message history
Show answer

Correct answer: B — The state was persisted by a checkpointer and the resuming call uses the same thread_id

The pause is state on disk, addressed by thread_id, so resumption needs both the checkpointer that wrote it and the identifier that finds it. The first option is the intuition that comes from thinking of an interrupt as a blocking call, and it is exactly the picture that makes overnight approval seem impossible. It is not a held process. It is a saved position.

Quiz

Gate or permission

An agent can write anywhere on a shared filesystem, and you want a human to sign off before it touches production paths. Which lever is the right primary one?

  1. APermissions, because path-scoped filesystem rules are exactly this problem
  2. Binterrupt_on, because a human decision is required rather than a static rule
  3. CNeither; this belongs in the system prompt as a standing instruction
  4. DBoth, but permissions alone are sufficient if the list is complete
Show answer

Correct answer: B — interrupt_on, because a human decision is required rather than a static rule

The requirement says a human signs off, which is a decision that cannot be expressed as a static rule, so it is an interrupt. Permissions are the strong distractor and they are genuinely part of a good answer, but a permission list encodes what is allowed rather than who decides, and it fails open where you forgot to write a rule. Use both, and know which one is doing which job.

Do

Design a gate list

Ten minutes with a scratch file. This is a judgement exercise, not a coding one.

  • List every tool an agent of yours can call, including the deep agent defaults.
  • Mark each one reversible or irreversible, judged by whether you could undo it in under a minute.
  • Put every irreversible one in a candidate interrupt_on list, then cross off any where the approval would arrive too often to be read.
Done whenMy final list is short enough that a human would actually read each prompt, and I can name at least one irreversible tool I deliberately left ungated along with the reason.
Sign in to track your progress →