Whetstone.
Building a Deep AgentMessages, threads, and checkpointers
Module 1, Lesson 616 min

Messages, threads, and checkpointers

Nothing in this lesson is new, and that is the lesson. A deep agent is a compiled StateGraph, so its conversation model is the runtime’s conversation model, unchanged.

Three things, cleanly separated

Messages are the conversation state itself, accumulated and merged by the runtime.

A thread_id identifies which conversation you are in. It is passed in the invocation config, not baked into the agent.

A checkpointer persists graph state after each step, so a later call can pick the conversation back up.

The Ark mapping, if you want a concrete one

Your own continuity architecture is the same shape and it is the fastest way to keep the roles straight.

state/session.json names which session is current. That is the thread_id: an address, not a store. The continuity letter is what actually survives the night, written down deliberately so the next session can read it. That is the checkpointer.

The failure modes map too. A session id pointing at nothing is a thread_id with no checkpointer. A letter nobody knows to read is a checkpointer with no thread_id. Both leave you starting from scratch, and neither raises an error, which is exactly what makes the class of bug annoying in both systems.

Where the parameter lives, and where the work happens

checkpointer is a create_agent parameter, inherited unchanged by create_deep_agent. It is not one of the six deep agent additions, which makes it a reliable distractor in “which of these is new” questions: it feels like heavyweight long-running-agent machinery, and deepagents is the long-running-agent product.

But the persistence is executed by LangGraph. The parameter is a form field on the harness. The behaviour belongs to the runtime.

Keep asking that question of every feature in this course. Where do I configure it, and where does it run? They are frequently different layers and the exam is built on the gap between them.

The consequence that matters most

Two things in the rest of this course only work because a checkpointer exists.

Human in the loop, the next lesson, is a persisted pause. The agent stops, state is written down, and a later call resumes it. If you picture the pause as a process blocking until someone answers, everything about the next lesson will feel wrong, because the human might answer an hour later from a different machine entirely.

Long-running work, which is the entire premise of the product. Four hours of agent time is four hours of things that can crash, and without persisted state a crash is a restart from nothing.

Practice

Try it yourself

Quiz

Second call, no checkpointer

An agent is built without a checkpointer. You invoke it, then invoke it again passing the same thread_id in the config. What does the second run see?

  1. AThe full history from the first run, because thread_id is what carries it
  2. BThe first run's final message only, replayed from the thread label
  3. CAn error, because passing thread_id without a checkpointer is invalid
  4. DNothing from the first run, because there was nowhere to write state
Show answer

Correct answer: D — Nothing from the first run, because there was nowhere to write state

A thread_id is a label, not a store. Without a checkpointer nothing was persisted, so there is nothing to load and the second run starts clean. The first option is the near-universal wrong answer, because thread_id is the thing you type when you want continuity, so it feels like the mechanism. It is the key. The checkpointer is the cupboard.

Recall

Key and cupboard

A distinction people collapse constantly, and collapsing it makes the persistence questions unanswerable.

What is the division of responsibility between thread_id and a checkpointer?

Reveal answer

The checkpointer is the persistence mechanism: it writes graph state after each step and can read it back. thread_id is the identifier that says which conversation's state to write to and read from. One is storage, the other is addressing. A checkpointer with no thread_id has nowhere specific to put things, and a thread_id with no checkpointer names a location in a store that does not exist.

Quiz

Whose feature is this

Checkpointing is what makes a long-running deep agent resumable. Which layer of the three-layer stack actually implements it?

  1. ALangGraph, because checkpointing is a property of the graph runtime
  2. BDeep Agents, because only deep agents run long enough to need it
  3. CThe middleware stack, since Filesystem middleware handles persistence
  4. DLangChain's create_agent, because that is where the parameter appears
Show answer

Correct answer: A — LangGraph, because checkpointing is a property of the graph runtime

Checkpointing belongs to the graph runtime. create_agent exposes a checkpointer parameter and create_deep_agent inherits it, but the persistence itself is executed by LangGraph. The last option is the sharpest distractor because it is half true: that is genuinely where you type the parameter. Where you configure a thing and where it executes are different questions, and this domain asks the second one.

Quiz

The prerequisite nobody mentions

You configure interrupt_on so a tool pauses for approval. The pause has to survive the gap between the agent stopping and a human answering. What does that require?

  1. ANothing extra, the interrupt holds the process open until an answer arrives
  2. BA checkpointer, so the paused state is persisted and can be resumed
  3. CA sandbox backend, so the paused tool call is isolated while it waits
  4. DA subagent, so the parent can continue while the child waits for a human
Show answer

Correct answer: B — A checkpointer, so the paused state is persisted and can be resumed

A human-in-the-loop pause is a persisted state that a later call resumes, which means it needs a checkpointer to exist in. The first option is the intuitive one if you picture the pause as a blocking call, and that picture is what breaks the moment the human answers an hour later from a different process. Holding a process open is not the mechanism; writing state down and picking it up again is.

Check

Map it onto something you already run

Take any system of yours that survives a restart and describe it in this vocabulary.

You should see

You named the thing that plays the role of the checkpointer, the thing that plays the role of the thread_id, and you can say what would break if you had one without the other.

Sign in to track your progress →