Whetstone.
MemoryMemory in LangGraph
Module 4, Lesson 118 min

Memory in LangGraph

Memory is a whole module in the official course and it is the part of this domain most likely to be under-studied, because it looks like an implementation detail and is actually two distinct components with different scopes, plus a database that is never optional, plus a cache that holds far less than anybody assumes.

Two components, two scopes

  • Checkpointer: short-term, thread-scoped. The state of one thread as it moves through a graph.
  • Store: long-term, cross-thread. Memory that outlives and spans individual conversations.

TTLs are configurable separately for each, which follows naturally from the scopes. A conversation’s working state and a user’s durable preferences do not want the same expiry policy, so they do not share one.

Both are declared in langgraph.json, in the where-state-goes family, which is a nice illustration of a recurring shape here: build-time declaration of a runtime behaviour.

Postgres is not optional

Read that list once more and notice how much of this domain’s vocabulary is sitting on it. Threads, runs, assistants: the entire runtime layer from module one lives in Postgres.

The examinable consequence is the one people get wrong. You can choose the MongoDB checkpointer backend. Doing so does not free you from Postgres. The checkpointer is one consumer of storage; everything on that list is still Postgres, always. It is a clean question with a very tempting wrong answer, because when people think “where does state go” they think checkpointer, so swapping the checkpointer feels like swapping the storage layer.

What Redis is for, which is less than you think

Redis holds ephemeral metadata. That is the entire remit. No user data. No run data. And specifically, streaming output is broadcasted but never stored.

A stream is a delivery mechanism, not a record. You watched all those tokens go past; none of them were written down as tokens. If you need to know afterwards what happened, you read the thread’s checkpointed state.

Durability modes

There are three, and they are literal strings worth recognising on sight:

The three durability modes
'sync' | 'async' | 'exit'

The default is 'async'. Read the three in order and they are a straight line from safest to fastest:

  • 'sync': changes are persisted before the next step starts.
  • 'async': changes are persisted while the next step executes. This is the default, and it is the throughput-versus-durability compromise.
  • 'exit': changes are persisted only when the graph exits. Fastest, and the one with the most to lose if the process dies mid-run.

Recognising the set matters as much as the default, so that an option offering 'strict' | 'relaxed' | 'none' gets binned instantly. Both halves are cheap marks.

In your own codebase, the half the ground school did not map

The ground school already mapped Ark onto the short-term half: state/session.json is the thread, the continuity writeback is a checkpoint, and state/incident-log.jsonl is the checkpoint history.

The store is the half that mapping left out, and it is sitting in the same repository. workspace/memory/ is cross-thread long-term memory by hand: continuity-anchor.md and MEMORY.md persist across every session boundary, are namespaced by topic rather than by conversation, and are read on boot regardless of which session is booting. The daily letter is the interesting middle case, because it is scoped to one day the way a checkpoint is scoped to one thread, and then it is deliberately promoted into the durable files when something on it turns out to matter.

That promotion step is exactly what a store write is: the decision that this piece of state should outlive the conversation that produced it. Ark makes the decision manually, in prose, every night. The store makes it an API call. Same architecture, one of them hand-rolled.

Practice

Try it yourself

Quiz

You choose the MongoDB checkpointer backend

A team wants to standardise on MongoDB and configures the MongoDB checkpointer backend. What is their Postgres situation?

  1. APostgres is no longer required, since checkpointing was the only component using it
  2. BPostgres is required only in Cloud deployments, since self-hosted uses your own store
  3. CPostgres becomes optional but remains recommended for query performance at scale
  4. DPostgres is still required: threads, runs, assistants, crons and the memory store all live there
Show answer

Correct answer: D — Postgres is still required: threads, runs, assistants, crons and the memory store all live there

LangSmith always requires PostgreSQL for threads, runs, assistants, crons, and the memory store. Choosing a different checkpointer backend swaps out one consumer of storage, not the platform's own database. The tempting wrong answer is the first one, because the checkpointer is the storage component people think about most, so replacing it feels like replacing the storage layer. It is not: the checkpointer is one item on a list, and the other items do not move.

Recall

Checkpointer against store

Two persistence components, two scopes. This distinction shows up in questions across several domains, not only this one.

What is the scope and lifespan of the checkpointer versus the store, and what is configurable separately for each?

Reveal answer

The checkpointer is short-term and thread-scoped: it holds the state of one thread as that thread progresses. The store is long-term and cross-thread: it holds memory that has to outlive and span individual conversations. TTLs are configurable separately for each, which follows directly from them having genuinely different lifespans. Short-term thread state and long-term cross-thread memory are the two phrases to keep.

Quiz

The durability modes

One of these sets names the actual durability modes.

Four candidate unions
type A = 'strict' | 'relaxed' | 'none'
type B = 'always' | 'onError' | 'never'
type C = 'sync' | 'async' | 'exit'
type D = 'immediate' | 'deferred' | 'batched'
  1. ASet A
  2. BSet B
  3. CSet C
  4. DSet D
Show answer

Correct answer: C — Set C

The modes are 'sync', 'async' and 'exit', and the default is 'async'. The other three sets are built to sound like configuration vocabulary you have met elsewhere, and set A is the most tempting because strict, relaxed and none is a very common shape in consistency and validation settings across many systems. These are exactly the sort of literal string values a semi-open-book exam is happy to test, because looking them up costs time you do not have.

Recall

The durability default

Three modes, one of them is what you get when you say nothing. The default is the half of this fact people skip.

What are the three durability modes, which one is the default, and what does each actually do?

Reveal answer

'sync', 'async' and 'exit', and the default is 'async'. sync persists changes before the next step starts. async persists changes while the next step executes, which is the default because it is the throughput-versus-durability compromise most runs want. exit persists only when the graph exits, which is the fastest and the one that loses the most if the process dies mid-run. Read the list in that order and it is a straight line from safest to fastest.

Recall

What Redis does not hold

Be precise about the negative here. The negative is considerably more examinable than the positive.

What does Redis hold in this architecture, and what does it explicitly not hold?

Reveal answer

Redis holds ephemeral metadata only. It explicitly does not hold user data or run data. Streaming output in particular is broadcasted but never stored, so a stream is a transient delivery mechanism rather than a persisted record. The practical consequence is that if you want a record of what was streamed, the durable copy comes from the thread's checkpointed state in Postgres, and nothing in the streaming path leaves a trace by itself.

Check

The restart drill

Reason it through in one pass, then check yourself against every box on your capstone drawing.

You should see

You can say that threads, runs, assistants, crons and the store all survive because they are in Postgres, that checkpointed thread state survives for the same reason, that ephemeral metadata in Redis does not survive and was never meant to, and that streamed tokens were never stored in the first place so there is nothing for a restart to lose. You can also state that a new revision does not empty any of the Postgres-backed state, because that state lives in the storage layer rather than in the deployed unit.

Sign in to track your progress →