Memory in deep agents
There are two things in this stack called memory and they are not the same thing. That sentence is the entire lesson, and it is exactly the sort of overlap a certification exam is built to exploit, because conflating two identically named concepts is a question that writes itself.
The one you already know
The LangGraph store is long-term, cross-thread memory living in the platform’s storage layer. You write to it and read from it through the Store API, it is organised by namespace, it has its own TTL configuration, and it is contrasted against the checkpointer’s short-term thread-scoped state. That is the previous two lessons.
The one deep agents add
A deep agent can be configured with a memory list. What that produces is always-loaded context, handled by MemoryMiddleware, and it is distinct from the LangGraph store.
That is architecturally different from the store in a way that goes beyond where the bytes live. The store is a place you ask. This is a thing that arrives.
The trade, stated plainly
Always-loaded context buys reliability: the agent has definitely seen it, so behaviour that depends on it is deterministic rather than contingent on a retrieval firing correctly. It costs tokens on every turn, forever, relevant or not.
Retrieved memory buys efficiency: nothing is paid for until something needs it. It costs contingency: the agent knows the fact only if the retrieval found it, and retrieval failures are quiet.
So the sorting rule is about size and universality rather than about importance. Small and always relevant goes in always-loaded context. Large and occasionally relevant goes behind search. A standing instruction about how to behave belongs in the first. Six months of conversation history belongs in the second, and putting it in the first is how you build an agent whose every reply costs a fortune and whose context window is full of last March.
What to check first when a question says memory
Ask which layer the question is about before you ask anything else. Platform storage, meaning the store and the checkpointer, sits underneath the deployment and is configured in langgraph.json. Agent context assembly, meaning middleware deciding what the model sees this turn, sits inside the agent and is configured where the agent is constructed.
They coexist happily in one deployment. They answer completely different questions. And an answer option that is entirely correct about one of them is, most of the time, the wrong answer to a question about the other.
Try it yourself
What deepagents memory actually is
A deep agent is configured with a memory list.
agent = create_deep_agent(
tools=[...],
memory=[...],
)Show answer
Correct answer: B — Always-loaded context handled by MemoryMiddleware, distinct from the LangGraph store
The deepagents memory list is always-loaded context, handled by MemoryMiddleware, and it is a different thing from the LangGraph store. The tempting wrong answer is the retrieval-tool option, because retrieval-augmented memory is by far the most common pattern in agent frameworks and the word memory has been used for it for years: the distinction here is that always-loaded means the model sees it every turn without deciding to, which changes both the behaviour and the token cost. The store option is the one an exam is most likely to use, because conflating two things both called memory is exactly the kind of question a certification writes.
Always-loaded against retrieved
The behavioural difference matters more than the plumbing, because it changes what the agent can be relied on to know.
What is the practical difference between context that is always loaded and memory that is retrieved on demand?
Reveal answer
Always-loaded context is in front of the model on every turn whether or not it is relevant, so the agent can be relied upon to have seen it, and you pay for it in tokens every single time. Retrieved memory is fetched only when something triggers the retrieval, so it costs nothing when unused, but the agent only knows it if the retrieval fired. The trade is reliability against cost: always-loaded guarantees presence, retrieval guarantees efficiency, and choosing wrongly gives you either an expensive agent or an agent that forgets things it demonstrably stored.
Which mechanism for which fact
One of these belongs in always-loaded context rather than in searchable long-term memory.
Show answer
Correct answer: D — The short standing instruction describing how this agent must behave for this tenant
A short standing instruction that must shape every single response is the textbook case for always-loaded context: it is small, it is universally relevant, and an agent that only sometimes sees it is an agent that only sometimes behaves correctly. The tempting wrong answer is the conversation transcript, because history feels like the most memory-shaped thing on the list and it is genuinely valuable: it is also large and only occasionally relevant, which is precisely the profile that wants retrieval. Ticket histories and documentation are both large corpora and both belong behind search.
Two things called memory, in one system
Name the confusion explicitly, because naming it is what stops it happening under time pressure.
A deployed deep agent can have both a LangGraph store and a deepagents memory configuration. What is each one for, and what should you check first when a question mentions memory?
Reveal answer
The LangGraph store is long-term cross-thread memory living in the platform's storage layer, written and read through the Store API, scoped by namespace, and it is the thing a checkpointer is contrasted against. The deepagents memory configuration is always-loaded context assembled by MemoryMiddleware and placed in front of the model each turn. When a question mentions memory, check first which layer it is talking about: platform storage or agent context assembly. They coexist in one deployment and they answer different questions, so an answer that is right about one is usually wrong about the other.