Running one, and choosing the model
Constructing a deep agent is one call. Running it is exactly what you already know how to do, and that is not a coincidence.
Running one
from deepagents import create_deep_agent
agent = create_deep_agent(
model="anthropic:claude-sonnet-4-6",
tools=[search, summarise],
system_prompt="You are a research assistant.",
)
result = agent.invoke({"messages": [{"role": "user", "content": "..."}]})create_deep_agent returns a compiled StateGraph. Not a wrapper around one, not a lookalike. The literal artifact you already know how to invoke, stream, checkpoint and rewind.
So every reflex you built in ground school still fires unchanged: .invoke(), .stream(), a checkpointer, a thread_id, state history. What you skipped was drawing the graph.
The model parameter
model is inherited straight from create_agent, and it takes two forms.
A provider-prefixed string, in the shape provider:model. The library resolves the provider and constructs the client for you. This is the form you want ninety percent of the time and it is what every example uses.
An initialised chat model instance. Reach for this when you need to configure the client yourself: temperature, a custom base URL, provider-specific options that a string has no room to express.
The string form is not a lesser version of the instance form. It is the same object, constructed for you from a shorter spelling.
The model is a budget, not just a brain
Choosing a model sets three things at once, and only the first is obvious.
Capability, which is what you were thinking about.
The context window, which is a property of the model and not something you configure. Module 3 measures input truncation against this number, so the model you pick silently moves that threshold.
Cost per turn, which a long-running agent multiplies by a number you did not choose in advance. A deep agent given a large corpus and four hours can take hundreds of model calls.
That last one is why the built-in catalogue has a whole family aimed at it. ModelCallLimitMiddleware and ToolCallLimitMiddleware are hard stops that count and refuse. ModelFallbackMiddleware and ModelRetryMiddleware handle a provider that is up but flaky.
The distinction between those two families is worth holding: limits count, retries re-invoke. A limit middleware cannot rescue a failed call and a retry middleware cannot bound your bill. They solve adjacent problems and the names are close enough to swap under pressure.
Where this goes next
You have the constructor, the return type and the model. The next three lessons fill in the three parameters you will actually spend your configuration time on: the system prompt, tools, and the MCP wiring that brings somebody else’s tools in.
Try it yourself
What comes back from the constructor
Worth knowing precisely, because the answer explains why every graph reflex you already have still fires.
Show answer
Correct answer: C — A compiled StateGraph
It returns a compiled StateGraph, exactly as create_agent does, because it is create_agent underneath. That is why invoke, stream, checkpointing, threads and time travel behave identically to anything you already built by hand. The uncompiled builder is the tempting answer, because that is the shape you get when you assemble a graph yourself, but the prebuilt constructors hand you the compiled artifact ready to run.
Two ways to hand over a model
A small mechanical fact that shows up inside larger questions rather than on its own.
What are the two forms the model parameter accepts, and what does the string form encode?
Reveal answer
Either a provider-prefixed identifier string, in the shape provider:model, or an already-initialised chat model instance. The string form encodes the provider and the model name together so the library can resolve and construct the client for you. Passing an instance is what you do when you need to configure the client yourself, for example to set temperature, a custom base URL, or provider-specific options that the string form has no way to express.
A flaky provider
Your model provider is intermittently returning 503s and the agent run dies partway through. You want the run to survive without the agent loop ever seeing the failure. Which built-in is the shape you need?
Show answer
Correct answer: D — ModelFallbackMiddleware
ModelFallbackMiddleware, with ModelRetryMiddleware as its sibling, is the family aimed at flaky infrastructure, and both work by wrapping the model call so they can re-invoke it. ModelCallLimitMiddleware is the tempting neighbour because the names are almost identical, but a limit middleware only counts and stops. It bounds cost. It cannot rescue a failed call, because counting is not re-invoking.
Whose number is it
max_input_tokens and the model's context window are two different quantities, and confusing them costs you a question in module 3.
Show answer
Correct answer: C — max_input_tokens is a knob you set, the context window is a property of the model
max_input_tokens is yours to configure, the context window belongs to the model. This matters because module 3 puts an 85 percent figure against each of them for two different mechanisms, and the only clean way to keep those apart is to remember that one denominator is a knob and the other is a fact about the model. The last option is the exact inversion, and it is what you will pick if you have half-remembered the pairing.
Sketch the invocation
No cloud account needed. This is a five-minute paper exercise about shape, not a build.
- Write a create_deep_agent call in a scratch file with a model string, a tools list and a system prompt.
- Underneath it, write the two lines that invoke it and then stream it.
- Add a one-line comment above each saying which layer of the three-layer stack that line's behaviour comes from.