Whetstone.
Building a Deep AgentRunning one, and choosing the model
Module 1, Lesson 217 min

Running one, and choosing the model

Constructing a deep agent is one call. Running it is exactly what you already know how to do, and that is not a coincidence.

Running one

Construct, then run
from deepagents import create_deep_agent

agent = create_deep_agent(
    model="anthropic:claude-sonnet-4-6",
    tools=[search, summarise],
    system_prompt="You are a research assistant.",
)

result = agent.invoke({"messages": [{"role": "user", "content": "..."}]})

create_deep_agent returns a compiled StateGraph. Not a wrapper around one, not a lookalike. The literal artifact you already know how to invoke, stream, checkpoint and rewind.

So every reflex you built in ground school still fires unchanged: .invoke(), .stream(), a checkpointer, a thread_id, state history. What you skipped was drawing the graph.

The model parameter

model is inherited straight from create_agent, and it takes two forms.

A provider-prefixed string, in the shape provider:model. The library resolves the provider and constructs the client for you. This is the form you want ninety percent of the time and it is what every example uses.

An initialised chat model instance. Reach for this when you need to configure the client yourself: temperature, a custom base URL, provider-specific options that a string has no room to express.

The string form is not a lesser version of the instance form. It is the same object, constructed for you from a shorter spelling.

The model is a budget, not just a brain

Choosing a model sets three things at once, and only the first is obvious.

Capability, which is what you were thinking about.

The context window, which is a property of the model and not something you configure. Module 3 measures input truncation against this number, so the model you pick silently moves that threshold.

Cost per turn, which a long-running agent multiplies by a number you did not choose in advance. A deep agent given a large corpus and four hours can take hundreds of model calls.

That last one is why the built-in catalogue has a whole family aimed at it. ModelCallLimitMiddleware and ToolCallLimitMiddleware are hard stops that count and refuse. ModelFallbackMiddleware and ModelRetryMiddleware handle a provider that is up but flaky.

The distinction between those two families is worth holding: limits count, retries re-invoke. A limit middleware cannot rescue a failed call and a retry middleware cannot bound your bill. They solve adjacent problems and the names are close enough to swap under pressure.

Where this goes next

You have the constructor, the return type and the model. The next three lessons fill in the three parameters you will actually spend your configuration time on: the system prompt, tools, and the MCP wiring that brings somebody else’s tools in.

Practice

Try it yourself

Quiz

What comes back from the constructor

Worth knowing precisely, because the answer explains why every graph reflex you already have still fires.

  1. AA bespoke DeepAgent class
  2. BAn uncompiled graph builder
  3. CA compiled StateGraph
  4. DA plain async function
Show answer

Correct answer: C — A compiled StateGraph

It returns a compiled StateGraph, exactly as create_agent does, because it is create_agent underneath. That is why invoke, stream, checkpointing, threads and time travel behave identically to anything you already built by hand. The uncompiled builder is the tempting answer, because that is the shape you get when you assemble a graph yourself, but the prebuilt constructors hand you the compiled artifact ready to run.

Recall

Two ways to hand over a model

A small mechanical fact that shows up inside larger questions rather than on its own.

What are the two forms the model parameter accepts, and what does the string form encode?

Reveal answer

Either a provider-prefixed identifier string, in the shape provider:model, or an already-initialised chat model instance. The string form encodes the provider and the model name together so the library can resolve and construct the client for you. Passing an instance is what you do when you need to configure the client yourself, for example to set temperature, a custom base URL, or provider-specific options that the string form has no way to express.

Quiz

A flaky provider

Your model provider is intermittently returning 503s and the agent run dies partway through. You want the run to survive without the agent loop ever seeing the failure. Which built-in is the shape you need?

  1. AModelCallLimitMiddleware
  2. BSummarizationMiddleware
  3. CLLMToolSelectorMiddleware
  4. DModelFallbackMiddleware
Show answer

Correct answer: D — ModelFallbackMiddleware

ModelFallbackMiddleware, with ModelRetryMiddleware as its sibling, is the family aimed at flaky infrastructure, and both work by wrapping the model call so they can re-invoke it. ModelCallLimitMiddleware is the tempting neighbour because the names are almost identical, but a limit middleware only counts and stops. It bounds cost. It cannot rescue a failed call, because counting is not re-invoking.

Quiz

Whose number is it

max_input_tokens and the model's context window are two different quantities, and confusing them costs you a question in module 3.

  1. ABoth are properties of the model, and neither one is configurable
  2. BBoth are settings you configure on the summarization middleware
  3. Cmax_input_tokens is a knob you set, the context window is a property of the model
  4. Dmax_input_tokens is a property of the model, the context window is a knob you set
Show answer

Correct answer: C — max_input_tokens is a knob you set, the context window is a property of the model

max_input_tokens is yours to configure, the context window belongs to the model. This matters because module 3 puts an 85 percent figure against each of them for two different mechanisms, and the only clean way to keep those apart is to remember that one denominator is a knob and the other is a fact about the model. The last option is the exact inversion, and it is what you will pick if you have half-remembered the pairing.

Do

Sketch the invocation

No cloud account needed. This is a five-minute paper exercise about shape, not a build.

  • Write a create_deep_agent call in a scratch file with a model string, a tools list and a system prompt.
  • Underneath it, write the two lines that invoke it and then stream it.
  • Add a one-line comment above each saying which layer of the three-layer stack that line's behaviour comes from.
Done whenMy invoke and stream lines are ordinary graph calls, and my comments attribute both to the LangGraph runtime rather than to deepagents.
Sign in to track your progress →