Whetstone.
Create AgentFoundational models
Module 1, Lesson 115 min

Foundational models

The model is the first positional argument to create_agent and it can be two completely different things. Getting that distinction straight now saves an argument with yourself later.

Two ways to name a model

String, or configured instance
from langchain.chat_models import init_chat_model
from langchain.agents import create_agent

# 1. A bare string. Fastest thing that works.
agent = create_agent("gpt-5.5", tools=[...])

# 2. An instance, when you need to configure anything at all.
model = init_chat_model(
    "google_genai:gemini-2.5-flash-lite",
    temperature=0,
    timeout=30,
    max_retries=10,
)
agent = create_agent(model, tools=[...])

The string form is provider:model, and the provider prefix is optional when the model name is unambiguous. Where a model identifier is ambiguous or vendor-namespaced, you pass model_provider separately instead, which is how Bedrock identifiers like us.anthropic.claude-sonnet-4-6 are wired up with model_provider="bedrock_converse".

The standard parameter surface

Five parameters recur across every provider, and one of them has a number worth remembering.

  • model, the identifier.
  • temperature, randomness of the output.
  • max_tokens, a cap on the response length.
  • timeout, seconds to wait before giving up.
  • max_retries, default 6, raised to roughly 10 to 15 on unreliable networks.

That retry count is a transport-level retry inside the model client. It is not the same mechanism as ModelRetryMiddleware, which sits in the agent’s middleware array and retries at a completely different layer. Two retry systems, stacked, and they multiply rather than replace each other.

Three ways to call it

invoke takes messages and returns the finished response. stream returns the same thing progressively as it is generated. batch sends many independent requests together for throughput.

Two more methods matter because create_agent is doing them for you under the hood: bind_tools attaches tools to a model, and with_structured_output constrains the response to a schema.

What create_agent does on your behalf
model_with_tools = model.bind_tools([get_weather])
model_with_structure = model.with_structured_output(Movie)

You will rarely call either directly once you are using the harness. Recognise them, because they show up constantly in older code and in questions about what the agent is doing internally.

The profile attribute

A model exposes what it can do through profile.

Capability data, not guesswork
model.profile["max_input_tokens"]
model.profile["tool_calling"]

This looks like trivia now. It is not: the compression middleware in module 3 can be configured as a fraction of the model’s maximum input, and that fraction is computed against exactly this number. When profile data is unavailable, you supply the maximum by hand, which is the fallback rather than the normal path.

Practice

Try it yourself

Quiz

Reading a provider-prefixed model string

Two ways of naming the same model.

Which of these is the provider-prefixed form
a = init_chat_model("claude-sonnet-4-6")
b = init_chat_model("google_genai:gemini-2.5-flash-lite")
c = init_chat_model("us.anthropic.claude-sonnet-4-6", model_provider="bedrock_converse")
  1. AOnly a, because it names the model without any provider
  2. BOnly b, because the prefix before the colon is the provider
  3. Cb and c, since both name a provider somewhere in the call
  4. DNone of them; the provider is always a separate argument
Show answer

Correct answer: B — Only b, because the prefix before the colon is the provider

The provider-prefixed form is provider then colon then model, which is b. In c the provider is supplied separately through model_provider, and the string itself is a Bedrock model identifier that happens to contain dots. The third option is the tempting one because c does clearly name a provider, but naming a provider and using the prefixed string form are two different mechanisms.

Recall

The retry default

One of the few standard parameters with a documented numeric default, which makes it exactly the kind of thing a question can pin down precisely.

What is the default value of max_retries on a chat model, and when would you raise it?

Reveal answer

The default is 6. You raise it, to somewhere around 10 to 15, when the network between you and the provider is unreliable, for example an agent running from a flaky connection or against a provider that rate-limits aggressively. Note that this is a transport-level retry on the model client itself, which is a different mechanism from ModelRetryMiddleware sitting in the agent's middleware array.

Quiz

String or instance

You need a model with temperature set to zero and a thirty second timeout, passed into create_agent. Which form do you have to use?

  1. AThe string form, since create_agent parses parameters out of it
  2. BEither, since create_agent forwards the same keyword arguments
  3. CNeither; temperature is set at invoke time, not at construction
  4. DAn instance, because a bare string carries no configuration
Show answer

Correct answer: D — An instance, because a bare string carries no configuration

A model string identifies a model and nothing else, so anything configured needs an instance you built yourself, usually via init_chat_model with the parameters set. The second option is genuinely tempting because create_agent has a large keyword surface and it feels like it should forward model parameters, but it does not, and there is no temperature parameter on create_agent at all.

Quiz

Where max_input_tokens comes from

A later lesson configures compression as a fraction of the model's maximum input. Where does that maximum come from?

  1. AFrom the model's profile attribute, which exposes capabilities like max_input_tokens
  2. BThe middleware carries a hardcoded table of limits, one entry per provider
  3. CIt must be supplied by hand every time, because nothing on the model exposes it
  4. DFrom counting the tokens already in the conversation, which sets the ceiling
Show answer

Correct answer: A — From the model's profile attribute, which exposes capabilities like max_input_tokens

Models expose a profile attribute carrying capability data, including max_input_tokens and tool_calling, and fraction-based configuration reads it from there. The third option is the tempting one because supplying the maximum by hand is a real fallback when profile data is unavailable, but it is the fallback rather than the mechanism.

Check

Name the three invocation methods

Close the lesson and write down the three ways of calling a model, with one sentence each on when you would pick it.

You should see

You named invoke, stream and batch, and your one-liners distinguish a complete response, a progressive one, and many independent requests sent together.

Sign in to track your progress →