Pydantic, the one that actually validates
This is the most important third-party library in Python for the work you are about to do. FastAPI is built on it, LangChain uses it for every tool schema and every structured output, and it is the only thing in this module that does anything at runtime.
What it is
A class whose annotations are executed rather than ignored.
from pydantic import BaseModel, Field
class SearchQuery(BaseModel):
query: str
limit: int = Field(default=10, ge=1, le=100)
include_archived: bool = False
q = SearchQuery.model_validate({"query": "python", "limit": "25"})
q.limit # 25, as an int. The string was coerced.
SearchQuery.model_validate({"query": "python", "limit": 500})
# ValidationError: limit -> Input should be less than or equal to 100Compare with the dataclass from the previous lesson, which would have stored the string "25" and accepted 500 without comment. Same-looking class, completely different runtime contract.
The core API is four methods, and the model_ prefix is deliberate so they cannot collide with your field names:
SearchQuery.model_validate(some_dict) # parse and validate
SearchQuery.model_validate_json(raw_text) # straight from a JSON string
q.model_dump() # -> dict
q.model_dump_json() # -> JSON string
SearchQuery.model_json_schema() # -> JSON SchemaCoercion, which will surprise you once
The default is lax mode, which converts where the conversion is unambiguous. For an int field, the documented rules are that a str is accepted if it is numeric only, a float is accepted if it has no fractional part, and a bool is accepted.
That is genuinely the right default at a boundary, where a form post or a query string delivers everything as text and you would otherwise write conversion code by hand. It is also genuinely surprising the first time it accepts something you expected it to reject.
from pydantic import ConfigDict, StrictInt
class Strict(BaseModel):
model_config = ConfigDict(strict=True)
attempts: int # now only a real int is accepted
class PartlyStrict(BaseModel):
attempts: StrictInt # just this field
name: str # still laxThe reason it is a hard dependency of every agent framework
This is worth understanding rather than accepting, because it explains a design decision that otherwise looks like taste.
An LLM tool call needs a JSON Schema describing the tool’s arguments, sent to the model provider. The model’s reply comes back as text, and something has to turn that text into a validated object or reject it.
Annotations alone cannot do either, because the interpreter ignores them. Pydantic reads them and produces both halves from one class definition:
class WeatherArgs(BaseModel):
city: str = Field(description="City name, e.g. 'London'")
units: Literal["c", "f"] = "c"
WeatherArgs.model_json_schema() # -> sent to the model as the tool schema
WeatherArgs.model_validate_json(raw) # <- validates what came backThe description on a Field is not a comment. It lands in the generated schema and the model reads it, which makes it prompt text disguised as a type annotation. Vague field descriptions produce bad tool calls, and this is one of the highest-leverage places to be specific in the entire agent stack.
This same machinery is what with_structured_output uses in LangChain: you hand it a model class, it generates the schema, constrains the response, and validates the result.
Validators
from pydantic import BaseModel, field_validator, model_validator
class Window(BaseModel):
start: int
end: int
@field_validator("start", "end")
@classmethod
def non_negative(cls, v: int) -> int:
if v < 0:
raise ValueError("must be non-negative")
return v
@model_validator(mode="after")
def ordered(self) -> "Window":
if self.start >= self.end:
raise ValueError("start must be before end")
return self@field_validator runs per field and sees one value. @model_validator(mode="after") runs once on the constructed instance and can see every field at once, which is what you need for any rule relating two fields. Both signal failure by raising ValueError, which Pydantic collects into a ValidationError rather than letting it escape.
That collection behaviour is the part worth noticing: a ValidationError carries every failure at once, each with a location path and an error type, so exc.errors() is directly shapeable into an API error response. It does not stop at the first problem.
Reading old answers
Pydantic v2 was a rewrite with a Rust core, and the API was renamed wholesale. Most search results and most model-generated Python still show v1.
| v1 | v2 |
|---|---|
.dict() |
.model_dump() |
.json() |
.model_dump_json() |
parse_obj() |
.model_validate() |
parse_raw() |
.model_validate_json() |
.copy() |
.model_copy() |
construct() |
.model_construct() |
schema() |
.model_json_schema() |
@validator |
@field_validator |
@root_validator |
@model_validator |
class Config: |
model_config = ConfigDict(...) |
The model_ prefix is the tell. One glance at a snippet tells you which era it is from, and that is worth more than memorising the table: you will not remember every row, but you will remember that .dict() means you are reading something old and should go looking for the current spelling.
TypeAdapter, for when there is no class
from pydantic import TypeAdapter
adapter = TypeAdapter(list[SearchQuery])
adapter.validate_python(raw_list)Any annotation can be validated, not just a BaseModel subclass. This is how you validate a bare list[int], a dict[str, Thing], or a TypedDict without inventing a wrapper model around it. Build the adapter once and reuse it; construction is the expensive part.
Where this goes
That is the type system, honestly described. Module 4 is the four idioms you will meet on your first day in an agent repo, all of which look like syntax and are actually protocols.
Try it yourself
Passing a string to an int field
A model declares attempts: int and you validate the payload {"attempts": "5"} with default settings. What do you get?
Show answer
Correct answer: B — An instance with attempts equal to the integer 5, because the default mode coerces a numeric string
Pydantic's default is lax mode, which coerces where the conversion is unambiguous: a numeric string becomes an int, a float becomes an int when it has no fractional part, and a bool becomes an int. This is genuinely useful at a JSON or form boundary where everything arrives as a string, and it is genuinely surprising the first time it silently accepts something you expected to be rejected. Strict mode, per-field or per-model, turns it off. Leaving the value as the string "5" is what a dataclass would do, which is exactly the contrast worth holding: the annotation is inert there and load-bearing here.
Reading a Stack Overflow answer
An answer you have found calls model.dict() and decorates a method with @validator. What does that tell you?
Show answer
Correct answer: C — It is v1. The v2 spellings are model_dump() and @field_validator, with a model_ prefix
The model_ prefix is the v2 tell, adopted precisely so that model methods cannot collide with your own field names. .dict(), .json(), parse_obj() and @validator are all v1. The compatibility shim described in the pydantic.v1-import option is real, and it works by importing from pydantic.v1 explicitly, which is a different thing from v1 code running unchanged. The reason this matters practically is that most Pydantic search results predate v2, so recognising the era of an answer in one glance saves you from writing code against an API you do not have.
Choosing Pydantic over a dataclass
You have an internal value object, constructed only by your own code from values that are already the right types. Which is right?
Show answer
Correct answer: D — A dataclass, because validating data you produced yourself buys nothing and costs work on every construction
Validation earns its cost at a boundary, where data arrives from somewhere you do not control. Inside the program you already own the invariants, so a dataclass says what you mean and does not re-check facts you established. Choosing the Pydantic model on safety grounds is the reflex worth resisting: safer-is-always-better leads to models everywhere, validation running on paths that cannot fail, and error handling written for exceptions that cannot be raised. The other two are simply false about dataclasses and about Pydantic construction respectively.
Why this library, specifically
The question worth being able to answer, because it explains a design decision you will otherwise find arbitrary.
Why does the whole LangChain and LangGraph ecosystem run on Pydantic rather than on plain type hints or dataclasses?
Reveal answer
Because these frameworks need the type information at RUNTIME, and Pydantic is the library that reads annotations and produces something executable from them. A tool definition has to become a JSON Schema that goes to the model provider, and a structured output has to be validated against that schema when the model's response comes back as text. Neither is possible with annotations alone, since the interpreter ignores them. Pydantic reads the annotations, generates the schema with model_json_schema(), and validates the response with model_validate(), so one class definition serves as the schema sent out and the parser applied on the way back in. That single round trip is why it is a hard dependency rather than a preference.
Break a model on purpose
Define a two-field model, validate a payload that is wrong in two different ways at once, and print the exception's .errors().
You got both failures in one exception rather than only the first, each carrying a location path and a type, and you can see how that structure would be turned into an API error response.