What is middleware
Middleware is where a surprising amount of this domain lives. Error handling, retries, approval gates, compression, model switching, tool filtering: none of those are parameters on create_agent, and all of them are middleware.
Six hooks
before_agent # around the whole agent invocation, entering
before_model # around a model call, entering
after_model # around a model call, leaving
after_agent # around the whole agent invocation, leaving
wrap_model_call # surrounds a model call
wrap_tool_call # surrounds a tool callTypeScript spells the same six beforeAgent, beforeModel, afterModel, afterAgent, wrapModelCall, wrapToolCall. Same hooks, same semantics, camelCase throughout, and that dual spelling runs through every name in this module.
Two shapes, not six things
The before_* and after_* hooks are observers and mutators at a boundary. They run, they can read and modify state, control carries on past them. They are one-shot: nothing is handed to them that they could choose to run.
The wrap_* hooks are surrounds. A wrap hook receives the call it is wrapping and decides what to do with it: run it, run it and then run it again, run something else instead, or not run it at all.
Three altitudes
- Agent level:
before_agent,after_agent. Once around the whole invocation. - Model-call level:
before_model,after_model,wrap_model_call. Once per turn, and a long-running agent may take hundreds. - Tool-call level:
wrap_tool_call. Once per individual tool invocation.
Putting per-turn work at agent level is a real bug and an extremely plausible exam scenario, because the code looks correct and simply never fires again after the first pass.
Two authoring surfaces
from langchain.agents.middleware import wrap_tool_call
from langchain.messages import ToolMessage
@wrap_tool_call
def handle_tool_errors(request, handler):
"""Convert tool exceptions into ToolMessages."""
try:
return handler(request)
except Exception as e:
return ToolMessage(content=f"Tool error: {e}",
tool_call_id=request.tool_call["id"])import { createMiddleware } from "langchain";
const timing = createMiddleware({
name: "timing",
wrapModelCall: async (request, handler) => handler(request),
});Python has decorators and a class-based API. TypeScript has neither. That asymmetry is the widest exposure a TypeScript developer has in this whole domain, and it is a dialect gap rather than a knowledge gap. Read SummarizationMiddleware and think “summarization middleware”, not “the Python one”.
Try it yourself
Name all six
Flat memorisation with no cleverness in it. These have to be reflex before the ordering lesson makes any sense at all.
Name the six middleware hooks in their Python spelling, and give the TypeScript spelling of each.
Reveal answer
before_agent and beforeAgent, before_model and beforeModel, after_model and afterModel, after_agent and afterAgent, wrap_model_call and wrapModelCall, wrap_tool_call and wrapToolCall. Four before-and-after hooks that observe and mutate around a boundary, and two wrap hooks that surround a call and can decide whether and how it runs.
Picking a hook for a retry
You want to catch a failing model call and retry it against a different provider, without the agent loop ever seeing the failure. Which hook shape do you need?
Show answer
Correct answer: D — wrap_model_call, because only a wrapping hook can re-invoke the call it holds
Only a wrap hook holds the call, so only a wrap hook can run it again. after_model is the tempting answer because that is where you would first see the error, but an after hook runs once, after the fact, with nothing left to invoke. That is the entire reason wrap hooks exist as a separate shape rather than as more before-and-after hooks.
Spot the bug
This middleware is meant to enforce a per-turn token budget on a long-running agent. It never fires more than once.
@before_agent
def enforce_budget(state, runtime):
if count_tokens(state["messages"]) > BUDGET:
raise BudgetExceeded()
return NoneShow answer
Correct answer: B — before_agent fires once per invocation, not once per turn
before_agent runs once around the whole invocation, so a hundred-turn run checks the budget exactly once, at the start, when nothing has been spent. Per-turn work belongs at model-call level in before_model. The last option is a plausible-sounding invention; returning None means no state change on this pass, and the hook still runs next time.
Authoring a custom middleware
You are writing a custom middleware and want to know which authoring styles are available in each language.
Show answer
Correct answer: A — Python has decorators and a class-based API; TypeScript only createMiddleware factories
Python gives you both a decorator style and a class-based middleware API. TypeScript has neither and authors with createMiddleware factories. The second option is the reasonable assumption, since parity is what you would expect from a framework maintained in both languages, and it is exactly the assumption that puts a TypeScript developer at a disadvantage in a Python-flavoured exam.
Reduce six to two
Write down the six hook names and then group them into the two shapes, with one sentence on what the second shape can do that the first cannot.
You grouped four observers against two surrounds, and your sentence is about re-invocation or substitution rather than about ordering.