Whetstone.
All courses
langsmithIntermediate

Production Monitoring for AI Agents with LangSmith

Monitor is exactly 10 of the 40 LCAE questions, and it is the only section with a single prep course behind it. It is also the most interface heavy domain on the exam: a LangSmith org is provisioned per candidate and several questions expect you to go and look. So this course teaches the interface alongside the concepts, organised the way the official course organises it, by what you are monitoring rather than by which feature does it. You will learn the greedy most-to-least-specific pricing walk and the three separate conditions that must all hold before any cost appears, why Insights and Online Evals are architecturally opposite, how to attribute latency down a tree whose parents contain their children, why user sentiment and security monitoring are compositions rather than products, what makes an online A/B split valid, and the exact metrics, windows and channels an alert rule can use. Where the docs and the launch blog disagree, this course says so instead of quietly picking one, and where a fact could not be sourced it says that too.

Start the course 17 lessons 257 min
PMCrash course
Curriculum

17 lessons across 6 modules

01Module 0: the objects you monitor2 lessons · 31 min
02Module 1: monitoring cost4 lessons · 61 min
03Module 2: monitoring product analytics3 lessons · 45 min
04Module 3: monitoring quality4 lessons · 60 min
05Module 4: monitoring security3 lessons · 46 min
06Capstone1 lessons · 14 min
Capstone

Wire one live agent end to end, from raw trace to a firing alert you would actually answer

You have one agent running in production and no honest idea whether it is fine. Fix that in one sitting, using only what this course covered, and finish with a monitor that can go red.

The build

Take any agent you already have, even a two node graph that calls one model. Then wire the monitoring surface onto it in this order, because each step depends on the one before it.

  1. Trace it. Wrap the entry point with traceable() and confirm a root run appears in a project you named yourself. Expand the tree, name the run type of every node, and note which of the seven never appear in your app.
  2. Thread it. Pass a stable thread_id in metadata across at least three invocations and find all three grouped in the threads view. Then do it once more with session_id instead and confirm it still groups.
  3. Attribute it. Decide your metadata vocabulary now, before you need it: tenant, environment, app version, feature, variant. Attach all five. Every question you will want to ask in month two is a group-by on one of these, and none of them can be added retroactively.
  4. Price it. Check all three cost conditions by hand: token counts present, ls_provider and ls_model_name present, and a matching active entry in the pricing map. Break exactly one on purpose and watch the cost column go blank rather than zero. That is the most instructive thirty seconds in this course.
  5. Time it. Take your slowest trace and walk the wall clock down the tree until you find the run that actually holds it. Note whether its siblings ran in parallel, which you can read off the timings alone.
  6. Judge it. Configure one online evaluator over a filter matching your traffic, with a sampling rate below 1 and a reference free criterion, and find the feedback it wrote back onto the sampled runs. Then check what it did to your data retention.
  7. Discover. If your plan allows it, schedule one Insights job and read the category tree. Notice that it told you what to look at and told you nothing about whether you are getting better. If your plan does not allow it, hand read twenty traces and build the tree yourself; the method is the point.
  8. Alarm. Create one threshold rule on a metric you chose deliberately, over a 5 or 15 minute window, pointed at a webhook you control. Break the agent on purpose and confirm the webhook fires. An alert you have never seen fire is a hypothesis, not a monitor.

The bar

You are done when you can point at your own project and answer five questions without opening the docs.

Which runs belong to which conversation. What last night cost, and whether that number is complete. What your most common failure category is, and whether that answer came from data or from your imagination. Whether your quality metric moved this week, and which evaluator and criterion produce it. Who gets woken when it breaks, on which metric, at what threshold, over which window, through which channel.

The three that will be uncomfortable

Most people can answer the first two immediately.

The third one usually reveals that nobody has ever looked, and the honest answer is a guess in the shape of a fact.

The fourth one usually reveals that there is no evaluator at all, which means there is no quality metric, which means the Feedback Score alert door is standing open with nothing behind it.

The fifth one usually reveals a rule somebody wrote once, with a number nobody can defend in both directions, on a channel nobody checks. If the answer to “who gets woken” is “an unread Slack channel”, that is the same as nobody.

Those three gaps are the real output of this capstone. They are also, almost exactly, the gaps the exam is testing, because the questions in this domain are overwhelmingly about the joins: what feedback is for, what an alert can actually watch, what makes a cost number trustworthy, and which instrument answers a question that has no hypothesis in it yet.