Insights against Online Evals, the distinction the exam is built on
This is the centre of the Monitor domain. Two features described in very similar words that are architecturally almost opposite, which is precisely the shape a question writer goes looking for.
Start from LangChain’s own sentence: Insights helps you identify what to improve; evals help you measure if improvements work.
The table
| Insights | Online Evals | |
|---|---|---|
| Shape | Unsupervised discovery: summarise, then cluster into hierarchical categories | Supervised measurement against a criterion you defined |
| Output | A report of categories and subcategories | Feedback written onto the run or thread |
| Trigger | Scheduled: daily, weekly or cron, interpreted in UTC | Continuous polling over a filter |
| Volume | Capped at 1,000 traces per run | Sampling rate from 0 to 1 |
| Plan | Plus and Enterprise only | Broadly available |
Every row is a potential question. The output row and the plan row are the two people get wrong most often.
How an online evaluator is triggered
Two settings do the work: a filter and a sampling rate.
The filter says which runs are candidates, expressed in the same language the runs table uses: this project, this run name, this metadata field, errors only. The sampling rate is a number from 0 to 1 saying what proportion of matching runs actually get judged, and it exists because every evaluated run is another model call that costs money and adds latency.
Then it runs continuously, polling for new matching runs as they arrive. It is not something you kick off. It is something that is on.
The one question that decides which tool
Do you already know what you are looking for?
If yes, you have a criterion, and a criterion is an evaluator. If no, you have a mess, and a mess is an Insights job. That is the whole decision procedure and it survives contact with nearly every scenario question you will be handed.
Try it yourself
Reproduce the comparison
If you learn one card from this course, make it this one. Five axes, and every single row has been a question somewhere.
Compare Insights and Online Evals on shape, output, trigger, volume, and plan availability.
Reveal answer
Shape: Insights is unsupervised discovery, summarise then cluster into hierarchical categories; Online Evals are supervised measurement against a criterion you wrote. Output: Insights produces a report of categories and subcategories; Online Evals write feedback onto the run or thread. Trigger: Insights is scheduled, daily, weekly or cron, in UTC; Online Evals poll continuously over a filter. Volume: Insights is capped at 1,000 traces per run; Online Evals use a sampling rate from 0 to 1. Plan: Insights is Plus and Enterprise only; Online Evals are broadly available. LangChain's own summary: Insights helps you identify what to improve, evals help you measure if improvements work.
A scenario with no hypothesis
Support says users are unhappy but nobody can say why. No theory, no failing test, no complaint that repeats.
Show answer
Correct answer: D — An Insights job over the last period, to see what the traffic clusters into
No hypothesis means nothing to measure, and an evaluator can only measure a criterion somebody wrote down. The satisfaction evaluator is the tempting answer because it sounds proactive and it does produce numbers, but you would be scoring satisfaction without knowing what is damaging it, so a low score tells you only what support already told you. Insights is the tool for the case where the question itself is missing. Once it names the categories, then you write the evaluator.
A scenario with a hypothesis
You shipped a prompt change on Tuesday intended to reduce ungrounded answers. Did it work?
Show answer
Correct answer: A — Run an online evaluator for groundedness across the change and compare feedback scores either side
A specific hypothesis and a specific criterion is the supervised case. The Insights option is the good distractor because category counts really would shift and you would see something, but Insights categories are regenerated per run and are not designed as a stable metric to diff, so you would be reading tea leaves rather than taking a measurement. Ungrounded answers are usually successful runs, so error rate is blind to them by construction.
What an online evaluator produces
The evaluator has run and formed a judgement about a sampled production run. Where does that judgement end up?
Show answer
Correct answer: B — As feedback attached to the run or thread that was evaluated
Online evaluators write feedback, onto the run or onto the thread, which is what makes the result filterable, chartable and alertable using machinery you already have. The output appending answer is the dangerous misconception rather than merely a wrong one: evaluation never mutates production data, and a system that did would make your traces a record of something that never happened. The report answer is the Insights answer, which is the confusion this whole lesson exists to kill.
Why online criteria have to be reference free
A structural consequence, not a design preference, and stating it that way is what makes it stick.
What does an offline evaluator receive that an online evaluator cannot, and what does that force about the criteria you can write?
Reveal answer
An offline evaluator runs against a dataset, so it receives both the run and the dataset example, which carries a reference output somebody wrote down in advance. An online evaluator runs against live production traffic, where no reference exists, because nobody knows the right answer to a question a real user just invented. It gets the run and nothing to compare it against. So online criteria must be reference free: groundedness, toxicity, did it answer the question, does this read like a frustrated user. Exact match against a known answer is structurally unavailable, and any option offering it as an online criterion is offering the wrong shape.