Whetstone.
Module 2: monitoring product analyticsInsights against Online Evals, the distinction the exam is built on
Module 2, Lesson 215 min

Insights against Online Evals, the distinction the exam is built on

This is the centre of the Monitor domain. Two features described in very similar words that are architecturally almost opposite, which is precisely the shape a question writer goes looking for.

Start from LangChain’s own sentence: Insights helps you identify what to improve; evals help you measure if improvements work.

The table

Insights Online Evals
Shape Unsupervised discovery: summarise, then cluster into hierarchical categories Supervised measurement against a criterion you defined
Output A report of categories and subcategories Feedback written onto the run or thread
Trigger Scheduled: daily, weekly or cron, interpreted in UTC Continuous polling over a filter
Volume Capped at 1,000 traces per run Sampling rate from 0 to 1
Plan Plus and Enterprise only Broadly available

Every row is a potential question. The output row and the plan row are the two people get wrong most often.

How an online evaluator is triggered

Two settings do the work: a filter and a sampling rate.

The filter says which runs are candidates, expressed in the same language the runs table uses: this project, this run name, this metadata field, errors only. The sampling rate is a number from 0 to 1 saying what proportion of matching runs actually get judged, and it exists because every evaluated run is another model call that costs money and adds latency.

Then it runs continuously, polling for new matching runs as they arrive. It is not something you kick off. It is something that is on.

The one question that decides which tool

Do you already know what you are looking for?

If yes, you have a criterion, and a criterion is an evaluator. If no, you have a mess, and a mess is an Insights job. That is the whole decision procedure and it survives contact with nearly every scenario question you will be handed.

Practice

Try it yourself

Recall

Reproduce the comparison

If you learn one card from this course, make it this one. Five axes, and every single row has been a question somewhere.

Compare Insights and Online Evals on shape, output, trigger, volume, and plan availability.

Reveal answer

Shape: Insights is unsupervised discovery, summarise then cluster into hierarchical categories; Online Evals are supervised measurement against a criterion you wrote. Output: Insights produces a report of categories and subcategories; Online Evals write feedback onto the run or thread. Trigger: Insights is scheduled, daily, weekly or cron, in UTC; Online Evals poll continuously over a filter. Volume: Insights is capped at 1,000 traces per run; Online Evals use a sampling rate from 0 to 1. Plan: Insights is Plus and Enterprise only; Online Evals are broadly available. LangChain's own summary: Insights helps you identify what to improve, evals help you measure if improvements work.

Quiz

A scenario with no hypothesis

Support says users are unhappy but nobody can say why. No theory, no failing test, no complaint that repeats.

  1. AAn online evaluator scoring user satisfaction on every production run
  2. BA threshold alert rule on the Feedback Score metric over a 15 minute window
  3. CAn offline evaluation run against your existing regression dataset
  4. DAn Insights job over the last period, to see what the traffic clusters into
Show answer

Correct answer: D — An Insights job over the last period, to see what the traffic clusters into

No hypothesis means nothing to measure, and an evaluator can only measure a criterion somebody wrote down. The satisfaction evaluator is the tempting answer because it sounds proactive and it does produce numbers, but you would be scoring satisfaction without knowing what is damaging it, so a low score tells you only what support already told you. Insights is the tool for the case where the question itself is missing. Once it names the categories, then you write the evaluator.

Quiz

A scenario with a hypothesis

You shipped a prompt change on Tuesday intended to reduce ungrounded answers. Did it work?

  1. ARun an online evaluator for groundedness across the change and compare feedback scores either side
  2. BSchedule a weekly Insights job and compare this week's categories against last week's
  3. CCheck whether the errors metric dropped in the days following Tuesday's deploy
  4. DRead twenty traces from either side of Tuesday by hand and form a judgement
Show answer

Correct answer: A — Run an online evaluator for groundedness across the change and compare feedback scores either side

A specific hypothesis and a specific criterion is the supervised case. The Insights option is the good distractor because category counts really would shift and you would see something, but Insights categories are regenerated per run and are not designed as a stable metric to diff, so you would be reading tea leaves rather than taking a measurement. Ungrounded answers are usually successful runs, so error rate is blind to them by construction.

Quiz

What an online evaluator produces

The evaluator has run and formed a judgement about a sampled production run. Where does that judgement end up?

  1. AIn a scheduled report you read in the Insights tab
  2. BAs feedback attached to the run or thread that was evaluated
  3. CAppended to the run's output so downstream consumers can see it
  4. DIn a dataset, as a new example for future offline evaluation
Show answer

Correct answer: B — As feedback attached to the run or thread that was evaluated

Online evaluators write feedback, onto the run or onto the thread, which is what makes the result filterable, chartable and alertable using machinery you already have. The output appending answer is the dangerous misconception rather than merely a wrong one: evaluation never mutates production data, and a system that did would make your traces a record of something that never happened. The report answer is the Insights answer, which is the confusion this whole lesson exists to kill.

Recall

Why online criteria have to be reference free

A structural consequence, not a design preference, and stating it that way is what makes it stick.

What does an offline evaluator receive that an online evaluator cannot, and what does that force about the criteria you can write?

Reveal answer

An offline evaluator runs against a dataset, so it receives both the run and the dataset example, which carries a reference output somebody wrote down in advance. An online evaluator runs against live production traffic, where no reference exists, because nobody knows the right answer to a question a real user just invented. It gets the run and nothing to compare it against. So online criteria must be reference free: groundedness, toxicity, did it answer the question, does this read like a frustrated user. Exact match against a known answer is structurally unavailable, and any option offering it as an online criterion is offering the wrong shape.

Sign in to track your progress →