Whetstone.
EvaluationEvaluating agents, offline and online
Module 2, Lesson 115 min

Evaluating agents, offline and online

Offline evals run against a dataset you built, so the right answer is sitting there to compare against. Online evals run against live production traffic, where there is no right answer, because nobody wrote one down.

That is the whole lesson. Everything below is consequences.

Offline: pre-deployment testing

Offline evaluation is what you do before you ship. You have a dataset of examples, each an input paired with a reference output, meaning the answer you decided was correct. You run your application over those inputs and evaluators score the outputs, often by comparing them against the reference.

This is the mode that answers whether version B is better than version A, because you hold the inputs fixed and vary the code. It is a regression suite that speaks in scores instead of booleans. Offline evaluators attach to a dataset.

Online: production monitoring

Online evaluation is what you do after you ship. It runs against real runs and threads flowing through a tracing project, continuously, on traffic you did not choose.

You configure it with a filter, deciding which runs qualify, and a sampling rate between 0 and 1, deciding what fraction of the qualifying ones actually get scored. You do not evaluate everything, because at production volume a judge on every run is a genuinely large model bill.

And there is no reference output anywhere in this picture. You can ask whether the output is valid JSON, whether it called a tool it should not have, whether a judge thinks it was helpful. You cannot ask whether it is correct, because correct compared to what.

The side effect nobody expects

Here is something with nothing to do with evaluation and everything to do with your invoice. Attaching an online evaluator automatically upgrades matching traces to extended data retention.

Read that again with your finance hat on. You wire up a cheap-looking heuristic evaluator on your highest-volume project, set the sampling rate generously because sampling more sounds more rigorous, and you have quietly changed the retention tier on a large slice of your traffic. It is a completely reasonable design decision, since an evaluated trace is one you will want to go back and look at, and it is a beautiful exam question precisely because it is a side effect rather than a feature.

Choosing between them, out loud

You do not pick one. They cover different failure modes and they feed each other.

Offline catches the change I just made broke the thing that used to work. Online catches reality contains inputs I never imagined. The bridge between them is the harvest: a production run scores badly online, you add it to a dataset, and now it is an offline test forever.

If an exam question gives you a scenario, the tell is usually the word production. Production traffic, live users, monitoring, drift, alerting: online. Regression, before deploying, comparing two versions, prompt A versus prompt B: offline.

Practice

Try it yourself

Recall

The one fact everything else falls out of

The highest-value card in the module, and worth answering before reading anything back.

What single fact about production traffic makes online evaluation structurally different from offline evaluation, and name two concrete consequences of it.

Reveal answer

Production runs have no reference output, because nobody wrote down the right answer; a real user typed a question and left. First consequence: the in-UI code evaluator loses its second parameter, becoming perform_eval(run) online where it is perform_eval(run, example) offline, because there is no example to pass. Second consequence: online evaluators can only score intrinsic properties of the run itself, such as whether it is valid JSON, whether it called the right tool, or whether a judge thinks it was helpful, and can never score correctness against a known answer. A third worth holding: online scores are a monitoring signal over time rather than a pass or fail gate before deploy.

Quiz

What each evaluator type attaches to

A configuration question that is really a comprehension question. Where does each kind of evaluator get wired up?

  1. AOffline evaluators attach to a tracing project, online evaluators attach to a dataset
  2. BBoth attach to a dataset, and the difference is a toggle on the evaluator itself
  3. COffline evaluators attach to a dataset, online evaluators attach to a tracing project
  4. DBoth attach to a tracing project, and the difference is whether a dataset is linked
Show answer

Correct answer: C — Offline evaluators attach to a dataset, online evaluators attach to a tracing project

Offline evaluation runs against curated examples, so it attaches to the dataset holding them. Online evaluation runs against live traffic, so it attaches to the tracing project receiving it. Option 0 is the same sentence with the nouns swapped, and it is the distractor they will use, because every word in it is correct and only the pairing is wrong. Read attachment questions twice. Option 1 is tempting if you assume the difference is a mode flag rather than a structural difference in which data exists.

Quiz

The billing side effect of an online evaluator

This one is not about evaluation at all, which is exactly why it makes a good exam question.

  1. AMatching traces are automatically excluded from your monthly trace quota
  2. BMatching traces are copied into a linked dataset so the scores can be replayed offline
  3. CMatching traces are downgraded to base retention to save storage
  4. DMatching traces are automatically upgraded to extended data retention
Show answer

Correct answer: D — Matching traces are automatically upgraded to extended data retention

Attaching an online evaluator auto-upgrades matching traces to extended data retention, which is a billing consequence you want to know about before setting a sampling rate of 1 on your busiest project. Option 2 is the tempting one because it sounds like a sensible cost optimisation and points the right lever the wrong way. Option 1 is worth rejecting explicitly: online evaluation does not put anything in a dataset by itself, that is an automation rule and a separate feature.

Recall

The two knobs on an online evaluator

Configuration detail that turns up as both a recall question and a cost question.

You attach an online evaluator to a busy tracing project. What two configuration controls decide which runs it actually scores, and what is the range on the numeric one?

Reveal answer

A filter, which decides which runs or threads qualify at all, keyed off metadata, name, tags and similar fields, and a sampling rate, a number from 0 to 1 deciding what fraction of qualifying runs get evaluated. The filter is about relevance, the sampling rate is about cost. Both matter more than they look: an LLM-as-judge at a sampling rate of 1 on high volume is a real model bill, and it also drags every matching trace into extended retention.

Quiz

Classify the scenario

A team wants to know whether changing the system prompt improved answer quality before rolling it out to users. Which setup fits?

  1. AOffline, an experiment over a dataset, comparing against the previous experiment as baseline
  2. BOnline, an evaluator on the tracing project at a low sampling rate to keep costs down
  3. COnline, with a tracing project filter restricted to runs tagged with the new prompt version
  4. DOffline, with the experiment attached to the tracing project so it scores real traffic too
Show answer

Correct answer: A — Offline, an experiment over a dataset, comparing against the previous experiment as baseline

Before rolling out means no production traffic exists for the new version yet, so the only thing you can hold fixed is a curated set of inputs, which is an offline experiment against a dataset with a baseline to compare to. Option 2 is the interesting distractor because it describes a perfectly sensible thing to do afterwards, during a staged rollout, and the exam likes questions where the wrong answer is the right answer to the next question. Option 3 is incoherent: offline and tracing project do not go together, which is the whole attachment rule.

Sign in to track your progress →