Evaluating agents, offline and online
Offline evals run against a dataset you built, so the right answer is sitting there to compare against. Online evals run against live production traffic, where there is no right answer, because nobody wrote one down.
That is the whole lesson. Everything below is consequences.
Offline: pre-deployment testing
Offline evaluation is what you do before you ship. You have a dataset of examples, each an input paired with a reference output, meaning the answer you decided was correct. You run your application over those inputs and evaluators score the outputs, often by comparing them against the reference.
This is the mode that answers whether version B is better than version A, because you hold the inputs fixed and vary the code. It is a regression suite that speaks in scores instead of booleans. Offline evaluators attach to a dataset.
Online: production monitoring
Online evaluation is what you do after you ship. It runs against real runs and threads flowing through a tracing project, continuously, on traffic you did not choose.
You configure it with a filter, deciding which runs qualify, and a sampling rate between 0 and 1, deciding what fraction of the qualifying ones actually get scored. You do not evaluate everything, because at production volume a judge on every run is a genuinely large model bill.
And there is no reference output anywhere in this picture. You can ask whether the output is valid JSON, whether it called a tool it should not have, whether a judge thinks it was helpful. You cannot ask whether it is correct, because correct compared to what.
The side effect nobody expects
Here is something with nothing to do with evaluation and everything to do with your invoice. Attaching an online evaluator automatically upgrades matching traces to extended data retention.
Read that again with your finance hat on. You wire up a cheap-looking heuristic evaluator on your highest-volume project, set the sampling rate generously because sampling more sounds more rigorous, and you have quietly changed the retention tier on a large slice of your traffic. It is a completely reasonable design decision, since an evaluated trace is one you will want to go back and look at, and it is a beautiful exam question precisely because it is a side effect rather than a feature.
Choosing between them, out loud
You do not pick one. They cover different failure modes and they feed each other.
Offline catches the change I just made broke the thing that used to work. Online catches reality contains inputs I never imagined. The bridge between them is the harvest: a production run scores badly online, you add it to a dataset, and now it is an offline test forever.
If an exam question gives you a scenario, the tell is usually the word production. Production traffic, live users, monitoring, drift, alerting: online. Regression, before deploying, comparing two versions, prompt A versus prompt B: offline.
Try it yourself
The one fact everything else falls out of
The highest-value card in the module, and worth answering before reading anything back.
What single fact about production traffic makes online evaluation structurally different from offline evaluation, and name two concrete consequences of it.
Reveal answer
Production runs have no reference output, because nobody wrote down the right answer; a real user typed a question and left. First consequence: the in-UI code evaluator loses its second parameter, becoming perform_eval(run) online where it is perform_eval(run, example) offline, because there is no example to pass. Second consequence: online evaluators can only score intrinsic properties of the run itself, such as whether it is valid JSON, whether it called the right tool, or whether a judge thinks it was helpful, and can never score correctness against a known answer. A third worth holding: online scores are a monitoring signal over time rather than a pass or fail gate before deploy.
What each evaluator type attaches to
A configuration question that is really a comprehension question. Where does each kind of evaluator get wired up?
Show answer
Correct answer: C — Offline evaluators attach to a dataset, online evaluators attach to a tracing project
Offline evaluation runs against curated examples, so it attaches to the dataset holding them. Online evaluation runs against live traffic, so it attaches to the tracing project receiving it. Option 0 is the same sentence with the nouns swapped, and it is the distractor they will use, because every word in it is correct and only the pairing is wrong. Read attachment questions twice. Option 1 is tempting if you assume the difference is a mode flag rather than a structural difference in which data exists.
The billing side effect of an online evaluator
This one is not about evaluation at all, which is exactly why it makes a good exam question.
Show answer
Correct answer: D — Matching traces are automatically upgraded to extended data retention
Attaching an online evaluator auto-upgrades matching traces to extended data retention, which is a billing consequence you want to know about before setting a sampling rate of 1 on your busiest project. Option 2 is the tempting one because it sounds like a sensible cost optimisation and points the right lever the wrong way. Option 1 is worth rejecting explicitly: online evaluation does not put anything in a dataset by itself, that is an automation rule and a separate feature.
The two knobs on an online evaluator
Configuration detail that turns up as both a recall question and a cost question.
You attach an online evaluator to a busy tracing project. What two configuration controls decide which runs it actually scores, and what is the range on the numeric one?
Reveal answer
A filter, which decides which runs or threads qualify at all, keyed off metadata, name, tags and similar fields, and a sampling rate, a number from 0 to 1 deciding what fraction of qualifying runs get evaluated. The filter is about relevance, the sampling rate is about cost. Both matter more than they look: an LLM-as-judge at a sampling rate of 1 on high volume is a real model bill, and it also drags every matching trace into extended retention.
Classify the scenario
A team wants to know whether changing the system prompt improved answer quality before rolling it out to users. Which setup fits?
Show answer
Correct answer: A — Offline, an experiment over a dataset, comparing against the previous experiment as baseline
Before rolling out means no production traffic exists for the new version yet, so the only thing you can hold fixed is a curated set of inputs, which is an offline experiment against a dataset with a baseline to compare to. Option 2 is the interesting distractor because it describes a perfectly sensible thing to do afterwards, during a staged rollout, and the exam likes questions where the wrong answer is the right answer to the next question. Option 3 is incoherent: offline and tracing project do not go together, which is the whole attachment rule.