Whetstone.
Moving Towards ProductionOnline evals in production
Module 3, Lesson 315 min

Online evals in production

You already know the mechanics from Module 2. This lesson is about operating them: what to point an online evaluator at, how much of it to score, and what to do when the number moves.

The two knobs, again

An online evaluator attaches to a tracing project and takes two controls.

A filter decides which runs or threads qualify at all, keyed off run name, tags, metadata, status and similar fields. This is where the metadata you attached at instrumentation time earns itself.

A sampling rate between 0 and 1 decides what fraction of qualifying runs actually get scored.

The division of labour is worth stating: the filter is about relevance, the sampling rate is about cost. A filter that is too broad gives you a number averaged across populations that have nothing to do with each other, which is worse than no number because it moves for reasons you cannot trace. Narrow with the filter first, then set sampling as low as you can tolerate.

And the signature, one more time

Attached to a tracing project
def perform_eval(run):
    ...

No example, because there is no reference output. The single most common way to break an online evaluator is to paste in one written for a dataset, where the body is entirely correct and the signature is not.

What you can actually ask

Intrinsic properties only. Is the output valid JSON against the schema. Did the agent call a tool it should not have. Did latency or token count exceed a budget. Does a judge consider this helpful, on-topic, or grounded in the context it was given.

Not correctness. There is no answer key. If a scenario asks whether a production response was correct, the honest reading is that it cannot be measured directly, and the nearest available moves are a judge scoring groundedness or a human labelling a sample in an annotation queue.

Reading an online score

Offline, a score is a verdict on a version. Online, a single score is close to meaningless and the trend is the whole signal.

Your helpfulness judge reads 0.78 today. Is that good? Unanswerable. Was it 0.78 last week and 0.71 the week before? Now you have something. Online evaluation is a monitoring instrument, and monitoring instruments are read as derivatives.

Which means two practical things. Establish a baseline period before you draw conclusions, because week one of any online evaluator is calibration rather than measurement. And be careful attributing movement, since your score can move because the agent got worse, because the traffic changed, or because the judge is nondeterministic and you are sampling few enough runs for that to show.

Finish the sentence

The question that turns an online evaluator from a chart into a mechanism is what happens when this moves.

If the answer is somebody will notice, you have built a lamp nobody is watching. The useful answers are concrete: an alert, or an automation rule that harvests the low-scoring runs into a dataset so the failure becomes a permanent offline test. That second one is the next lesson, and it is the reason online evaluation is worth doing at all.

Practice

Try it yourself

Quiz

Spot the bug

This code evaluator was written for a dataset and is being pasted into the UI, attached to a tracing project, to flag responses that leaked an email address.

One line stops this working online. Which?
def perform_eval(run, example):
    leaked = "@" in run.outputs.get("answer", "")
    return {"key": "no_pii", "score": 0 if leaked else 1}
  1. AThe return dictionary needs a comment field when attached to a tracing project
  2. BThe function must be renamed when the evaluator is attached to a tracing project
  3. CThe signature takes an example parameter, and there is no example online
  4. Drun.outputs is not populated on production runs, only on runs from a dataset
Show answer

Correct answer: C — The signature takes an example parameter, and there is no example online

The body is fine and never touches example, which is exactly what makes this realistic: the code is correct and the signature is not. Online there is no reference output, so there is no example object to pass, and the parameter has to go. Option 1 is the good distractor because a name change feels like the sort of thing that would differ between two attachment points, and it is the one rule that stays constant: perform_eval in both worlds and both languages. Option 3 is wrong and worth rejecting, since the run is precisely what an online evaluator is given.

Quiz

The sampling rate

You are configuring an online evaluator and want it to score roughly a quarter of qualifying runs. What do you set, and what is the valid range?

  1. A25, on a scale from 0 to 100
  2. B4, meaning one run in four
  3. C0.25, on a scale from 0 to 1
  4. D0.25, but only if a filter is also set, since sampling requires a filter
Show answer

Correct answer: C — 0.25, on a scale from 0 to 1

The sampling rate is a number from 0 to 1, so a quarter is 0.25. Option 0 is the tempting one because percentages are the more common human unit and plenty of tools do take 0 to 100, which is why the range is worth holding as a fact rather than an assumption. Option 3 invents a dependency: a filter and a sampling rate are two independent controls, one about relevance and one about cost, and either can be left permissive.

Recall

What can and cannot be scored online

A boundary worth being able to draw quickly, because scenario questions usually turn on it rather than on any configuration detail.

Give three properties an online evaluator can score and one it structurally cannot, and say why.

Reveal answer

It can score anything intrinsic to the run: whether the output is valid JSON against a schema, whether the agent called a tool it should not have, whether latency or token count exceeded a budget, whether a judge considers the answer helpful, on-topic, or grounded in the context it was given. What it cannot score is correctness against a known answer, because there is no reference output in production: a user asked a question and left without supplying the right answer. If a scenario asks you to check whether a production response was correct, the honest reading is that it cannot be done directly, and the nearest available move is a judge scoring groundedness or a human labelling a sample.

Recall

Two costs of turning the sampling rate up

One of these is obvious and one catches people, and the second is the more likely exam question.

You raise the sampling rate on an LLM-as-judge attached to your busiest tracing project. Name both costs you have just incurred.

Reveal answer

The obvious one is model spend: every sampled run is an extra model call, and at production volume a generous sampling rate on a high-traffic project is a real and recurring bill. The one that catches people is retention, because attaching an online evaluator automatically upgrades matching traces to extended data retention, so raising coverage also raises how much trace data you are storing at the higher tier. The second is a side effect rather than a feature, which is exactly what makes it a good exam question and a bad surprise on an invoice.

Check

Design one, end to end

On paper, for an agent you know. No account needed.

You should see

You have named the property, decided code-based or judge and defended it, written the filter in words, chosen a sampling rate and justified it against both costs, and stated what you would do when the score moves. That last part is the one people skip, and an online evaluator whose alarm has no action behind it is a chart nobody opens.

Sign in to track your progress →