Whetstone.
Module 3: monitoring qualityA/B testing, offline experiments and online splits
Module 3, Lesson 316 min

A/B testing, offline experiments and online splits

Two things get called A/B testing in this space and they are not the same instrument. Getting the pair straight is worth more than any mechanism, because most of the wrong answers here come from applying the offline shape to an online question.

Offline first, because it is nearly free

An offline experiment runs both variants over a dataset with reference outputs you wrote, and compares the scores. It is repeatable, it costs no user exposure, and you can run it fifty times before lunch.

Its limitation is exactly its strength inverted: it measures your variants against examples you curated, so it is only as representative as your imagination was on the day you built the dataset. That is why it rejects bad variants but never confirms good ones.

Online, and the three things that decide whether the result means anything

The variant must be in metadata, at trace time. Runs are immutable. If nobody attached the arm while the traffic was flowing, the experiment did not happen, however much data you collected. This is module 0’s metadata rule with a deadline on it.

The randomisation unit must match the measurement unit. Split per request in a multi turn app and a single conversation gets a blend of both arms, so any thread level score, and sentiment is a thread level score, is measuring a mixture. Derive the variant deterministically from the user or conversation identifier so it is stable for the life of the thread.

Both arms must be measured identically. One online evaluator whose filter matches both arms guarantees the same criterion at the same sampling rate. Two separately configured evaluators is tidier to look at and gives you two things that can drift.

Reading the result

The surface for this is a dashboard chart with Group by set to your variant metadata field. Group by accepts a run tag or a metadata key, and it splits one chart into one series per value, which is exactly the shape an A/B read wants: both arms, same metric, same axes, no second chart to eyeball against the first.

Two practical details that catch people out. Group by shows the top five values by frequency by default and can be raised to twenty, which matters the moment your variant field has more than two values or you left an old arm running. And if you want both trace-level and LLM-level charts grouped the same way, the metadata has to be attached to both the root run and the child runs, not just the trace, or the LLM charts come back ungrouped.

Then compare the aggregated feedback score per arm, and be honest about how much data each arm actually has. Sampling shrinks it, and a thread level criterion shrinks it again, because a hundred runs might only be twenty conversations.

Decide the stopping condition before you start. Watching two numbers until one of them is ahead is not an experiment, it is a slot machine.

Practice

Try it yourself

Quiz

What has to be true before the traces are written

You want to compare two prompt variants on live traffic and analyse the result next week. What single thing must already be in place while the traffic is running?

  1. AA variant identifier attached to every run's metadata at trace time
  2. BA dashboard with two charts, one filtered to each of the two variants
  3. CAn alert rule on feedback score, so a regression pages you mid experiment
  4. DA dataset containing labelled examples drawn from both variants
Show answer

Correct answer: A — A variant identifier attached to every run's metadata at trace time

You can only split a population by a field somebody attached while the runs were being created, and runs are immutable records. Without a variant field, next week's analysis is impossible and no amount of clever filtering recovers it. The dashboard option is tempting because it is what the finished analysis looks like, and it is easy to confuse the artefact you want with the precondition that makes it possible. The alert rule is a sensible safety net and does nothing for attribution.

Quiz

Spot the bug in the split

This assigns a variant on every incoming request in a multi turn chat app.

Variant assignment
const variant = Math.random() < 0.5 ? "control" : "treatment";

const result = await traceable(runTurn, {
  name: "chatTurn",
  metadata: { thread_id: conversationId, variant },
})(userMessage);
  1. AMath.random() is not cryptographically secure, so the variant split is predictable
  2. BThe variant should be a tag rather than metadata, so as written it cannot be filtered
  3. CAssignment is per request, so a conversation flips between variants and thread level scores mean nothing
  4. DA 50/50 split is too aggressive and should start at 5% of traffic before ramping
Show answer

Correct answer: C — Assignment is per request, so a conversation flips between variants and thread level scores mean nothing

The randomisation unit is wrong. A user in a ten turn conversation gets roughly five of each variant, so a thread scored by a sentiment or task success evaluator is scoring a blend rather than an arm. Derive the variant deterministically from the conversation or user identifier instead, so it is stable for the life of the thread. The predictability answer is the plausible one because unseeded randomness genuinely is a real concern in other contexts, and it is irrelevant here: nobody is attacking your prompt experiment.

Recall

Offline experiment against online split

Two different instruments with two different jobs, and reaching for the wrong one is usually a symptom of not knowing which question you are asking yet.

Describe the difference between running an offline experiment over a dataset and running an online A/B split on production traffic. When is each the right instrument?

Reveal answer

An offline experiment runs both variants over a fixed dataset with known reference outputs, and compares scores in an experiment comparison view. It is repeatable, cheap, has no user impact, and it measures the variants against examples you curated, so it is only as representative as your dataset. Use it before you ship, to reject obviously worse variants. An online A/B split routes real traffic to both variants, tags each run with its arm, and measures with reference free online evaluators plus whatever explicit feedback arrives. It measures real distribution and real users, costs real user exposure, and takes real time to accumulate. Use it after the offline pass, to confirm on traffic you did not curate.

Quiz

Why a before and after comparison is not an A/B test

You deployed variant B on Tuesday and compare last week's feedback scores against this week's. B looks better.

  1. AIt is a valid A/B test, since both variants ran on real production traffic in turn
  2. BIt is valid provided traffic volume was similar across the two weeks being compared
  3. CIt is invalid because feedback scores cannot be compared across different time ranges
  4. DIt is not an A/B test at all: everything else changed too, so the variant is confounded with the week
Show answer

Correct answer: D — It is not an A/B test at all: everything else changed too, so the variant is confounded with the week

A/B means the two arms run concurrently on the same traffic distribution, which is what removes time as a variable. A before and after comparison confounds your change with the news cycle, a marketing push, a different mix of users, a provider model update and any other deploy that week. The volume answer is the strongest distractor because equal volume genuinely does matter for statistical power, so it is a real concern attached to the wrong problem: it addresses noise, not confounding.

Quiz

One evaluator or two

You are scoring both arms with the same reference free criterion at a sampling rate of 0.2.

  1. ATwo evaluators, one filtered to each arm, each configured at a rate of 0.2
  2. BOne evaluator whose filter matches both arms, so both are sampled at the same rate by one configuration
  3. COne evaluator at 1.0 on the treatment arm and 0.2 on control, since treatment is the one under test
  4. DNeither, because a sampled evaluation cannot support a comparison between arms
Show answer

Correct answer: B — One evaluator whose filter matches both arms, so both are sampled at the same rate by one configuration

One variant agnostic filter is the simplest way to guarantee both arms are sampled identically, which is what keeps the comparison fair. Two separately configured evaluators is the tempting option because it looks tidier and more explicit, and it introduces a real risk: two configurations can drift apart, in rate, in criterion wording, or in when somebody paused one of them. Deliberately over sampling one arm biases nothing about the score itself but does give the arms different precision, which is a trap in the opposite direction.

Check

Design one split on paper

No infrastructure, one scratch file, ten minutes. This is the whole lesson compressed into a design you could hand to someone.

You should see

You have written down the variant field name, the identifier the split is derived from and why it is stable across a conversation, the reference free criterion, the sampling rate, the granularity you are scoring at, and the one sentence that says how you will decide the experiment is over.

Sign in to track your progress →