A/B testing, offline experiments and online splits
Two things get called A/B testing in this space and they are not the same instrument. Getting the pair straight is worth more than any mechanism, because most of the wrong answers here come from applying the offline shape to an online question.
Offline first, because it is nearly free
An offline experiment runs both variants over a dataset with reference outputs you wrote, and compares the scores. It is repeatable, it costs no user exposure, and you can run it fifty times before lunch.
Its limitation is exactly its strength inverted: it measures your variants against examples you curated, so it is only as representative as your imagination was on the day you built the dataset. That is why it rejects bad variants but never confirms good ones.
Online, and the three things that decide whether the result means anything
The variant must be in metadata, at trace time. Runs are immutable. If nobody attached the arm while the traffic was flowing, the experiment did not happen, however much data you collected. This is module 0’s metadata rule with a deadline on it.
The randomisation unit must match the measurement unit. Split per request in a multi turn app and a single conversation gets a blend of both arms, so any thread level score, and sentiment is a thread level score, is measuring a mixture. Derive the variant deterministically from the user or conversation identifier so it is stable for the life of the thread.
Both arms must be measured identically. One online evaluator whose filter matches both arms guarantees the same criterion at the same sampling rate. Two separately configured evaluators is tidier to look at and gives you two things that can drift.
Reading the result
The surface for this is a dashboard chart with Group by set to your variant metadata field. Group by accepts a run tag or a metadata key, and it splits one chart into one series per value, which is exactly the shape an A/B read wants: both arms, same metric, same axes, no second chart to eyeball against the first.
Two practical details that catch people out. Group by shows the top five values by frequency by default and can be raised to twenty, which matters the moment your variant field has more than two values or you left an old arm running. And if you want both trace-level and LLM-level charts grouped the same way, the metadata has to be attached to both the root run and the child runs, not just the trace, or the LLM charts come back ungrouped.
Then compare the aggregated feedback score per arm, and be honest about how much data each arm actually has. Sampling shrinks it, and a thread level criterion shrinks it again, because a hundred runs might only be twenty conversations.
Decide the stopping condition before you start. Watching two numbers until one of them is ahead is not an experiment, it is a slot machine.
Try it yourself
What has to be true before the traces are written
You want to compare two prompt variants on live traffic and analyse the result next week. What single thing must already be in place while the traffic is running?
Show answer
Correct answer: A — A variant identifier attached to every run's metadata at trace time
You can only split a population by a field somebody attached while the runs were being created, and runs are immutable records. Without a variant field, next week's analysis is impossible and no amount of clever filtering recovers it. The dashboard option is tempting because it is what the finished analysis looks like, and it is easy to confuse the artefact you want with the precondition that makes it possible. The alert rule is a sensible safety net and does nothing for attribution.
Spot the bug in the split
This assigns a variant on every incoming request in a multi turn chat app.
const variant = Math.random() < 0.5 ? "control" : "treatment";
const result = await traceable(runTurn, {
name: "chatTurn",
metadata: { thread_id: conversationId, variant },
})(userMessage);Show answer
Correct answer: C — Assignment is per request, so a conversation flips between variants and thread level scores mean nothing
The randomisation unit is wrong. A user in a ten turn conversation gets roughly five of each variant, so a thread scored by a sentiment or task success evaluator is scoring a blend rather than an arm. Derive the variant deterministically from the conversation or user identifier instead, so it is stable for the life of the thread. The predictability answer is the plausible one because unseeded randomness genuinely is a real concern in other contexts, and it is irrelevant here: nobody is attacking your prompt experiment.
Offline experiment against online split
Two different instruments with two different jobs, and reaching for the wrong one is usually a symptom of not knowing which question you are asking yet.
Describe the difference between running an offline experiment over a dataset and running an online A/B split on production traffic. When is each the right instrument?
Reveal answer
An offline experiment runs both variants over a fixed dataset with known reference outputs, and compares scores in an experiment comparison view. It is repeatable, cheap, has no user impact, and it measures the variants against examples you curated, so it is only as representative as your dataset. Use it before you ship, to reject obviously worse variants. An online A/B split routes real traffic to both variants, tags each run with its arm, and measures with reference free online evaluators plus whatever explicit feedback arrives. It measures real distribution and real users, costs real user exposure, and takes real time to accumulate. Use it after the offline pass, to confirm on traffic you did not curate.
Why a before and after comparison is not an A/B test
You deployed variant B on Tuesday and compare last week's feedback scores against this week's. B looks better.
Show answer
Correct answer: D — It is not an A/B test at all: everything else changed too, so the variant is confounded with the week
A/B means the two arms run concurrently on the same traffic distribution, which is what removes time as a variable. A before and after comparison confounds your change with the news cycle, a marketing push, a different mix of users, a provider model update and any other deploy that week. The volume answer is the strongest distractor because equal volume genuinely does matter for statistical power, so it is a real concern attached to the wrong problem: it addresses noise, not confounding.
One evaluator or two
You are scoring both arms with the same reference free criterion at a sampling rate of 0.2.
Show answer
Correct answer: B — One evaluator whose filter matches both arms, so both are sampled at the same rate by one configuration
One variant agnostic filter is the simplest way to guarantee both arms are sampled identically, which is what keeps the comparison fair. Two separately configured evaluators is the tempting option because it looks tidier and more explicit, and it introduces a real risk: two configurations can drift apart, in rate, in criterion wording, or in when somebody paused one of them. Deliberately over sampling one arm biases nothing about the score itself but does give the arms different precision, which is a trap in the opposite direction.
Design one split on paper
No infrastructure, one scratch file, ten minutes. This is the whole lesson compressed into a design you could hand to someone.
You have written down the variant field name, the identifier the split is derived from and why it is stable across a conversation, the reference free criterion, the sampling rate, the granularity you are scoring at, and the one sentence that says how you will decide the experiment is over.