Evaluator alignment, and the feature it is not
You wrote a judge. It gives scores. You have no idea whether those scores are any good, and until this feature existed the only available response was a shrug.
The alignment score is the shrug replaced by a number. It also has a sibling feature that looks similar, works completely differently, and is the likeliest thing in this domain to be taught wrong.
The definition
An alignment score is the percentage of examples where the evaluator’s judgment matches that of the human expert.
Learn it close to that wording. It is agreement, not accuracy. For a subjective property like tone or helpfulness there is no objective ground truth to be accurate against, so the closest available standard is a competent human’s verdict, and the question becomes whether your judge reaches the same conclusions they do.
That reframing is the whole feature. Your judge stops being an oracle and becomes an instrument with a measured agreement rate, which is the only basis on which you should be shipping decisions off its scores.
The workflow, as clicks
Select your experiments. You start from experiments that already exist, choosing the ones whose examples you want in your alignment set.
Human-label in an annotation queue. The part everyone wants to skip and cannot. You, or a domain expert, apply verdicts by hand. Slow, boring, and the entire foundation of the number you are about to compute.
Add to Reference Dataset. Your human labels get promoted into the reference dataset alignment will measure against.
Iterate in the Evaluator Playground, then Start Alignment. Tune the judge prompt, hit Start Alignment, read your score, go round again.
The two numbers: at least 20 examples, balanced across both 0 and 1 labels.
Overfitting, and the documented answer
The obvious failure mode: you tweak the prompt, the score rises, you tweak again, it rises more, and eventually you have a judge exquisitely tuned to twenty specific examples and useless on the twenty-first.
The docs’ remedy is direct. Add more labelled examples and re-test. Not a regularisation trick, not a holdout split with a clever name. More labelled data, then run it again. Which is annoying, because labelling is the expensive step, and that is precisely why it works: it cannot be faked by prompt gymnastics.
The feature this is not
Few-shot self-improving evaluators have a completely different rhythm. You are looking at results for some other reason, you see a score that is plainly wrong, and you correct it inline. That correction is then automatically injected into the judge prompt at a {{Few-shot examples}} placeholder. Next time the judge runs, your correction sits in its context as a worked example.
No labelling session. No annotation queue. No score. You did not schedule anything; you fixed one thing you happened to notice, and the fix persisted into the prompt.
| Align Evals | Few-shot self-improving | |
|---|---|---|
| Trigger | A deliberate session you sit down for | Noticing a wrong score in passing |
| Human does | Labels a balanced set in an annotation queue | Corrects one score inline |
| Produces | An alignment score, a percentage | A modified judge prompt |
| Mechanism | Compares judge verdicts to human labels | Injects corrections at a placeholder |
| Minimum effort | About 20 balanced labelled examples | Seconds |
Align Evals measures. Few-shot self-improvement modifies. One is an instrument reading, the other is a nudge. They are complementary, not alternatives, and the give-away in any question is whether it mentions a number.
What the score does not tell you
It tells you your judge agrees with one human on this labelled set. Not that the human was right, not that a second expert would have labelled the same way, not that agreement holds on inputs unlike the ones you gathered.
None of that makes the number useless. It makes it a number with a scope, and knowing the scope of your instrument is roughly the entire theme of this course.
Try it yourself
Define the alignment score
Close to verbatim if you can. This is the definitional card for the whole focus area.
What exactly is an alignment score?
Reveal answer
The percentage of examples where the evaluator's judgment matches that of the human expert. It is agreement between your LLM-as-judge and a human labeller, measured over a set of examples a human has actually labelled. It is not accuracy against ground truth in the usual sense, it is concordance with a specific human's verdicts, which is the right framing because for subjective properties those human labels are the closest thing to ground truth that exists.
How many labelled examples, and shaped how
You are setting up Align Evals. What do the docs recommend for the labelled set?
Show answer
Correct answer: D — At least 20 examples, balanced across both 0 and 1 labels
At least 20 examples, balanced across both 0 and 1 labels. Balance is the part that matters and the part people skip: an evaluator tested only on failures can score beautifully by calling everything a failure, and one tested only on passes can score beautifully by calling everything a pass. Option 0 sounds like sensible ML practice and gets the balance requirement exactly backwards. Option 2 is what most people actually do, and it is how you end up with a lopsided set by accident.
The Align Evals workflow, in order
Four steps, and this is a UI workflow with a provisioned org waiting for you at exam time, so learn it as clicks rather than as prose.
Walk through the Align Evals workflow from a set of experiments to an alignment score.
Reveal answer
One, select the experiments you want to draw examples from. Two, human-label those examples in an annotation queue, applying 0 and 1 verdicts by hand. Three, Add to Reference Dataset, promoting your human labels into the reference set that alignment will be measured against. Four, iterate the judge prompt in the Evaluator Playground, then Start Alignment to score the judge against the human labels. The loop then repeats: read the score, adjust the prompt, re-run alignment.
Which feature uses the few-shot placeholder
A judge prompt contains a placeholder written as a double-braced Few-shot examples token. Which feature populates it?
Show answer
Correct answer: B — Few-shot self-improving evaluators, which inject your inline score corrections
The placeholder belongs to few-shot self-improving evaluators: when you correct a score inline, that correction is automatically injected into the judge prompt at that placeholder, so the judge learns from it on subsequent runs. Option 0 is the distractor and it is the conflation this entire lesson exists to prevent. Align Evals uses human labels to compute an agreement score; it does not push them into the prompt for you. Two features, two mechanisms, and the exam is well aware they get confused.
Pick the feature for the scenario
You are about to make your judge the gate on a release, and you need to be able to state how much it can be trusted. Which feature do you reach for?
Show answer
Correct answer: C — Align Evals, because it produces a measurable agreement rate against human labels
You need a number you can quote, and only Align Evals produces one. Option 0 is genuinely tempting because few-shot correction does plausibly improve the judge, but it gives you no measurement, so you still cannot say how much to trust it. Improving something and measuring something are different acts, and this question separates them. Option 3 reduces noise in the experiment, which is real and tells you nothing about whether the judge is any good.
Two columns, from memory
Draw the comparison table yourself without looking.
Four rows filled for both features: what triggers it, what the human does, what it produces, and what it changes. If you produce four clean rows with no cell borrowed from the wrong column, you have the distinction, and it is the one most likely to be asked about in this focus area.