LLM-as-judge
Two families of evaluator, and the exam wants to know you can put a given job in the right one. That part is quick. The part worth your attention is what a judge actually is, because it changes what its score means.
The two families
A code-based evaluator is a function. It looks at the run, applies deterministic logic, returns a score. No model call anywhere. Everything you would normally write as an assertion lives here: valid JSON, parses against the schema, under the length limit, contains the required disclaimer, called the tool it was supposed to, returned an ID that exists in the database, finished in under two seconds. Fast, free, deterministic. The score means exactly what the code says and nothing else.
An LLM-as-judge is a prompt. You describe the property you care about, hand the model the inputs and the output, and offline optionally the reference output, and ask it to score. This is the only thing that works for properties with no closed-form definition: helpfulness, tone, whether an answer is grounded in retrieved context rather than invented, whether a summary preserved the key point. You cannot regex your way to is this dismissive of the customer.
The framing that pays off next lesson
A code-based evaluator is correct by construction. An LLM-as-judge is a hypothesis.
Once you have read a code evaluator and agreed with its logic, its score means what you think it means, forever. A judge produces a number with an unknown error rate. It might be lenient, it might be harsh, it might be secretly scoring on length because verbose answers look thorough. You cannot tell by staring at the score, and the score does not come with an error bar.
That is not a reason to avoid judges. It is the reason evaluator alignment is its own feature, its own workflow and its own focus area, and the entire purpose of the next lesson is turning a hypothesis into an instrument with a measured agreement rate.
What a judge is handed
Offline, the judge can see the inputs, your application’s output, and the reference output where one exists. That last one is worth stating precisely, because it looks like it contradicts the rule from the dataset lesson and does not.
Your application must never see the reference output. Feeding the thing under test the answer key makes every score a lie. The evaluator seeing it is the entire point. Same object, two consumers, opposite rules, and confusing them is a genuine misconception the exam can build a question on.
Online, there is no reference output to offer, so a judge can only assess intrinsic quality: is this helpful, is this on-topic, does this read as grounded in the context it was given.
Where they run
Both families run in both worlds, which is worth stating plainly because it is a natural place to get confused. A code-based evaluator can be offline or online. A judge can be offline or online. What changes between offline and online is not the evaluator family, it is what data the evaluator receives.
Try it yourself
Pick the right evaluator for the job
Four properties you might want to score. Which one genuinely needs an LLM-as-judge rather than a code-based evaluator?
Show answer
Correct answer: C — The response is polite and does not sound dismissive of the customer
Politeness and tone are fuzzy, subjective, and have no closed-form definition, which is the exact profile of a judge job. The other three are deterministic string or structural checks: parse it, count it, substring match it. Option 3 is the tempting one because it looks semantic, but presence of a known order ID is a substring check against the input, and doing it with a model would be slower, more expensive, nondeterministic and worse. The general rule: if you can write the check, write the check.
What you pay for a judge
An LLM-as-judge is not free in any dimension, and one of the four costs is the reason the next lesson exists.
List the concrete downsides of using an LLM-as-judge where a code-based evaluator would have worked.
Reveal answer
Money, since every evaluation is an extra model call and at production sampling rates that compounds. Latency, which matters for online evaluation at volume. Nondeterminism, since the judge can score the same output differently across runs, giving your measuring instrument its own noise floor. And trust: a code evaluator is correct by construction once you have read it, whereas a judge has to be aligned against human labels before its number means anything. That last one is the big one and it is why evaluator alignment exists as its own feature and its own focus area.
Why an unaligned judge is worse than nothing
A framing worth getting into your bones early, because the next lesson depends entirely on it.
Why is an unaligned LLM-as-judge score dangerous, arguably more dangerous than having no evaluator at all?
Reveal answer
Because it produces a number, and numbers get believed. An unaligned judge might be systematically lenient, systematically harsh, or correlated with something irrelevant like output length, and none of that is visible from looking at the score. You then make shipping decisions on it and confidently move in the wrong direction. Having no evaluator at least leaves you honestly uncertain. This is exactly what the alignment score fixes: it measures the percentage of examples where the judge's verdict matches a human expert's, turning the judge from an oracle you hope is right into an instrument with a known error rate.
What a judge is handed
An LLM-as-judge is running offline against a dataset whose examples have reference outputs. What can the judge prompt be given?
Show answer
Correct answer: C — The inputs, the output, and optionally the reference output
Offline, a judge can see the inputs, your application's output, and the reference output if one exists, which is what lets it grade against ground truth rather than only on intrinsic quality. Option 0 sounds like methodological care and confuses two different things: your application must never see the reference output, because that would be feeding it the answers, but the evaluator seeing it is the entire point. Option 1 is the same confusion in a milder form. Online, of course, there is no reference output to offer.
Audit your own instinct
Take any agent you have built and list five things you would want to evaluate about it.
For each of the five you can say code-based or judge and defend it, and at least three turn out to be code-based. If all five came out as judge jobs, you are reaching for the model where a substring check, a schema parse or a tool-call assertion would do, which is the most common evaluation design mistake there is.