Whetstone.
EvaluationPairwise evaluations
Module 2, Lesson 715 min

Pairwise evaluations

Most evaluation asks how good is this. Pairwise evaluation asks a different question: of these two, which is better. That sounds like a smaller question and it is frequently a much easier one to answer reliably, which is the entire reason the feature exists.

Why absolute scoring struggles

Ask a judge to score a summary out of ten and you are asking it to hold a stable internal rubric across hundreds of calls. It will not. The same summary scores 6 on Tuesday and 8 on Wednesday, and the drift is not random noise you can average away, it moves with phrasing, with length, with what happened to be in context.

Now show the same judge two summaries and ask which is better. That is a comparison, not a calibration. There is no scale to hold steady, no rubric to remember, just a relative preference. Models are noticeably better at this, and so are humans, which is why every wine tasting, every A/B test and every chess rating is built on comparisons rather than absolute scores.

What a pairwise evaluation is

You take two experiments run over the same dataset, and for each example you show a judge both outputs and ask which is preferred. Across the dataset that aggregates into a preference rate: version B beat version A on 34 of 50 examples.

When to reach for it

When better is easier to judge than good. Tone, style, helpfulness, readability: anything where you can instantly tell which of two drafts you prefer and would struggle to defend a number.

When you are choosing between two candidates rather than measuring one. Two prompts, two models, two retrieval strategies. The decision is comparative, so measure it comparatively.

When writing reference outputs is impractical. For open-ended generation, authoring the one correct answer is often harder than the task itself. Preference judging sidesteps that entirely.

And when not to: anything you need to track over time, anything you need to report as a level rather than a direction, and anything online, where only one output exists.

Do not confuse it with the comparison view

They look similar on screen and do genuinely different things.

The comparison view takes scores your evaluators already produced independently and colours them red or green relative to a baseline experiment. Its input is numbers.

A pairwise evaluation is itself an evaluation: a judge sees both outputs together and expresses a preference that did not exist until it was asked. Its input is outputs.

How you actually run one

The documented entry point is evaluate() itself. Pass a two-tuple of existing experiments instead of a target function, and it performs a pairwise comparison rather than a normal experiment:

Python: the documented form
from langsmith import evaluate

evaluate(
    ("experiment-1", "experiment-2"),
    evaluators=[ranked_preference],
    randomize_order=True,
    max_concurrency=4,
)

evaluate_comparative() is the underlying function and is still exported, so you will meet both spellings. Its full signature:

Python: the underlying function
from langsmith.evaluation import evaluate_comparative

evaluate_comparative(
    experiments,              # positional-only: a two-tuple of experiment names or ids
    /,
    evaluators,
    experiment_prefix=None,
    description=None,
    max_concurrency=5,
    client=None,
    metadata=None,
    load_nested=False,        # load child runs, not just top-level roots
    randomize_order=False,
)

Three details worth banking, because each is the kind of thing a question hangs on. experiments is positional-only, so it cannot be passed by keyword. max_concurrency defaults to 5 here, which is not the same as evaluate()’s own default. And a comparative evaluator receives a list of runs plus the shared example, rather than the single run an ordinary evaluator gets.

The hygiene, and the one bit of it the SDK does for you

Position bias is handled in-product. Models tend to favour whichever candidate they see first, or last, and randomize_order exists precisely for that: it shuffles the order the outputs are presented in for each comparison. Note the default is False, so it is opt-in. Turning it on is the single highest-value thing you can do to a pairwise run, and forgetting it is the most common way a preference result ends up measuring presentation order.

Length bias is on you. Longer answers read as more thorough, and nothing in the SDK controls for it. If your winner is systematically the longer output, you may have measured verbosity. This one is standard practice from the wider evaluation literature rather than a LangSmith feature.

Both are reminders that a preference judge is still a judge, so it deserves the same treatment as any other: sample it, get human labels, and check whether it agrees with them.

Practice

Try it yourself

Quiz

When pairwise beats absolute scoring

You have two versions of a summarisation agent and a judge that has never agreed with itself twice about what a 7 out of 10 summary looks like. Which situation most favours a pairwise evaluation?

  1. AYou need to know which of the two versions is better, and better is easier to judge than good
  2. BYou need one absolute quality number on a 0 to 1 scale to report to a stakeholder
  3. CYou need to detect a regression on one specific example against a baseline experiment
  4. DYou need to score live production traffic, where only one output per request exists
Show answer

Correct answer: A — You need to know which of the two versions is better, and better is easier to judge than good

Pairwise exists for exactly the case where a judge is unreliable on an absolute scale but reliable at picking a winner, which is a very common asymmetry with subjective properties. Option 1 is the tempting one because reporting to a stakeholder is precisely when you feel the pull towards a single tidy number, and pairwise does not give you one: it gives you a preference rate between two specific things. Option 3 rules itself out structurally, since a comparison needs two outputs and production gives you one.

Recall

What a pairwise run actually produces

The output shape is the thing that decides which questions this feature can answer, so it is worth being able to state before reasoning about anything else.

A pairwise evaluation compares two experiments over the same dataset. What does it produce, and what can it not tell you?

Reveal answer

It produces a relative preference between the two experiments over the shared examples: for each example the judge picks which output is better, and across the dataset that aggregates into a win rate for one experiment over the other. What it cannot tell you is whether either output is any good in absolute terms. A version can win every comparison while both versions are unusable, because the only question ever asked was which of these two, and there is no bar anywhere in the procedure. If you need an absolute standard you need a normal evaluator with a rubric, or a human.

Quiz

Pairwise evaluation versus the comparison view

Both involve two experiments side by side. Which statement correctly separates them?

  1. AThey are one feature under two names, pairwise being what the comparison view is called in the SDK
  2. BThe comparison view runs offline over datasets and pairwise evaluation runs online over traffic
  3. CThe comparison view is offline and pairwise evaluation is online
  4. DPairwise evaluation requires reference outputs, whereas the comparison view does not
Show answer

Correct answer: B — The comparison view runs offline over datasets and pairwise evaluation runs online over traffic

The comparison view takes scores that already exist and colours them relative to a baseline, so its input is numbers your evaluators produced independently. A pairwise evaluation is an evaluation in its own right: a judge sees both outputs together and expresses a preference. Option 0 is the conflation worth guarding against, because they look identical on screen and are doing different things. Option 3 has it backwards in spirit: preference judging is the technique you reach for precisely when good ground truth is hardest to write down.

Recall

The hygiene a preference judge needs

One half of this is a parameter the SDK gives you and the other half is yours to handle, so be clear which is which.

You are asking a model to pick the better of two outputs. What biases should you control for, and which of them does LangSmith control for you?

Reveal answer

Position bias is the main one: models have a documented tendency to favour whichever candidate appears first or last. LangSmith handles this with the randomize_order parameter, which shuffles the presentation order per comparison, and it defaults to False, so it is opt-in and easy to forget. Length bias is the second, since longer answers often read as more thorough regardless of quality, and nothing in the SDK controls for it, so checking whether the winner is systematically the longer output is your job. Both are reasons a pairwise result deserves the same scepticism as any other judge output, which means it deserves the same treatment: get human labels on a sample and check the judge agrees with them.

Check

Confirm the signature at the source

A short docs-navigation drill against a page you should be able to reach without hunting.

You should see

You can state the pairwise entry point without looking: evaluate() with a two-tuple of existing experiments in the target position, handing off to evaluate_comparative() underneath. Then find the reference page and confirm two things the prose above asserts, that randomize_order defaults to False and that max_concurrency defaults to 5, so you have seen them on the page as well as here. The habit being drilled is checking a default rather than assuming it, since defaults are exactly what changes between versions.

Sign in to track your progress →