Whetstone.
EvaluationRunning experiments, and reading them honestly
Module 2, Lesson 315 min

Running experiments, and reading them honestly

An experiment is one run of your application over one dataset, scored by one or more evaluators, recorded so you can compare it to the next one. The API is small. The interesting parts are the parameter almost nobody sets and the statistics that are not there.

The call

You give evaluate() three things: the thing under test, the data to test it on, and the evaluators to score with.

Python
from langsmith import evaluate

results = evaluate(
    my_app,
    data="my-dataset",
    evaluators=[correctness, conciseness],
    num_repetitions=3,
)
TypeScript
import { evaluate } from "langsmith/evaluation";

const results = await evaluate(myApp, {
  data: "my-dataset",
  evaluators: [correctness, conciseness],
  numRepetitions: 3,
});

Note what changed beyond the obvious. num_repetitions became numRepetitions, and data went from optional in Python to required in TypeScript. Both are documented differences and both are exactly the shape of question that tests whether you read the language toggle.

Python also has aevaluate() for async evaluation and TypeScript has no equivalent: no aEvaluate, no evaluateAsync, no flag. Worth naming explicitly, because the natural response to not finding something in the docs is to assume you searched badly.

The parameter almost nobody sets

Repetitions defaults to 1.

Sit with that. Every experiment you have run without thinking about it ran each example exactly once. You drew one sample per example from a nondeterministic system, compared it against another single sample, and concluded which version was better.

Raise it and each example runs multiple times. LangSmith then shows the average and the standard deviation across repetitions, which is frequently humbling: a two-point improvement looks less exciting once you can see the same code varies by five points between runs.

Reading the comparison view

You pick a baseline experiment and everything is coloured relative to it. Red means regression, green means improvement, per example and in aggregate, which lets you scan a long dataset and land on the handful of rows where your change did damage.

That drill-down is the actual value. An aggregate moving from 0.82 to 0.85 tells you almost nothing about what to do next. Six red rows that turn out to be the same category of question tells you where to go.

Colour is relative, always. A red cell is not a bad score, it is a score that went down. You can be red at 0.9 because your baseline was 0.95. The comparison view answers whether this change helped, never whether this is good. Which also makes the baseline a load-bearing and trivially abusable choice: compare against your worst historical run and the screen turns satisfyingly green. Say which baseline you used whenever you report a result.

What you do not get

No confidence intervals. No significance test. No minimum sample size guidance. Nothing warns you that twelve examples is too few, and the tool will happily draw you a lovely green comparison across all twelve.

I am naming the absence as carefully as the presence, because the gap is where people get hurt. A team that believes there is a significance test somewhere in the interface will assume any comparison they are shown has been vetted by one. There is no such vetting.

So decide with judgement and the spread in front of you: raise repetitions before you care about a result, compare the gap against the standard deviation, read the rows that moved, and grow the dataset, which is the crudest and most reliable lever there is.

In the exam, be suspicious of the most rigorous-sounding option. References to p-values, confidence intervals, statistical significance or recommended sample sizes describe capabilities the documentation does not contain, and the paper is semi-open-book against that documentation.

Practice

Try it yourself

Quiz

Repetitions, casing and default

You want each example evaluated three times in a single TypeScript experiment.

Which line is correct, and what was it before you touched it?
const results = await evaluate(myApp, {
  data: "my-dataset",
  evaluators: [correctness],
  /* ??? */
});
  1. Anum_repetitions: 3, up from a default of 1
  2. BnumRepetitions: 3, up from a default of 3
  3. Crepetitions: 3, up from a default of 0
  4. DnumRepetitions: 3, up from a default of 1
Show answer

Correct answer: D — numRepetitions: 3, up from a default of 1

TypeScript uses numRepetitions, Python uses num_repetitions, and the default is 1 in both. Option 0 is the distractor and it is the Python casing, which is what you will have seen most often in docs and blog posts. A default of 1 is worth holding onto because it means every experiment you have run without thinking about this ran each example exactly once, with no measure of run-to-run variance at all.

Quiz

The data parameter

A TypeScript evaluate() call omits the data parameter. What happens?

  1. AIt is an error, data is required in TypeScript even though it is optional in Python
  2. BIt is fine, data is optional in TypeScript exactly as it is in Python
  3. CIt is fine, data defaults to the dataset used by the most recent experiment
  4. DIt is an error, because data is a required parameter in Python and TypeScript alike
Show answer

Correct answer: A — It is an error, data is required in TypeScript even though it is optional in Python

data is optional in Python and required in TypeScript. Option 3 is the tempting one for anyone reasoning that an evaluation obviously needs to know what to evaluate against, which is a sensible instinct that happens to be wrong about Python. Option 2 invents a convenience default, and inventing plausible conveniences is a reliable way to lose marks in this domain, because the real APIs are less accommodating than you would guess.

Quiz

What red means in the comparison view

You are looking at an experiment compared against a baseline. A cell is red. What does that tell you?

  1. AThe evaluator raised an exception while scoring that example
  2. BThe score is below an absolute pass threshold you configured
  3. CThat example regressed relative to the baseline experiment
  4. DYour application errored on that example and produced no output to score
Show answer

Correct answer: C — That example regressed relative to the baseline experiment

Red means regression and green means improvement, both measured against the baseline experiment you chose. Option 1 is the distractor worth understanding: the colouring is relative, never absolute. A red cell can be a score of 0.9 if the baseline scored 0.95, and a green cell can be a dismal 0.2 if the baseline managed 0.1. Colour tells you direction of change, never quality of result.

Quiz

What the tooling actually computes for you

Which of these does the documented LangSmith experiment tooling give you?

  1. AA 95 percent confidence interval on each metric across repetitions
  2. BAn average and a standard deviation across repetitions
  3. CA significance test between an experiment and its baseline
  4. DA recommended minimum sample size for your dataset
Show answer

Correct answer: B — An average and a standard deviation across repetitions

Average and standard deviation across repetitions is what you get, and repetitions is the only documented mechanism for handling run-to-run noise. Options 0, 2 and 3 are all things a rigorous evaluation setup would ideally have and none of them appear in the documentation. They are tempting precisely because they are what you would expect a serious evaluation product to provide, which is what makes this a good question: it tests whether you know the tool or whether you are reasoning from what a tool like this ought to do.

Recall

What a score difference licenses you to conclude

The judgement call this lesson builds to, and the one that generalises furthest beyond the exam.

Experiment B scores 0.86 against a baseline of 0.83, on a 40-example dataset, with repetitions left at the default. What can you legitimately conclude, and what would you need in order to conclude more?

Reveal answer

Almost nothing. With repetitions at 1 you have a single sample per example from a nondeterministic system, so you cannot separate a real three-point gain from run-to-run variance, and there is no confidence interval or significance test in the tooling to resolve it statistically either. To conclude more you would raise repetitions, compare the gap against the reported standard deviation, and drill into the individual rows that moved to check they moved for the reason you intended. Judgement, informed by spread, on inspected examples, is genuinely the state of the art in the documented tooling.

Check

Report a result without overclaiming

Write one sentence reporting an experiment result the way you would put it in a pull request.

You should see

The sentence names the baseline, names the metric, gives the average, gives the standard deviation if you ran repetitions, and states the number of examples. If your sentence is accuracy improved to 87% with none of that context, rewrite it, because that sentence is unfalsifiable and it is the kind that gets a feature shipped on noise.

Sign in to track your progress →