Running experiments, and reading them honestly
An experiment is one run of your application over one dataset, scored by one or more evaluators, recorded so you can compare it to the next one. The API is small. The interesting parts are the parameter almost nobody sets and the statistics that are not there.
The call
You give evaluate() three things: the thing under test, the data to test it on, and the evaluators to score with.
from langsmith import evaluate
results = evaluate(
my_app,
data="my-dataset",
evaluators=[correctness, conciseness],
num_repetitions=3,
)import { evaluate } from "langsmith/evaluation";
const results = await evaluate(myApp, {
data: "my-dataset",
evaluators: [correctness, conciseness],
numRepetitions: 3,
});Note what changed beyond the obvious. num_repetitions became numRepetitions, and data went from optional in Python to required in TypeScript. Both are documented differences and both are exactly the shape of question that tests whether you read the language toggle.
Python also has aevaluate() for async evaluation and TypeScript has no equivalent: no aEvaluate, no evaluateAsync, no flag. Worth naming explicitly, because the natural response to not finding something in the docs is to assume you searched badly.
The parameter almost nobody sets
Repetitions defaults to 1.
Sit with that. Every experiment you have run without thinking about it ran each example exactly once. You drew one sample per example from a nondeterministic system, compared it against another single sample, and concluded which version was better.
Raise it and each example runs multiple times. LangSmith then shows the average and the standard deviation across repetitions, which is frequently humbling: a two-point improvement looks less exciting once you can see the same code varies by five points between runs.
Reading the comparison view
You pick a baseline experiment and everything is coloured relative to it. Red means regression, green means improvement, per example and in aggregate, which lets you scan a long dataset and land on the handful of rows where your change did damage.
That drill-down is the actual value. An aggregate moving from 0.82 to 0.85 tells you almost nothing about what to do next. Six red rows that turn out to be the same category of question tells you where to go.
Colour is relative, always. A red cell is not a bad score, it is a score that went down. You can be red at 0.9 because your baseline was 0.95. The comparison view answers whether this change helped, never whether this is good. Which also makes the baseline a load-bearing and trivially abusable choice: compare against your worst historical run and the screen turns satisfyingly green. Say which baseline you used whenever you report a result.
What you do not get
No confidence intervals. No significance test. No minimum sample size guidance. Nothing warns you that twelve examples is too few, and the tool will happily draw you a lovely green comparison across all twelve.
I am naming the absence as carefully as the presence, because the gap is where people get hurt. A team that believes there is a significance test somewhere in the interface will assume any comparison they are shown has been vetted by one. There is no such vetting.
So decide with judgement and the spread in front of you: raise repetitions before you care about a result, compare the gap against the standard deviation, read the rows that moved, and grow the dataset, which is the crudest and most reliable lever there is.
In the exam, be suspicious of the most rigorous-sounding option. References to p-values, confidence intervals, statistical significance or recommended sample sizes describe capabilities the documentation does not contain, and the paper is semi-open-book against that documentation.
Try it yourself
Repetitions, casing and default
You want each example evaluated three times in a single TypeScript experiment.
const results = await evaluate(myApp, {
data: "my-dataset",
evaluators: [correctness],
/* ??? */
});Show answer
Correct answer: D — numRepetitions: 3, up from a default of 1
TypeScript uses numRepetitions, Python uses num_repetitions, and the default is 1 in both. Option 0 is the distractor and it is the Python casing, which is what you will have seen most often in docs and blog posts. A default of 1 is worth holding onto because it means every experiment you have run without thinking about this ran each example exactly once, with no measure of run-to-run variance at all.
The data parameter
A TypeScript evaluate() call omits the data parameter. What happens?
Show answer
Correct answer: A — It is an error, data is required in TypeScript even though it is optional in Python
data is optional in Python and required in TypeScript. Option 3 is the tempting one for anyone reasoning that an evaluation obviously needs to know what to evaluate against, which is a sensible instinct that happens to be wrong about Python. Option 2 invents a convenience default, and inventing plausible conveniences is a reliable way to lose marks in this domain, because the real APIs are less accommodating than you would guess.
What red means in the comparison view
You are looking at an experiment compared against a baseline. A cell is red. What does that tell you?
Show answer
Correct answer: C — That example regressed relative to the baseline experiment
Red means regression and green means improvement, both measured against the baseline experiment you chose. Option 1 is the distractor worth understanding: the colouring is relative, never absolute. A red cell can be a score of 0.9 if the baseline scored 0.95, and a green cell can be a dismal 0.2 if the baseline managed 0.1. Colour tells you direction of change, never quality of result.
What the tooling actually computes for you
Which of these does the documented LangSmith experiment tooling give you?
Show answer
Correct answer: B — An average and a standard deviation across repetitions
Average and standard deviation across repetitions is what you get, and repetitions is the only documented mechanism for handling run-to-run noise. Options 0, 2 and 3 are all things a rigorous evaluation setup would ideally have and none of them appear in the documentation. They are tempting precisely because they are what you would expect a serious evaluation product to provide, which is what makes this a good question: it tests whether you know the tool or whether you are reasoning from what a tool like this ought to do.
What a score difference licenses you to conclude
The judgement call this lesson builds to, and the one that generalises furthest beyond the exam.
Experiment B scores 0.86 against a baseline of 0.83, on a 40-example dataset, with repetitions left at the default. What can you legitimately conclude, and what would you need in order to conclude more?
Reveal answer
Almost nothing. With repetitions at 1 you have a single sample per example from a nondeterministic system, so you cannot separate a real three-point gain from run-to-run variance, and there is no confidence interval or significance test in the tooling to resolve it statistically either. To conclude more you would raise repetitions, compare the gap against the reported standard deviation, and drill into the individual rows that moved to check they moved for the reason you intended. Judgement, informed by spread, on inspected examples, is genuinely the state of the art in the documented tooling.
Report a result without overclaiming
Write one sentence reporting an experiment result the way you would put it in a pull request.
The sentence names the baseline, names the metric, gives the average, gives the standard deviation if you ran repetitions, and states the number of examples. If your sentence is accuracy improved to 87% with none of that context, rewrite it, because that sentence is unfalsifiable and it is the kind that gets a feature shipped on noise.