Whetstone.
LangSmith: Tracing, Datasets, and EvalsDatasets and evaluate(): a real accuracy score
Module 8, Lesson 225 min

Datasets and evaluate(): a real accuracy score

Datasets are labelled examples. A dataset is a set of { inputs, outputs } pairs. Build one from a real run: chronicler’s cached day results are already input (messages) plus expected output (A/B/C decisions). Upload those as a dataset and you have ground truth.

evaluate() runs an experiment. You give it a target function, a dataset, and evaluators, functions that score one output against its expected value.

Running an evaluation experiment
import { evaluate } from "langsmith/evaluation";

await evaluate(
  (input) => runClassifier(input),          // target
  {
    data: "chronicler-day-2026-07-01",       // dataset name
    evaluators: [
      ({ run, example }) => ({
        key: "category_match",
        score: run.outputs.category === example.outputs.category ? 1 : 0,
      }),
    ],
  },
);

Each evaluator returns { key, score }. LangSmith aggregates them into an experiment with a real number: classification accuracy across the dataset. Verify the exact evaluate() TS signature against the LangSmith reference at build time, the evaluator argument shape in particular is worth confirming.

Practice

Try it yourself

Recall

Why chronicler's cached day is a good dataset source

This has two halves, and the second is easy to skip. A dataset needs inputs and a trusted expected output.

Why chronicler's cached day specifically, and where does "expected output" come from?

Reveal answer

The cached day already has real messages (inputs) paired with the classifications that were actually produced and cached (which, once you've spot-checked them, serve as a trusted expected output / ground truth). It's real production data rather than synthetic examples, which makes the eval meaningful -- you're scoring against decisions that already survived scrutiny in the real pipeline.

Quiz

Evaluator vs target function

In evaluate(target, { data, evaluators }), what is the difference between the target function and an evaluator?

  1. AThey are interchangeable; evaluate() calls both of them the exact same way
  2. BThe target function returns a score, the evaluator returns the output
  3. CThere's no target function, only evaluators
  4. DThe target is the code under test; the evaluator scores its output against the expected one
Show answer

Correct answer: D — The target is the code under test; the evaluator scores its output against the expected one

evaluate() runs the target function against every example in the dataset to produce actual outputs, then runs each evaluator to compare those outputs against expected values and produce a {key, score}. The target is what you're testing; evaluators are how you grade it.

Do

Trace, dataset, and eval: classifier plus vault-tidy

Instrument the module 1 and 2 builds, then extend the same lens to vault-tidy.

  • Set LANGSMITH_TRACING=true and LANGSMITH_API_KEY, and confirm the m1 classifier spike and m2 todoist agent both appear as traces in LangSmith.
  • Wrap one raw-SDK call (or a small standalone script using @anthropic-ai/sdk) with wrapAnthropic and confirm it also traces.
  • Build a dataset from chronicler's cached day results: messages as inputs, cached A/B/C decisions as expected outputs.
  • Run evaluate() with a category_match evaluator and get a real accuracy score for the m1 classifier.
  • Point the same lens at vault-tidy: take a sample of notes, define what a 'correct' frontmatter fix looks like, and score vault-tidy's actual fixes against that sample.
  • Write up the vault-tidy baseline as a real measured number, not a vibe.
Done whenTraces are visible in LangSmith for both the graph run and the wrapAnthropic-wrapped call, one dataset and one experiment produced a real accuracy score, and a written vault-tidy quality baseline exists.
Check

vault-tidy quality is measured, not vibed

Confirm the vault-tidy eval produced a real baseline.

You should see

A concrete accuracy or quality score exists for vault-tidy's frontmatter fixes on a real sample of notes, with a written explanation of how expected outputs were constructed for a task that has no obvious ground truth.

Sign in to track your progress →