Whetstone.
Test agentsTest agents
Focus area: Test20 min

Test agents

The Test domain is LangSmith evaluation. This second set covers dataset splits and versions, the evaluator techniques (human annotation queues, pairwise, the two LLM-judge modes), the feedback an evaluator emits, when online evaluation fits, adding examples from a trace, the example inputs field, and pinning a CI benchmark to a dataset version. Ten questions.

Practice

Try it yourself

Quiz

What dataset splits are for

What are dataset splits in LangSmith used for?

  1. AThey store the reference outputs separately from the example inputs for safety
  2. BThey are named subsets that segment a dataset''s examples into separate groups
  3. CThey version the dataset automatically each time an example is added or changed
  4. DThey group live production traces into a single multi-turn session for review
Show answer

Correct answer: B — They are named subsets that segment a dataset''s examples into separate groups

Splits are named subsets that segment examples into groups (ML-style, category-based, or staged rollout), and one example may sit in multiple splits. Docs: docs.langchain.com/langsmith/evaluation-concepts (Test).

Quiz

Reproducing an old evaluation

You edit examples in a dataset, then later need to reproduce an earlier evaluation exactly. What does LangSmith provide?

  1. ADataset versions, created automatically whenever examples change, which you can pin
  2. BA checkpointer snapshot of the dataset that the agent restores on the next run
  3. CA deployment revision that captures the dataset state at the moment you deployed
  4. DA thread that records each change as its own turn, which you can replay in order later on
Show answer

Correct answer: A — Dataset versions, created automatically whenever examples change, which you can pin

LangSmith automatically creates dataset versions when examples change; you can tag versions and pin a run to a specific one. Docs: docs.langchain.com/langsmith/evaluation-concepts (Test).

Quiz

Humans scoring outputs

You want humans to manually review application outputs and their traces and score them. Which LangSmith evaluation approach is that?

  1. AA code or heuristic evaluator that applies deterministic rules to each output
  2. BAn LLM-as-judge evaluator scoring each output against a written rubric
  3. CHuman evaluation via annotation queues, in single-run or pairwise form
  4. DAn online evaluator that scores incoming production runs automatically
Show answer

Correct answer: C — Human evaluation via annotation queues, in single-run or pairwise form

Human evaluation is manual review of outputs and traces, supported by annotation queues (single-run and pairwise types). Docs: docs.langchain.com/langsmith/evaluation-concepts (Test).

Quiz

Who runs a pairwise comparison

A pairwise evaluation compares the outputs of two application versions. What can actually perform that comparison?

  1. AOnly an LLM-as-judge, because a pairwise choice always needs a language model
  2. BOnly a deterministic code evaluator comparing the two outputs field by field
  3. COnly a human reviewer, since ranking two answers is inherently subjective work
  4. DA heuristic, an LLM, or a human, depending on how you configure the comparison
Show answer

Correct answer: D — A heuristic, an LLM, or a human, depending on how you configure the comparison

Pairwise evaluators compare outputs from two application versions using heuristics, LLMs, or humans. Docs: docs.langchain.com/langsmith/evaluation-concepts (Test).

Quiz

Two modes of an LLM judge

An LLM-as-judge evaluator can run in two modes. What distinguishes them?

  1. AOne runs offline and the other online; that timing is the only real difference
  2. BReference-free scores the output alone; reference-based compares it to a stored reference output
  3. COne is deterministic while the other is not; the judge model makes it repeatable
  4. DOne mode is allowed to use few-shot examples while the other is strictly forbidden from using any at all
Show answer

Correct answer: B — Reference-free scores the output alone; reference-based compares it to a stored reference output

An LLM-as-judge can be reference-free (scoring the output on its own) or reference-based (scoring it against a stored reference output). Docs: docs.langchain.com/langsmith/evaluation-concepts (Test).

Quiz

What an evaluator produces

When an evaluator scores a run or an example, what does it actually produce?

  1. AFeedback: a key, a score or value, and an optional comment attached to the run
  2. BA new example that is appended to the dataset, with the evaluator score stored as its reference
  3. CA deployment revision recording the evaluated version of the application
  4. DA reference output written back onto the example for the next evaluation
Show answer

Correct answer: A — Feedback: a key, a score or value, and an optional comment attached to the run

An evaluator emits feedback containing a key, a score or value, and an optional comment. Docs: docs.langchain.com/langsmith/evaluation-concepts (Test).

Quiz

When online evaluation fits

Which task is online evaluation best suited to?

  1. APre-deployment regression testing of a change against a labelled dataset
  2. BBacktesting a new prompt over historical examples with known good answers
  3. CUnit-testing the agent''s output shape in CI before a release goes out
  4. DReal-time monitoring and anomaly detection on live production runs
Show answer

Correct answer: D — Real-time monitoring and anomaly detection on live production runs

Online evaluation runs on live runs and threads (inputs and outputs only) for production monitoring and anomaly detection; the other three are offline, pre-deployment uses. Docs: docs.langchain.com/langsmith/evaluation-concepts (Test).

Quiz

Turning a trace into an example

You spot a real production trace that would make a great regression case. How do you turn it into a dataset example?

  1. AAttach a checkpointer to the run so its state is saved into the dataset directly
  2. BAdd a reference output to the run and it enrolls into the dataset automatically
  3. CSelect the run in the tracing project and use Add to Dataset to add it
  4. DCreate a deployment revision from the run, which snapshots it as an example
Show answer

Correct answer: C — Select the run in the tracing project and use Add to Dataset to add it

In the tracing project you can multi-select runs and click Add to Dataset to add them as examples. Docs: docs.langchain.com/langsmith/manage-datasets (Test).

Quiz

The inputs field of an Example

In a LangSmith dataset Example, what is the inputs field?

  1. AA dictionary of input variables passed to your application when it runs
  2. BThe reference answer the evaluator compares the application output against
  3. CThe feedback score recorded for the example after an evaluation finishes
  4. DThe model and prompt configuration the application should use for this case
Show answer

Correct answer: A — A dictionary of input variables passed to your application when it runs

inputs is a dictionary of input variables passed to your application; the reference output (separate and optional) is used only inside evaluators. Docs: docs.langchain.com/langsmith/evaluation-concepts (Test).

Quiz

Pinning a CI benchmark

Why would a CI evaluation pipeline pin itself to a specific dataset version?

  1. ATo lower token cost, because older dataset versions contain fewer examples
  2. BTo keep the benchmark stable, running against a fixed set of examples as the dataset evolves
  3. CTo enable online evaluation, which the docs say can only attach to a pinned dataset version
  4. DTo share the dataset publicly, because only a pinned version can be made visible to people outside the workspace
Show answer

Correct answer: B — To keep the benchmark stable, running against a fixed set of examples as the dataset evolves

Pinning to a dataset version keeps a CI benchmark reproducible: it runs against a fixed set of examples even as the dataset changes later. Docs: docs.langchain.com/langsmith/evaluation-concepts (Test).

Sign in to track your progress →