Test agents
The Test domain is LangSmith evaluation. This second set covers dataset splits and versions, the evaluator techniques (human annotation queues, pairwise, the two LLM-judge modes), the feedback an evaluator emits, when online evaluation fits, adding examples from a trace, the example inputs field, and pinning a CI benchmark to a dataset version. Ten questions.
Try it yourself
What dataset splits are for
What are dataset splits in LangSmith used for?
Show answer
Correct answer: B — They are named subsets that segment a dataset''s examples into separate groups
Splits are named subsets that segment examples into groups (ML-style, category-based, or staged rollout), and one example may sit in multiple splits. Docs: docs.langchain.com/langsmith/evaluation-concepts (Test).
Reproducing an old evaluation
You edit examples in a dataset, then later need to reproduce an earlier evaluation exactly. What does LangSmith provide?
Show answer
Correct answer: A — Dataset versions, created automatically whenever examples change, which you can pin
LangSmith automatically creates dataset versions when examples change; you can tag versions and pin a run to a specific one. Docs: docs.langchain.com/langsmith/evaluation-concepts (Test).
Humans scoring outputs
You want humans to manually review application outputs and their traces and score them. Which LangSmith evaluation approach is that?
Show answer
Correct answer: C — Human evaluation via annotation queues, in single-run or pairwise form
Human evaluation is manual review of outputs and traces, supported by annotation queues (single-run and pairwise types). Docs: docs.langchain.com/langsmith/evaluation-concepts (Test).
Who runs a pairwise comparison
A pairwise evaluation compares the outputs of two application versions. What can actually perform that comparison?
Show answer
Correct answer: D — A heuristic, an LLM, or a human, depending on how you configure the comparison
Pairwise evaluators compare outputs from two application versions using heuristics, LLMs, or humans. Docs: docs.langchain.com/langsmith/evaluation-concepts (Test).
Two modes of an LLM judge
An LLM-as-judge evaluator can run in two modes. What distinguishes them?
Show answer
Correct answer: B — Reference-free scores the output alone; reference-based compares it to a stored reference output
An LLM-as-judge can be reference-free (scoring the output on its own) or reference-based (scoring it against a stored reference output). Docs: docs.langchain.com/langsmith/evaluation-concepts (Test).
What an evaluator produces
When an evaluator scores a run or an example, what does it actually produce?
Show answer
Correct answer: A — Feedback: a key, a score or value, and an optional comment attached to the run
An evaluator emits feedback containing a key, a score or value, and an optional comment. Docs: docs.langchain.com/langsmith/evaluation-concepts (Test).
When online evaluation fits
Which task is online evaluation best suited to?
Show answer
Correct answer: D — Real-time monitoring and anomaly detection on live production runs
Online evaluation runs on live runs and threads (inputs and outputs only) for production monitoring and anomaly detection; the other three are offline, pre-deployment uses. Docs: docs.langchain.com/langsmith/evaluation-concepts (Test).
Turning a trace into an example
You spot a real production trace that would make a great regression case. How do you turn it into a dataset example?
Show answer
Correct answer: C — Select the run in the tracing project and use Add to Dataset to add it
In the tracing project you can multi-select runs and click Add to Dataset to add them as examples. Docs: docs.langchain.com/langsmith/manage-datasets (Test).
The inputs field of an Example
In a LangSmith dataset Example, what is the inputs field?
Show answer
Correct answer: A — A dictionary of input variables passed to your application when it runs
inputs is a dictionary of input variables passed to your application; the reference output (separate and optional) is used only inside evaluators. Docs: docs.langchain.com/langsmith/evaluation-concepts (Test).
Pinning a CI benchmark
Why would a CI evaluation pipeline pin itself to a specific dataset version?
Show answer
Correct answer: B — To keep the benchmark stable, running against a fixed set of examples as the dataset evolves
Pinning to a dataset version keeps a CI benchmark reproducible: it runs against a fixed set of examples even as the dataset changes later. Docs: docs.langchain.com/langsmith/evaluation-concepts (Test).