Test agents
The Test domain is LangSmith evaluation: datasets and examples, evaluator types, and reading experiments without overfitting to noise. Ten questions.
Try it yourself
What a Dataset is
In LangSmith, a Dataset is best described as:
Show answer
Correct answer: D — A collection of Examples used to evaluate an application
A Dataset is a collection of Examples used to evaluate an application. Docs: docs.langchain.com/langsmith/evaluate-llm-application (Test).
Reference outputs
Reference outputs on an Example are:
Show answer
Correct answer: A — Used only inside evaluators, and never passed to the app
Reference outputs are used only inside evaluators and are never passed to the app. Docs: docs.langchain.com/langsmith/evaluate-llm-application (Test).
Online-capable evaluators
Which evaluator type can run on live production traffic (online evaluation)?
Show answer
Correct answer: B — Reference-free evaluators (safety, format, quality heuristics)
Reference-free evaluators run both offline and online; reference-based ones need a gold answer and are offline only. Docs: docs.langchain.com/langsmith (Test).
Deterministic JSON check
You want a deterministic check that the output is valid JSON of the right shape. Best evaluator type?
Show answer
Correct answer: C — A code / heuristic evaluator that checks the shape directly
Deterministic shape and format checks are code / heuristic evaluators. Docs: docs.langchain.com/langsmith/evaluate-llm-application (Test).
When pairwise fits
When is pairwise evaluation the right choice?
Show answer
Correct answer: D — When directly scoring one output is hard but comparing two is straightforward
Pairwise fits when comparing two outputs is easier than scoring one absolutely. Docs: docs.langchain.com/langsmith (Test).
What an Experiment is
An Experiment in LangSmith represents:
Show answer
Correct answer: A — The results of evaluating a specific application version on a dataset
An Experiment is the result of evaluating one specific application version on a dataset. Docs: docs.langchain.com/langsmith/evaluate-llm-application (Test).
Evaluator alignment
"Evaluator alignment" refers to:
Show answer
Correct answer: B — Making an automated evaluator agree with human judgement before trusting it at scale
Evaluator alignment means tuning an automated judge (e.g. LLM-as-judge) to agree with humans before scaling. The Study Pack marks the exact LangSmith UI or API name [unverified] at source; the tested concept (align the judge to humans, then scale) is what to know. Docs: docs.langchain.com/langsmith (Test).
Offline vs online
Offline evaluation differs from online evaluation primarily because offline:
Show answer
Correct answer: C — Runs on datasets/examples and can use reference outputs, pre-deployment
Offline evaluation runs on datasets/examples with reference outputs, pre-deployment. Docs: docs.langchain.com/langsmith/evaluate-llm-application (Test).
Reading noisy scores
A single example's score jumps around between runs. The exam-endorsed reading is:
Show answer
Correct answer: D — Look at aggregate metrics across the dataset rather than overfitting to one noisy example
Read aggregate metrics across the dataset; do not overfit to one noisy example (the guide's "without overfitting to noise"). Docs: docs.langchain.com/langsmith (Test).
Few-shot and judges
Few-shot examples are most associated with improving which evaluator type?
Show answer
Correct answer: A — An LLM-as-judge evaluator scoring subjective quality
Few-shot examples improve LLM-as-judge performance. Docs: docs.langchain.com/langsmith (Test).