Test agents
Ten Test questions for LCAE Mock Exam C. Study them one at a time here, or take the whole timed paper in exam mode.
Try it yourself
Where a reference output cannot exist
In which situation can a reference output NOT exist?
Show answer
Correct answer: C — For a live production run a user left after asking their question
A reference output is the correct answer written in advance, attached to an example, and seen only by evaluators. Live production traffic has no example object and nobody wrote the answer down, so online evaluation never has a reference output; offline can. Docs: docs.langchain.com/langsmith/evaluation-concepts (Test).
Classifying an online safety check
You want a safety check that flags toxic responses, running continuously against live production traffic. Is it offline or online, and does it use a reference output?
Show answer
Correct answer: C — Online, and reference-free
"Production / live traffic / continuous" means online, which never has a reference output; a toxicity check scores the output on its own terms, so it is reference-free and can run in both worlds. Docs: docs.langchain.com/langsmith/evaluation-concepts (Test).
The perform_eval seam
An evaluator is written inline in the LangSmith UI and attached to a tracing project (i.e. online). What must it be named, and what argument(s) does it receive?
Show answer
Correct answer: C — perform_eval; it receives (run) only
In-UI evaluators must be named perform_eval so the platform can find them. Attached to a dataset (offline) it takes run and example; attached to a tracing project (online) there is no example, so it takes run alone. One parameter present offline, absent online. Docs: docs.langchain.com/langsmith/code-evaluator (Test).
A correctly-formed SDK evaluator
This experiment passes an evaluator named correctness into evaluate:
from langsmith import evaluate
results = evaluate(
my_app,
data="regression-set",
evaluators=[correctness],
)Which definition of correctness is a correctly-formed Python SDK
evaluator?
Show answer
Correct answer: A — def correctness(inputs, outputs, reference_outputs): ...
Python SDK evaluators take three positional arguments, inputs, outputs, reference_outputs, and have no name rule (they are found by reference). The destructured single-object form is TypeScript, and perform_eval(run, ...) is the in-UI surface, not an SDK evaluator. Docs: docs.langchain.com/langsmith/evaluate-llm-application (Test).
One score moved between identical runs
You run the same experiment twice with no code change and one example's score differs. What is the documented interpretation, and the fix?
Show answer
Correct answer: C — It is noise from model/judge variance; read the aggregate and raise repetitions from the default of 1
Model and judge nondeterminism produce per-row variance; the documented move is to read aggregate metrics, not react to one row. The lever is repetitions, and the default is 1, so by default there is no averaging at all. Docs: docs.langchain.com/langsmith/evaluation-concepts (Test).
Keeping a benchmark comparable
Between two experiments your dataset gained 10 examples. Are the two experiments directly comparable, and what lets a CI benchmark stay fixed as the dataset grows?
Show answer
Correct answer: B — No; pin the evaluation to a dataset version
Changing the examples changes the exam, so the two are not directly comparable. LangSmith auto-creates a dataset version whenever examples change, and you can pin an evaluation to a version so a benchmark runs against a fixed set. Splits are named subsets, not versions. Docs: docs.langchain.com/langsmith/manage-datasets (Test).
Comparison view versus pairwise
Which statement correctly distinguishes the experiment comparison view from a pairwise evaluation?
Show answer
Correct answer: B — The comparison view colours each example's existing scores against a baseline; a pairwise evaluation has a judge pick the better of two outputs and yields a win rate
The comparison view takes scores your evaluators already produced and colours them against a baseline. Pairwise puts two outputs side by side and a judge (heuristic, LLM, or human) picks; its output is a win rate, which is a preference, not a quality score. Docs: docs.langchain.com/langsmith/evaluate-pairwise (Test).
Trusting an LLM judge
Before gating a release on an LLM-as-judge, you align it. An alignment score measures what, and how many labels do you need?
Show answer
Correct answer: B — Agreement with human labels; at least 20, balanced across pass and fail
Alignment is the percentage of examples where the evaluator matches the human expert, agreement rather than objective accuracy. Use at least 20 balanced labels; an all-failures set is aced by a judge that calls everything a failure, so balance forces discrimination. Docs: docs.langchain.com/langsmith/align-evaluator (Test).
Getting a production run into a dataset
Which of the following are valid ways to add a real production run to a LangSmith dataset? Select all that apply.
Show answer
Correct answer: A — Multi-select from the Runs table, an automation rule that matches runs, and sending a run from an annotation queue
UI routes include Runs-table multi-select, automation rules, the annotation queue (hotkey D), the Playground, and the Examples tab; SDK routes are create-examples, CSV upload, and dataframe upload (Python only). Runs are immutable, so you cannot edit a trace into an example. Docs: docs.langchain.com/langsmith/manage-datasets (Test).
Correcting a score inline
While reviewing results, someone corrects one wrong evaluator score inline, with no queue and no session. What feature is this, and does it produce an alignment percentage?
Show answer
Correct answer: B — Few-shot self-improving evaluators; the correction is injected into the judge prompt and it produces no number
Few-shot self-improving evaluators inject the inline correction into the judge prompt at a few-shot placeholder; it modifies the prompt and produces no number. If a question mentions a percentage, it is Align Evals. Docs: docs.langchain.com/langsmith/align-evaluator (Test).