Whetstone.
Test agentsTest agents
Focus area: Test22 min

Test agents

Ten Test questions for LCAE Mock Exam C. Study them one at a time here, or take the whole timed paper in exam mode.

Practice

Try it yourself

Quiz

Where a reference output cannot exist

In which situation can a reference output NOT exist?

  1. AIn a curated dataset built before deployment
  2. BFor an example a reviewer hand-writes for a regression set
  3. CFor a live production run a user left after asking their question
  4. DIn an offline experiment run against a dataset
Show answer

Correct answer: C — For a live production run a user left after asking their question

A reference output is the correct answer written in advance, attached to an example, and seen only by evaluators. Live production traffic has no example object and nobody wrote the answer down, so online evaluation never has a reference output; offline can. Docs: docs.langchain.com/langsmith/evaluation-concepts (Test).

Quiz

Classifying an online safety check

You want a safety check that flags toxic responses, running continuously against live production traffic. Is it offline or online, and does it use a reference output?

  1. AOffline, and reference-based
  2. BOffline, and reference-free
  3. COnline, and reference-free
  4. DOnline, and reference-based
Show answer

Correct answer: C — Online, and reference-free

"Production / live traffic / continuous" means online, which never has a reference output; a toxicity check scores the output on its own terms, so it is reference-free and can run in both worlds. Docs: docs.langchain.com/langsmith/evaluation-concepts (Test).

Quiz

The perform_eval seam

An evaluator is written inline in the LangSmith UI and attached to a tracing project (i.e. online). What must it be named, and what argument(s) does it receive?

  1. AAny name; it receives (run, example)
  2. Bperform_eval; it receives (run, example)
  3. Cperform_eval; it receives (run) only
  4. Devaluate; it receives (run) only
Show answer

Correct answer: C — perform_eval; it receives (run) only

In-UI evaluators must be named perform_eval so the platform can find them. Attached to a dataset (offline) it takes run and example; attached to a tracing project (online) there is no example, so it takes run alone. One parameter present offline, absent online. Docs: docs.langchain.com/langsmith/code-evaluator (Test).

Quiz

A correctly-formed SDK evaluator

This experiment passes an evaluator named correctness into evaluate:

python
from langsmith import evaluate

results = evaluate(
    my_app,
    data="regression-set",
    evaluators=[correctness],
)

Which definition of correctness is a correctly-formed Python SDK evaluator?

  1. Adef correctness(inputs, outputs, reference_outputs): ...
  2. Bdef correctness({inputs, outputs, referenceOutputs}): ...
  3. Cdef perform_eval(run, example): ...
  4. Ddef correctness(run): ...
Show answer

Correct answer: A — def correctness(inputs, outputs, reference_outputs): ...

Python SDK evaluators take three positional arguments, inputs, outputs, reference_outputs, and have no name rule (they are found by reference). The destructured single-object form is TypeScript, and perform_eval(run, ...) is the in-UI surface, not an SDK evaluator. Docs: docs.langchain.com/langsmith/evaluate-llm-application (Test).

Quiz

One score moved between identical runs

You run the same experiment twice with no code change and one example's score differs. What is the documented interpretation, and the fix?

  1. AIt is a regression; roll back the change
  2. BIt is noise; delete the flapping example
  3. CIt is noise from model/judge variance; read the aggregate and raise repetitions from the default of 1
  4. DIt is a bug in the evaluator; rewrite it
Show answer

Correct answer: C — It is noise from model/judge variance; read the aggregate and raise repetitions from the default of 1

Model and judge nondeterminism produce per-row variance; the documented move is to read aggregate metrics, not react to one row. The lever is repetitions, and the default is 1, so by default there is no averaging at all. Docs: docs.langchain.com/langsmith/evaluation-concepts (Test).

Quiz

Keeping a benchmark comparable

Between two experiments your dataset gained 10 examples. Are the two experiments directly comparable, and what lets a CI benchmark stay fixed as the dataset grows?

  1. AYes, comparable; splits keep the set fixed
  2. BNo; pin the evaluation to a dataset version
  3. CYes; repetitions keep the set fixed
  4. DNo; snapshot the set with a deployment revision
Show answer

Correct answer: B — No; pin the evaluation to a dataset version

Changing the examples changes the exam, so the two are not directly comparable. LangSmith auto-creates a dataset version whenever examples change, and you can pin an evaluation to a version so a benchmark runs against a fixed set. Splits are named subsets, not versions. Docs: docs.langchain.com/langsmith/manage-datasets (Test).

Quiz

Comparison view versus pairwise

Which statement correctly distinguishes the experiment comparison view from a pairwise evaluation?

  1. AThey are two names for the same feature
  2. BThe comparison view colours each example's existing scores against a baseline; a pairwise evaluation has a judge pick the better of two outputs and yields a win rate
  3. CThe comparison view produces a win rate; pairwise colours existing scores
  4. DPairwise evaluation can only ever use a human judge
Show answer

Correct answer: B — The comparison view colours each example's existing scores against a baseline; a pairwise evaluation has a judge pick the better of two outputs and yields a win rate

The comparison view takes scores your evaluators already produced and colours them against a baseline. Pairwise puts two outputs side by side and a judge (heuristic, LLM, or human) picks; its output is a win rate, which is a preference, not a quality score. Docs: docs.langchain.com/langsmith/evaluate-pairwise (Test).

Quiz

Trusting an LLM judge

Before gating a release on an LLM-as-judge, you align it. An alignment score measures what, and how many labels do you need?

  1. AAccuracy against objective truth; 50 labels
  2. BAgreement with human labels; at least 20, balanced across pass and fail
  3. CThe judge's speed over the dataset; 1000 labels
  4. DCorrelation with token count; 10 labels
Show answer

Correct answer: B — Agreement with human labels; at least 20, balanced across pass and fail

Alignment is the percentage of examples where the evaluator matches the human expert, agreement rather than objective accuracy. Use at least 20 balanced labels; an all-failures set is aced by a judge that calls everything a failure, so balance forces discrimination. Docs: docs.langchain.com/langsmith/align-evaluator (Test).

Quiz

Getting a production run into a dataset

Which of the following are valid ways to add a real production run to a LangSmith dataset? Select all that apply.

  1. AMulti-select from the Runs table, an automation rule that matches runs, and sending a run from an annotation queue
  2. BOnly by exporting the run to CSV in the browser and re-importing it
  3. CBy editing the run's output inline in the trace view so it becomes an example
  4. DOnly through the SDK; there is no UI route at all
Show answer

Correct answer: A — Multi-select from the Runs table, an automation rule that matches runs, and sending a run from an annotation queue

UI routes include Runs-table multi-select, automation rules, the annotation queue (hotkey D), the Playground, and the Examples tab; SDK routes are create-examples, CSV upload, and dataframe upload (Python only). Runs are immutable, so you cannot edit a trace into an example. Docs: docs.langchain.com/langsmith/manage-datasets (Test).

Quiz

Correcting a score inline

While reviewing results, someone corrects one wrong evaluator score inline, with no queue and no session. What feature is this, and does it produce an alignment percentage?

  1. AAlign Evals; yes, it reports an alignment percentage
  2. BFew-shot self-improving evaluators; the correction is injected into the judge prompt and it produces no number
  3. CAn annotation queue; yes, it reports a percentage
  4. DDataset versioning; no
Show answer

Correct answer: B — Few-shot self-improving evaluators; the correction is injected into the judge prompt and it produces no number

Few-shot self-improving evaluators inject the inline correction into the judge prompt at a few-shot placeholder; it modifies the prompt and produces no number. If a question mentions a percentage, it is Align Evals. Docs: docs.langchain.com/langsmith/align-evaluator (Test).

Sign in to track your progress →