Whetstone.
Test agentsTest agents
Focus area: Test20 min

Test agents

The Test domain is LangSmith evaluation: datasets and examples, evaluator types, and reading experiments without overfitting to noise. Ten questions.

Practice

Try it yourself

Quiz

What a Dataset is

In LangSmith, a Dataset is best described as:

  1. AA deployed revision of a graph running in production
  2. BA group of traces bound together by one thread_id
  3. CA single input and output pair captured from exactly one production run
  4. DA collection of Examples used to evaluate an application
Show answer

Correct answer: D — A collection of Examples used to evaluate an application

A Dataset is a collection of Examples used to evaluate an application. Docs: docs.langchain.com/langsmith/evaluate-llm-application (Test).

Quiz

Reference outputs

Reference outputs on an Example are:

  1. AUsed only inside evaluators, and never passed to the app
  2. BRequired for every single online evaluation to run
  3. CJust another name for the feedback scores on a run
  4. DShown to the agent at runtime as additional grounding context
Show answer

Correct answer: A — Used only inside evaluators, and never passed to the app

Reference outputs are used only inside evaluators and are never passed to the app. Docs: docs.langchain.com/langsmith/evaluate-llm-application (Test).

Quiz

Online-capable evaluators

Which evaluator type can run on live production traffic (online evaluation)?

  1. AReference-based correctness checks measured against a stored gold answer
  2. BReference-free evaluators (safety, format, quality heuristics)
  3. CExact-match evaluators comparing output to the gold answer
  4. DAny evaluator, provided it is supplied a reference output
Show answer

Correct answer: B — Reference-free evaluators (safety, format, quality heuristics)

Reference-free evaluators run both offline and online; reference-based ones need a gold answer and are offline only. Docs: docs.langchain.com/langsmith (Test).

Quiz

Deterministic JSON check

You want a deterministic check that the output is valid JSON of the right shape. Best evaluator type?

  1. AAn LLM-as-judge scoring the structure qualitatively for validity
  2. BA pairwise comparison against another version's output
  3. CA code / heuristic evaluator that checks the shape directly
  4. DA human annotation queue review of each output
Show answer

Correct answer: C — A code / heuristic evaluator that checks the shape directly

Deterministic shape and format checks are code / heuristic evaluators. Docs: docs.langchain.com/langsmith/evaluate-llm-application (Test).

Quiz

When pairwise fits

When is pairwise evaluation the right choice?

  1. AWhen there are no reference outputs available anywhere for the dataset being scored
  2. BWhen you want to instrument production traces for cost
  3. CWhen you need a precise numeric latency measurement
  4. DWhen directly scoring one output is hard but comparing two is straightforward
Show answer

Correct answer: D — When directly scoring one output is hard but comparing two is straightforward

Pairwise fits when comparing two outputs is easier than scoring one absolutely. Docs: docs.langchain.com/langsmith (Test).

Quiz

What an Experiment is

An Experiment in LangSmith represents:

  1. AThe results of evaluating a specific application version on a dataset
  2. BA deployment environment such as cloud versus hybrid
  3. CA configured graph together with its system prompt
  4. DA single production run of one deployed assistant executing on a thread
Show answer

Correct answer: A — The results of evaluating a specific application version on a dataset

An Experiment is the result of evaluating one specific application version on a dataset. Docs: docs.langchain.com/langsmith/evaluate-llm-application (Test).

Quiz

Evaluator alignment

"Evaluator alignment" refers to:

  1. AAligning the base model's internal weights during a supervised fine-tuning run on labelled data
  2. BMaking an automated evaluator agree with human judgement before trusting it at scale
  3. CAligning trace timestamps consistently across serving regions
  4. DMatching a deployment revision back to its source graph
Show answer

Correct answer: B — Making an automated evaluator agree with human judgement before trusting it at scale

Evaluator alignment means tuning an automated judge (e.g. LLM-as-judge) to agree with humans before scaling. The Study Pack marks the exact LangSmith UI or API name [unverified] at source; the tested concept (align the judge to humans, then scale) is what to know. Docs: docs.langchain.com/langsmith (Test).

Quiz

Offline vs online

Offline evaluation differs from online evaluation primarily because offline:

  1. AOnly works against already-deployed assistants
  2. BRuns on live production runs and threads and has no reference outputs at all
  3. CRuns on datasets/examples and can use reference outputs, pre-deployment
  4. DCannot make use of code-based evaluators at all
Show answer

Correct answer: C — Runs on datasets/examples and can use reference outputs, pre-deployment

Offline evaluation runs on datasets/examples with reference outputs, pre-deployment. Docs: docs.langchain.com/langsmith/evaluate-llm-application (Test).

Quiz

Reading noisy scores

A single example's score jumps around between runs. The exam-endorsed reading is:

  1. ADelete the noisy example from the dataset entirely
  2. BImmediately roll back the current deployment to the previous known-good revision just to be safe
  3. CTrust the single example and ship the fix straight away
  4. DLook at aggregate metrics across the dataset rather than overfitting to one noisy example
Show answer

Correct answer: D — Look at aggregate metrics across the dataset rather than overfitting to one noisy example

Read aggregate metrics across the dataset; do not overfit to one noisy example (the guide's "without overfitting to noise"). Docs: docs.langchain.com/langsmith (Test).

Quiz

Few-shot and judges

Few-shot examples are most associated with improving which evaluator type?

  1. AAn LLM-as-judge evaluator scoring subjective quality
  2. BAn exact-match evaluator against a gold answer
  3. CA latency measurement recorded per run
  4. DA code / heuristic evaluator running deterministic rules
Show answer

Correct answer: A — An LLM-as-judge evaluator scoring subjective quality

Few-shot examples improve LLM-as-judge performance. Docs: docs.langchain.com/langsmith (Test).

Sign in to track your progress →