Whetstone.
Monitor agentsMonitor agents
Focus area: Monitor22 min

Monitor agents

Ten Monitor questions for LCAE Mock Exam C. Study them one at a time here, or take the whole timed paper in exam mode.

Practice

Try it yourself

Quiz

Building the nested trace tree

Runs are posted asynchronously and can arrive out of order. Which two fields let the UI draw the correct nested tree, and why is a shared trace id not enough?

  1. Atrace_id and the run's start timestamp
  2. Bparent_run_id and dotted_order
  3. Cthread_id and session_id
  4. Drun_type and parent_run_id only
Show answer

Correct answer: B — parent_run_id and dotted_order

parent_run_id gives the edges; dotted_order encodes each run's full path from the root as a sortable string, so the frontend renders order and indentation by sorting a flat list. A shared trace_id only says the runs belong together, and timestamps cannot recover depth. Docs: docs.langchain.com/langsmith/observability-concepts (Monitor).

Quiz

Two thread keys at once

A conversation's runs carry both thread_id = abc and session_id = xyz in metadata. Which value does LangSmith group the thread by?

  1. Axyz, because the legacy session_id key wins
  2. Babc, because thread_id is read first
  3. CNeither; you must create a thread object explicitly
  4. DBoth, producing two separate threads
Show answer

Correct answer: B — abc, because thread_id is read first

Two keys are accepted for thread grouping, thread_id and session_id, and the backend reads thread_id first, against the backward-compatibility instinct. (Separately, session_id in the SDK/REST sense means the tracing project id.) Docs: docs.langchain.com/langsmith/observability-concepts (Monitor).

Quiz

Reading latency down the tree

Sorting the runs table by latency, the slowest run is the root trace at 7.1s. Its only child, a reranker step, shows 6.9s, and that reranker's own child (an LLM call) shows 0.2s. Where is the time actually being spent?

  1. AIn the root trace, since it has the highest latency
  2. BIn the reranker step, which has about 6.7s of its own time beyond its child
  3. CIn the LLM call at the bottom
  4. DIt cannot be determined from these numbers
Show answer

Correct answer: B — In the reranker step, which has about 6.7s of its own time beyond its child

A parent's latency includes its children, so the root is always near the top and least useful. Walk down subtracting what children account for: the reranker's 6.9s minus its 0.2s child leaves about 6.7s of its own work, so it is the slow leaf. Docs: docs.langchain.com/langsmith/dashboards (Monitor).

Quiz

Thumbs against a judge

For the same week of traffic, end-user thumbs feedback reads 94% positive while an online LLM-groundedness judge averages 0.61. Which number is wrong?

  1. AThe thumbs; tune them down to match the judge
  2. BThe judge; tune it up until it agrees with the thumbs
  3. CNeither: thumbs are self-selected and sparse, while the judge is dense but only as good as its criterion
  4. DBoth are invalid without a reference output
Show answer

Correct answer: C — Neither: thumbs are self-selected and sparse, while the judge is dense but only as good as its criterion

Explicit feedback is trustworthy but self-selected and sparse; an LLM judge is dense but only as good as its criterion. They measure different things, and tuning the judge to agree destroys your only broad-coverage signal. Docs: docs.langchain.com/langsmith/online-evaluations (Monitor).

Quiz

A blank cost cell

Run A shows a cost; run B is blank. Same model, same day, and both show token counts. Run B's metadata carries the key ls_model_nme. Which of the three cost conditions failed?

  1. ACondition 1 (token counts), because usage metadata is missing
  2. BCondition 2, the model-identity key ls_model_name, which is misspelled and matches nothing
  3. CCondition 3, because there is no active pricing entry for that date
  4. DNone; cost tracking is plan-gated and B is on a lower plan
Show answer

Correct answer: B — Condition 2, the model-identity key ls_model_name, which is misspelled and matches nothing

Cost needs token counts, the ls_provider and ls_model_name metadata keys, and an active regex-matching pricing entry. Tokens are visible (rules out condition 1); a dated pricing gap would blank both same-day runs (rules out condition 3); the misspelled ls_model_nme is present, plausible, and matches nothing. Misspelled is worse than missing. Docs: docs.langchain.com/langsmith/monitoring-costs (Monitor).

Quiz

Instrumenting your own function

This helper never appears in the trace tree, and its time is silently absorbed into its parent:

python
def fetch_docs(query: str) -> list[str]:
    return search(query)

Which change makes it show up as its own run in the LangSmith trace?

  1. ADecorate it with @traceable
  2. BRename it so it starts with run_
  3. CPass it to create_agent as the response_format
  4. DNothing is needed; LangChain auto-instruments every function you write
Show answer

Correct answer: A — Decorate it with @traceable

LangChain and LangGraph components auto-instrument, but your own functions need the traceable decorator (Python) or wrapper (TypeScript), the trace context manager (Python-only), or the RunTree API. An undecorated helper is invisible. Docs: docs.langchain.com/langsmith/observability-concepts (Monitor).

Quiz

What is true of alerts

Which of the following are TRUE of LangSmith alerts? Select all that apply.

  1. AThe only rule type is a threshold, windows are 5 or 15 minutes, and Feedback Score is an alertable metric
  2. BAlerts support anomaly detection with baseline learning
  3. CThere is a native email channel alongside Slack and PagerDuty
  4. DAlert windows can be set to any custom duration
Show answer

Correct answer: A — The only rule type is a threshold, windows are 5 or 15 minutes, and Feedback Score is an alertable metric

Alerts are threshold-only over 5- or 15-minute windows; the five metrics are Run Count, Cost, Errors, Feedback Score, and Latency; channels are Slack, PagerDuty, Dynatrace, and Webhook (email goes out through a Webhook). No anomaly detection, no native email, no custom windows. Docs: docs.langchain.com/langsmith/alerts (Monitor).

Quiz

Where each output goes

Where does the OUTPUT of Insights go, versus the output of an online evaluator?

  1. ABoth write feedback onto the run
  2. BInsights writes feedback; an online evaluator produces a report
  3. CInsights produces a report of categories you read; an online evaluator writes feedback onto the run or thread
  4. DBoth produce a scheduled report you read
Show answer

Correct answer: C — Insights produces a report of categories you read; an online evaluator writes feedback onto the run or thread

Insights is unsupervised discovery producing a report of categories and subcategories; an online evaluator is supervised measurement writing feedback, which is what makes it filterable, chartable, and alertable. Output and plan are the two rows people get wrong. Docs: docs.langchain.com/langsmith/insights (Monitor).

Quiz

A cap and its mirror image

How many traces does a single Insights job read, and which cap is it commonly confused with?

  1. A25,000 traces, confused with runs per Insights job
  2. B1,000 traces per Insights job, confused with the 25,000-runs-per-trace cap
  3. C100,000 traces, confused with the context-editing threshold
  4. DUnlimited, with no related cap
Show answer

Correct answer: B — 1,000 traces per Insights job, confused with the 25,000-runs-per-trace cap

Insights is capped at 1,000 traces per run, and the mirror-image number is 25,000 runs per trace. Both numbers right and attached to the wrong things is the most common way to be wrong. Docs: docs.langchain.com/langsmith/insights (Monitor).

Quiz

Harvesting thumbs-downs

Users leave a thumbs-down about once a day. Should an automation add those runs straight to your evaluation dataset?

  1. AYes; harvested failures are the best regression tests
  2. BNo: route them through an annotation queue first, because thumbs-downs include noise and the harvested runs have no reference output for a correctness evaluator
  3. CYes, but only when delivered through a Webhook
  4. DNo; thumbs-downs cannot trigger an automation at all
Show answer

Correct answer: B — No: route them through an annotation queue first, because thumbs-downs include noise and the harvested runs have no reference output for a correctness evaluator

An automation can add-to-dataset, add-to-annotation-queue, or trigger an online evaluator. Straight-to-dataset poisons it: thumbs-downs include misunderstandings and slow replies, and the runs carry no reference output, so route through an annotation queue where a human validates the failure and writes the reference output. Docs: docs.langchain.com/langsmith/automations (Monitor).

Sign in to track your progress →