Whetstone.
Monitor agentsMonitor agents
Focus area: Monitor20 min

Monitor agents

The Monitor domain is LangSmith observability: the trace data model, grouping traces into threads, reading traces, cost and feedback, alerting, dashboards, and Online Evals versus Insights. Ten questions.

Practice

Try it yourself

Quiz

Run as a span

In the LangSmith trace model, a Run is analogous to what in OpenTelemetry?

  1. AA trace spanning one whole operation
  2. BA span within a trace
  3. CA project containing many traces
  4. DA dataset of evaluation examples
Show answer

Correct answer: B — A span within a trace

A Run is analogous to a span; a Trace is the collection of runs for one top-level operation. Docs: docs.langchain.com/langsmith (Monitor).

Quiz

Grouping into threads

How are multiple traces grouped into a single multi-turn session?

  1. ABy the evaluator feedback attached to each run
  2. BBy the single shared trace ID that spans the traces
  3. CBy a thread_id metadata key (forming a Thread)
  4. DBy the deployment revision that produced them
Show answer

Correct answer: C — By a thread_id metadata key (forming a Thread)

Traces group into a Thread via the thread_id metadata key. Docs: docs.langchain.com/langsmith (Monitor).

Quiz

What a trajectory is

A Trajectory is:

  1. AA category-based split of an evaluation dataset
  2. BThe control plane's internal routing table
  3. CA single deployment served simultaneously across multiple geographic serving regions
  4. DA flat, ordered list of messages showing the path an agent took start to finish
Show answer

Correct answer: D — A flat, ordered list of messages showing the path an agent took start to finish

A trajectory is a flat, ordered list of messages showing the agent's full path, rendered in the Messages view. Docs: docs.langchain.com/langsmith (Monitor).

Quiz

Manual instrumentation

Which is the correct way to manually instrument a function for tracing?

  1. AThe @traceable decorator (or the trace context manager / RunTree API)
  2. BThe response_format parameter on the agent
  3. CA HumanInTheLoopMiddleware configured to wrap and record the function's calls
  4. DA pairwise evaluator attached to the run
Show answer

Correct answer: A — The @traceable decorator (or the trace context manager / RunTree API)

@traceable, the trace context manager, and the RunTree API are the manual instrumentation tools. Docs: docs.langchain.com/langsmith (Monitor).

Quiz

Online Evals vs Insights

The exam frames Online Evals vs Insights as:

  1. ATwo effectively identical features that merely happen to carry two different product names inside the LangSmith UI
  2. BOnline Evals score individual live runs; Insights surface aggregate trends/patterns across many traces
  3. COnline Evals are offline, while Insights are online
  4. DBoth run only pre-deployment against datasets
Show answer

Correct answer: B — Online Evals score individual live runs; Insights surface aggregate trends/patterns across many traces

Online Evals score individual live runs; Insights find aggregate trends across many traces. The Study Pack marks the exact Insights mechanics [unverified] at source; the tested distinction is the concept stated here. Docs: docs.langchain.com/langsmith (Monitor).

Quiz

Reading a trace

You open one trace to debug why a single operation ran slow. What are you reading?

  1. AA tree of runs: the runs recorded for that one operation
  2. BA dataset of examples with their stored reference outputs
  3. CA single deployment revision and its bundled secrets
  4. DAn experiment that compares two application versions across a dataset
Show answer

Correct answer: A — A tree of runs: the runs recorded for that one operation

A trace is the collection of runs for a single operation, shaped as a tree of runs, and it is the surface you reach for to debug why one operation failed or ran slow. Docs: docs.langchain.com/langsmith/observability-concepts (Monitor).

Quiz

How per-run cost is computed

LangSmith shows token cost per run automatically. What has to be in place for those costs to be calculated?

  1. AA reference output stored on every example in the linked dataset
  2. BA pairwise evaluator attached to each production run
  3. CA thread_id that first groups the runs into one session
  4. DA model pricing map from model names to per-token prices
Show answer

Correct answer: D — A model pricing map from model names to per-token prices

Costs are computed from token counts using the model pricing map, which maps model names to per-token prices; the UI splits spend into input, output, and other. Docs: docs.langchain.com/langsmith/cost-tracking (Monitor).

Quiz

Capturing user sentiment

You want to record an end user's thumbs-up or thumbs-down against the run that produced each answer. In LangSmith that is captured as:

  1. AA revision, versioning the deployment each time a user reacts
  2. BFeedback: a tag and a score bound to the run by its run ID
  3. CA trajectory, the flat ordered list of the run's messages
  4. DAn online evaluator's reference output stored on the example
Show answer

Correct answer: B — Feedback: a tag and a score bound to the run by its run ID

Each feedback entry is a tag and a score, bound to a run by its run ID, and can be continuous or discrete (categorical); end-user sentiment such as thumbs up or down is recorded this way. Docs: docs.langchain.com/langsmith/observability-concepts (Monitor).

Quiz

What alerts can watch

Which set describes metrics LangSmith can raise a threshold alert on?

  1. AOnly the total number of runs, and nothing else
  2. BDataset size and the count of stored reference outputs
  3. CRun count, error rate, latency, feedback score, and cost
  4. DOnly offline experiment scores against a stored gold answer
Show answer

Correct answer: C — Run count, error rate, latency, feedback score, and cost

Alerts fire when a threshold is breached on run count, errors (count or rate), latency, feedback score, or cost, over a 5 or 15 minute window, and deliver to Slack, PagerDuty, or a webhook. Docs: docs.langchain.com/langsmith/alerts (Monitor).

Quiz

Prebuilt vs custom dashboards

What does LangSmith give you for monitoring one project's production performance over time?

  1. AA single global dashboard shared across every project
  2. BCustom dashboards only, which you must build before any chart appears
  3. CNothing until you attach an evaluator to the project's live runs
  4. DA prebuilt dashboard per project, and custom dashboards too
Show answer

Correct answer: D — A prebuilt dashboard per project, and custom dashboards too

Every tracing project gets a prebuilt dashboard covering trace count, error rates, and token usage, and you can also assemble custom dashboards from configurable charts. Docs: docs.langchain.com/langsmith/dashboards (Monitor).

Sign in to track your progress →