Whetstone.
Monitor agentsMonitor agents
Focus area: Monitor20 min

Monitor agents

The Monitor domain is LangSmith observability. This second set covers the trace ID and the 25,000-run cap, projects, auto-instrumentation via integrations, tags versus metadata, what a feedback score can be, alert channels and windows, the cost breakdown, Insights for aggregate patterns, what counts as one run, and the prebuilt dashboard. Ten questions.

Practice

Try it yourself

Quiz

What binds runs into a trace

What binds individual runs together into a single trace, and is there a size limit?

  1. AA shared thread_id metadata key, and a trace may hold unlimited runs
  2. BThe feedback score attached to each run, with a cap of 1,000 runs per trace
  3. CA unique trace ID, and each trace is limited to a maximum of 25,000 runs
  4. DThe assistant id that produced them, with no documented limit on run count
Show answer

Correct answer: C — A unique trace ID, and each trace is limited to a maximum of 25,000 runs

Runs are bound to a trace by a unique trace ID, and each trace is limited to a maximum of 25,000 runs. Docs: docs.langchain.com/langsmith/observability-concepts (Monitor).

Quiz

What a project is

What is a tracing Project in LangSmith?

  1. AA container for all the traces related to a single application or service
  2. BA sequence of traces that together make up one whole multi-turn conversation session
  3. CA collection of examples with reference outputs used to score the app
  4. DA single deployed graph together with its configured prompt and tools
Show answer

Correct answer: A — A container for all the traces related to a single application or service

A project is a container for all the traces related to a single application or service. Docs: docs.langchain.com/langsmith/observability-concepts (Monitor).

Quiz

Tracing without decorators

You build on LangChain, LangGraph, or the OpenAI SDK and want tracing without decorating each function. What gives you that?

  1. AYou must still wrap every single traced function with the @traceable decorator by hand
  2. BYou configure an online evaluator, which begins recording the traces for you
  3. CYou pass a thread_id on each call, which turns on tracing for the project
  4. DAn integration (auto-instrumentation) captures inputs, outputs, and metadata for you
Show answer

Correct answer: D — An integration (auto-instrumentation) captures inputs, outputs, and metadata for you

An integration is the equivalent of auto-instrumentation: with a supported framework it captures inputs, outputs, and metadata with no manual code changes. Docs: docs.langchain.com/langsmith/observability-concepts (Monitor).

Quiz

Filtering and grouping runs

You want to categorize, filter, and group runs in the LangSmith UI. Which two run attributes are meant for that?

  1. AReference outputs and dataset splits attached to each run when it is recorded
  2. BTags (strings) and metadata (key-value pairs), both usable to filter and group
  3. CA checkpointer and a store, which the UI reads to organize the run list for you
  4. DThe trace ID alone, which is the only value the UI can filter or group runs by
Show answer

Correct answer: B — Tags (strings) and metadata (key-value pairs), both usable to filter and group

Tags are strings and metadata is key-value pairs attached to runs; both let you filter and group runs in the UI. Docs: docs.langchain.com/langsmith/observability-concepts (Monitor).

Quiz

What a feedback score can be

A feedback score attached to a run in LangSmith can be:

  1. AOnly a boolean pass or fail on the run, with nothing more granular than that allowed
  2. BOnly a number between 0 and 1, since scores are always normalized that way
  3. CContinuous or discrete (categorical), and its tag can be reused across runs
  4. DOnly a value set by an online evaluator, never one a human enters by hand
Show answer

Correct answer: C — Continuous or discrete (categorical), and its tag can be reused across runs

Feedback can be continuous or discrete (categorical), and tags can be reused across runs within an organization. Docs: docs.langchain.com/langsmith/observability-concepts (Monitor).

Quiz

Where an alert delivers

When a LangSmith alert threshold is breached, where can it deliver and over what window?

  1. ATo Slack, PagerDuty, or a webhook, evaluated over a 5- or 15-minute window
  2. BTo email only, and only once per day when the daily batch job runs
  3. CTo a dataset, which stores the breaching runs for later offline evaluation
  4. DTo a dashboard chart only, where the breach is highlighted but not sent out
Show answer

Correct answer: A — To Slack, PagerDuty, or a webhook, evaluated over a 5- or 15-minute window

Alerts fire on a threshold breach over a 5- or 15-minute window and deliver to Slack, PagerDuty, or a webhook. Docs: docs.langchain.com/langsmith/alerts (Monitor).

Quiz

How cost is broken down

LangSmith computes token cost per run. How does the UI break that spend down?

  1. AInto prompt and completion tokens only, with no other categories shown
  2. BInto input, output, and other, once a model pricing map is configured
  3. CInto a single total figure per project, not split per run at all
  4. DInto per-tool and per-thread subtotals derived from the run''s metadata
Show answer

Correct answer: B — Into input, output, and other, once a model pricing map is configured

Costs are computed from token counts via the model pricing map, and the UI splits spend into input, output, and other. Docs: docs.langchain.com/langsmith/cost-tracking (Monitor).

Quiz

Finding patterns across traces

You want to discover recurring failure patterns across thousands of production traces, without scoring each run one by one. Which LangSmith feature fits?

  1. AOnline Evals, which attach a reference-free evaluator to score each live run
  2. BAn alert, which fires the moment one run breaches a metric threshold you set
  3. CA dataset split, which groups the failing traces into a named subset for review
  4. DInsights, which surfaces aggregate trends and patterns across many traces
Show answer

Correct answer: D — Insights, which surfaces aggregate trends and patterns across many traces

Insights surfaces aggregate trends and patterns across many traces; Online Evals score individual live runs, the opposite grain. Docs: docs.langchain.com/langsmith/observability-concepts (Monitor).

Quiz

What counts as one run

Within a trace, each run records one unit of work. Which of these is such a run?

  1. AA single LLM call, a tool invocation, or a retrieval step
  2. BA whole multi-turn session between the user and the agent
  3. CThe deployed application as a whole, across all of its users
  4. DA dataset of examples paired with their reference outputs
Show answer

Correct answer: A — A single LLM call, a tool invocation, or a retrieval step

A run is a single unit of work such as an LLM call, a tool invocation, or a retrieval; a trace is the collection of runs for one operation. Docs: docs.langchain.com/langsmith/observability-concepts (Monitor).

Quiz

What the prebuilt dashboard shows

Every tracing project gets a prebuilt dashboard. What does it show out of the box?

  1. AThe reference-output coverage of the linked evaluation dataset
  2. BThe version history of each assistant on the deployment
  3. CTrace count, error rates, and token usage for the project
  4. DThe dataset split sizes and the count of stored examples
Show answer

Correct answer: C — Trace count, error rates, and token usage for the project

Every tracing project gets a prebuilt dashboard covering trace count, error rates, and token usage, and you can also build custom dashboards. Docs: docs.langchain.com/langsmith/dashboards (Monitor).

Sign in to track your progress →