Whetstone.
ObservationTracing with LangSmith
Module 1, Lesson 216 min

Tracing with LangSmith

Tracing is the mechanism that turns an application into something you can look at. It is also where the exam’s favourite piece of trivia lives, the two environment-variable prefixes, so that gets stated outright below rather than left as a thing to look up.

What actually gets recorded

One top-level request produces one trace, which is a tree of runs. Each run carries its own inputs, outputs, start and end time, status, and where relevant token counts and cost.

The tree structure is the point. A flat log tells you an error happened. A tree tells you the error happened in the third tool call, which was invoked by the planning step, which had already been retried once. Latency rolls up, so a nine-second root run is a question about which child it spent those seconds waiting on.

Above traces sit threads, grouping successive turns of one conversation so you can read the exchange rather than a single turn out of context.

How instrumentation happens

Two sources, and most applications have both.

Framework components are instrumented already. LangChain and LangGraph pieces emit runs without you writing anything, which is why a graph often produces a rich trace on the first try.

Your own functions need decorating. In Python that is the @traceable decorator; in TypeScript it is the traceable() wrapper. Whatever runs inside a traced function nests underneath it.

Python
from langsmith import traceable

@traceable
def summarise(ticket: str) -> str:
    ...
TypeScript
import { traceable } from "langsmith/traceable";

const summarise = traceable(async (ticket: string) => {
  // ...
});

Undecorated helpers are invisible, and their time is silently attributed to whatever called them. That is the usual reason a trace looks less useful than you hoped: not that tracing failed, but that the interesting step was never a run.

The environment variables

Tracing is switched on by environment variables. The current set is prefixed LANGSMITH_, and these four are the ones worth knowing without looking:

Variable What it does
LANGSMITH_TRACING Set to true to switch tracing on
LANGSMITH_API_KEY Authenticates to LangSmith
LANGSMITH_PROJECT Which tracing project runs land in; defaults to default
LANGSMITH_ENDPOINT The API URL, which you only set for self-hosted or EU deployments

Beyond these four the names get rarer and more specific, and those genuinely are a lookup rather than a memory task. Practise finding the page for the long tail; do not spend revision time on it.

Metadata is the part people skip

Runs can carry metadata and tags: a version string, a tenant, an environment, a feature flag, whatever you might want to slice by later.

This feels like premature bookkeeping and it is not, because every selection mechanism downstream is a filter. Online evaluators are configured with a filter plus a sampling rate. Automation rules fire on a filter. Both of those read fields that instrumentation put there.

The asymmetry is what makes it worth doing early: adding metadata today helps traffic from today onwards, and the traces you most want to slice are always the ones you have already collected. When somebody asks in six weeks whether the regression is confined to one customer, the answer is either one filter or a fortnight of guessing.

What to walk in knowing

Runs nest into traces, traces group into threads, all of it lands in a tracing project. Framework components self-instrument; your own functions need @traceable or traceable(). Configuration is by environment variable: LANGSMITH_TRACING, LANGSMITH_API_KEY, LANGSMITH_PROJECT and LANGSMITH_ENDPOINT, with the older LANGCHAIN_ spellings still honoured as a fallback. Metadata and tags are what make filters possible, and filters are what Modules 2 and 3 are built on.

Practice

Try it yourself

Quiz

Reading a traced function

A colleague has instrumented one helper in an otherwise untraced service.

What will the trace tree look like?
from langsmith import traceable

@traceable
def summarise(ticket: str) -> str:
    context = fetch_history(ticket)      # not decorated
    return model.invoke(prompt(context)) # LangChain model
  1. ANothing appears at all, because the service entrypoint above summarise is not decorated
  2. Bsummarise appears as a run, the model call nests under it, fetch_history does not appear
  3. CAll three appear as sibling runs at the top level of the trace, with no nesting
  4. DOnly the model call appears, because the LangChain client emits runs and decorators do not
Show answer

Correct answer: B — summarise appears as a run, the model call nests under it, fetch_history does not appear

The decorated function becomes a run, and work performed inside it nests underneath. LangChain and LangGraph components are instrumented already, so the model call shows up without you doing anything; a plain undecorated Python function does not, so fetch_history is invisible and its latency is silently attributed to its parent. Option 3 is the tempting one for anyone who thinks of tracing as something the model client does, and it is exactly backwards: the decorator is what gives you the tree structure to hang the model call on.

Recall

Why you add metadata you do not yet need

This is the piece of tracing setup people skip, and the cost of skipping it does not appear until two modules later.

You are instrumenting a service. Why is it worth attaching metadata and tags to runs before you have any use for them, and which two later features consume them?

Reveal answer

Because you cannot filter on a field you never recorded, and every later selection mechanism is a filter. Online evaluators are configured with a filter plus a sampling rate, so an evaluator that should only score checkout-flow traffic needs something on the run that says checkout. Automation rules are also filter-driven, so a rule that harvests failures from one customer segment needs that segment recorded. Add version, feature, tenant or environment at instrumentation time; adding them later only helps traffic from that point onwards, and the traces you most want to slice are always the ones you have already collected.

Quiz

Two prefixes, and which one wins

A service you inherit sets LANGCHAIN_API_KEY. A service you are writing sets LANGSMITH_API_KEY. What is the relationship between the two prefixes?

  1. AOnly the LANGCHAIN names are read; the LANGSMITH names were proposed and never shipped in the SDK
  2. BBoth are read; LANGSMITH is the current prefix and takes precedence, with LANGCHAIN as a fallback
  3. CNeither is read, because tracing credentials are passed to the Client constructor in code
  4. DOnly one prefix works per SDK version, so the two services need different SDK versions
Show answer

Correct answer: B — Both are read; LANGSMITH is the current prefix and takes precedence, with LANGCHAIN as a fallback

Both prefixes are honoured. The SDK looks each variable up under LANGSMITH first and falls back to LANGCHAIN, so the current name wins where both are set and the inherited service keeps tracing correctly. The distinction the exam cares about is superseded against removed: the older names are superseded, and an answer that calls them dead is as wrong as one that calls them current. Option 3 is the plausible miss, because incompatible-by-version genuinely is how a lot of renames work, and it is not how this one works.

Recall

What the tree is made of

A structural card. Getting the containment right is what makes a trace readable at a glance.

Describe the structure of a trace, from the top-level request down, and say where a multi-turn conversation fits.

Reveal answer

One top-level request produces one trace, which is a tree of runs. The root run is the entrypoint, and every nested unit of work, each chain step, tool call and model call, is a child run underneath it, carrying its own inputs, outputs, timing, status and token counts. Threads sit above traces: a thread groups the traces from successive turns of the same conversation, so you can read the exchange rather than one turn in isolation. The practical consequence is that latency and cost roll up, so a slow parent is usually a question about which child it is waiting on.

Check

Trace something tiny

Offline-friendly if the wifi is bad, since the point is the shape rather than the upload.

You should see

You can write down, for a two-step script, exactly which runs you expect to appear and how they nest, before running anything. If your prediction and the actual tree disagree, the interesting case is almost always an undecorated helper whose time got absorbed into its parent, which is the most common reason a trace looks less informative than expected.

Sign in to track your progress →