Monitoring latency and errors
Latency and errors are the two metrics everybody already thinks they understand, which is why they are worth a lesson. The concepts are familiar. The trap is that a trace is a tree, and tree shaped timing does not behave like the flat timing you are used to reading in an APM.
A parent’s latency contains its children
This is the whole lesson in one sentence, and almost every latency mistake comes from forgetting it.
Sort a runs table by latency descending and you get root runs. Of course you do: the root is the wall clock of everything. It is the least informative number in the trace and it is always at the top of the list.
The skill is walking down the tree until the time stops being explained by something beneath it. At each level, subtract what the children account for. What is left is what that run spent on itself, and the run where the unexplained remainder is large is the run that is slow.
Errors are per run, and the root can lie by telling the truth
Run status is recorded on each run independently. A tool that raised, was caught, and was retried successfully produces an errored tool run inside a successful trace. Nothing is corrupt. Both statements are true records of what happened.
The operational consequence is sharp: if your error monitoring only counts root level failures, a dependency failing half the time behind a retry is completely invisible to you, right up until the retry budget runs out during an incident. When you filter for errors, be deliberate about whether you mean failed traces or runs that errored anywhere.
The three surfaces, and where they lead
On the run, latency and status are facts about one case. In the runs table, they are a column and a filter, which is how you isolate a population. On dashboards, they are aggregates over a window, which is how you notice something you were not looking for.
Try it yourself
Find the step that is actually slow
One trace, five runs, latencies as recorded. The children listed under the root are its only children.
root chain 8.4s
├─ tool retriever 0.3s
├─ llm call-1 0.4s
└─ chain rerank 7.1s
└─ llm score-docs 6.9sShow answer
Correct answer: C — The score-docs llm run, at 6.9s, because it is the deepest run that accounts for the time
A parent's latency includes its children, so the root is always near the top when you sort by latency and is almost never the culprit. Walk down until the time stops being explained by something beneath it: score-docs holds 6.9 of the rerank chain's 7.1, so rerank itself contributes about 0.2 and the leaf is the answer. The rerank option is the good distractor because 7.1 is where the number first looks alarming, and stopping at the first alarming number is exactly the habit this exercise is trying to break.
A successful trace containing errors
A trace's root run shows status success. Inside it, one tool run shows status error. The user received a sensible answer.
Show answer
Correct answer: B — The tool failed, the agent caught it and recovered, and both facts are recorded truthfully
Run status is per run. A caught and handled tool failure is recorded as an errored tool run inside a successful trace, which is the honest record of what happened. The propagation answer is the tempting one because that is how exceptions behave in a single call stack, and the instinct is normally correct in code. It is wrong here, and it matters operationally: if you only ever count root level errors, a tool that fails half the time behind a retry is completely invisible to you.
Where latency and errors live
Both metrics appear on three surfaces with different jobs, and the third one is what turns a chart into a phone call.
Name the three places latency and error information surfaces in LangSmith, and say what each is for.
Reveal answer
One, on the individual run, as its recorded latency and its status, which is what you read in the trace detail view when debugging one case. Two, in the runs table, as a sortable column and a filterable status, which is how you isolate a population such as everything slower than ten seconds in the last hour. Three, on dashboards and as alert metrics, aggregated over a window, which is how you notice a regression you were not looking for. Latency and Errors are two of the five metrics an alert rule can watch, so this is the module that connects directly to module 4.
Children that add up to more than their parent
A parent chain records 3.0s. Its four children record 2.6s, 2.4s, 2.5s and 2.7s, summing to 10.2s.
Show answer
Correct answer: C — The children ran in parallel, so the parent's wall clock is roughly the slowest child rather than the sum
Parallel execution is the ordinary explanation and the one that should come first: wall clock for a fan out is bounded below by the slowest branch, not by the sum. The timeout answer is the interesting distractor because it would produce a similar looking record and it is a real failure mode in other systems, but a timed out parent would not be recorded as a clean success. Noticing that children can exceed their parent is also the fastest way to tell a parallel step from a sequential one without reading any code.
Attribute your own worst latency
Pick the slowest trace in any project you have and do this before reading any code.
You can name the single run responsible for most of the wall clock, say whether its siblings ran in parallel or in sequence, and state how much time the parent itself accounts for once its children are subtracted.