Analyzing your agent
You have traces arriving. Now you have to use them, which is a different skill from producing them, and it is mostly the skill of narrowing down before reading.
This lesson describes the interface structurally. Button labels and column names shift between releases; the shape does not, and a LangSmith org is provisioned for you at exam time, so learn it as a place you move around in.
Filter first, always
The Runs view is a table of runs and traces with filters over it. Time window, run name, status, tags, metadata fields, latency, token count, feedback score.
The instinct is to sort by something and start reading. The better move is to narrow to a slice that matches the complaint, then read a handful in full. Ten traces read properly beat four hundred skimmed, and the filters are how you get from four hundred to ten.
This is where the metadata from the last lesson pays for itself. A complaint that says slow since Tuesday for enterprise customers is one filter if you recorded a tenant field and an afternoon of guessing if you did not.
Reading a single trace
Once you are inside one, there are four questions and they have an order.
Where did the time go. Durations are per run and they roll up, so a slow root is a question about which child it waited on. This is usually the fastest route to a cause because it points at a specific step before you have understood anything.
What did the tools return. Not what they were called with, though that matters: what came back. A confidently wrong answer is very often a retrieval step that returned nothing and a model that improvised over the silence, and the final output gives no hint of it.
What did the model actually see. Not what you believe the template produces. Truncation, a missing variable, a system prompt that never got substituted: all of it is visible here and nowhere else.
What failed and got retried. A step that errored, retried, and succeeded is invisible in the output and extremely visible in your latency and your bill.
Feedback lives on runs
Scores attach to runs, and they come from three places: evaluators, humans labelling in an annotation queue, and end users sending signals from your own application.
They land in the same place, which is the design decision that makes everything in Module 3 possible. One automation rule can react to a thumbs-down from a user or a low score from a judge, because by the time the rule sees it, both are just feedback on a run.
Nothing derives a quality score from latency or token count. Those are recorded, not judged.
Finish the investigation properly
You have found the bad trace and understood it. The move that makes the work pay off more than once is to add the run to a dataset before you close the ticket.
Do it now, while the incident is annoying and the judgement about what good would have looked like is still in your head. It is one action from the run view. Later, that example is a permanent regression test, and the next person who refactors that prompt gets a red score instead of an outage.
That is the harvest edge of the lifecycle loop, performed by hand. Module 3 is the same move performed while you sleep.
Try it yourself
What you interrogate a trace for
A LangSmith org is provisioned per candidate and some exam items are answered by doing, so this is worth holding as a procedure rather than a concept.
You open a single slow, wrong trace. What four things do you check, and in what order?
Reveal answer
First, where the time went: read the per-run durations and find which child the root was waiting on, because a slow parent is almost always a question about one child. Second, the tool calls: which tools were invoked, with what arguments, and what came back, since a plausible-looking wrong answer is very often a tool that returned nothing and a model that improvised over the gap. Third, what the model actually saw, rather than what you believe the prompt template produces, because truncation and missing context show up here and nowhere else. Fourth, status and retries, since a step that failed and was retried silently is a latency and cost story that never reaches the final output.
Where analysis actually starts
You have ten thousand traces and a report that the agent is sometimes wrong. What is the first move in the Runs view?
Show answer
Correct answer: A — Filter to a slice defined by the complaint, then read a handful in full
Analysis is filtering before it is reading. The Runs view exists to be narrowed by time window, status, name, tags, metadata and feedback score, and the whole point of instrumenting metadata was to make that narrowing possible. Option 1 is genuinely tempting because sorting by duration is a real and useful move, but it answers a latency question rather than a correctness one, and it is how people end up investigating a slow trace that was perfectly correct. Reading a handful in full beats skimming hundreds.
Where a feedback score can come from
Feedback attached to a run can originate from several places. Which of these is NOT a source of feedback in LangSmith?
Show answer
Correct answer: D — The tracing SDK inferring a score from run latency
Nothing infers a quality score from latency. Feedback arrives from evaluators, from humans in annotation queues, and from end users through your own application, and all three land in the same place, which is why one automation rule can react to any of them. Option 3 is built to be plausible because latency genuinely is recorded on every run, so it feels like the sort of thing the platform might helpfully derive; the point of the question is that recorded and scored are different acts.
What analysis is ultimately for
The connection this module has been building towards, and the reason this lesson sits before the evaluation module rather than after it.
You have found and understood a bad trace. What is the move that makes the investigation pay off more than once, and why does it belong in the same sitting?
Reveal answer
Add the run to a dataset before you close the ticket. That converts a one-off investigation into a permanent offline test, so the next person to touch that prompt gets a red score rather than a repeat incident. It belongs in the same sitting because you will not come back: the run is already open, the add is one action from the run view, and the judgement about what good would have looked like is in your head right now and will not be next week. This is the harvest edge of the lifecycle loop, done by hand; Module 3's automations are the same move done unattended.
Read one trace all the way down
Pick any traced request you have, ideally one that was slightly wrong rather than dramatically broken.
You can state where the time went, what each tool returned, and what the model was actually shown, and at least one of those three surprised you. If nothing surprised you, the trace was probably too simple to be worth the exercise; a two-step happy path teaches nothing, and the traces worth reading are the ones with a retry or an empty tool result in them.