The vocabulary, and finding things fast
Every later question assumes this vocabulary. It takes ten minutes and it prevents the specific failure of reading a question three times because two words in it mean nearly the same thing to you.
The observation nouns
A run is one unit of work: a single model call, a single tool call, a single chain step. It has inputs, outputs, a start and end time, a status, and usually token counts.
A trace is the whole tree of runs produced by one top-level request. One user question, one trace, many runs nested inside it.
A thread groups traces that belong to the same multi-turn conversation, so you can read a whole exchange rather than a single turn.
A tracing project is the container all of that lands in. It is the online-side counterpart to a dataset, and remembering that pairing saves you on attachment questions later.
The evaluation nouns
A dataset is a list of examples. An example is inputs, optionally a reference output, and metadata.
An experiment is one run of your application over one dataset, scored by one or more evaluators, recorded so you can compare it against the next one.
Feedback is a score attached to a run, whether it came from an evaluator, a human in an annotation queue, or an end user clicking thumbs down. That last source matters more than it sounds: user feedback is what automation rules most often trigger on.
What the docs are for during the exam
You get docs.langchain.com and smith.langchain.com. No general search, no AI. That has two consequences and they pull in opposite directions.
Memorising signatures is nearly worthless. Anything you could look up in ninety seconds is a bad exam question, so the paper will not lean on it.
Navigating fast is genuinely scoreable. The questions that survive are ones where a permitted tab confirms an answer you already half know. Knowing that the evaluation reference has a language toggle, and reaching for it rather than translating a Python name into camelCase, is worth real marks in this domain specifically, because Python and TypeScript are not at parity here.
The navigation habit worth building
Before the exam, spend twenty minutes clicking, not reading. Find the evaluation reference. Find the page describing online evaluators. Find where dataset creation methods are listed. Find the language toggle and flip it.
You are building a rough index of where things live, not learning content. Under time pressure the difference between a candidate who knows a page exists and one who is searching from scratch is about ninety seconds per question, and ten questions in this section makes that fifteen minutes.
Try it yourself
Run, trace, thread
Three words that get used loosely in conversation and precisely in the product. Which set of definitions is right?
Show answer
Correct answer: C — A run is one unit of work, a trace is the tree of runs for one request, a thread is a conversation
A run is a single unit of work, one model call or one tool call or one chain step. A trace is the whole tree of runs produced by one top-level request. A thread groups traces belonging to the same multi-turn conversation. Option 0 is the good distractor because it inverts run and trace, which is the mistake almost everyone makes out loud, and it lands because in casual speech people say run when they mean the whole request. Option 3 is worth rejecting deliberately: runs and traces being distinct is what makes the trace tree readable at all.
Two minutes, docs open, where do you look
The exam is semi-open-book with docs.langchain.com and smith.langchain.com permitted. A question asks whether a particular evaluation parameter exists in the TypeScript SDK. What is the fastest reliable move?
Show answer
Correct answer: B — Open the evaluation reference page and switch the language toggle to TypeScript
The docs carry per-language toggles, and this domain has real asymmetries where a Python capability has no TypeScript twin at all. Going to the reference page and flipping the toggle answers existence questions directly. Option 2 is the trap this entire course keeps warning about: assuming parity and camelCasing the name yields a confident answer that is sometimes simply wrong. Option 0 is not unreasonable but search results frequently land you on the Python page, which is how people convince themselves a Python-only function exists in TypeScript.
Reference output, precisely
This term carries more weight than any other in the glossary, because half the distinctions in the course reduce to whether one exists.
What is a reference output, which part of the system is allowed to see it, and in which of the two evaluation worlds does it not exist at all?
Reveal answer
A reference output is the answer you decided was correct for a given example, written down in advance by you, a domain expert, or a human labeller. It is handed to evaluators only: during an experiment your application receives the example's inputs and nothing else, so it runs in the dark exactly as it would in production. It does not exist online, because a production user asked a question and left without supplying the correct answer. Nearly every offline-versus-online distinction in this course, including a function signature you will be asked about, falls out of that absence.
Which of these is not a thing
Which of the following is NOT a first-class object you will meet in LangSmith?
Show answer
Correct answer: C — Evaluation gate
Evaluation gate is invented. It sounds right because gating a release on evaluation results is a real practice, and this question type usually works exactly that way: take a genuine concept and attach it to a product noun that does not exist. Annotation queues, tracing projects and experiments are all real objects with their own views. When you meet a which-is-not question in this domain, check whether the odd one out is a real idea wearing a fake product name.
Draw the object map from memory
Small, offline, no account needed. This is the map you want in your head before the clock starts, and drawing it is faster than rereading it.
- On one side write the observation objects, run, trace, thread, tracing project, and draw the containment arrows between them.
- On the other side write the evaluation objects, dataset, example, reference output, experiment, evaluator, feedback score.
- Draw the two arrows that connect the sides, one for harvesting a run into a dataset and one for attaching an evaluator to a tracing project.
- Mark which side of the map has reference outputs on it and which does not.