Creating datasets, and the six routes in
A dataset is a list of examples. An example is inputs, optionally a reference output, and metadata. The data model takes ninety seconds, so the interesting questions are all about how the pieces get used and how examples get in.
The three parts
Inputs are what your application receives. They should look exactly like real production input, because that is the point.
Reference outputs are what you decided good looks like. Somebody wrote them down: you, a domain expert, a human labeller in an annotation queue, or a previous version of the system whose output you blessed. This is the part that exists offline and cannot exist online.
Metadata is everything else. Splits, tags, provenance, which incident this example commemorates. Useful later for filtering an experiment down to a slice.
The sentence worth learning
Straight from the docs, and it settles an intuition people get wrong:
Reference outputs do not get passed to your application, they are only used in evaluators.
During an experiment your app is fed the inputs and nothing else. It runs exactly as it would in production, in the dark, with no idea it is being examined. Then the evaluator gets handed the app’s output and the reference output, and compares them.
If you ever design something where the app needs the reference output, you have written an evaluator and put it in the wrong file.
The six UI routes
Bulk multi-select from the Runs table. Filter the table down to something interesting, tick several runs, add them at once. The workhorse for harvesting a batch of failures after an incident.
A single run, via Add to and then Dataset. You are looking at one trace, you notice it is wrong or interestingly right, you add it there and then.
Automation rules. A rule that fires on runs matching a filter and adds them without you present. The documented example triggers on low feedback scores, which is genuinely lovely machinery: users tell you something was bad and your regression suite grows itself overnight.
Annotation queue. A reviewer working a queue can send a run to a dataset as part of labelling it. There is a D hotkey for it, which is exactly the kind of specific detail that makes a clean exam question. Remember D for dataset.
The Playground. You are iterating on a prompt, you get an output you like, you save it as an example. This is how a reference output gets authored by hand without anyone typing JSON.
The Examples tab, plus Example button. Inside the dataset, the manual route. Type the inputs, type the reference output. What you use for the first ten examples before any traffic exists.
Group them so they stick: two from the Runs table, two from a human-in-the-loop surface, one that runs without you, one from inside the dataset itself.
The three SDK routes
create_examples for programmatic creation. upload_csv for a file, which is how a spreadsheet a domain expert filled in gets in without retyping. upload_dataframe for a pandas DataFrame, and this one is Python-only, which is not favouritism, it is that TypeScript has no DataFrame to hand you.
The trap in the question shape
A which-of-these-is-NOT question does not usually invent something absurd. It takes a real feature and attaches it to a plausible but fictional UI path, or takes a real path and puts it in the wrong menu. The distractor is almost always made of true parts assembled wrongly. Check the path, not just the capability.
Try it yourself
Where reference outputs never go
One sentence from the docs worth learning close to verbatim, because it settles a question people get wrong on intuition.
During an experiment, does your application receive the reference output for each example, and what is it used for instead?
Reveal answer
No. Reference outputs do not get passed to your application, they are only used in evaluators. Your app sees the example's inputs and nothing else, exactly as it would in production. The reference output is handed to the evaluator afterwards so it can compare what your app produced against what you said good looked like. This has to be true or the whole thing is circular: an application that can see the right answer is not being tested, it is being given the answers.
Which of these is NOT a way to add an example
The classic shape for this topic. One of these four is not a documented route for getting an example into a dataset.
Show answer
Correct answer: D — Exporting a trace to CSV and re-importing it from the dataset settings page
There is a CSV route, but it is the SDK's upload_csv, not an export-then-reimport flow hanging off dataset settings. Option 3 is built to feel plausible by combining a real capability, CSV, with an invented UI path, which is exactly how this question type is usually constructed. The other three are documented routes. When you meet a which-is-not question here, check whether the odd one out is a real feature described in a place it does not live.
All six UI routes, cold
The whole set. Take a breath and list them before checking, because this is the most list-shaped topic in the section.
Name the six documented UI routes for adding examples to a LangSmith dataset.
Reveal answer
One, bulk multi-select from the Runs table. Two, a single run using Add to and then Dataset. Three, automation rules, which add matching runs automatically, for example any run with a low feedback score. Four, an annotation queue, where a reviewer sends the run to a dataset while labelling it, with D as the hotkey. Five, the Playground, saving a run you just produced there. Six, the Examples tab inside the dataset itself, with the plus Example button, which is the type-it-in-by-hand route. The grouping that makes them stick: two from the Runs table, two from a human review surface, one automatic, one manual from inside the dataset.
The SDK methods, and the language trap
Which statement about the SDK routes for creating examples is correct?
Show answer
Correct answer: D — upload_dataframe is Python-only
upload_dataframe is Python-only, which makes sense the moment you say it aloud, since a DataFrame is a pandas object and there is no equivalent in the TypeScript ecosystem. Option 0 is the answer you give if you assume SDK parity, an assumption this domain punishes repeatedly. Options 1 and 2 pick the wrong method: CSV upload and create_examples are not the Python-only ones.
An example with no reference output
Somebody adds fifty production runs to a dataset. None of them have reference outputs, because production runs never do. What is true of that dataset?
Show answer
Correct answer: B — It is usable, but only by evaluators that score intrinsic properties of the output
Reference outputs are optional on an example, so the dataset is fine, and any evaluator scoring an intrinsic property works: valid JSON, under the token budget, correct tool called, no email address leaked. What cannot work is a comparison against ground truth, because there is none. Option 2 is the tempting overcorrection: optional on the example is not the same as unnecessary to the evaluator, and a correctness evaluator handed no reference has nothing to compare against. Option 3 describes a good idea rather than a requirement.
Rebuild the six routes in a scratch file
Small, offline, no account needed. A retrieval exercise rather than a build.
- Open a scratch file and write the six UI routes from memory, no peeking.
- Beside each one write the situation where it is the right choice, in under ten words.
- Add the three SDK methods underneath and mark which one is Python-only.
- Compare against the lesson and note only what you missed.