Whetstone.
CapstoneCapstone quiz: the whole chain under exam conditions
Capstone, Quiz14 min

Capstone quiz: the whole chain under exam conditions

Ten of the forty LCAE questions come from this domain, and every one of them is somewhere on a single chain. This quiz walks that chain, and the items are deliberately weighted toward the joins between modules rather than the interiors, because the joins are where the confusable pairs live.

The confusable pairs, collected

These are the ones that cost marks, and they cost marks because each member of the pair is individually easy.

25,000 against 1,000. Runs inside a trace, traces inside an Insights job. Different features, different units, both round.

Insights against Online Evals. Discovery against measurement, report against feedback, scheduled against continuous, capped against sampled, plan gated against broadly available.

thread_id against session_id. Both accepted, thread_id wins, and the instinct that says the legacy key should take priority is the wrong instinct here.

ls_provider and ls_model_name against provider and model. The prefix is the whole question.

Email against the four real channels. The most catchable trap in the domain precisely because it is the least clever one.

What the course would not tell you

Three things in this material are compositions rather than products: user sentiment, security monitoring, and online A/B testing. Each one is assembled from feedback, evaluators, filters, metadata and alert rules.

That is not a gap in LangSmith. It is a design: the platform ships primitives and one open door, Feedback Score, through which any judgement you can express becomes something a threshold rule can watch. Recognising a composition question is worth more than any individual fact, because the wrong answers to those questions all take the same form, a tidy dedicated feature that does not exist.

Practice

Try it yourself

Quiz

Two numbers that live next to each other

Under time pressure these two limits get swapped, and a question will put them in the same options list on purpose. Pick the pairing that is correct.

  1. A1,000 runs per trace, and 25,000 traces per Insights job
  2. B25,000 runs per trace, and 25,000 traces per Insights job
  3. C1,000 runs per trace, and 1,000 traces per Insights job
  4. D25,000 runs per trace, and 1,000 traces per Insights job
Show answer

Correct answer: D — 25,000 runs per trace, and 1,000 traces per Insights job

25,000 runs inside one trace. 1,000 traces inside one Insights job. They belong to different features and different units, which is exactly why they blur: both are round limits attached to the word trace. The first option is the mirror image, and it is the one you produce if you remember that there are two numbers, remember both numbers correctly, and attach them to the wrong things, which is the most common way to be wrong here.

Quiz

Spot the invented feature

Three of these are real LangSmith capabilities described in this course. One does not exist.

  1. AA hierarchical category report produced on a UTC schedule
  2. BA dedicated user sentiment dashboard populated from run metadata
  3. CFeedback attached to a thread rather than to an individual run
  4. DA threshold alert rule on an aggregated Feedback Score metric
Show answer

Correct answer: B — A dedicated user sentiment dashboard populated from run metadata

There is no first class sentiment feature and no metadata field that lights one up. Sentiment is composed from explicit feedback, an LLM judge writing feedback, and Insights clustering. The other three are real: the first is Insights, the third is thread level feedback granularity, and the fourth is the Feedback Score alert metric. The general pattern is worth more than this item: when an option offers a tidy dedicated product for a messy human concept, check that it exists.

Recall

Trace one signal from run to phone

The synthesis card for the entire course. If you can walk this chain without hesitating, the domain is done.

Follow a single user's frustration from the moment it happens to the moment somebody is paged. Name every object and mechanism it passes through.

Reveal answer

The turn is recorded as a trace, a tree of runs built from parent_run_id and dotted_order. The turn joins its siblings as a thread, because a thread_id was attached in metadata. An online evaluator, running continuously over a filter at some sampling rate, judges the thread against a reference free criterion and writes its verdict as feedback on the thread. That feedback is aggregated as the Feedback Score metric. An alert rule, the only type being a threshold, compares that aggregate over a 5 or 15 minute window and fires to one of Slack, PagerDuty, Dynatrace or Webhook. If the destination is email, it goes out through Webhook, because email is not a native channel.

Quiz

Four scenarios, one tool each

Which pairing is wrong?

  1. ANo hypothesis about why users are unhappy, so run an Insights job over last week
  2. BA specific groundedness hypothesis to test, so run an online evaluator and compare feedback scores
  3. CCost is blank on some runs, so walk the three pricing conditions in order
  4. DUsers complain the assistant forgets context, so open individual traces and check each one for errors
Show answer

Correct answer: D — Users complain the assistant forgets context, so open individual traces and check each one for errors

A cross turn memory defect is invisible per trace by construction: every turn answered correctly with the context it was given, so nothing errored and nothing looks wrong. The threads view is the only surface where the space between turns exists as a visible object. The other three pairings are correct. This item is here because the wrong option describes what people actually do, and it is the most expensive habit in the domain.

Quiz

Ninety seconds, docs open, and a half remembered number

One question left. It turns on the aggregation window options for an alert rule, and you also half remember a channel count from something you read once that does not match what you revised. The exam is semi open book: docs.langchain.com and smith.langchain.com are available, general web search and AI assistance are not.

  1. AThe LangSmith alerting docs page, and answer with its figures rather than the half remembered ones
  2. BThe alerting launch blog post, since it is the primary announcement of the feature
  3. CThe LangSmith SDK client reference, since alert rules are configured programmatically
  4. DA general web search for the current alert metric and notification channel lists
Show answer

Correct answer: A — The LangSmith alerting docs page, and answer with its figures rather than the half remembered ones

The docs page carries the windows, the five metrics and the four channels in one place, and the docs are the source to trust: the launch blog lists three metrics and two channels and is almost certainly the older artefact. The blog option is the trap this whole course has been priming you for, because a primary announcement genuinely does sound authoritative. General web search is not available to you, which is the deeper point: this exam rewards knowing which documentation section owns a topic far more than it rewards memorising a signature, because the lookup-able facts are the ones you can look up.

Do

Capstone dry run: answer the five questions from memory

This is the bar the capstone project sets, compressed into something you can do on a train. Write the five answers out before you open anything.

  • Pick one real project. Write down, without looking, which runs belong to which conversation and what mechanism makes that true.
  • Write down what it cost last night, and name which of the three cost conditions you are least confident is satisfied.
  • Write down what your most common failure category is, and say whether that answer came from data or from your imagination.
  • Write down whether your quality metric moved this week, and name the evaluator and criterion that produce it. If there is no evaluator, write that instead; it is the honest answer and it is the gap.
  • Write down who gets woken when it breaks, through which channel, on which metric, at what threshold, over which window.
  • Now open the org and check all five. Mark each one right, wrong, or unanswerable. The unanswerable ones are the real output of this exercise.
Done whenI have five written answers, five verdicts against reality, and a list of the questions my current setup cannot answer at all.
Sign in to track your progress →