Monitoring user sentiment, which is a composition and not a product
Here is the honest version, and it is more useful than the tidy version: LangSmith has no dedicated user sentiment feature.
There is no sentiment toggle, no sentiment tab, no field that lights up a dashboard. The official course has a lesson under this exact name, so the topic is genuinely examinable. It is examinable as a composition, and being able to describe it as one is the thing that demonstrates you understand the primitives rather than the marketing.
The three ingredients
Explicit feedback. Your app collects a thumb, a star, a “was this helpful”, and writes it as feedback under a key you choose. This is the highest quality signal you will ever have, because a human actually said it, and it is desperately sparse, because approximately nobody clicks the thumb.
An LLM judge. An online evaluator scores runs or threads against a criterion you wrote, something like “does this exchange show the user becoming frustrated”, and writes its verdict back as feedback exactly like a human rating would. Coverage is whatever your sampling rate says. Quality is whatever your criterion deserves.
Insights clustering. Insights groups traffic into categories without being told what to look for, so it can surface a frustration pattern nobody thought to write an evaluator for. Scheduled, capped and plan gated, so treat it as a periodic study rather than a gauge.
Score threads, not turns
Frustration is a sequence property. The reply that finally broke the user was probably a perfectly reasonable reply; the problem was that it was the fourth one.
So point your sentiment evaluator at threads where you can. Feedback attaches at run or thread granularity, and this is the case that choice was made for. Getting it wrong is a design mistake rather than a syntax error, which means nothing errors and the numbers just quietly mean less than you think.
Where it ends up
Once sentiment is a feedback score, everything else works on it for free. It is filterable in the runs table, chartable on a dashboard, and it is one of the five metrics an alert rule can watch. That last one is the point of the whole exercise: a quality regression becomes a page rather than a discovery three weeks later.
Try it yourself
Where is the sentiment feature
You are asked to start monitoring user sentiment on a live app tomorrow. What do you actually go and do?
Show answer
Correct answer: D — Compose it from explicit feedback, an LLM judge scoring threads, and Insights for the patterns nobody named
There is no standalone sentiment product in LangSmith, and that is the distinction being tested. Sentiment is assembled from primitives you already know: feedback, evaluators, Insights, threads. The alert rule option is the sharpest distractor because alert rules genuinely do read a Feedback Score metric, so you really can be paged about sentiment, once you have made it into a feedback score. Creating the rule is the last step, not the mechanism. A sentiment feature and an ls_sentiment field are invented surfaces.
Three sources of sentiment signal
Name them and say what each is bad at. The weaknesses are the lesson, because they are what force you to use more than one.
Sentiment monitoring is composed from three kinds of signal. Name all three and give the main weakness of each.
Reveal answer
One, explicit user feedback, a thumbs up or a rating written under a feedback key you choose. It is the most trustworthy signal because the human said it, and the sparsest, because almost nobody clicks the thumb. Two, an LLM judge running as an online evaluator, scoring runs or threads against a criterion you wrote. It gives coverage on every sampled interaction, and it is only ever as good as your criterion, so it measures what you thought to ask about. Three, Insights clustering, which surfaces recurring frustration patterns nobody wrote an evaluator for. It is the only one that can find an unknown unknown, and it is scheduled, capped and plan gated, so it is not a live signal.
Score the turn or score the thread
You point an LLM judge at production traffic with the criterion "does this show the user becoming frustrated". Run level or thread level?
Show answer
Correct answer: C — Thread level, because frustration lives in the sequence and most turns in a bad conversation look neutral
Frustration is almost never located in one response. It is the third rephrasing, the second time the assistant asked for something already given. Score turns and most of a genuinely bad conversation scores neutral, so you get a diluted, noisy signal. The higher resolution argument is the tempting one because finer granularity usually is better, and here it actively destroys the phenomenon you are measuring. Online evaluators can write to a thread, and a thread score is a judgement about the whole exchange rather than an average of its parts.
Why you need both human and judge signal
Your explicit thumbs data says satisfaction is 94%. Your LLM judge says 61%. Both were computed correctly.
Show answer
Correct answer: B — Explicit feedback is self selected and sparse, so 94% describes the people who bothered to rate
The classic shape is that people who click the thumb are the ones who were pleased or the ones already engaged enough to interact with your UI, so explicit feedback skews high and covers almost nobody. Tuning the judge until it agrees is the seductive wrong move, and it is worth naming because teams do it: you would be calibrating a dense signal against a biased sparse one and destroying the only broad coverage you have. Sparse and trustworthy against dense and inferred is the whole point of holding both.
Pick your feedback key now
A tiny decision with disproportionate consequences downstream. Two minutes.
You have chosen a single feedback key name for user satisfaction in your own app and can explain why renaming it in six months would fragment every chart, every filter and every alert threshold that reads it.