Whetstone.
Module 2: monitoring product analyticsAnalyze trends with Insights
Module 2, Lesson 116 min

Analyze trends with Insights

Every other tool in this domain measures something you already decided to care about. Insights is the one instrument that tells you what you should have been looking at, and that difference is the centre of this module.

LangChain’s own one liner is the cleanest statement of the pair, and it is worth memorising verbatim: Insights helps you identify what to improve; evals help you measure if improvements work.

What it actually does

Two stages, and the two stage shape explains almost everything else about the feature.

Summarise. Every trace in scope gets condensed on its own. This is the Summarization model’s job, and it runs once per trace, so this stage scales with volume.

Cluster. Those summaries are grouped into hierarchical categories and subcategories. This is the Thinking model’s job, and it is the stage that has to hold the whole set in mind at once.

The output is a report: a tree of categories with counts, describing what your traffic is actually made of. Not a score. Not feedback on a run. A document you read.

Its two hard constraints

Capped at 1,000 traces per run. This is a sample of your traffic, not a census, and the cap follows directly from the clustering stage having to hold the whole set at once.

Plus and Enterprise plans only. This is the only plan gated feature in the Monitor domain. Tracing is not gated, online evals are not gated, alerts are not gated. Insights is. Any scenario question that mentions a smaller plan and asks what is available to you has exactly one exclusion, and this is it.

Scheduled, and scheduled in UTC

Insights runs daily, weekly, or on a cron expression, interpreted in UTC. Set it and it produces a fresh report each period.

That is a completely different rhythm from online evaluators, which poll continuously. Insights is a periodic study of what happened. It cannot tell you what is on fire right now, and reaching for it during an incident is a category error rather than a slow answer.

The UTC detail is small and it bites: a daily job you think of as “overnight” fires on UTC’s midnight, not yours, so during British Summer Time your report boundary sits an hour off where you assumed it was.

Practice

Try it yourself

Quiz

Insights asks you to configure two models

Setting up an Insights job asks for two separate model configurations. What are the two roles?

  1. AA primary model and a fallback used only when the primary is rate limited or errors
  2. BAn embedding model to vectorise the traces and a chat model to name the clusters it finds
  3. CA cheap model for sampling traces and an expensive model for writing the final report
  4. DA Summarization model that condenses each trace, and a Thinking model that does the clustering
Show answer

Correct answer: D — A Summarization model that condenses each trace, and a Thinking model that does the clustering

Two roles, two separately configured models: Summarization runs once per trace, Thinking does the clustering across the summaries. The embedding answer is genuinely tempting because that is how you would build unsupervised clustering from scratch, it is how a lot of comparable tooling works, and it therefore feels like informed reasoning rather than a guess. It is not how this feature is configured, and knowing that is a strong signal you have actually met it.

Recall

The two stage pipeline

Saying the shape out loud explains the two models and the trace cap at the same time, which is why this card is worth more than either fact alone.

Describe what an Insights job does to your traces, in order, and explain why it needs two stages rather than one.

Reveal answer

First it summarises each trace individually, which is the Summarization model's job and scales linearly with the number of traces. Then it clusters those summaries into hierarchical categories and subcategories, which is the Thinking model's job and is the stage that has to reason across the whole set at once. Two stages because no model can hold thousands of raw traces in one context and compare them, so you compress each one first and let the reasoning step work on something that fits. That also explains the 1,000 trace cap, since the clustering stage has a finite budget for how much it can hold and compare.

Quiz

The two hard constraints

Two limits on Insights, both very quotable, and one of them is the only plan gate in this entire domain.

  1. ACapped at 1,000 traces per run, available on Plus and Enterprise plans only
  2. BCapped at 25,000 traces per run, available on every plan including Free
  3. CCapped at 1,000 traces per run, available on every plan including Free
  4. DNo trace cap at all, and restricted to Enterprise plans only
Show answer

Correct answer: A — Capped at 1,000 traces per run, available on Plus and Enterprise plans only

1,000 traces per Insights run, and Plus and Enterprise only. The 25,000 answer is the run cap inside a single trace from module 0, placed here deliberately because those two numbers live next to each other in memory and swap under time pressure. The third option gets the number right and drops the plan gate, which is the half most people forget, and it is the half a scenario question about a smaller plan will hinge on.

Quiz

When does an Insights job run

Production went sideways ten minutes ago and somebody suggests running Insights to find out what is happening.

  1. AGood idea, Insights polls continuously and will pick the incident up within a few minutes
  2. BInsights runs on a schedule, daily, weekly or cron in UTC, so it is a periodic study not a live signal
  3. CInsights runs on demand only and has no scheduling at all, so trigger it right now
  4. DInsights runs automatically whenever a project crosses a trace count threshold you set
Show answer

Correct answer: B — Insights runs on a schedule, daily, weekly or cron in UTC, so it is a periodic study not a live signal

Scheduled, with the schedule interpreted in UTC, which is its own quiet trap for anyone reasoning in local time about when a daily job actually fires. The continuous polling answer is the trap worth naming: it describes online evaluators, not Insights, and confusing those two rhythms is the most common error in this whole domain. For a live incident, you want the runs table and an error filter.

Quiz

Two minutes, docs open, find the Insights limits

A question hinges on whether your workspace plan can run Insights at all, and you want to confirm the cap and the plan gate rather than trust memory.

  1. AThe LangSmith self hosting docs, under deployment sizing and workspace plan limits
  2. BThe LangChain how-to guides, under summarisation chains and clustering patterns
  3. CThe LangSmith observability docs, on the Insights page, plus the pricing or plans page for the gate
  4. DThe LangSmith evaluation docs, under evaluator configuration and feedback scores
Show answer

Correct answer: C — The LangSmith observability docs, on the Insights page, plus the pricing or plans page for the gate

Insights is an observability feature, so its own page carries the cap and the model configuration, and the plan availability is confirmed on the plans page. The evaluation section is the near miss that costs people time, because Insights and online evals sit adjacent in mental models even though they live in different parts of the documentation. Summarisation chains are a LangChain building block and have nothing to do with this feature despite sharing the word.

Check

Test one of your own questions

Take a real question you have about your own production traffic and run the single test that decides the tool.

You should see

You can say whether Insights could answer it, and the test you used was whether you already know what you are looking for. If you do, it is not an Insights question, and the next lesson tells you what it is instead.

Sign in to track your progress →