Whetstone.
Moving Towards ProductionThe Insights Agent
Module 3, Lesson 215 min

The Insights Agent

Every evaluator you own is answering a question you already asked. That is a real limitation and it is invisible from inside, because a dashboard full of green scores looks exactly the same whether your traffic matches your assumptions or not.

The Insights Agent is the instrument for the other job: finding out what is in there.

Unsupervised discovery

You point it at a tracing project. It works in two stages.

Summarise. A Summarization model runs per trace and reduces each one to a short description of what happened.

Cluster. A Thinking model takes those summaries and groups them into hierarchical categories: themes, with sub-themes underneath.

You did not supply the categories. That is the entire point. The output is a picture of what your traffic actually consists of, including the third of it that turns out to be a use case nobody designed for.

The two-model design is worth understanding rather than merely memorising, because it explains everything else about the feature. The per-trace work is deliberately cheap summarisation; the reasoning-heavy clustering happens once, over the summaries. And it explains the cap, since one model call per trace is what makes a limit necessary.

The numbers

Capped at 1,000 traces per run. Scheduled, daily, weekly, or by cron expression, with cron interpreted in UTC. Plus and Enterprise plans only. Two configured models, a Thinking model for clustering and a Summarization model per trace.

Insights against evals

The docs put the division of labour cleanly, and it is worth learning close to the wording:

Insights helps you identify what to improve; evals help you measure if improvements work.

Insights is discovery and produces themes. Evals are measurement and produce numbers. Neither substitutes for the other and the failure modes are symmetrical: expect a score from Insights and you will conclude it is broken, expect a category from an evaluator and you will conclude the same.

The practical sequence is a handoff. Insights surfaces that eleven per cent of your traffic is people trying to cancel a subscription through a support agent that was never given a cancellation tool. That is not a score and could not have been, because you had no cancellation evaluator to score it with. Now you write one, and now you can measure whether the fix worked.

Where it fits in the loop

This is the instrument that keeps your dataset honest.

Left alone, an offline dataset drifts away from reality: it encodes the failures you knew about on the day you built it. Insights is the periodic re-survey that tells you the traffic has moved, which categories are growing, and which of them nothing in your test suite currently touches.

Use it as a prompt for the harvest rather than as a report to read. A cluster you did not expect is a signal to go and pull examples from it into a dataset, and that is the point at which discovery turns into a regression test.

Practice

Try it yourself

Quiz

How many traces per run

The Insights Agent is capped per run, and the number is exactly the sort of quotable figure a semi-open-book exam likes.

  1. A100 traces
  2. B1,000 traces
  3. C10,000 traces
  4. DUnlimited, bounded only by your retention window
Show answer

Correct answer: B — 1,000 traces

The cap is 1,000 traces per run. Option 3 is the tempting one because an unsupervised clustering feature feels like it ought to consume everything you have, and the cap is what forces you to think about which slice you point it at. That consequence is the real content of the fact: on a project doing far more than 1,000 traces in a scheduling period, the cap means you are sampling whether or not you meant to, so the filter you choose is doing real work.

Recall

The two models it uses

A configuration detail that also explains the shape of the output, which is why it is worth knowing rather than merely quotable.

The Insights Agent is configured with two models. Name both and say what each one does.

Reveal answer

A Summarization model, which runs per trace and reduces each one to a short description, and a Thinking model, which clusters those summaries into hierarchical categories. That two-stage design is why the output is a hierarchy of themes rather than a pile of traces: the expensive per-trace work is a cheap summarisation, and the reasoning-heavy work happens once over the summaries. It also explains the 1,000-trace cap, since the summarisation stage costs a model call per trace.

Quiz

Who can run it

Which statement about Insights Agent availability and scheduling is correct?

  1. AAvailable on all plans, and runs continuously as new traces arrive in the project
  2. BEnterprise only, and runs once per calendar month on a fixed schedule
  3. CAvailable on all plans, but only when triggered manually from a tracing project
  4. DPlus and Enterprise only, scheduled daily, weekly or by cron expression in UTC
Show answer

Correct answer: D — Plus and Enterprise only, scheduled daily, weekly or by cron expression in UTC

It is a Plus and Enterprise feature, scheduled daily, weekly or by cron expression, with cron interpreted in UTC. Option 0 is the tempting one because continuous processing is what you would assume from something described as an agent watching your traces, and getting this wrong changes how you would design around it: scheduled means results are a periodic report rather than a live signal, so it is not the mechanism you reach for when you need to know within minutes. The UTC detail is the kind of thing that quietly matters when a cron lands on a boundary in your own timezone.

Recall

Insights against evals, in one sentence each

The docs draw this contrast explicitly, and it is close to verbatim worth learning because it settles which instrument answers which question.

State the division of labour between the Insights Agent and evaluators.

Reveal answer

Insights helps you identify what to improve; evals help you measure if improvements work. Insights is unsupervised discovery, so it can surface a category of traffic or failure you never thought to look for, and it produces themes rather than scores. Evaluators are supervised measurement, so they can only ever score the property you specified, and they produce numbers you can track. The practical sequence is that Insights tells you where to point an evaluator, and the evaluator tells you whether the fix landed.

Quiz

What you get back

An Insights run completes over a tracing project. What is the output?

  1. AA pass or fail verdict on each of the traces it processed
  2. BA single quality score for the tracing project over the scheduled period
  3. CA set of hierarchical categories that the traces were clustered into
  4. DA ranked list of the slowest and most expensive traces in the period
Show answer

Correct answer: C — A set of hierarchical categories that the traces were clustered into

The output is a hierarchy of categories: it summarises each trace, then clusters those summaries into themes with sub-themes. Option 1 is the distractor, and it is the one to reject deliberately, because the whole value of the feature is that it produces no score at all. If you expect a number from Insights you will conclude it is broken, and if you expect categories from an evaluator you will conclude the same. Option 3 describes what the Runs view already gives you by sorting.

Sign in to track your progress →