Module 2 lab and quiz: what is my traffic actually made of
The official course puts one teaching lesson in this module, “Analyze Trends with Insights”, and then a lab and a quiz. That ratio is itself information: this module is short on facts and long on judgement, and judgement is what a lab is for.
The habit this lab builds
Commit to your guess before you look. Almost everyone believes they know what their users do. The gap between the guessed distribution and the real one is the entire value of unsupervised discovery, and you can only see the gap if you wrote the guess down first.
That is not a study trick, it is the argument for the feature. If your guess were reliable, you would skip discovery and go straight to writing evaluators. Insights earns its keep exactly in proportion to how wrong you were.
Why hand simulating the two stages is worth ten minutes
Summarising twenty traces by hand and then grouping the summaries teaches you two things no amount of reading does.
First, the summarisation stage is lossy on purpose, and you feel which detail you dropped. That is why a category tree can be right about the shape of your traffic and useless for debugging any individual case.
Second, the clustering stage is doing something you can actually do, which demystifies it. It is not magic and it is not embeddings. It is a model looking at a pile of short descriptions and naming the piles.
Try it yourself
A workspace on the wrong plan
A startup on the free tier wants to understand what their users are actually asking for. Nobody has a hypothesis. What can they do today?
Show answer
Correct answer: C — Compose a rough substitute from what is not gated, then upgrade when the categories justify it
Insights is Plus and Enterprise only, so it is off the table, but the plan gate does not gate the underlying problem. Tracing, filtering, feedback and online evaluators are all available, so you can hand read a sample, tag recurring shapes as metadata, and get a crude category count. The second option is the tempting one because it is technically correct about the gate and it feels rigorous to say so, but answering "nothing useful" to a product question is almost never the right professional move. Offline evaluation over a dataset measures a criterion you already have, which is the thing this team explicitly lacks.
Match the rhythm to the question
Four questions about the same production app. One of them is the only one Insights is the right instrument for.
Show answer
Correct answer: A — Which recurring themes is our traffic made of that nobody has named yet
Only the first has no hypothesis in it, and no hypothesis is the Insights signature. Groundedness week over week is a defined criterion, so it is an online evaluator plus a chart. Errors in the last fifteen minutes is a live operational question for the runs table or an alert. Yesterday's spend is a cost dashboard question. The groundedness option is the distractor worth pausing on, because it involves comparing periods, and "comparing periods" feels like trend analysis, which sounds like the word Insights.
Insights in one card
Consolidation card for the module. Three pairs, and each pair explains the other two.
Name the two models an Insights job configures, the two stages of its pipeline, and its two hard constraints. Then say how the stages explain the constraints.
Reveal answer
Models: a Summarization model and a Thinking model. Stages: summarise each trace individually, then cluster the summaries into hierarchical categories and subcategories. Constraints: capped at 1,000 traces per run, and available on Plus and Enterprise plans only. The clustering stage has to reason across the entire set at once, which is what bounds the trace count; the summarisation stage exists precisely so that the clustering stage is comparing short summaries rather than raw traces it could never fit. The plan gate follows from the cost of running two models over a thousand traces on a schedule.
What a sampling rate of 0.1 does
You set an online evaluator's sampling rate to 0.1 on a filter matching roughly 10,000 runs a day.
Show answer
Correct answer: D — It evaluates approximately 10% of matching runs, so roughly 1,000 a day
Sampling rate is a value from 0 to 1 controlling the proportion of matching runs that get judged. The evaluate everything and discard option is the tempting misread and it is the expensive interpretation, which defeats the entire reason sampling exists: judge calls cost money and latency. Taking the first ten percent of the day would bias the sample toward one part of the day, which is exactly what random sampling avoids.
The daily job that fires at the wrong time
You schedule a daily Insights job for what you think of as midnight, sit in London in August, and find the report boundary is not where you expected.
Show answer
Correct answer: B — Schedules are interpreted in UTC, so during British Summer Time the job fires at 01:00 local
Schedules are interpreted in UTC, so any local time offset shifts when the job actually fires and therefore which traces land either side of the boundary. The last option is the sharpest distractor because the 1,000 cap is real and it genuinely does mean a busy day is sampled rather than covered, so it is a true statement that is not the cause of this particular symptom. Two real facts, only one of which explains the clock.
Lab: name your own categories before the machine does
The official lab schedules a real Insights job. The valuable half of that exercise works offline, and doing it in this order is what makes the result mean something.
- Before looking at any data, write down the five categories you believe your production traffic falls into, with a guessed percentage for each. Commit to the numbers.
- Now open a project and read twenty traces at random. Tally each one against your five categories, and open a sixth bucket called "none of the above".
- Count the "none of the above" bucket. That number is the honest measure of how much of your own product you cannot currently see.
- Write down what a Summarization model would have condensed each of those unclassified traces into, in one line each. You are hand simulating stage one.
- Group your one line summaries into categories. You have now hand run stage two, and you know what the clustering step is actually doing.
- If you have a Plus or Enterprise workspace, schedule the real job and compare its category tree against yours. If you do not, write down which of your categories you would turn into an online evaluator first, and why that one.