Whetstone.
Module 2: monitoring product analyticsModule 2 lab and quiz: what is my traffic actually made of
Module 2, Lab and Quiz14 min

Module 2 lab and quiz: what is my traffic actually made of

The official course puts one teaching lesson in this module, “Analyze Trends with Insights”, and then a lab and a quiz. That ratio is itself information: this module is short on facts and long on judgement, and judgement is what a lab is for.

The habit this lab builds

Commit to your guess before you look. Almost everyone believes they know what their users do. The gap between the guessed distribution and the real one is the entire value of unsupervised discovery, and you can only see the gap if you wrote the guess down first.

That is not a study trick, it is the argument for the feature. If your guess were reliable, you would skip discovery and go straight to writing evaluators. Insights earns its keep exactly in proportion to how wrong you were.

Why hand simulating the two stages is worth ten minutes

Summarising twenty traces by hand and then grouping the summaries teaches you two things no amount of reading does.

First, the summarisation stage is lossy on purpose, and you feel which detail you dropped. That is why a category tree can be right about the shape of your traffic and useless for debugging any individual case.

Second, the clustering stage is doing something you can actually do, which demystifies it. It is not magic and it is not embeddings. It is a model looking at a pile of short descriptions and naming the piles.

Practice

Try it yourself

Quiz

A workspace on the wrong plan

A startup on the free tier wants to understand what their users are actually asking for. Nobody has a hypothesis. What can they do today?

  1. ASchedule an Insights job, since unnamed discovery is exactly the problem they have
  2. BNothing useful, since this is an Insights problem and Insights is gated above the free tier
  3. CCompose a rough substitute from what is not gated, then upgrade when the categories justify it
  4. DRun offline evaluations against a regression dataset until a pattern emerges
Show answer

Correct answer: C — Compose a rough substitute from what is not gated, then upgrade when the categories justify it

Insights is Plus and Enterprise only, so it is off the table, but the plan gate does not gate the underlying problem. Tracing, filtering, feedback and online evaluators are all available, so you can hand read a sample, tag recurring shapes as metadata, and get a crude category count. The second option is the tempting one because it is technically correct about the gate and it feels rigorous to say so, but answering "nothing useful" to a product question is almost never the right professional move. Offline evaluation over a dataset measures a criterion you already have, which is the thing this team explicitly lacks.

Quiz

Match the rhythm to the question

Four questions about the same production app. One of them is the only one Insights is the right instrument for.

  1. AWhich recurring themes is our traffic made of that nobody has named yet
  2. BIs our groundedness score higher this week than it was last week
  3. CHow many runs errored in the last fifteen minutes on this project
  4. DWhat did we spend on the summarisation step over yesterday's traffic
Show answer

Correct answer: A — Which recurring themes is our traffic made of that nobody has named yet

Only the first has no hypothesis in it, and no hypothesis is the Insights signature. Groundedness week over week is a defined criterion, so it is an online evaluator plus a chart. Errors in the last fifteen minutes is a live operational question for the runs table or an alert. Yesterday's spend is a cost dashboard question. The groundedness option is the distractor worth pausing on, because it involves comparing periods, and "comparing periods" feels like trend analysis, which sounds like the word Insights.

Recall

Insights in one card

Consolidation card for the module. Three pairs, and each pair explains the other two.

Name the two models an Insights job configures, the two stages of its pipeline, and its two hard constraints. Then say how the stages explain the constraints.

Reveal answer

Models: a Summarization model and a Thinking model. Stages: summarise each trace individually, then cluster the summaries into hierarchical categories and subcategories. Constraints: capped at 1,000 traces per run, and available on Plus and Enterprise plans only. The clustering stage has to reason across the entire set at once, which is what bounds the trace count; the summarisation stage exists precisely so that the clustering stage is comparing short summaries rather than raw traces it could never fit. The plan gate follows from the cost of running two models over a thousand traces on a schedule.

Quiz

What a sampling rate of 0.1 does

You set an online evaluator's sampling rate to 0.1 on a filter matching roughly 10,000 runs a day.

  1. AIt waits until 10 runs have accumulated, then evaluates them as a batch
  2. BIt evaluates the first 10% of the day's matching runs and then stops
  3. CIt evaluates every matching run but stores only 10% of the results
  4. DIt evaluates approximately 10% of matching runs, so roughly 1,000 a day
Show answer

Correct answer: D — It evaluates approximately 10% of matching runs, so roughly 1,000 a day

Sampling rate is a value from 0 to 1 controlling the proportion of matching runs that get judged. The evaluate everything and discard option is the tempting misread and it is the expensive interpretation, which defeats the entire reason sampling exists: judge calls cost money and latency. Taking the first ten percent of the day would bias the sample toward one part of the day, which is exactly what random sampling avoids.

Quiz

The daily job that fires at the wrong time

You schedule a daily Insights job for what you think of as midnight, sit in London in August, and find the report boundary is not where you expected.

  1. AThe schedule drifted because the previous run overran its window
  2. BSchedules are interpreted in UTC, so during British Summer Time the job fires at 01:00 local
  3. CInsights schedules are approximate and fire within a few hours of the requested time
  4. DThe report covers the previous 1,000 traces rather than a fixed period, so boundaries move with volume
Show answer

Correct answer: B — Schedules are interpreted in UTC, so during British Summer Time the job fires at 01:00 local

Schedules are interpreted in UTC, so any local time offset shifts when the job actually fires and therefore which traces land either side of the boundary. The last option is the sharpest distractor because the 1,000 cap is real and it genuinely does mean a busy day is sampled rather than covered, so it is a true statement that is not the cause of this particular symptom. Two real facts, only one of which explains the clock.

Do

Lab: name your own categories before the machine does

The official lab schedules a real Insights job. The valuable half of that exercise works offline, and doing it in this order is what makes the result mean something.

  • Before looking at any data, write down the five categories you believe your production traffic falls into, with a guessed percentage for each. Commit to the numbers.
  • Now open a project and read twenty traces at random. Tally each one against your five categories, and open a sixth bucket called "none of the above".
  • Count the "none of the above" bucket. That number is the honest measure of how much of your own product you cannot currently see.
  • Write down what a Summarization model would have condensed each of those unclassified traces into, in one line each. You are hand simulating stage one.
  • Group your one line summaries into categories. You have now hand run stage two, and you know what the clustering step is actually doing.
  • If you have a Plus or Enterprise workspace, schedule the real job and compare its category tree against yours. If you do not, write down which of your categories you would turn into an online evaluator first, and why that one.
Done whenI have a pre committed list of five categories, a count of traces that fit none of them, and either a real Insights category tree to compare against or a named first evaluator with a reason.
Sign in to track your progress →