Whetstone.
Module 3: monitoring qualityModule 3 lab and quiz: the quality signal nobody is watching
Module 3, Lab and Quiz14 min

Module 3 lab and quiz: the quality signal nobody is watching

Quality is the module where monitoring stops being operations and starts being product. Cost tells you what you spent. Errors tell you what crashed. Neither of them can see the thing that actually loses you customers, which is a system that works perfectly and answers badly.

The failure class this module exists for

Successful, fast, correctly priced, and wrong.

That is the shape of nearly every serious agent defect: an ungrounded answer, a hallucinated citation, a tool call that succeeded against the wrong record, a conversation that quietly stopped carrying context. None of them error. None of them are slow. Some of them are cheaper than the correct behaviour, which means a naive cost dashboard will read the regression as an optimisation.

The offline half of the lab

You do not need an org to do the valuable part. The step that changes how people work is the first one: writing down the ways your app can fail while looking healthy.

Most engineers can list two immediately and then stall, and the stall is informative. It usually means the third failure mode is one nobody has ever looked for, which makes it the one that has been happening for months.

What to carry into module 4

Everything you just built is inert until something reads it. A feedback score sitting on a run that nobody charts and nobody alarms on is a diary entry, not a monitor. The next module is where the chain finishes, and it finishes with a threshold, a window, and somebody’s phone.

Practice

Try it yourself

Quiz

The regression nothing caught

A prompt change made answers subtly less grounded. Error rate is unchanged, latency is unchanged, cost is slightly down, and users have not complained loudly enough to notice.

  1. ALatency percentiles, because degraded reasoning shows up as faster and shallower answers
  2. BCost, since the slightly cheaper number is itself the evidence that quality moved
  3. CA Feedback Score from an online groundedness evaluator, the one alert metric of the five that sees this
  4. DRun count, since better grounded answers need fewer retries and fewer runs overall
Show answer

Correct answer: C — A Feedback Score from an online groundedness evaluator, the one alert metric of the five that sees this

Ungrounded answers are successful, fast, cheap runs. Four of the five alert metrics are structurally blind to them, and Feedback Score is the one that is not, which is exactly why it exists in that list. The cost option is the interesting distractor because the cheaper number really is correlated here and you would notice it, but a cost drop is consistent with a dozen benign causes and tells you nothing about quality on its own.

Quiz

Which run do you optimise

Trace timings
root      chain    12.0s
  ├─ chain  plan       1.1s
  ├─ chain  research  10.4s
  │    ├─ tool  search-a   4.9s
  │    ├─ tool  search-b   5.1s
  │    └─ tool  search-c   4.8s
  └─ llm    compose    0.4s
  1. AThe research chain as a whole, since its searches overlap and no single one holds the time
  2. BAny one of the three searches, since removing one saves about five seconds
  3. CThe root chain at 12.0s, since it holds the total for the trace
  4. DNothing can be optimised by removing work; the three searches total 14.8s inside a 10.4s parent, so the timings must be misread
Show answer

Correct answer: A — The research chain as a whole, since its searches overlap and no single one holds the time

The three searches total 14.8s inside a 10.4s parent, so they overlap: they ran in parallel and the parent's wall clock is near the slowest branch plus overhead. Removing one search therefore saves almost nothing, which kills the second option, and that is the tempting one because dividing 10.4 by three is the reflex. The work to attack is the research chain as a whole: fewer parallel branches will not help, but a faster backend or a cheaper strategy will. The last option mistakes parallelism for a data error.

Recall

The quality module in one card

Consolidation card. Three lessons, one thread running through them, and the thread is what the exam actually tests.

State the rule for attributing latency in a trace, the rule for reading error status, what user sentiment is composed from, and the one property that makes an online A/B split valid.

Reveal answer

Latency: a parent's latency includes its children, so walk down until the time stops being explained by something beneath it; children that sum to more than their parent ran in parallel. Errors: status is per run, so an errored tool run can sit inside a successful trace, and counting only root failures hides every handled failure. Sentiment: composed from explicit feedback, an LLM judge as an online evaluator, and Insights clustering, scored at thread granularity because frustration is a sequence property. A/B: both arms must run concurrently with a stable variant identifier in metadata, or the change is confounded with time.

Quiz

The two numbers disagree

Explicit thumbs data says 91% satisfied across 300 ratings. Your LLM judge says 58% across 40,000 sampled threads. Nothing is misconfigured.

  1. ATake the thumbs figure, because a human said it and human signal outranks inference
  2. BTake the judge figure, because 40,000 sampled threads is vastly more data than 300
  3. CAverage the two figures, weighted by their respective sample sizes, into one number
  4. DReport both, because they describe different populations, and treat the gap itself as the finding
Show answer

Correct answer: D — Report both, because they describe different populations, and treat the gap itself as the finding

Three hundred self selected ratings describe the people who chose to rate you; forty thousand sampled threads describe your traffic. Neither is the truth about the other. The sample size answer is the strongest distractor because it applies a correct statistical instinct to two populations that are not the same population, which is precisely when sample size stops being the deciding argument. Averaging them produces a number describing nobody.

Quiz

Pick the granularity for each criterion

Four criteria. Which one is the clearest case for scoring at thread level rather than run level?

  1. ADid this response cite a source from the retrieved context
  2. BDid the user ultimately get what they came for
  3. CDid this call use the correct tool for the request
  4. DWas this response under the length limit
Show answer

Correct answer: B — Did the user ultimately get what they came for

Only the second is a property of the whole exchange and cannot be located in any single turn: most individual turns of a successful conversation are not themselves the success. The other three are judgements about one response and score cleanly per run. Tool correctness is the near miss, because a wrong tool choice often reveals itself over several turns, but the choice itself happened in one run and is judgeable there.

Do

Lab: find the quality signal you are not watching

The official lab is hands on in a provisioned org. This version is readable and reasonable offline, and the output is a design you could ship on Monday.

  • Write down the three ways your own app can fail that produce a successful, fast, normally priced run. Be specific; "bad answer" does not count.
  • For each one, write which of the five alert metrics could detect it. Most will come back with only Feedback Score, and that is the point of the exercise.
  • Pick the one that would embarrass you most in front of a customer. Write a reference free criterion for it in one sentence, judgeable from the run or thread alone with no reference answer available.
  • Decide the granularity, run or thread, and write the sentence justifying it.
  • Pick a sampling rate and write the arithmetic: matching runs per day, times the rate, equals judged runs per day. If that number is under about fifty, say what you will do about it.
  • Now take the slowest trace you can find and attribute its wall clock down the tree until you reach the run that actually holds the time.
Done whenI have three silent failure modes, one written reference free criterion with a granularity and a justification, an explicit sampling arithmetic, and one latency attribution walked to a leaf.
Sign in to track your progress →