Whetstone.
Module 4: monitoring securityModule 4 lab and quiz: the alert you would actually answer
Module 4, Lab and Quiz14 min

Module 4 lab and quiz: the alert you would actually answer

The chain finishes here. A trace records what happened, a thread groups it into a conversation, cost prices it, evaluators judge it, and an alert rule turns one of those numbers into a phone call. Everything before this module was potential energy.

The rule that gets muted is worse than no rule

This is the only genuinely hard part of alerting and it is not on the exam, so it is worth ten seconds of attention purely because it will matter to you professionally.

A threshold you chose too low pages you during ordinary variance. You mute it. The mute persists. Six weeks later the metric moves for real and nothing happens, and the incident review records that you had monitoring, which is true and useless.

Defending a number in both directions is the discipline that prevents this. Too low and you will silence it. Too high and it will never fire. If you can only argue one side, you picked a number that felt safe rather than one that works.

Low volume changes the arithmetic

A five minute window on a service doing forty runs an hour holds about three runs. A single failure is a third of the window, and any percentage based threshold reads that as a catastrophe.

The instinct to reach for the shorter window because faster is better inverts on low traffic. Volume is the denominator, so less traffic means each event carries more weight, not less. That is one of the few places in this domain where the arithmetic contradicts the reflex.

Practice

Try it yourself

Quiz

Match the incident to the metric

A downstream provider starts returning 429s. Your agent retries, mostly succeeds, and users see slightly slower answers. Which alert metric fires first and most reliably?

  1. AFeedback Score, because degraded answers under retry will score lower
  2. BLatency, because retries add wall clock to runs that still succeed
  3. CErrors, because the provider's 429 responses are recorded as failed runs
  4. DCost, because every retry multiplies the token spend on that request
Show answer

Correct answer: B — Latency, because retries add wall clock to runs that still succeed

Retried requests still succeed, so the root traces are successful and the extra time lands squarely on latency. The errors option is the strong distractor because the failed attempts genuinely are recorded as errored runs, so errors will move too, but only if your rule counts runs that errored anywhere rather than failed traces. Cost rises slowly and lags. Feedback Score has nothing to see; the answers are fine, just late.

Quiz

Choosing between the two windows

You are writing an errors threshold rule for a low volume internal tool that handles about forty runs an hour.

  1. A5 minutes, because faster detection always beats slower detection on an internal tool
  2. B5 minutes, because a low volume service cannot generate enough noise for that to matter
  3. C15 minutes, because at three runs per five minute window one failure looks like a catastrophic error rate
  4. DNeither works, since a low volume service cannot be threshold alerted meaningfully at all
Show answer

Correct answer: C — 15 minutes, because at three runs per five minute window one failure looks like a catastrophic error rate

At forty runs an hour a five minute window holds about three runs, so one failure is a third of the window and any percentage based threshold screams. The second option gets the volume reasoning exactly backwards, and it is tempting because "low volume means low noise" sounds right until you think about the denominator: low volume means each event carries more weight, not less. The last option overstates the constraint; a longer window and an absolute count both work fine.

Recall

Module 4 in one card

Consolidation card. The three lists plus the one thing that is not on any of them.

Name the five alert metrics, the rule type, the window options and the four channels. Then name the notification mechanism that is not a channel, and say how you deliver it.

Reveal answer

Metrics: Run Count, Cost, Errors, Feedback Score, Latency. Rule type: one only, threshold. Windows: 5 or 15 minutes. Channels: Slack, PagerDuty, Dynatrace, Webhook. Email is not a native channel, and you deliver it by pointing a Webhook at an email service such as SendGrid or Mailgun. Webhook is the general escape hatch for anything not on the list: Teams, Discord, a ticketing system, your own handler. The docs give five and four; the launch blog gives three and two and is older.

Quiz

Getting paged about a policy violation

You want to be alerted when your agent produces output that violates your content policy. What is the chain, end to end?

  1. AAn online evaluator judges output against the policy and writes feedback, then a Feedback Score rule fires
  2. BA content policy rule type on the alerting screen, pointed at the tracing project
  3. CInsights clusters the violations into a category and the category count triggers an alert
  4. DThe errors metric, since policy violations are recorded as failed runs by the guardrail
Show answer

Correct answer: A — An online evaluator judges output against the policy and writes feedback, then a Feedback Score rule fires

Judgement becomes feedback, feedback becomes a metric, the metric crosses a threshold and something fires. It is the same composition as sentiment and the same composition as injection detection, because there is no dedicated product for any of them. The Insights option is the best distractor because Insights really would find the pattern, but it produces a report on a schedule, not a metric, and category counts are not alertable. A policy violation that your system handled gracefully is not an error.

Quiz

The compliance question you did not know you were answering

Your data policy says user content is retained for the base period and no longer. A colleague attaches an online groundedness evaluator to a filter matching all production traffic.

  1. ANo impact, since evaluators only read traces and do not change how they are stored
  2. BNo impact, since retention is set per workspace and overrides anything an evaluator does
  3. CMatched traces are upgraded automatically to extended data retention, so the policy is now breached
  4. DOnly the sampled traces are affected, so the breach is proportional to the sampling rate
Show answer

Correct answer: C — Matched traces are upgraded automatically to extended data retention, so the policy is now breached

Attaching an online evaluator automatically upgrades matching traces to extended data retention, which turns a quality decision into a storage and compliance one. The sampling option is the sharpest distractor because it sounds precise and reasonable, and reasoning about which traces the upgrade actually covers is genuinely worth doing, but the mechanism is driven by the evaluator's match rather than by whether a given trace happened to be judged. Confirm the exact scope in the docs before you write it into a policy document.

Do

Lab: write the four rules you would actually keep

The official lab configures real rules in a provisioned org. This version produces the same artefact on paper, and the constraint is the teaching device.

  • Write four alert rules for your own app, using four different metrics from the five. For each, give the metric, the number, the window and the channel.
  • Beside each number, write the sentence that defends it against being too low. Then write the sentence that defends it against being too high. If you cannot write both, the number is a guess.
  • For any rule where you wanted email, write the webhook target instead, and name the service that will actually send the mail.
  • Now delete the weakest rule and write one line saying why. Four rules you would answer beats six you would mute.
  • Take the Feedback Score rule specifically and trace it backwards: which evaluator writes that score, what criterion does it use, at what granularity, and at what sampling rate. If any link is missing, the rule is watching an empty metric.
  • Finally, write down what happens at 3am when the rule fires and the person paged does not know the system. If the alert text does not tell them where to look, add that to the payload design.
Done whenI have three surviving rules with defended numbers, a real webhook target for anything email shaped, and a Feedback Score rule whose entire upstream chain I can name.
Sign in to track your progress →