Whetstone.
Module 4: monitoring securitySetting up alerts, five metrics, one rule type, and the email trap
Module 4, Lesson 217 min

Setting up alerts, five metrics, one rule type, and the email trap

Alerting is the least conceptually interesting part of this domain and the most quotable, which makes it dense in exam value per minute of study. Three short lists, one famous trap. Learn it and move on.

Five metrics

An alert rule can be built on exactly five metrics: Run Count, Cost, Errors, Feedback Score, and Latency.

Four of those are properties of the machine. Feedback Score is the interesting one, because it is the bridge from everything else in this course to somebody’s phone actually ringing. Your online evaluator writes feedback. Your sentiment composition produces feedback. Your injection detector produces feedback. An alert rule reads the aggregated score. That is how a quality regression becomes a page rather than a discovery three weeks later.

One rule type

There is a single alert type: threshold. You choose a metric, a number and a comparison, and the alert fires when the aggregated value crosses the number.

That is it. No anomaly detection, no baseline learning, no trend or rate of change rules, no composite conditions.

Two window sizes

The metric is aggregated over a window before comparison, and the window is 5 minutes or 15 minutes. Not an arbitrary duration, not a free text field.

Shorter means faster detection and more noise from ordinary variance. Longer means calmer alerts and a slower page. With only two options this is a genuinely easy decision, which is presumably why it was constrained.

Four channels, and the one that is missing

Slack, PagerDuty, Dynatrace, and Webhook.

Slack for the team channel. PagerDuty for the on call rotation. Dynatrace for shops that already centralise observability there. Webhook for absolutely everything else.

There is no native email channel, and it is worth being annoyed about this for a moment so that it sticks. Email is the most universal notification mechanism there has ever been, so an option saying “configure the email channel” reads as boring and obviously correct. That is precisely why it works as a distractor: nothing about it trips your suspicion, because nothing about it seems clever.

To send email, use the Webhook channel pointed at an email service such as SendGrid or Mailgun. Webhook is the general escape hatch, not an email workaround: Teams, Discord, a ticketing system, a bespoke internal handler, all the same pattern.

Practice

Try it yourself

Quiz

The email question

Your team wants alert emails going to a distribution list. Read all four options before answering; one of them is built to look like the safe choice.

  1. AThe Slack channel, since Slack forwards channel messages to email automatically
  2. BThe email channel, entering the distribution list address in the recipient field
  3. CA webhook channel pointed at an email service like SendGrid, because email is not a native channel
  4. DEmail alerts require an Enterprise plan, so upgrade the workspace before configuring
Show answer

Correct answer: C — A webhook channel pointed at an email service like SendGrid, because email is not a native channel

Email is not a native notification channel. The four are Slack, PagerDuty, Dynatrace and Webhook. The email channel option is the highest quality distractor in this domain, because email is the most obvious notification mechanism in existence and every other tool you have used has it, so the option reads as unremarkable and you tick it without friction. To get email you send a webhook to something that sends email. The Enterprise option invents a plan gate; the only plan gate in this domain is Insights.

Recall

The alerting facts, as three lists

Pure enumeration, three short lists, and enumeration is the cheapest mark per minute of study anywhere on this exam.

Name the five metrics an alert rule can watch, the rule types available, and the aggregation window options.

Reveal answer

Metrics: Run Count, Cost, Errors, Feedback Score, Latency. Rule types: exactly one, a threshold rule, where the aggregated metric crosses a number you set. There is no anomaly detection, no baseline learning, no trend or rate of change rule and no composite condition combining two metrics. Windows: 5 minutes or 15 minutes, and nothing else. Feedback Score is the interesting entry in the metric list, because it is the bridge from every quality and security composition in this course to something that can actually wake a human.

Quiz

Alerting on unusual latency

You want to be paged when latency is unusually high compared with a normal Tuesday. What can you configure?

  1. AAn anomaly rule that learns the baseline, plus a threshold rule as a backstop
  2. BA trend rule that fires on the rate of change of the latency metric over time
  3. CA composite rule combining the latency and error count metrics in one condition
  4. DOnly a threshold rule, so you choose a latency number yourself and alert when it is crossed
Show answer

Correct answer: D — Only a threshold rule, so you choose a latency number yourself and alert when it is crossed

There is exactly one alert type and it is threshold. Anomaly detection is the distractor with the best story, because "learns your normal baseline" is standard in observability products and is genuinely what this scenario wants. It does not exist here, and neither do trend or composite rules. Notice how the question is phrased to make the sophisticated answer feel like the informed one; that phrasing is the trap, not the content.

Quiz

The aggregation window options

An alert rule aggregates its metric over a window before comparing it against the threshold. What are the choices?

  1. A5 minutes or 15 minutes
  2. BAny duration from 1 minute to 24 hours
  3. C1 minute, 5 minutes, or 1 hour
  4. DFixed at 5 minutes and not configurable
Show answer

Correct answer: A — 5 minutes or 15 minutes

Two choices, 5 or 15 minutes. The arbitrary duration option is the tempting one, because free window configuration is normal in monitoring tools and nothing about the feature signals otherwise until you look. The fixed at 5 minutes option is wrong in the direction people guess when they have only ever accepted the default, which makes it a good distractor for anyone who has seen the screen once and not read it.

Recall

The two sources disagree

A live documentation disagreement, which is worth more to you than the fact on its own, because it tells you why an options list might look wrong.

The LangSmith docs and the alerting launch blog give different counts for metrics and channels. State both counts, say which you answer with, and say why.

Reveal answer

The docs describe five metrics and four channels. The launch blog lists three metrics and two channels. The blog is almost certainly the older artefact, written when the feature shipped in a smaller form, so answer with the docs: five and four. The exam is semi open book with docs.langchain.com available, which is itself a signal about which source is treated as authoritative. The practical value of knowing the disagreement exists is that if you meet an options list built around three metrics, you will recognise where it came from instead of spending thirty seconds you do not have doubting yourself.

Check

Pick a threshold you would accept being woken for

One rule, for your own app, written out in full. The number is the hard part and it is meant to be.

You should see

You have a metric, a number and a window, and you can defend the number against both directions of failure: low enough to catch a real incident, high enough that you would not mute the rule after the third false page. A rule that gets muted is worse than no rule, because it looks like coverage.

Sign in to track your progress →