Production operations · Intermediate
Alerting that people act on
An alert rule is a classifier: score it on what it catches, what it misses, how late it speaks and how often it cries wolf.
Educational material only. Not investment advice.
Contents
Abstract
An alert asks someone to act, so each one should say what is wrong, what to do and when it clears. When incidents are rare, even a good detector’s alerts are mostly false; Bayes’ rule shows that the rate of false alarms on quiet checks decides how many can be believed. On a synthetic metric with known incidents, a run length and hysteresis raise the share of useful alerts from about a sixth to about two fifths while catching almost as many incidents, at a cost of twenty minutes to half an hour of delay. Alerts go out through a durable outbox with receipts and escalation, a heartbeat catches a monitor that has stopped, and each rule is kept or retired on its record.
Key takeaways
- An alert needs someone to act now; everything else belongs in a digest, and each alert names what is wrong, why it matters, what to do and when it clears.
- When incidents are rare, the false alarm rate on quiet checks sets the share of alerts that are real.
- A run length and hysteresis cut false and repeated alerts from a bare threshold; the price is delay, and a fast alert needs a faster metric.
- One incident, one alert, one recovery message, with open episodes stored durably; a watcher elsewhere makes a silent monitor an alert.
- Deliver from a durable outbox with retries, receipts and escalation, and time acknowledgements separately.
- Keep or retire each rule on its precision, recall, delay and action rate.
Before you start
- Monitoring a live book, for the checks whose results become alerts
- Conditional probability, for Bayes’ rule
On the morning Knight Capital lost more than 460 million dollars, its systems sent 97 emails warning of the fault before the market opened. They had not been designed as alerts, and the staff who received them generally did not read them.1 An alert rule is a classifier, and it can be scored like one: how often it is right, how much it misses, and how late it speaks.
An alert names an action
The daily checks in Monitoring a live book produce a report; only some of its results should interrupt a person. The test is simple: does this need someone to act, now? If yes, it is an alert. If it can wait for the next session, it belongs in a digest. If nobody would do anything differently, it belongs in neither. Google’s site reliability engineers put it in two rules: every page should be actionable, and every page should require intelligence, because a page that merits only a robotic response should have been automated.2,3 They also advise alerting on symptoms, what is broken for the user, rather than on every cause that might explain it.
| Field | Example | Why it is there |
|---|---|---|
| What is wrong | Holdings differ from targets in two instruments | A symptom in the reader’s terms, not a stack trace |
| Why it matters | The book is well over its intended exposure | Lets the reader judge urgency without opening a dashboard |
| What to do | Runbook: compare the broker’s orders with the journal; cancel duplicates | An alert without an action is a report |
| Owner and severity | On call; act before the next session | Someone specific, with a deadline |
| Evidence | The check’s witness: what it read, and when | So the reader can trust it without rerunning it |
| When it clears | When holdings equal targets for one full check | So the recovery is announced, not assumed |
An alert in this shape can be acted on from a phone. The examples are illustrative.
The arithmetic of alert fatigue
Most alerts from a reasonable detector are false, and the reason is arithmetic. If an incident is under way on a fraction b of checks, a check made during an incident alerts with probability s, and a quiet check alerts with probability f, the share of alerts that are real is
- Step 1 of 4
A month of checks
A check every 15 minutes makes 2,880 checks a month. With an incident under way on 1 in 1,000 checks, about 2.9 of them fall during an incident, and a detector with s at 90% alerts on 2.6. Those are the alerts we want.
- Step 2 of 4
The quiet checks
The other 2,877 checks are quiet. If 1 quiet check in 100 alerts, they send 28.8 false alerts a month, about 11 for every real one, and only 8% of alerts are real.
- Step 3 of 4
Perfect detection
A detector that never misses raises the real alerts from 2.6 to 2.9 a month. The share that is real moves from 8% to 9%, because the false alarms are untouched.
- Step 4 of 4
Quieter checks
If only 1 quiet check in 1,000 alerts, 47% of alerts are real; at 1 in 10,000, 90%. When incidents are rare, the false alarm rate on quiet checks decides how many alerts can be believed.
Arithmetic, not a simulation: Bayes’ rule with a check every 15 minutes, as in the laboratory below; the shares are those of the base-rate table in alerting.py. Without script, or with reduced motion, the figure shows its final state.
Axelsson made the same point about intrusion detection: when the base rate is low, the false alarm rate is the limiting factor.4 Hospitals learned it at a cost: Sendelbach and Funk report that between 72% and 99% of clinical alarms are false, and they and Cvach describe staff who, overwhelmed, silence or ignore them.5,6
Of the alerts, 8% come from a real incident. Checked every 15 minutes, that is about 2.6 alerting checks during incidents and 28.8 false ones a month.
Thresholds, run length and hysteresis
Most alert rules compare a metric with a threshold, and a bare threshold is a poor rule. Noise that crosses it for one sample fires an alert; a real problem that hovers around it fires one alert after another as it crosses back and forth. Two old ideas fix most of this. A run length asks for the condition to hold for several consecutive samples before firing, trading delay for noise; Page’s cumulative-sum schemes are the principled version of the same trade.7 A threshold with hysteresis clears only when the metric falls some distance below the level that set it, which is how Schmitt’s trigger stopped a noisy signal from chattering in 1938.8 In the listing, the run length applies to clearing as well.
def detect(x, threshold: float, run: int = 1, hysteresis: float = 0.0):
"""Alert episodes as (start, end). An alert fires once the metric has been at or above
`threshold` for `run` consecutive samples, and clears only after it has been below
`threshold - hysteresis` for `run` consecutive samples. One episode is one alert: no
repeats while it stays on. `end` is None for an alert still on when the data ends."""
out, on, count, start = [], False, 0, 0
for t, v in enumerate(x):
if not on:
count = count + 1 if v >= threshold else 0
if count >= run:
on, start, count = True, t, 0
else:
count = count + 1 if v < threshold - hysteresis else 0
if count >= run:
out.append((start, t))
on, count = False, 0
if on:
out.append((start, None))
return out64 alerts in six weeks, about 46 a month. 10 tell someone about a new incident; 39 are false and 15 repeat an incident already alerted. 10 of 12 incidents caught, on average 6 minutes after they began; 2 missed.
Each point is a threshold, with the run length and hysteresis above. Moving left cuts noise; moving down cuts misses. A run length or hysteresis that pulls the solid curve towards the corner, against the dashed curve of a bare threshold, is buying both at once. No setting reaches the corner, with no noise and no misses: some incidents are smaller than the noise.
A bare threshold at 2 standard deviations fires 64 alerts in six weeks, about 46 a month. Only 10 of them tell anyone about a new incident: 39 are false and 15 repeat an incident already alerted, a precision of 16% (useful alerts over all alerts, so repeats count against it). The same threshold with a run length of 2 samples and a hysteresis of 1 sd sends 26 (38% useful) and still catches 10 of the 12 incidents. The price is time: the first alert arrives on average 39 minutes after an incident begins, against 6 minutes. Nudging the threshold to 2.25 standard deviations sends 19 alerts and still catches 10. Raising the threshold alone is the blunt fix: at 3.5 standard deviations 5 incidents are missed, and 11 of its 18 alerts are still repeats, because a higher threshold does nothing for flapping; hysteresis does, and so does a run length applied to clearing (at 3.5 standard deviations, a run of 2 cuts the repeats to 5 and a hysteresis of 1 sd to 3). An alert that must be fast needs a metric sampled faster, not a shorter run.
One path with 12 incidents is a small sample. Over 200 synthetic paths of the same model, the tuned rule’s precision beats the bare threshold’s on every one of them (40% against 16% on average). The misses are less clear-cut: it catches 11.1 incidents of 12 against 11.6, fewer on 45% of the paths, and its first alert comes 22 minutes later on average.
The scoring hides one subtlety: an alert already on when an incident begins tells nobody anything new, so it is scored false and the incident missed. Long, low-threshold alerts lose incidents this way. Choosing a point on the curve is a statement about which mistake costs more, and it should be written down.
One incident, one alert
One incident should produce one alert. An episode, from the moment a condition fires to the moment it clears, is the unit: while it lasts, the alert is not sent again. When several checks fire for one cause (a broker outage makes holdings, exposure and P&L all unavailable), a de-duplication key groups them under the symptom that matters, so the reader gets one message that lists the others. Every alert ends with a recovery message saying that the condition cleared, when, and how long it lasted; without it, a fixed problem and a monitor that stopped looking read the same. The recovery is sent only when the condition has cleared: an alert still on when the data ends has none. Open episodes are stored durably, or a restart sends the alert again and never sends its recovery.
def messages(episodes, name: str, minutes_per_sample: int = 15) -> list:
"""One alert per episode and one recovery message when it clears, keyed so a resend is
recognised. The open episodes must be stored durably, or a restart re-sends the alert and
never sends its recovery."""
out = []
for a, b in episodes:
key = f"{name}:{a}"
out.append((key, "fire", f"{name}: condition met at sample {a}"))
if b is not None: # still on: no recovery until it clears
out.append((key, "clear", f"{name}: cleared at sample {b}, after {(b - a) * minutes_per_sample} minutes"))
return outdef heartbeat_ok(last_beat_received: float, now: float, every_s: float, missed: int = 2, slack_s: float = 30.0) -> bool:
"""A dead man's switch, run by a watcher on other infrastructure. Both times are the
watcher's own clock (when it received the last beat, and now), so the monitor's clock
cannot mislead it. It alerts once `missed` beats in a row are overdue, with some slack.
Silence from the monitor is then an alert, never an all-clear."""
return now - last_beat_received <= missed * every_s + slack_sSeverity and escalation
Severity decides who is interrupted and how fast; escalation decides what happens when nobody answers. Both are properties of the alert rule, decided in advance.
| Level | Means | Delivered | If nobody acknowledges |
|---|---|---|---|
| Act now | Money is at risk before the next decision: wrong positions, orders out of control, the kill switch tripped | Phone, with sound, at any hour | Second channel after 15 minutes, second person after 30 |
| Act before the next session | The next session will be wrong unless fixed: stale data, a failed snapshot, a check unavailable | Message, in the evening | Promoted to act now an hour before the session |
| Digest | Worth knowing: slippage above the model, a slow feed, a warning | The daily report | Nothing; reviewed weekly |
The timings illustrate a policy; they are not a recommendation for any particular operation. Rules should know the market calendar: a stale price on a holiday is not an incident.
Delivery you can prove
An alert that was generated but not delivered is worse than none, because the record says someone was told. We write every alert to a durable outbox before sending it and deliver from the outbox: retries with exponential backoff until the channel returns a receipt, and after the last retry the next channel. Hohpe and Woolf describe the pattern as guaranteed delivery through a persistent channel;9 it delivers at least once, so every alert carries an id and the receiving end ignores a duplicate. A receipt says the message arrived; only an acknowledgement says a person has it, and the two need separate timers.
class Outbox:
"""Alerts are written to a durable table first and delivered from it. Each pass tries every
pending alert that is due, on its current channel; a failure (or an exception) schedules a
retry with exponential backoff, and after `max_tries` the alert moves to the next channel
at once. A delivered alert has a receipt; an alert that exhausts every channel is marked
undeliverable, which is itself something to alert on by another route. A person's
acknowledgement is a separate matter, with its own escalation timer."""
def __init__(self, channels, path: str = ":memory:", max_tries: int = 3, base_delay_s: float = 30.0):
self.channels, self.max_tries, self.base = channels, max_tries, base_delay_s
self.db = sqlite3.connect(path)
self.db.execute("CREATE TABLE IF NOT EXISTS outbox (id TEXT PRIMARY KEY, msg TEXT, channel INT DEFAULT 0,"
" tries INT DEFAULT 0, due REAL, state TEXT DEFAULT 'pending', receipt_at REAL)")
def put(self, alert_id: str, message: str, now: float) -> None:
"""Idempotent: the same id twice is one alert."""
self.db.execute("INSERT OR IGNORE INTO outbox (id, msg, due) VALUES (?, ?, ?)", (alert_id, message, now))
self.db.commit()
def flush(self, now: float) -> None:
due = self.db.execute("SELECT id, msg, channel, tries FROM outbox WHERE state = 'pending' AND due <= ?"
" ORDER BY due", (now,)).fetchall()
for alert_id, msg, ch, tries in due:
try:
ok = bool(self.channels[ch](alert_id, msg)) # True only with a receipt
except Exception: # noqa: BLE001
ok = False
if ok:
self.db.execute("UPDATE outbox SET state = 'delivered', receipt_at = ? WHERE id = ?", (now, alert_id))
continue
tries += 1
if tries < self.max_tries:
self.db.execute("UPDATE outbox SET tries = ?, due = ? WHERE id = ?",
(tries, now + self.base * 2 ** (tries - 1), alert_id))
elif ch + 1 < len(self.channels): # escalate to the next channel now
self.db.execute("UPDATE outbox SET channel = ?, tries = 0, due = ? WHERE id = ?", (ch + 1, now, alert_id))
else:
self.db.execute("UPDATE outbox SET state = 'undeliverable' WHERE id = ?", (alert_id,))
self.db.commit()| Channels | Outcome | Channel used | Delivered after |
|---|---|---|---|
| Primary up | delivered | primary | 0 s |
| Primary flaky | delivered | primary | 90 s |
| Primary times out | delivered | secondary | 100 s |
| Both channels down | undeliverable | — | never |
From delivery_demo in alerting.py: one alert, three tries per channel with backoff from 30 seconds, the outbox flushed every 10 seconds.
The delivery path needs its own test, and the simplest is to use it: a scheduled test alert, sent through every channel at a fixed time, whose receipt is itself checked. A channel that has quietly stopped working is then found on a calm afternoon.
Measuring and retiring alerts
Every alert rule is a hypothesis: this condition, at this level, needs a person. Its history tests the hypothesis. For each rule we keep what fired, whether it was real, whether anyone acted, and how long the action took. From that record come the rule’s precision, its recall against incidents found by other means, and its action rate. A rule whose alerts are rarely acted on is either wrongly tuned or wrongly classified, and it is changed, demoted to the digest, or deleted. Google’s workbook derives alert rules from the error budget a service can afford, pairing a long window with a short one so that a fast burn pages quickly, a slow burn is still caught, and a brief blip that has already stopped does not page; the same reasoning applies to a tracking difference or a data delay.10
| Measure | From the history | Acts on |
|---|---|---|
| Precision | Alerts that told someone something new, out of all alerts | The threshold, the run length |
| Recall | Incidents alerted, out of incidents found any way | Missing rules, thresholds set too high |
| Delay | From the start of an incident to its alert | The run length, the sampling rate |
| Action rate | Alerts someone acted on | Rules that belong in the digest |
| Repeats | Alerts per incident beyond the first | Hysteresis, de-duplication |
Some incidents never raise an alert at all and are found in the next day’s reconciliation. Those are the misses that matter most, and they are the subject of Live versus backtest reconciliation.
References
- U.S. Securities and Exchange Commission (2013). In the matter of Knight Capital Americas LLC (Exchange Act Release 34-70694). Administrative proceeding. www.sec.gov/files/litigation/admin/2013/34-70694.pdf
- Ewaschuk, R. (2016). Monitoring distributed systems. In B. Beyer, C. Jones, J. Petoff, & N. R. Murphy (Eds.), Site reliability engineering: How Google runs production systems (ch. 6). O’Reilly Media. sre.google/sre-book/monitoring-distributed-systems/
- Ewaschuk, R. (n.d.). My philosophy on alerting. Google Docs. docs.google.com/document/d/199PqyG3UsyXlwieHaqbGiWVa8eMWi8zzAn0YfcApr8Q/edit
- Axelsson, S. (2000). The base-rate fallacy and the difficulty of intrusion detection. ACM Transactions on Information and System Security, 3(3), 186–205. https://doi.org/10.1145/357830.357849
- Sendelbach, S., & Funk, M. (2013). Alarm fatigue: A patient safety concern. AACN Advanced Critical Care, 24(4), 378–386. https://doi.org/10.4037/nci.0b013e3182a903f9
- Cvach, M. (2012). Monitor alarm fatigue: An integrative review. Biomedical Instrumentation & Technology, 46(4), 268–277. https://doi.org/10.2345/0899-8205-46.4.268
- Page, E. S. (1954). Continuous inspection schemes. Biometrika, 41(1–2), 100–115. https://doi.org/10.1093/biomet/41.1-2.100
- Schmitt, O. H. (1938). A thermionic trigger. Journal of Scientific Instruments, 15(1), 24–26. https://doi.org/10.1088/0950-7671/15/1/305
- Hohpe, G., & Woolf, B. (2004). Enterprise integration patterns: Designing, building, and deploying messaging solutions. Addison-Wesley.
- Thurgood, S. (2018). Alerting on SLOs. In B. Beyer, N. R. Murphy, D. K. Rensin, K. Kawahara, & S. Thorne (Eds.), The site reliability workbook: Practical ways to implement SRE (ch. 5). O’Reilly Media. sre.google/workbook/alerting-on-slos/