Skip to content

Production operations · Intermediate

Alerting that people act on

An alert rule is a classifier: score it on what it catches, what it misses, how late it speaks and how often it cries wolf.

15 min read10 referencesIntroduction to the series

Educational material only. Not investment advice.

Contents
  1. 01An alert names an action
  2. 02The arithmetic of alert fatigue
  3. 03Thresholds, run length and hysteresis
  4. 04One incident, one alert
  5. 05Severity and escalation
  6. 06Delivery you can prove
  7. 07Measuring and retiring alerts
  8. —References

Abstract

An alert asks someone to act, so each one should say what is wrong, what to do and when it clears. When incidents are rare, even a good detector’s alerts are mostly false; Bayes’ rule shows that the rate of false alarms on quiet checks decides how many can be believed. On a synthetic metric with known incidents, a run length and hysteresis raise the share of useful alerts from about a sixth to about two fifths while catching almost as many incidents, at a cost of twenty minutes to half an hour of delay. Alerts go out through a durable outbox with receipts and escalation, a heartbeat catches a monitor that has stopped, and each rule is kept or retired on its record.

Key takeaways

  • An alert needs someone to act now; everything else belongs in a digest, and each alert names what is wrong, why it matters, what to do and when it clears.
  • When incidents are rare, the false alarm rate on quiet checks sets the share of alerts that are real.
  • A run length and hysteresis cut false and repeated alerts from a bare threshold; the price is delay, and a fast alert needs a faster metric.
  • One incident, one alert, one recovery message, with open episodes stored durably; a watcher elsewhere makes a silent monitor an alert.
  • Deliver from a durable outbox with retries, receipts and escalation, and time acknowledgements separately.
  • Keep or retire each rule on its precision, recall, delay and action rate.

Before you start

  • Monitoring a live book, for the checks whose results become alerts
  • Conditional probability, for Bayes’ rule

On the morning Knight Capital lost more than 460 million dollars, its systems sent 97 emails warning of the fault before the market opened. They had not been designed as alerts, and the staff who received them generally did not read them.1 An alert rule is a classifier, and it can be scored like one: how often it is right, how much it misses, and how late it speaks.

An alert names an action

The daily checks in Monitoring a live book produce a report; only some of its results should interrupt a person. The test is simple: does this need someone to act, now? If yes, it is an alert. If it can wait for the next session, it belongs in a digest. If nobody would do anything differently, it belongs in neither. Google’s site reliability engineers put it in two rules: every page should be actionable, and every page should require intelligence, because a page that merits only a robotic response should have been automated.2,3 They also advise alerting on symptoms, what is broken for the user, rather than on every cause that might explain it.

FieldExampleWhy it is there
What is wrongHoldings differ from targets in two instrumentsA symptom in the reader’s terms, not a stack trace
Why it mattersThe book is well over its intended exposureLets the reader judge urgency without opening a dashboard
What to doRunbook: compare the broker’s orders with the journal; cancel duplicatesAn alert without an action is a report
Owner and severityOn call; act before the next sessionSomeone specific, with a deadline
EvidenceThe check’s witness: what it read, and whenSo the reader can trust it without rerunning it
When it clearsWhen holdings equal targets for one full checkSo the recovery is announced, not assumed

An alert in this shape can be acted on from a phone. The examples are illustrative.

The arithmetic of alert fatigue

Most alerts from a reasonable detector are false, and the reason is arithmetic. If an incident is under way on a fraction b of checks, a check made during an incident alerts with probability s, and a quiet check alerts with probability f, the share of alerts that are real is

P(incident | alert)  =  s·b / ( s·b + f·(1 − b) )
the precision of an alert, by Bayes’ rule
Step by stepA month of checks, and why most alerts are false
One month, a check every 15 minutes, an incident under way on 1 in 1,000 checks2,880checks a month2.9during an incident2,877quiet2.6real alerts (90%)28.8false alerts (1 in 100)Alerts a month: real and false90% detection, 1 in 100 quiet checks alert8% real100% detection, 1 in 100 quiet checks alert9% real90% detection, 1 in 1,000 quiet checks alert47% real90% detection, 1 in 10,000 quiet checks alert90% real
  1. Step 1 of 4

    A month of checks

    A check every 15 minutes makes 2,880 checks a month. With an incident under way on 1 in 1,000 checks, about 2.9 of them fall during an incident, and a detector with s at 90% alerts on 2.6. Those are the alerts we want.

  2. Step 2 of 4

    The quiet checks

    The other 2,877 checks are quiet. If 1 quiet check in 100 alerts, they send 28.8 false alerts a month, about 11 for every real one, and only 8% of alerts are real.

  3. Step 3 of 4

    Perfect detection

    A detector that never misses raises the real alerts from 2.6 to 2.9 a month. The share that is real moves from 8% to 9%, because the false alarms are untouched.

  4. Step 4 of 4

    Quieter checks

    If only 1 quiet check in 1,000 alerts, 47% of alerts are real; at 1 in 10,000, 90%. When incidents are rare, the false alarm rate on quiet checks decides how many alerts can be believed.

Arithmetic, not a simulation: Bayes’ rule with a check every 15 minutes, as in the laboratory below; the shares are those of the base-rate table in alerting.py. Without script, or with reduced motion, the figure shows its final state.

Axelsson made the same point about intrusion detection: when the base rate is low, the false alarm rate is the limiting factor.4 Hospitals learned it at a cost: Sendelbach and Funk report that between 72% and 99% of clinical alarms are false, and they and Cvach describe staff who, overwhelmed, silence or ignore them.5,6

Figure 1The base-rate arithmetic of alerts

Of the alerts, 8% come from a real incident. Checked every 15 minutes, that is about 2.6 alerting checks during incidents and 28.8 false ones a month.

Alerts that are real
8%
P(incident | alert), Bayes’ rule
Real a month
2.6
90% of the checks made during incidents
False a month
28.8
2,880 checks a month
A hundred alerts
Real incidentFalse alarm
Arithmetic, not a simulation: a check every 15 minutes, with the rates set above. It opens at the prose’s first example. Move the false alarm rate from 1 in 100 to 1 in 1,000, then set detection to 100% and see how little that changes.

Thresholds, run length and hysteresis

Most alert rules compare a metric with a threshold, and a bare threshold is a poor rule. Noise that crosses it for one sample fires an alert; a real problem that hovers around it fires one alert after another as it crosses back and forth. Two old ideas fix most of this. A run length asks for the condition to hold for several consecutive samples before firing, trading delay for noise; Page’s cumulative-sum schemes are the principled version of the same trade.7 A threshold with hysteresis clears only when the metric falls some distance below the level that set it, which is how Schmitt’s trigger stopped a noisy signal from chattering in 1938.8 In the listing, the run length applies to clearing as well.

site_research/fieldnotes/alerting.py
def detect(x, threshold: float, run: int = 1, hysteresis: float = 0.0):
    """Alert episodes as (start, end). An alert fires once the metric has been at or above
    `threshold` for `run` consecutive samples, and clears only after it has been below
    `threshold - hysteresis` for `run` consecutive samples. One episode is one alert: no
    repeats while it stays on. `end` is None for an alert still on when the data ends."""
    out, on, count, start = [], False, 0, 0
    for t, v in enumerate(x):
        if not on:
            count = count + 1 if v >= threshold else 0
            if count >= run:
                on, start, count = True, t, 0
        else:
            count = count + 1 if v < threshold - hysteresis else 0
            if count >= run:
                out.append((start, t))
                on, count = False, 0
    if on:
        out.append((start, None))
    return out
Figure 2Tuning an alert on a metric with known incidents

64 alerts in six weeks, about 46 a month. 10 tell someone about a new incident; 39 are false and 15 repeat an incident already alerted. 10 of 12 incidents caught, on average 6 minutes after they began; 2 missed.

Alerts a month
46
64 in six weeks: 39 false, 15 repeats
Precision
16%
alerts that told someone something new
Incidents missed
2 of 12
recall 83%
Time to alert
6 min
after an incident starts, on average
The metric (in standard deviations), incidents shaded, alerts on top−20246day 1day 8day 15day 22day 29day 36day 42
Noise against misses, as the threshold moves024681012050100150alerts a monthincidents missed2.00 sd

Each point is a threshold, with the run length and hysteresis above. Moving left cuts noise; moving down cuts misses. A run length or hysteresis that pulls the solid curve towards the corner, against the dashed curve of a bare threshold, is buying both at once. No setting reaches the corner, with no noise and no misses: some incidents are smaller than the noise.

True alertFalse alertIncident (known, synthetic); outlined if missedThreshold; dashed, where it clearsBare threshold, for comparison
Synthetic: six weeks of a noisy, autocorrelated health metric sampled every 15 minutes, with 12 incidents of known start, length and size. An alert is true if it starts during an incident or within an hour after it; only the first alert on an incident counts towards precision. Start from a bare threshold of 2 sd, add a run length of 2 and a hysteresis of 1, then raise the threshold and watch the misses climb.

A bare threshold at 2 standard deviations fires 64 alerts in six weeks, about 46 a month. Only 10 of them tell anyone about a new incident: 39 are false and 15 repeat an incident already alerted, a precision of 16% (useful alerts over all alerts, so repeats count against it). The same threshold with a run length of 2 samples and a hysteresis of 1 sd sends 26 (38% useful) and still catches 10 of the 12 incidents. The price is time: the first alert arrives on average 39 minutes after an incident begins, against 6 minutes. Nudging the threshold to 2.25 standard deviations sends 19 alerts and still catches 10. Raising the threshold alone is the blunt fix: at 3.5 standard deviations 5 incidents are missed, and 11 of its 18 alerts are still repeats, because a higher threshold does nothing for flapping; hysteresis does, and so does a run length applied to clearing (at 3.5 standard deviations, a run of 2 cuts the repeats to 5 and a hysteresis of 1 sd to 3). An alert that must be fast needs a metric sampled faster, not a shorter run.

One path with 12 incidents is a small sample. Over 200 synthetic paths of the same model, the tuned rule’s precision beats the bare threshold’s on every one of them (40% against 16% on average). The misses are less clear-cut: it catches 11.1 incidents of 12 against 11.6, fewer on 45% of the paths, and its first alert comes 22 minutes later on average.

The scoring hides one subtlety: an alert already on when an incident begins tells nobody anything new, so it is scored false and the incident missed. Long, low-threshold alerts lose incidents this way. Choosing a point on the curve is a statement about which mistake costs more, and it should be written down.

One incident, one alert

One incident should produce one alert. An episode, from the moment a condition fires to the moment it clears, is the unit: while it lasts, the alert is not sent again. When several checks fire for one cause (a broker outage makes holdings, exposure and P&L all unavailable), a de-duplication key groups them under the symptom that matters, so the reader gets one message that lists the others. Every alert ends with a recovery message saying that the condition cleared, when, and how long it lasted; without it, a fixed problem and a monitor that stopped looking read the same. The recovery is sent only when the condition has cleared: an alert still on when the data ends has none. Open episodes are stored durably, or a restart sends the alert again and never sends its recovery.

site_research/fieldnotes/alerting.py
def messages(episodes, name: str, minutes_per_sample: int = 15) -> list:
    """One alert per episode and one recovery message when it clears, keyed so a resend is
    recognised. The open episodes must be stored durably, or a restart re-sends the alert and
    never sends its recovery."""
    out = []
    for a, b in episodes:
        key = f"{name}:{a}"
        out.append((key, "fire", f"{name}: condition met at sample {a}"))
        if b is not None:                                   # still on: no recovery until it clears
            out.append((key, "clear", f"{name}: cleared at sample {b}, after {(b - a) * minutes_per_sample} minutes"))
    return out
site_research/fieldnotes/alerting.py
def heartbeat_ok(last_beat_received: float, now: float, every_s: float, missed: int = 2, slack_s: float = 30.0) -> bool:
    """A dead man's switch, run by a watcher on other infrastructure. Both times are the
    watcher's own clock (when it received the last beat, and now), so the monitor's clock
    cannot mislead it. It alerts once `missed` beats in a row are overdue, with some slack.
    Silence from the monitor is then an alert, never an all-clear."""
    return now - last_beat_received <= missed * every_s + slack_s

Severity and escalation

Severity decides who is interrupted and how fast; escalation decides what happens when nobody answers. Both are properties of the alert rule, decided in advance.

LevelMeansDeliveredIf nobody acknowledges
Act nowMoney is at risk before the next decision: wrong positions, orders out of control, the kill switch trippedPhone, with sound, at any hourSecond channel after 15 minutes, second person after 30
Act before the next sessionThe next session will be wrong unless fixed: stale data, a failed snapshot, a check unavailableMessage, in the eveningPromoted to act now an hour before the session
DigestWorth knowing: slippage above the model, a slow feed, a warningThe daily reportNothing; reviewed weekly

The timings illustrate a policy; they are not a recommendation for any particular operation. Rules should know the market calendar: a stale price on a holiday is not an incident.

Figure 3The life of an alert
Alert lifecycle: condition, alert rule with run length, hysteresis and de-duplication, digest for the rest, outbox, delivery with retries and receipt, next channel, acknowledgement or escalation, recovery message; a heartbeat watcher alerts when the monitor stopsreceiptCheck resultfrom the monitorAlert rulerun, hysteresis, dedupDigesteverything elseOutboxdurableDeliveryretries, then next channelAcknowledgedor escalated on a timerRecovery messagecleared, and for how longHeartbeatevery monitor passWatcher, elsewhereown channel to a person
A condition becomes an alert only after the run-length rule and de-duplication; the rest goes to the digest. The alert is written to an outbox and delivered with retries until a receipt comes back, moving to the next channel if none does. A person’s acknowledgement has its own timer, and a missing one escalates. When the condition clears, a recovery message closes the episode. A watcher elsewhere alerts if the monitor’s heartbeat stops.

Delivery you can prove

An alert that was generated but not delivered is worse than none, because the record says someone was told. We write every alert to a durable outbox before sending it and deliver from the outbox: retries with exponential backoff until the channel returns a receipt, and after the last retry the next channel. Hohpe and Woolf describe the pattern as guaranteed delivery through a persistent channel;9 it delivers at least once, so every alert carries an id and the receiving end ignores a duplicate. A receipt says the message arrived; only an acknowledgement says a person has it, and the two need separate timers.

site_research/fieldnotes/alerting.py
class Outbox:
    """Alerts are written to a durable table first and delivered from it. Each pass tries every
    pending alert that is due, on its current channel; a failure (or an exception) schedules a
    retry with exponential backoff, and after `max_tries` the alert moves to the next channel
    at once. A delivered alert has a receipt; an alert that exhausts every channel is marked
    undeliverable, which is itself something to alert on by another route. A person's
    acknowledgement is a separate matter, with its own escalation timer."""

    def __init__(self, channels, path: str = ":memory:", max_tries: int = 3, base_delay_s: float = 30.0):
        self.channels, self.max_tries, self.base = channels, max_tries, base_delay_s
        self.db = sqlite3.connect(path)
        self.db.execute("CREATE TABLE IF NOT EXISTS outbox (id TEXT PRIMARY KEY, msg TEXT, channel INT DEFAULT 0,"
                        " tries INT DEFAULT 0, due REAL, state TEXT DEFAULT 'pending', receipt_at REAL)")

    def put(self, alert_id: str, message: str, now: float) -> None:
        """Idempotent: the same id twice is one alert."""
        self.db.execute("INSERT OR IGNORE INTO outbox (id, msg, due) VALUES (?, ?, ?)", (alert_id, message, now))
        self.db.commit()

    def flush(self, now: float) -> None:
        due = self.db.execute("SELECT id, msg, channel, tries FROM outbox WHERE state = 'pending' AND due <= ?"
                              " ORDER BY due", (now,)).fetchall()
        for alert_id, msg, ch, tries in due:
            try:
                ok = bool(self.channels[ch](alert_id, msg))          # True only with a receipt
            except Exception:                                        # noqa: BLE001
                ok = False
            if ok:
                self.db.execute("UPDATE outbox SET state = 'delivered', receipt_at = ? WHERE id = ?", (now, alert_id))
                continue
            tries += 1
            if tries < self.max_tries:
                self.db.execute("UPDATE outbox SET tries = ?, due = ? WHERE id = ?",
                                (tries, now + self.base * 2 ** (tries - 1), alert_id))
            elif ch + 1 < len(self.channels):                        # escalate to the next channel now
                self.db.execute("UPDATE outbox SET channel = ?, tries = 0, due = ? WHERE id = ?", (ch + 1, now, alert_id))
            else:
                self.db.execute("UPDATE outbox SET state = 'undeliverable' WHERE id = ?", (alert_id,))
        self.db.commit()
The outbox: one table, each pass trying what is due, backoff after a failure, the next channel after the last try, and an undeliverable state that is itself reported another way.
ChannelsOutcomeChannel usedDelivered after
Primary updeliveredprimary0 s
Primary flakydeliveredprimary90 s
Primary times outdeliveredsecondary100 s
Both channels downundeliverable—never

From delivery_demo in alerting.py: one alert, three tries per channel with backoff from 30 seconds, the outbox flushed every 10 seconds.

The delivery path needs its own test, and the simplest is to use it: a scheduled test alert, sent through every channel at a fixed time, whose receipt is itself checked. A channel that has quietly stopped working is then found on a calm afternoon.

Measuring and retiring alerts

Every alert rule is a hypothesis: this condition, at this level, needs a person. Its history tests the hypothesis. For each rule we keep what fired, whether it was real, whether anyone acted, and how long the action took. From that record come the rule’s precision, its recall against incidents found by other means, and its action rate. A rule whose alerts are rarely acted on is either wrongly tuned or wrongly classified, and it is changed, demoted to the digest, or deleted. Google’s workbook derives alert rules from the error budget a service can afford, pairing a long window with a short one so that a fast burn pages quickly, a slow burn is still caught, and a brief blip that has already stopped does not page; the same reasoning applies to a tracking difference or a data delay.10

MeasureFrom the historyActs on
PrecisionAlerts that told someone something new, out of all alertsThe threshold, the run length
RecallIncidents alerted, out of incidents found any wayMissing rules, thresholds set too high
DelayFrom the start of an incident to its alertThe run length, the sampling rate
Action rateAlerts someone acted onRules that belong in the digest
RepeatsAlerts per incident beyond the firstHysteresis, de-duplication

Some incidents never raise an alert at all and are found in the next day’s reconciliation. Those are the misses that matter most, and they are the subject of Live versus backtest reconciliation.

References

  1. U.S. Securities and Exchange Commission (2013). In the matter of Knight Capital Americas LLC (Exchange Act Release 34-70694). Administrative proceeding. www.sec.gov/files/litigation/admin/2013/34-70694.pdf
  2. Ewaschuk, R. (2016). Monitoring distributed systems. In B. Beyer, C. Jones, J. Petoff, & N. R. Murphy (Eds.), Site reliability engineering: How Google runs production systems (ch. 6). O’Reilly Media. sre.google/sre-book/monitoring-distributed-systems/
  3. Ewaschuk, R. (n.d.). My philosophy on alerting. Google Docs. docs.google.com/document/d/199PqyG3UsyXlwieHaqbGiWVa8eMWi8zzAn0YfcApr8Q/edit
  4. Axelsson, S. (2000). The base-rate fallacy and the difficulty of intrusion detection. ACM Transactions on Information and System Security, 3(3), 186–205. https://doi.org/10.1145/357830.357849
  5. Sendelbach, S., & Funk, M. (2013). Alarm fatigue: A patient safety concern. AACN Advanced Critical Care, 24(4), 378–386. https://doi.org/10.4037/nci.0b013e3182a903f9
  6. Cvach, M. (2012). Monitor alarm fatigue: An integrative review. Biomedical Instrumentation & Technology, 46(4), 268–277. https://doi.org/10.2345/0899-8205-46.4.268
  7. Page, E. S. (1954). Continuous inspection schemes. Biometrika, 41(1–2), 100–115. https://doi.org/10.1093/biomet/41.1-2.100
  8. Schmitt, O. H. (1938). A thermionic trigger. Journal of Scientific Instruments, 15(1), 24–26. https://doi.org/10.1088/0950-7671/15/1/305
  9. Hohpe, G., & Woolf, B. (2004). Enterprise integration patterns: Designing, building, and deploying messaging solutions. Addison-Wesley.
  10. Thurgood, S. (2018). Alerting on SLOs. In B. Beyer, N. R. Murphy, D. K. Rensin, K. Kawahara, & S. Thorne (Eds.), The site reliability workbook: Practical ways to implement SRE (ch. 5). O’Reilly Media. sre.google/workbook/alerting-on-slos/