Back to the blog
  • monitoring
  • on-call

Every alert should be worth waking someone up for

Alert fatigue isn’t a discipline problem. It’s a design problem. Here’s the review we run to cut the noise without going blind.

Most on-call engineers can tell you which alerts they ignore. They’ll say it with a shrug: “that one fires every Tuesday”, “that one resolves itself”, “that one’s been broken since the migration”. Nobody decided to ignore them. The team just learned, one false page at a time, that the pager lies.

That’s the real cost of alert noise. Lost sleep matters. The worse part is the day a real incident fires and gets the same shrug.

Symptoms, not causes

The most common source of noise we see is alerting on causes instead of symptoms.

CPU at 85%, a pod restart, a queue depth above some number picked two years ago: these are potential causes of user pain. Sometimes they matter. Often they don’t. A service can run at 90% CPU all day and serve every request under budget.

Symptoms are what users actually feel:

  • requests failing
  • requests getting slow
  • work not getting done (jobs stuck, messages not processed)

If you page on symptoms, you page when it matters. Causes belong on dashboards, where the on-call engineer looks after being paged, to work out why.

The four questions

For every alert that can page a human, we ask four questions. If the answer to any of them is “no”, the alert gets fixed, downgraded to a ticket, or deleted.

  1. Is a user affected, or about to be? If not, it isn’t urgent.
  2. Does it need a human now? If it can wait until morning, it’s a ticket, not a page.
  3. Is there a clear first action? An alert without a runbook link is a puzzle, not an alert.
  4. Would we notice if it didn’t exist? If it has never led to a real fix, it’s decoration.

Page on error budget burn

Once you have an SLO for a user journey, you can stop guessing thresholds. Alert on how fast you’re burning through the error budget instead. A fast burn pages; a slow burn opens a ticket.

Here’s a simplified multi-window burn-rate alert, in Prometheus syntax, for a 99.9% availability SLO:

- alert: CheckoutErrorBudgetFastBurn
  expr: |
    (
      sum(rate(http_requests_total{service="checkout",code=~"5.."}[1h]))
      / sum(rate(http_requests_total{service="checkout"}[1h]))
    ) > (14.4 * 0.001)
    and
    (
      sum(rate(http_requests_total{service="checkout",code=~"5.."}[5m]))
      / sum(rate(http_requests_total{service="checkout"}[5m]))
    ) > (14.4 * 0.001)
  labels:
    severity: page
  annotations:
    summary: Checkout is burning its error budget fast
    runbook: https://runbooks.internal/checkout/error-budget

The long window avoids paging on a blip. The short window makes the alert stop quickly once things recover. The runbook annotation is not optional.

Running the review

You don’t need a quarter-long project to fix this. A first pass looks like this:

Step What you do
Export Pull every alert that paged in the last 90 days, with fire count and who acked it.
Sort Order by fire count. The top of the list is usually a handful of alerts making most of the noise.
Ask Run the four questions on each, with the people who actually carry the pager in the room.
Act Fix, downgrade to ticket, or delete. Write down why.
Repeat Put a short alert review in the on-call handover, every rotation.

Deleting alerts feels risky. It’s worth saying out loud: an alert everyone ignores already provides zero protection. Removing it just makes that honest.

What good looks like

You’re in a good place when the on-call engineer trusts the pager. When it goes off, they get up. It’s almost always real, and it almost always tells them where to start.

Getting there is mostly alert design work. It’s also one of the first things we look at in an audit.

[OPEN] Free audit

If everyone ignores your alerts, your on-call isn’t protecting anything.

A free first audit. We find what’s breaking and tell you where to start.

Request a free audit