Part 2 · 2 chapters · ~16 min
Alerting
Paging on error budget burn with multi-window, multi-burn-rate alerts: fast and medium burns page, slow burns ticket, and causes move to dashboards; then a paging policy that every page must pass, runbooks, measuring fatigue, the weekly review, and the usual culprits.
4
Symptom-based alerting with burn rates
page on budget burn, with two windows
- Raw thresholds page on blips and miss slow burns.
- One window is either slow or noisy. Use two.
- Fast burn: 14.4× over 1 hour and 5 minutes. Page.
- Medium burn: 6× over 6 hours and 30 minutes. Page.
- Slow burn: 1× over 3 days and 6 hours. Ticket.
- Causes go to dashboards. The pager carries budget burn plus a few hard failures.
code
# Prometheus: the fast-burn page for a 99.9% SLO
groups:
- name: transfer-availability-slo
rules:
- record: slo:transfer_errors:ratio_rate1h
expr: sum(rate(lb_requests_total{route="/v1/transfers",code=~"5.."}[1h])) / sum(rate(lb_requests_total{route="/v1/transfers"}[1h]))
- record: slo:transfer_errors:ratio_rate5m
expr: sum(rate(lb_requests_total{route="/v1/transfers",code=~"5.."}[5m])) / sum(rate(lb_requests_total{route="/v1/transfers"}[5m]))
- alert: TransferBudgetFastBurn
expr: slo:transfer_errors:ratio_rate1h > (14.4 * 0.001) and slo:transfer_errors:ratio_rate5m > (14.4 * 0.001)
labels: { severity: page, team: payments }
annotations:
summary: "Transfers losing >2% of monthly error budget per hour"
runbook: "https://runbooks.acme.ng/payments/transfer-availability"MULTI-WINDOW, MULTI-BURN-RATE ALERTS
paging on budget burn with two windows per alert, so pages are fast for big problems, quiet for small ones, and reset quickly
swipe the figure sideways, or tap expand for full screen
1/6
raw thresholds
Why not "error rate > 1% for 5 minutes": a threshold on the raw rate pages on short blips that spend almost no budget, and misses a 0.5% error rate that quietly burns the month's budget in a week. The question an alert should answer is "are we losing budget fast enough that a human must act?"
5
Paging policy and alert fatigue
every page earns its place
- Four questions: urgent, actionable, user-visible, novel.
- Severities route to page, urgent ticket, ticket or dashboard.
- A runbook behind every page.
- Measure fatigue: actionable share, after-hours pages, repeat offenders.
- Review weekly: delete, tune, automate or fix.
- Common culprits and their standard fixes.
code
# Alertmanager: one page per incident, not twenty
route:
group_by: [team, slo]
group_wait: 30s
group_interval: 5m
repeat_interval: 3h
receiver: default
routes:
- matchers: ['severity="page"']
receiver: pagerduty-payments
- matchers: ['severity="ticket"']
receiver: jira-payments
inhibit_rules:
- source_matchers: ['alertname="PaymentsDown"'] # the big one
target_matchers: ['team="payments"'] # silences the symptoms it causes
equal: [cluster]PAGING POLICY AND ALERT FATIGUE
what deserves a page, the cost of a bad one, and the weekly review that keeps the pager honest
swipe the figure sideways, or tap expand for full screen
1/6
four questions
The four questions for every page: does it require action now (not in the morning)? Can the on-call engineer do something about it? Does it reflect user-visible pain, now or imminently? Is it novel, or is it the same automated fix every time (which should be automated, not paged)? Any "no" means it is not a page.