Part 5 · 1 chapters · ~8 min
On-Call Culture
Rotation sizes, follow-the-sun, page budgets and alert hygiene, handovers, escalation policies, incident roles (commander, communications, operations), compensation, blameless post-incident reviews, action-item follow-through, and on-call as a measure of system health.
6
Sustainable rotations
| incident role | job |
|---|---|
| incident commander | coordinates, decides, keeps the timeline; does not debug |
| operations lead | investigates and mitigates with responders |
| communications lead | status page, support, executives, regulators if money is affected |
| scribe | records actions and times for the review |
code
weekly on-call review (30 min): pages this week: 14 (target ≤ 2 per shift) → top source: payouts-dlq-depth (9 pages, 0 actionable) action: change payouts-dlq-depth to a ticket; page only on age of oldest DLQ message > 30 min open incident action items: 6 (2 overdue → escalate in staff meeting)
ON-CALL THAT PEOPLE CAN SUSTAIN
practices from SRE organisations
swipe the figure sideways, or tap expand for full screen
1/5
rotation
Too few people on a rotation means burnout and attrition. Google's SRE book recommends at least eight engineers for a single-site rotation, or six per site for two sites.
enough people8 single-site, 6 per site dual