Part 6 · 10 chapters · ~50 min

On-call: the human side of reliability

Every mechanism in the previous parts eventually resolves to a person being woken up. This part is about that person: how the rotation should be designed so they stay, what the page budget means as a limit rather than a metric, why the incident commander must not debug, and why blameless is a mechanism rather than a tone of voice.

57

What on-call is actually for

the question

“Why do we need someone awake at 3am if the system is automated?”

the honest answer, in three parts
  1. To limit damage from what automation cannot handle. Novel failures, ambiguous states, and decisions requiring judgement about money.
  2. To provide a feedback loop. The person woken is the person motivated to fix the cause, which is why on-call should be held by the team that writes the code.
  3. Not to compensate for a system that cannot look after itself. If on-call exists to restart things nightly, on-call is a symptom.
worked numbers
the test of a healthy rotation:

  most nights, nothing happens
  when something does, it is genuinely novel
  the responder has a runbook and authority to act
  the fix lands within the following week

an unhealthy rotation: the same page every week,
a runbook that says "restart it", and no time to fix the cause.
the framing that matters to the team
On-call is a cost the organisation pays, not a duty the engineer owes. That framing changes the decisions that follow: it means pages are budgeted, the rotation is compensated, the load is measured, and reducing pages is work that gets scheduled. Teams that treat on-call as an obligation accumulate toil until people leave, and the leaving is usually attributed to something else.
58

Rotation design, and the humane constraints

the constraints, which are not negotiable if you want people to stay
  1. A minimum of six people, ideally eight. With four, everyone is on-call a full week every month, which is not sustainable.
  2. Primary and secondary, with the secondary for escalation and for when the primary does not acknowledge.
  3. Compensated. Time off in lieu, payment, or both. Unpaid on-call is a hidden pay cut that falls hardest on people with caring responsibilities.
  4. Recovery time. Someone paged at 3am does not start at 9am. This has to be stated, or people will do it anyway and burn out quietly.
  5. Predictable and swappable. Published months ahead, with a simple swap process.
  6. Follow-the-sun if you can. Two timezones removes night pages entirely, and it is worth real organisational effort.
the equity point that is usually missed
On-call falls unequally. An engineer with young children, a long commute or a health condition pays a much higher price for the same rotation. Flexibility about swaps, genuine compensation, and not treating night availability as a proxy for commitment are what stop on-call quietly selecting for a narrow demographic. This is a management responsibility and Part 10 returns to it.
59

Severity levels that mean something

SevMeansResponseExample
Sev 1Money is wrong or at riskPage immediately, all handsInvariant violation, funds unaccounted
Sev 2Core function unavailable for manyPage immediatelyTransfers failing, card auth down
Sev 3Degraded, or a subset affectedPage in hoursOne rail down, elevated latency
Sev 4Minor or cosmeticTicketA dashboard is wrong
Sev 5InformationalTicketA certificate expires in 30 days
the bank-specific distinction
Sev 1 is reserved for correctness, not for availability. Transfers being down is bad and it is Sev 2: customers retry, money is safe, and the system recovers. An invariant violation means money may have been created or destroyed, which is unrecoverable in a different way and warrants a different response. Part 13 made the same point about error budgets: availability gets one, correctness does not.
what makes severity levels work
  1. Defined by customer impact, never by which component failed.
  2. Anyone can declare. Requiring permission to declare an incident delays every response and nobody does it.
  3. Easy to downgrade. If declaring high is punished, people declare low and the response is inadequate.
  4. Each level has a defined response, written down, so declaring a Sev 2 means specific things happen without anyone deciding.
60

The page budget, and treating it as a limit

worked numbers
target: fewer than 2 pages per on-call night.

above that:
  · responders stop reading pages carefully
  · the real one is missed
  · sleep debt accumulates and judgement degrades
  · people leave, and cite something else in the exit interview

alert fatigue is a DETECTION failure, not a morale problem.
what to do when the budget is breached
  1. Treat it as an incident in itself. Sustained breach triggers a review, exactly like an error budget breach.
  2. Categorise last month's pages: actionable and correct, actionable but should have been a ticket, not actionable, duplicate.
  3. Delete or demote everything not actionable. This is usually a third of them and it is the fastest win available.
  4. Fix the top repeat cause. One recurring page is usually a large fraction of the total.
  5. If the budget cannot be met, stop feature work until it can. Agreed in advance, per Part 0, so it is not a negotiation at the point of pain.
61

Handover, and the shift report

what a handover must transfer
  1. Anything currently degraded, and what is known about it.
  2. Anything deployed recently that might still bite, with links.
  3. Any suppressed or silenced alert, and when the silence expires. A forgotten silence is how a real incident goes undetected.
  4. Any in-flight manual intervention, such as a stuck payment being resolved by hand.
  5. Anything expected: a planned migration, a provider maintenance window, a known batch running long.
the failure handover prevents
The incident that was half-diagnosed by someone who has now gone to sleep. Without a handover, the next responder starts from zero on a problem that already has forty minutes of investigation behind it. Ten minutes of written handover is worth an hour of re-investigation, and it is the single cheapest process improvement in on-call.
62

Incident command: roles during an incident

RoleDoesDoes not
Incident commanderCoordinates, decides, tracks stateDebug. The moment the IC is in a terminal, coordination stops
Operations leadMakes the changes, hands onCommunicate externally
Communications leadUpdates stakeholders, status pageDebug
ScribeTimestamps everythingAnything else
Subject expertsAnswer questions, investigateAct without the IC knowing
why explicit roles matter more than they seem
  1. Without an IC, nobody is tracking the whole picture and three people investigate the same thing while nothing gets decided.
  2. The IC must not debug. The most common failure: the most knowledgeable person takes command and immediately disappears into the problem, leaving the incident uncoordinated.
  3. The scribe is undervalued. A timestamped log makes the postmortem accurate rather than reconstructed, and reconstruction is where blame creeps in.
  4. For small incidents one person holds several roles, and that is fine as long as the roles are named so the responsibilities are not dropped.
63

Communicating during an incident, internally and out

internal
  1. One channel, and say which. Splitting across three means nobody has the full picture.
  2. Regular updates even when nothing has changed. Silence is read as chaos, and people start asking, which costs the responders time.
  3. State what is known, what is not, and what happens next. Speculation in a shared channel becomes fact within twenty minutes.
  4. Protect the responders. The comms lead exists so executives asking for updates do not interrupt the people fixing it.
external
  1. Acknowledge early. Customers already know. Silence damages trust more than the outage.
  2. Say what is affected in their terms, not "the posting service is degraded" but "transfers to other banks are failing; your balance and card are unaffected".
  3. Never promise a time you do not know. "Next update in 30 minutes" is a promise you can keep; "fixed within the hour" is one you cannot.
  4. Regulatory obligations may have started. For a payment system, operational incidents often require notification within hours, and the clock starts at detection.
  5. Never speculate about cause externally while the incident is live.
the point about money specifically
If money is affected, say so, and say what you are doing about it. Customers forgive an outage; they do not forgive discovering later that their money was at risk and nobody told them. The Part 6 suspense account design means you can usually say "your money is in a known place and will complete or be returned", which is a genuinely reassuring and genuinely true thing to be able to say.
64

The postmortem, and making blamelessness real

what makes a postmortem useful
  1. A timeline with timestamps, from the scribe, not from memory.
  2. Impact quantified. How many customers, how much money, for how long. Vague impact produces vague prioritisation of the follow-ups.
  3. Contributing factors, plural. Never "root cause": real incidents have several, and picking one usually picks a person.
  4. What went well. Genuinely, because the detection or the rollback that worked is a control worth preserving.
  5. Follow-up actions with owners and dates, tracked like any other work.
  6. Published widely. The learning is the output, and a postmortem only the involved team reads has wasted most of its value.
worked numbers
blameless is a mechanism, not a tone of voice.

saying "we are blameless" and then asking
  "why did you deploy without checking?"
is blame with better manners.

the mechanism: ask what made the wrong action reasonable
  what did the dashboard show them?
  what did the runbook say?
  what did the tooling permit without warning?

the system allowed it. that is the finding.
the test
If the action item is "be more careful" or "add a review step", the postmortem failed. Those are not mechanisms and they do not survive contact with a busy week. "The deploy tool now blocks when the canary correctness metric moves" is a mechanism. The difference is whether the same mistake is still possible tomorrow.
65

Follow-up actions that actually get done

worked numbers
the usual fate of postmortem actions:

  week 1: urgent, everyone agrees
  week 3: deprioritised for a launch
  month 6: the same incident

an action list with no completion mechanism is a wish list,
and repeating an incident you already analysed is the worst outcome
available, because it means the cost was paid and nothing was bought.
what makes them land
  1. Fewer, prioritised. Three actions that get done beat twelve that do not. Ruthlessly cut to the ones that prevent recurrence.
  2. Each has one named owner, not a team.
  3. They go into the normal backlog with normal priority rules, not a separate list that nobody grooms.
  4. A Sev 1 or Sev 2 action blocks feature work for that team until done. Agreed in advance.
  5. Reviewed at the monthly scorecard, per Part 0. Visibility is most of the mechanism.
  6. Repeat incidents are tracked as a metric. A recurrence is a failure of the process, and it should be visible as one.
66

Measuring on-call health

MetricTargetSignals
Pages per night< 2Alert quality and system stability
Pages outside hours< 1 per shiftThe actual human cost
Actionable rate> 80%Alert precision. Below this, fatigue is setting in
Time to acknowledge< 5 minWhether the rotation is functioning
Repeat incidents< 10%Whether follow-ups are landing
Runbook coverage100% of pagesPreparedness
Toil share< 30%How much of on-call is manual work that should be automated
the metric to watch hardest
Actionable rate. Everything else can look acceptable while this quietly falls, and when it does, responders begin pattern-matching pages as noise. The first missed real incident after a period of low actionable rate is not bad luck; it is the predictable consequence, and it is visible months in advance if anyone is looking.