Part 4 · 1 chapters · ~12 min
Postmortems
The blameless postmortem: a UTC timeline with what was known, impact in user and budget terms, contributing factors rather than a root cause, language about what the system allowed, three to five owned and dated action items, and the review loop that closes them, with a complete worked example.
7
Blameless practice and action items that get done
learning, written down, acted on
- Timeline in UTC, recording what was known at each moment.
- Impact in user terms and in error budget.
- Contributing factors rather than a single root cause.
- Blameless language: describe what the system allowed.
- Three to five action items, owned and dated: prevent, detect, mitigate, process.
- Close the loop: track them, review them, look for patterns, share the write-up.
code
# POSTMORTEM: Transfers failing after 4.18 deploy (SEV2) · 2026-10-06 Owner: Bayo · Reviewers: payments, platform · Status: actions in progress ## Summary A synchronous call to the fee service, added in 4.18 without a timeout, failed when the fee service slowed under load. 3,412 transfers failed and were auto-reversed over 41 min. ## Impact 3,412 transfers · 41 min · 18% of monthly budget · ₦0 lost · 212 contacts ## Detection burn-rate page 6 min after impact began (fast-burn window) ## Timeline (UTC) 13:58 deploy · 14:04 page · 14:09 SEV2 · 14:16 rollback · 14:21 baseline ## Contributing factors - New synchronous dependency on the fee service in the transfer path - Fee client had no timeout (default: none) - Canary skipped: change labelled config-only, contained code - Fee service had no latency SLI; its slowdown was invisible ## What went well rollback took 5 min; auto-reversal protected every customer ## Action items | type | action | owner | due | | prevent | default 2 s timeout on all outbound clients (template) | Chidi | 20 Oct | | detect | latency SLI + alert on fee service | Efe | 17 Oct | | mitigate | auto-rollback when fast-burn fires within 30 min of a deploy | Dami | 31 Oct | | process | canary skip requires a second approver | Bayo | 13 Oct |
THE BLAMELESS POSTMORTEM
from a timeline to contributing factors to action items that actually close
swipe the figure sideways, or tap expand for full screen
1/6
timeline
The timeline: reconstructed from the incident channel, deploy logs, alerts and dashboards, in UTC, with what was known at each moment: 13:58 deploy 4.18 begins; 14:04 burn-rate page fires; 14:09 commander declares SEV2; 14:16 rollback started; 14:21 error rate back to baseline. Detection time, time to mitigate and time to resolve fall out of it.