Part 4 · 1 chapters · ~12 min

Postmortems

The blameless postmortem: a UTC timeline with what was known, impact in user and budget terms, contributing factors rather than a root cause, language about what the system allowed, three to five owned and dated action items, and the review loop that closes them, with a complete worked example.

7

Blameless practice and action items that get done

learning, written down, acted on
  1. Timeline in UTC, recording what was known at each moment.
  2. Impact in user terms and in error budget.
  3. Contributing factors rather than a single root cause.
  4. Blameless language: describe what the system allowed.
  5. Three to five action items, owned and dated: prevent, detect, mitigate, process.
  6. Close the loop: track them, review them, look for patterns, share the write-up.
code
# POSTMORTEM: Transfers failing after 4.18 deploy (SEV2) · 2026-10-06
Owner: Bayo · Reviewers: payments, platform · Status: actions in progress

## Summary
A synchronous call to the fee service, added in 4.18 without a timeout, failed when the
fee service slowed under load. 3,412 transfers failed and were auto-reversed over 41 min.

## Impact            3,412 transfers · 41 min · 18% of monthly budget · ₦0 lost · 212 contacts
## Detection         burn-rate page 6 min after impact began (fast-burn window)
## Timeline (UTC)    13:58 deploy · 14:04 page · 14:09 SEV2 · 14:16 rollback · 14:21 baseline
## Contributing factors
- New synchronous dependency on the fee service in the transfer path
- Fee client had no timeout (default: none)
- Canary skipped: change labelled config-only, contained code
- Fee service had no latency SLI; its slowdown was invisible
## What went well    rollback took 5 min; auto-reversal protected every customer
## Action items
| type     | action                                              | owner | due    |
| prevent  | default 2 s timeout on all outbound clients (template) | Chidi | 20 Oct |
| detect   | latency SLI + alert on fee service                   | Efe   | 17 Oct |
| mitigate | auto-rollback when fast-burn fires within 30 min of a deploy | Dami | 31 Oct |
| process  | canary skip requires a second approver               | Bayo  | 13 Oct |
THE BLAMELESS POSTMORTEM
from a timeline to contributing factors to action items that actually close
swipe the figure sideways, or tap expand for full screen
1/6
timeline
The timeline: reconstructed from the incident channel, deploy logs, alerts and dashboards, in UTC, with what was known at each moment: 13:58 deploy 4.18 begins; 14:04 burn-rate page fires; 14:09 commander declares SEV2; 14:16 rollback started; 14:21 error rate back to baseline. Detection time, time to mitigate and time to resolve fall out of it.