Part 11 · 1 chapters · ~11 min

Readiness and On-call Health

Production readiness reviews for services and for frontend launches, measuring on-call load and keeping rotations healthy, and turning a quarter of incidents into a reliability roadmap argued in the business's units, with a worked readiness checklist for the non-indebtedness letter.

16

Before launch, during on-call, across quarters

code
PRODUCTION READINESS: Letter of Non-Indebtedness (self-serve)
[x] SLO: 99.5% of letter requests produce a correct PDF within 60 s (30-day window)
[x] alerts: burn rate 14.4x over 1 h and 6x over 6 h; queue age > 10 min
[x] runbook: stuck generation, wrong balance shown, PDF service down
[x] dependencies: loan ledger (read), PDF service, email; each with timeout and fallback state
[x] degraded: ledger slow → "we will email your letter"; never issue a letter on stale data
[x] flag: per-tenant rollout with kill switch; old manual path kept for 30 days
[x] client: error tracking with source maps; conversion and drop-off dashboard per step
[ ] restore drill for generated-letter archive (scheduled before GA)
[x] on-call: owning team rotation; support briefed with macros
on-call metrichealthyact when
actionable pages per shifta handful or fewerconsistently above the agreed limit
out-of-hours pagesraremore than one a week for a month
pages with no action needednear zeroany: fix or delete the alert
people in rotationsix to eight or morefewer than five
a staff artifact you can produce
Take one quarter of incidents and action items from your team, group them by theme, estimate what each theme cost, and propose two projects. That one page is a better staff signal than any individual fix, because it shows you can see the system and choose.
PRODUCTION READINESS AND ON-CALL HEALTH
the checklist before launch, the health of the rotation after it, and reliability as a staff engineer's roadmap
swipe the figure sideways, or tap expand for full screen
1/6
readiness review
The readiness review: before a new service or major feature launches, the owning team answers a short checklist with a reliability engineer: SLOs defined, dashboards and alerts in place, runbooks written, dependencies and their failure behaviour known, capacity tested, rollback rehearsed, on-call staffed, data backed up and restorable.