Part 11 · 1 chapters · ~11 min
Readiness and On-call Health
Production readiness reviews for services and for frontend launches, measuring on-call load and keeping rotations healthy, and turning a quarter of incidents into a reliability roadmap argued in the business's units, with a worked readiness checklist for the non-indebtedness letter.
16
Before launch, during on-call, across quarters
code
PRODUCTION READINESS: Letter of Non-Indebtedness (self-serve) [x] SLO: 99.5% of letter requests produce a correct PDF within 60 s (30-day window) [x] alerts: burn rate 14.4x over 1 h and 6x over 6 h; queue age > 10 min [x] runbook: stuck generation, wrong balance shown, PDF service down [x] dependencies: loan ledger (read), PDF service, email; each with timeout and fallback state [x] degraded: ledger slow → "we will email your letter"; never issue a letter on stale data [x] flag: per-tenant rollout with kill switch; old manual path kept for 30 days [x] client: error tracking with source maps; conversion and drop-off dashboard per step [ ] restore drill for generated-letter archive (scheduled before GA) [x] on-call: owning team rotation; support briefed with macros
| on-call metric | healthy | act when |
|---|---|---|
| actionable pages per shift | a handful or fewer | consistently above the agreed limit |
| out-of-hours pages | rare | more than one a week for a month |
| pages with no action needed | near zero | any: fix or delete the alert |
| people in rotation | six to eight or more | fewer than five |
a staff artifact you can produce
Take one quarter of incidents and action items from your team, group them by theme, estimate what each theme cost, and propose two projects. That one page is a better staff signal than any individual fix, because it shows you can see the system and choose.
PRODUCTION READINESS AND ON-CALL HEALTH
the checklist before launch, the health of the rotation after it, and reliability as a staff engineer's roadmap
swipe the figure sideways, or tap expand for full screen
1/6
readiness review
The readiness review: before a new service or major feature launches, the owning team answers a short checklist with a reliability engineer: SLOs defined, dashboards and alerts in place, runbooks written, dependencies and their failure behaviour known, capacity tested, rollback rehearsed, on-call staffed, data backed up and restorable.