Part 10 · 1 chapters · ~11 min
Disaster Recovery and Game Days
Recovery point and recovery time objectives agreed with the business, backups that are separate, immutable and actually restored, timed restore drills as the real RTO, failover with fencing, game days with a hypothesis and a stop button, and continuous chaos experiments with bounded blast radius.
15
Recovery objectives, restores and practising failure
| system | RPO | RTO | mechanism | proof |
|---|---|---|---|---|
| ledger | 0 | 15 min | synchronous replica in another zone; async to another region | quarterly failover drill |
| customer documents | 1 h | 4 h | versioned object storage replicated cross-region | monthly sample restore |
| analytics | 24 h | 2 days | nightly snapshot | semi-annual restore |
| frontend assets | 0 | minutes | immutable builds in two buckets, CDN with origin failover | redeploy of a previous build |
code
GAME DAY: bank adapter unavailable (staging, 45 minutes)
hypothesis transfers are accepted as pending within 2 s; no duplicates; users see "pending";
the bank-adapter alert fires within 5 min; the runbook gets on-call to the breaker dashboard
inject network policy blocks egress from bank-adapter pods, 14:00
observers SRE (timeline), frontend (UI states), product (customer comms)
abort if any ledger mismatch, or staging is shared with a demo
results pending state correct; alert fired at 7 min (threshold too slow: action item);
UI showed a spinner for 30 s before "pending" (frontend timeout too long: action item);
runbook link pointed to an old dashboard (fixed during the exercise)why game days find frontend bugs
Backend teams run them for their own systems, and the client is where the hypothesis most often fails: a timeout that is too long, a missing pending state, a retry that double-submits. Ask to be an observer.
DISASTER RECOVERY AND GAME DAYS
RPO and RTO, backups that are actually restored, failover, and practising failure on purpose
swipe the figure sideways, or tap expand for full screen
1/6
RPO and RTO
RPO and RTO: a ledger might need RPO zero (no committed transaction lost) and RTO 15 minutes; an analytics store might accept RPO 24 hours and RTO two days. The numbers set the cost: RPO zero needs synchronous replication; a long RTO can be met with restores from backup.