Part 10 · 1 chapters · ~11 min

Disaster Recovery and Game Days

Recovery point and recovery time objectives agreed with the business, backups that are separate, immutable and actually restored, timed restore drills as the real RTO, failover with fencing, game days with a hypothesis and a stop button, and continuous chaos experiments with bounded blast radius.

15

Recovery objectives, restores and practising failure

systemRPORTOmechanismproof
ledger015 minsynchronous replica in another zone; async to another regionquarterly failover drill
customer documents1 h4 hversioned object storage replicated cross-regionmonthly sample restore
analytics24 h2 daysnightly snapshotsemi-annual restore
frontend assets0minutesimmutable builds in two buckets, CDN with origin failoverredeploy of a previous build
code
GAME DAY: bank adapter unavailable (staging, 45 minutes)
hypothesis   transfers are accepted as pending within 2 s; no duplicates; users see "pending";
             the bank-adapter alert fires within 5 min; the runbook gets on-call to the breaker dashboard
inject       network policy blocks egress from bank-adapter pods, 14:00
observers    SRE (timeline), frontend (UI states), product (customer comms)
abort if     any ledger mismatch, or staging is shared with a demo
results      pending state correct; alert fired at 7 min (threshold too slow: action item);
             UI showed a spinner for 30 s before "pending" (frontend timeout too long: action item);
             runbook link pointed to an old dashboard (fixed during the exercise)
why game days find frontend bugs
Backend teams run them for their own systems, and the client is where the hypothesis most often fails: a timeout that is too long, a missing pending state, a retry that double-submits. Ask to be an observer.
DISASTER RECOVERY AND GAME DAYS
RPO and RTO, backups that are actually restored, failover, and practising failure on purpose
swipe the figure sideways, or tap expand for full screen
1/6
RPO and RTO
RPO and RTO: a ledger might need RPO zero (no committed transaction lost) and RTO 15 minutes; an analytics store might accept RPO 24 hours and RTO two days. The numbers set the cost: RPO zero needs synchronous replication; a long RTO can be met with restores from backup.