Part 8 · 2 chapters · ~12 min

Reliability at Scale

High-availability topologies and failover automation per engine, backup strategy at terabyte scale with restore time as the real metric, disaster recovery with RPO and RTO and tested restores, chaos testing the database tier, and runbooks for the incidents that actually happen.

19

HA, backups and DR at scale

engineHA and failoverbackup at scale
PostgresPatroni + etcd, or managed (RDS Multi-AZ, Cloud SQL HA)pgBackRest / WAL-G: parallel, incremental, PITR
MySQLOrchestrator, InnoDB Cluster, Vitess, or managedXtraBackup + binlogs for PITR
MongoDBreplica set electionssnapshots + oplog
Cassandrano failover: replicas per token rangeper-node snapshots + incremental, Medusa

At 5 TB, a single-threaded logical dump and restore can take more than a day. Restore time is the real metric: parallel physical backups, incremental strategies, and restore drills that measure the RTO you actually have (SRE part 10). An untested backup does not exist.

20

Chaos tests and the incidents that happen

incidentfirst moves in the runbook
connections exhaustedcheck pooler waiting clients and pg_stat_activity states; kill idle-in-transaction; find the slow query holding connections
replication lag climbingcheck long queries on the replica, heavy writes (bulk jobs), network; pause the bulk job
disk fillingWAL retained by slots? temp files from a bad query? bloat? Act on the cause, add space as a stopgap
CPU at 100%top queries by total time right now; a plan change after ANALYZE or a deploy?
lock pile-uppg_blocking_pids; the migration waiting for ACCESS EXCLUSIVE; cancel it
primary failureconfirm automation promoted; verify apps reconnected; check data loss window

Chaos testing the data tier: kill the primary in staging and time failover; inject replica lag and watch read-your-writes; fill a disk in a sandbox; throttle IO. Each experiment has a hypothesis and a stop condition.