Reliability at Scale
High-availability topologies and failover automation per engine, backup strategy at terabyte scale with restore time as the real metric, disaster recovery with RPO and RTO and tested restores, chaos testing the database tier, and runbooks for the incidents that actually happen.
HA, backups and DR at scale
| engine | HA and failover | backup at scale |
|---|---|---|
| Postgres | Patroni + etcd, or managed (RDS Multi-AZ, Cloud SQL HA) | pgBackRest / WAL-G: parallel, incremental, PITR |
| MySQL | Orchestrator, InnoDB Cluster, Vitess, or managed | XtraBackup + binlogs for PITR |
| MongoDB | replica set elections | snapshots + oplog |
| Cassandra | no failover: replicas per token range | per-node snapshots + incremental, Medusa |
At 5 TB, a single-threaded logical dump and restore can take more than a day. Restore time is the real metric: parallel physical backups, incremental strategies, and restore drills that measure the RTO you actually have (SRE part 10). An untested backup does not exist.
Chaos tests and the incidents that happen
| incident | first moves in the runbook |
|---|---|
| connections exhausted | check pooler waiting clients and pg_stat_activity states; kill idle-in-transaction; find the slow query holding connections |
| replication lag climbing | check long queries on the replica, heavy writes (bulk jobs), network; pause the bulk job |
| disk filling | WAL retained by slots? temp files from a bad query? bloat? Act on the cause, add space as a stopgap |
| CPU at 100% | top queries by total time right now; a plan change after ANALYZE or a deploy? |
| lock pile-up | pg_blocking_pids; the migration waiting for ACCESS EXCLUSIVE; cancel it |
| primary failure | confirm automation promoted; verify apps reconnected; check data loss window |
Chaos testing the data tier: kill the primary in staging and time failover; inject replica lag and watch read-your-writes; fill a disk in a sandbox; throttle IO. Each experiment has a hypothesis and a stop condition.