10 parts · 22 chapters
Scaling Databases
Scaling a database is a sequence of decisions, each with a cost you keep paying. The order matters: measure, fix queries, scale up, pool connections, cache, add replicas, and only then partition or shard. This course teaches that sequence as a discipline and links into the engine courses rather than re-teaching them.
Ten parts: knowing your limits; query and schema scaling; connections and concurrency; caching; read scaling; write scaling and sharding with Vitess and Citus; multi-region and the NewSQL systems; event-driven and streaming data; reliability at scale; and a written decision framework as the capstone.
limits firstGolden signals, capacity models, and scaling up before out.
connectionsPooling architectures, transaction pooling, serverless storms, admission control.
cachingCache-aside to write-behind, invalidation, stampedes, Redis, the dual-write problem.
readsReplicas, lag-aware routing, read-your-writes, CQRS.
writesSharding strategies, shard keys, live resharding, Vitess and Citus internals.
globalMulti-region topologies, Spanner, CockroachDB, Aurora, Neon, data residency.
00
Know Your Limits First
Golden signals and the ladder · Capacity modelling and the single-node ceiling
2 ch · ~12 min01Query and Schema Scaling
The slow-query pipeline and index strategy · Schema for scale: summaries, counters and tiers · Big deletes and online schema change
3 ch · ~18 min02Connection and Concurrency Scaling
The connection problem and pooling architectures · Pooling modes and what breaks
2 ch · ~12 min03Caching
Topology and patterns · Invalidation, stampedes and Redis · Consistency between cache and database
3 ch · ~18 min04Read Scaling
Replicas, routing and read-your-writes · Geography, fan-out and CQRS
2 ch · ~12 min05Write Scaling and Sharding
Decompose first, then choose a strategy and a key · Resharding live, cross-shard queries and transactions · Vitess, Citus and where sharding lives
3 ch · ~18 min06Multi-Region and Global
Topologies and conflict resolution
1 ch · ~8 min07Event-Driven and Streaming Data
CDC, Debezium and the outbox · Kafka as a log, stream processing and the OLAP split
2 ch · ~12 min08Reliability at Scale
HA, backups and DR at scale · Chaos tests and the incidents that happen
2 ch · ~12 min09The Scaling Decision Framework
The decision tree · Capstone: a scaling RFC
2 ch · ~12 minBuilt on the engine coursesAssumes MySQL Internals, PostgreSQL Internals and SQL; links to Cassandra and MongoDB, Redis and Kafka where they come up, and to Distributed Systems for the theory.