Read Scaling
Read replicas and the routing layer, replica-lag-aware routing and read-your-writes guarantees, geographic replicas and the latency map, read amplification and fan-out control, and CQRS with its synchronisation mechanism.
Replicas, routing and read-your-writes
| routing approach | how | trade-off |
|---|---|---|
| in the app | two pools (primary, replicas); the repository chooses per query | explicit and testable; every query needs a decision |
| in a proxy | ProxySQL query rules, Pgpool, Vitess vtgate | transparent; hard to express read-your-writes |
| in the driver | multi-host connection strings with target_session_attrs | simple failover; not load balancing by itself |
Geography, fan-out and CQRS
Geographic replicas put read copies near users (Lagos, London, Virginia): reads drop from 150 ms to 10 ms, writes still cross the ocean to the primary. Read amplification appears when one page triggers dozens of queries, or one query fans out to every shard; control it with batching, denormalised read models and limits on page sizes.
CQRS: separate the write model from the read model
writes → transfers (normalised, constrained, on the primary)
→ outbox row in the same transaction
→ CDC / outbox relay → Kafka
→ consumer builds account_activity_view (denormalised, per screen) in a read store
reads → account_activity_view (fast, eventually consistent, rebuildable from the log)CQRS earns its complexity when the read shape differs a lot from the write shape (dashboards, search, feeds) or read volume dwarfs writes. The read model is derived and disposable: if it breaks, replay the log and rebuild it.