Part 9 · 2 chapters · ~12 min
The Scaling Decision Framework
A written decision tree from symptom to diagnosis, intervention, cost and risk, and the capstone: a complete scaling plan for a realistic system in RFC form, defensible in review.
21
The decision tree
The framework is the course in one picture: measure, diagnose the actual bottleneck, choose the cheapest intervention that fixes it, price its cost and risk, and know how you will verify and undo it.
THE SCALING DECISION TREE
symptom, diagnosis, intervention, cost, risk
swipe the figure sideways, or tap expand for full screen
1/5
symptom
Start from a user-visible symptom and a measurement, never from a technology ("we should use Cassandra").
start from a measured symptomnot from a technology
22
Capstone: a scaling RFC
code
RFC: scaling the loans ledger to 10× (template, filled in for a lending product)
1 context 2M monthly users, 16.7k queries/s peak, Postgres 17 primary + 2 replicas, 1.6 TB
2 measurements top 5 queries = 71% of DB time; peak CPU 78%; pool waits at month end; lag < 200 ms
3 goals p99 < 150 ms at 10× peak; RPO 0, RTO 15 min; no cross-country data movement (NDPA)
4 options A tune + scale up + PgBouncer + Redis B + partition by month + CQRS read models
C + regional sharding by country (Citus or app-level)
5 analysis cost, risk, team effort, reversibility, operational load for each (a trade table)
6 decision A now (6 weeks), B next quarter, C only if a third country launches
7 migration expand/contract steps, backfills, verification queries, dual-run, cutover, rollback
8 risks register (Systems Engineering P9): pooler as SPOF, cache staleness on balances (never cached)
9 success metrics and the date we will check them