Part 9 · 2 chapters · ~12 min

The Scaling Decision Framework

A written decision tree from symptom to diagnosis, intervention, cost and risk, and the capstone: a complete scaling plan for a realistic system in RFC form, defensible in review.

21

The decision tree

The framework is the course in one picture: measure, diagnose the actual bottleneck, choose the cheapest intervention that fixes it, price its cost and risk, and know how you will verify and undo it.

THE SCALING DECISION TREE
symptom, diagnosis, intervention, cost, risk
symptomp99 latency upCPU saturatedconnections exhaustedIO / cache missesreplica lagfix top queriesadd a poolerscale up RAMshard or decompose
swipe the figure sideways, or tap expand for full screen
1/5
symptom
Start from a user-visible symptom and a measurement, never from a technology ("we should use Cassandra").
start from a measured symptomnot from a technology
22

Capstone: a scaling RFC

code
RFC: scaling the loans ledger to 10× (template, filled in for a lending product)

1 context      2M monthly users, 16.7k queries/s peak, Postgres 17 primary + 2 replicas, 1.6 TB
2 measurements top 5 queries = 71% of DB time; peak CPU 78%; pool waits at month end; lag < 200 ms
3 goals        p99 < 150 ms at 10× peak; RPO 0, RTO 15 min; no cross-country data movement (NDPA)
4 options      A tune + scale up + PgBouncer + Redis   B + partition by month + CQRS read models
               C + regional sharding by country (Citus or app-level)
5 analysis     cost, risk, team effort, reversibility, operational load for each (a trade table)
6 decision     A now (6 weeks), B next quarter, C only if a third country launches
7 migration    expand/contract steps, backfills, verification queries, dual-run, cutover, rollback
8 risks        register (Systems Engineering P9): pooler as SPOF, cache staleness on balances (never cached)
9 success      metrics and the date we will check them