Part 7 · 1 chapters · ~8 min
Queueing Theory and Little's Law
Arrival rate, service time and utilisation, the M/M/1 queue and the 1/(1-ρ) curve, multiple servers (M/M/c) and why pooling helps, variability makes queues worse (Kingman), Little's law for sizing pools and concurrency limits, load shedding and admission control, and the universal scalability law.
8
Queues explode near 100%
code
// M/M/1: mean time in system W = S / (1 - ρ), ρ = λS for (const rho of [0.5, 0.7, 0.8, 0.9, 0.95, 0.99]) console.log(rho, 10 / (1 - rho)); // 20 33 50 100 200 1000 ms // Little's law: L = λ × W (holds for any stable system, any distribution) const inFlight = 2000 /* req/s */ * 0.050 /* s */; // 100 concurrent requests // a pool of 20 DB connections with 25 ms queries supports at most 20 / 0.025 = 800 queries/s // Kingman: waiting ∝ ρ/(1-ρ) × (Ca² + Cs²)/2 × S → burstier arrivals or more variable service = longer queues
Practical rules: keep steady-state utilisation of latency-sensitive resources well below the knee (often 60-70%); pool servers rather than giving each a separate queue (one queue to many servers beats many queues); cut variability (split slow and fast work into different queues); and when overloaded, reject early (load shedding) instead of letting queues grow.
LATENCY VERSUS UTILISATION (M/M/1)
service time 10 ms; mean time in system = S / (1 - ρ)
swipe the figure sideways, or tap expand for full screen
1/4
half busy
At 50% utilisation a 10 ms request takes 20 ms on average: half of it waiting in the queue.
ρ 0.5: 2× service timewaiting already matters