Part 7 · 1 chapters · ~8 min

Queueing Theory and Little's Law

Arrival rate, service time and utilisation, the M/M/1 queue and the 1/(1-ρ) curve, multiple servers (M/M/c) and why pooling helps, variability makes queues worse (Kingman), Little's law for sizing pools and concurrency limits, load shedding and admission control, and the universal scalability law.

8

Queues explode near 100%

code
// M/M/1: mean time in system W = S / (1 - ρ),  ρ = λS
for (const rho of [0.5, 0.7, 0.8, 0.9, 0.95, 0.99]) console.log(rho, 10 / (1 - rho));   // 20 33 50 100 200 1000 ms

// Little's law: L = λ × W (holds for any stable system, any distribution)
const inFlight = 2000 /* req/s */ * 0.050 /* s */;     // 100 concurrent requests
// a pool of 20 DB connections with 25 ms queries supports at most 20 / 0.025 = 800 queries/s
// Kingman: waiting ∝ ρ/(1-ρ) × (Ca² + Cs²)/2 × S  → burstier arrivals or more variable service = longer queues

Practical rules: keep steady-state utilisation of latency-sensitive resources well below the knee (often 60-70%); pool servers rather than giving each a separate queue (one queue to many servers beats many queues); cut variability (split slow and fast work into different queues); and when overloaded, reject early (load shedding) instead of letting queues grow.

LATENCY VERSUS UTILISATION (M/M/1)
service time 10 ms; mean time in system = S / (1 - ρ)
ρ = 0.520 msρ = 0.733 msρ = 0.850 msρ = 0.9100 msρ = 0.95200 msρ = 0.991,000 ms
swipe the figure sideways, or tap expand for full screen
1/4
half busy
At 50% utilisation a 10 ms request takes 20 ms on average: half of it waiting in the queue.
ρ 0.5: 2× service timewaiting already matters