Part 1 · 1 chapters · ~8 min
Percentiles and Tails
Why means mislead for skewed data, percentiles and how to compute them, histograms versus summaries, why percentiles cannot be averaged across servers, tail amplification under fan-out (1 - 0.99^n), hedged requests, SLOs on percentiles, and HDR histograms and t-digest.
2
Percentiles, and why tails multiply
code
const pct = (a: number[], p: number) => { const b = [...a].sort((x, y) => x - y); return b[Math.min(b.length - 1, Math.floor(p / 100 * b.length))]; };
// 100,000 simulated requests: mean 49.4 · p50 40.1 · p90 63.9 · p99 413.3 · p99.9 945.5 (ms)
// fan-out: probability a request waits on at least one p99-slow backend
1 - 0.99 ** 10 // 0.096
1 - 0.99 ** 100 // 0.634Never average percentiles: the mean of ten servers' p99s is not the fleet p99. Aggregate histograms (Prometheus histograms, HDR Histogram, t-digest, DDSketch) and compute percentiles from the merged distribution. Hedged requests send a second copy of a slow request after the p95 delay and take whichever returns first, cutting tail latency for a small amount of extra load.
ONE LATENCY DISTRIBUTION, FIVE SUMMARIES
100,000 simulated requests: log-normal body + 1% slow tail
swipe the figure sideways, or tap expand for full screen
1/4
the mean
The mean, 49.4 ms, describes no actual request: it sits between the typical 40 ms and the slow tail it averages in.
49.4 ms describes nobodyaverages hide shape