Benchmark Statistics
Sources of benchmark noise, warm-up and JIT effects, repetitions and interleaving, medians and percentiles over means, intervals and significance tests (Mann-Whitney U, bootstrap), tools that do this for you (hyperfine, benchstat, criterion, JMH, mitata), coordinated omission in load tests, and reporting results honestly.
Is it really faster?
# hyperfine: warm-up, many runs, mean ± σ, and a relative comparison hyperfine --warmup 3 --runs 30 'node old.js' 'node new.js' # Go: run each benchmark 10 times and compare distributions with a significance test go test -bench=. -count=10 > old.txt # then on the new code → new.txt benchstat old.txt new.txt # shows the change with a p-value, or "~" when not significant # Rust: criterion reports confidence intervals and detects regressions between runs cargo bench
Simulated: with 5% run-to-run noise and 10 runs per side, identical code showed an apparent difference of 5% or more in 2.6% of 10,000 comparisons. Combined with publication bias (only "wins" get posted), unverified micro-benchmarks are a reliable source of folklore. The Systems Performance course covers the measurement environment; load tests need constant-arrival-rate generators to avoid coordinated omission (Backend Disciplines P2).