Part 10 · 1 chapters · ~10 min
Measuring Performance
How to get performance numbers you can trust: JIT warm-up measured at 250 µs cold against 6 µs warm on this machine, variance and repeated runs, percentiles instead of averages, sampling profilers, reading flame graphs, and a one-change-at-a-time method on the hardware users actually carry.
16
Numbers you can trust
code
// the warm-up measurement from this chapter (node, Apple M3 Pro)
function work(arr: number[]) { let s = 0; for (let i = 0; i < arr.length; i++) s += arr[i] * 2 + (arr[i] >> 1); return s; }
const arr = Array.from({ length: 10_000 }, (_, i) => i);
const us: number[] = [];
for (let k = 0; k < 2000; k++) {
const t = process.hrtime.bigint(); work(arr); us.push(Number(process.hrtime.bigint() - t) / 1000);
}
// first call ≈ 250 µs · median of calls 1-10 ≈ 33 µs · median of calls 1000-2000 ≈ 6 µs (p99 ≈ 7 µs)code
# profiling a Node script and opening the result in DevTools node --cpu-prof --cpu-prof-dir=prof script.mjs # writes a .cpuprofile # DevTools → Performance → Load profile → Bottom-up for self time, flame chart for shape # in the browser: Performance panel → CPU 4x slowdown → record the interaction # then: what is wide that I did not expect?
| mistake | effect | fix |
|---|---|---|
| timing one cold run | measures the interpreter, not the optimised code | warm up, or report cold separately on purpose |
| reporting the average | hides the tail users feel | median plus p95 and p99 |
| dead code in a microbenchmark | the JIT removes the work you meant to measure | use the result (sum it, return it) |
| optimising without a profile | effort on the wrong function | profile first; fix the widest unexpected bar |
| testing only on a fast laptop | problems invisible until users report them | CPU throttling and a real low-end phone |
MEASURING PERFORMANCE PROPERLY
warm-up, variance, percentiles, profilers and flame graphs: how to get numbers you can trust
swipe the figure sideways, or tap expand for full screen
1/6
warm-up
Warm-up: JavaScript engines interpret first, then compile hot functions with optimising compilers. Measured on this machine, a 10,000-element loop function took about 250 µs on its first call, a median of about 33 µs over the first ten calls, and about 6 µs once warm. Benchmark the state your users will see, and say which one you measured.