Part 10 · 1 chapters · ~10 min

Measuring Performance

How to get performance numbers you can trust: JIT warm-up measured at 250 µs cold against 6 µs warm on this machine, variance and repeated runs, percentiles instead of averages, sampling profilers, reading flame graphs, and a one-change-at-a-time method on the hardware users actually carry.

16

Numbers you can trust

code
// the warm-up measurement from this chapter (node, Apple M3 Pro)
function work(arr: number[]) { let s = 0; for (let i = 0; i < arr.length; i++) s += arr[i] * 2 + (arr[i] >> 1); return s; }
const arr = Array.from({ length: 10_000 }, (_, i) => i);
const us: number[] = [];
for (let k = 0; k < 2000; k++) {
  const t = process.hrtime.bigint(); work(arr); us.push(Number(process.hrtime.bigint() - t) / 1000);
}
// first call ≈ 250 µs · median of calls 1-10 ≈ 33 µs · median of calls 1000-2000 ≈ 6 µs (p99 ≈ 7 µs)
code
# profiling a Node script and opening the result in DevTools
node --cpu-prof --cpu-prof-dir=prof script.mjs     # writes a .cpuprofile
# DevTools → Performance → Load profile → Bottom-up for self time, flame chart for shape

# in the browser: Performance panel → CPU 4x slowdown → record the interaction
# then: what is wide that I did not expect?
mistakeeffectfix
timing one cold runmeasures the interpreter, not the optimised codewarm up, or report cold separately on purpose
reporting the averagehides the tail users feelmedian plus p95 and p99
dead code in a microbenchmarkthe JIT removes the work you meant to measureuse the result (sum it, return it)
optimising without a profileeffort on the wrong functionprofile first; fix the widest unexpected bar
testing only on a fast laptopproblems invisible until users report themCPU throttling and a real low-end phone
MEASURING PERFORMANCE PROPERLY
warm-up, variance, percentiles, profilers and flame graphs: how to get numbers you can trust
swipe the figure sideways, or tap expand for full screen
1/6
warm-up
Warm-up: JavaScript engines interpret first, then compile hot functions with optimising compilers. Measured on this machine, a 10,000-element loop function took about 250 µs on its first call, a median of about 33 µs over the first ten calls, and about 6 µs once warm. Benchmark the state your users will see, and say which one you measured.