Part 9 · 3 chapters · ~18 min
Performance Engineering
Benchmarking honestly, CPU profiling with --cpu-prof, --prof and DevTools, Linux perf with JIT frames, memory profiling and the five common leaks, production-safe inspection, diagnostic reports, diagnostics_channel and trace events, optimisation techniques that matter, and when native code or WASM is justified.
29
Benchmarking and CPU profiling
code
# load: autocannon for HTTP, mitata or tinybench for functions npx autocannon -c 100 -d 30 -p 1 http://localhost:3000/transfers # report p50, p97.5, p99, req/s # profile while under load node --cpu-prof --cpu-prof-dir=./prof server.js # writes .cpuprofile on exit node --prof server.js && node --prof-process isolate-*.log > ticks.txt # V8 tick profiler, includes C++ # Linux perf with JavaScript frames resolved node --perf-basic-prof server.js & perf record -F 99 -g -p $! -- sleep 30 && perf script | stackcollapse-perf.pl | flamegraph.pl > flame.svg
honest benchmarking
Warm up first (the Computers course measured 250 µs cold against 6 µs warm for one function), run long enough to include GC, report percentiles not averages, compare against a baseline run on the same machine, and repeat the whole run: a difference smaller than run-to-run variance is not a difference.
READING A NODE FLAME GRAPH
a request handler profile, widest bars first
swipe the figure sideways, or tap expand for full screen
1/5
sampling
node --cpu-prof samples the JS stack about every millisecond and writes a .cpuprofile. Open it in Chrome DevTools (Performance, load profile) or speedscope. Overhead is low enough for staging and short production captures.
a stack sample every ~1 ms--cpu-prof, or the inspector in production
30
Memory, leaks and production diagnostics
| the five common leaks | reproduction | fix |
|---|---|---|
| unbounded module-level cache | const cache = new Map() keyed by request data | LRU with max size and TTL |
| event listeners never removed | emitter.on per request on a long-lived emitter | once, off in cleanup, AbortSignal listeners |
| closures captured by timers | setInterval referencing request objects, never cleared | clear on completion; unref long-lived timers |
| promises that never settle | awaiting a callback that never fires; each holds its closure | timeouts on every await of external work |
| global request context | storing per-request data on a singleton | AsyncLocalStorage, scoped to the request |
code
# production-safe tools
node --inspect=127.0.0.1:9229 server.js # never expose 0.0.0.0; reach it via an SSH tunnel or kubectl port-forward
kill -USR1 <pid> # enable the inspector on a running process
node --report-on-fatalerror --report-uncaught-exception server.js # JSON diagnostic report: stacks, heap, handles, env
process.report.writeReport(); # on demand
// diagnostics_channel: zero-cost hooks libraries publish to (undici, http, pg via instrumentations)
import dc from 'node:diagnostics_channel';
dc.subscribe('undici:request:create', ({ request }) => console.log('out', request.origin, request.path));31
Optimisation techniques, and when to go native
| technique | why it works (part 1) |
|---|---|
| stable object shapes (classes, same property order) | monomorphic inline caches, fewer deopts |
| avoid megamorphic hot functions | normalise inputs at the boundary into one class |
| avoid repeated string concatenation in loops | build arrays and join, or write to streams |
| JSON: smaller payloads, schema serialisers | stringify cost is linear in size; schemas skip type checks |
| reuse buffers and objects in hot loops | less allocation means less GC |
| move CPU work to workers | the loop stays responsive, latency for others stays low |
When native code or WASM is justified: only when a profile shows a pure computation (not IO, not serialisation of data you could avoid sending) dominating CPU, the JavaScript is already reasonably optimised, and a benchmark of the native version shows a gain that survives the boundary-crossing cost. Write the benchmark first; it is usually where the idea dies.