Part 9 · 3 chapters · ~18 min

Performance Engineering

Benchmarking honestly, CPU profiling with --cpu-prof, --prof and DevTools, Linux perf with JIT frames, memory profiling and the five common leaks, production-safe inspection, diagnostic reports, diagnostics_channel and trace events, optimisation techniques that matter, and when native code or WASM is justified.

29

Benchmarking and CPU profiling

code
# load: autocannon for HTTP, mitata or tinybench for functions
npx autocannon -c 100 -d 30 -p 1 http://localhost:3000/transfers   # report p50, p97.5, p99, req/s

# profile while under load
node --cpu-prof --cpu-prof-dir=./prof server.js        # writes .cpuprofile on exit
node --prof server.js && node --prof-process isolate-*.log > ticks.txt   # V8 tick profiler, includes C++

# Linux perf with JavaScript frames resolved
node --perf-basic-prof server.js &
perf record -F 99 -g -p $! -- sleep 30 && perf script | stackcollapse-perf.pl | flamegraph.pl > flame.svg
honest benchmarking
Warm up first (the Computers course measured 250 µs cold against 6 µs warm for one function), run long enough to include GC, report percentiles not averages, compare against a baseline run on the same machine, and repeat the whole run: a difference smaller than run-to-run variance is not a difference.
READING A NODE FLAME GRAPH
a request handler profile, widest bars first
(root)100% of sampleshttp request handler92%middleware chain12%JSON.stringify(response)41%pg query + parse rows30%(garbage collector)9%
swipe the figure sideways, or tap expand for full screen
1/5
sampling
node --cpu-prof samples the JS stack about every millisecond and writes a .cpuprofile. Open it in Chrome DevTools (Performance, load profile) or speedscope. Overhead is low enough for staging and short production captures.
a stack sample every ~1 ms--cpu-prof, or the inspector in production
30

Memory, leaks and production diagnostics

the five common leaksreproductionfix
unbounded module-level cacheconst cache = new Map() keyed by request dataLRU with max size and TTL
event listeners never removedemitter.on per request on a long-lived emitteronce, off in cleanup, AbortSignal listeners
closures captured by timerssetInterval referencing request objects, never clearedclear on completion; unref long-lived timers
promises that never settleawaiting a callback that never fires; each holds its closuretimeouts on every await of external work
global request contextstoring per-request data on a singletonAsyncLocalStorage, scoped to the request
code
# production-safe tools
node --inspect=127.0.0.1:9229 server.js        # never expose 0.0.0.0; reach it via an SSH tunnel or kubectl port-forward
kill -USR1 <pid>                               # enable the inspector on a running process
node --report-on-fatalerror --report-uncaught-exception server.js   # JSON diagnostic report: stacks, heap, handles, env
process.report.writeReport();                  # on demand

// diagnostics_channel: zero-cost hooks libraries publish to (undici, http, pg via instrumentations)
import dc from 'node:diagnostics_channel';
dc.subscribe('undici:request:create', ({ request }) => console.log('out', request.origin, request.path));
31

Optimisation techniques, and when to go native

techniquewhy it works (part 1)
stable object shapes (classes, same property order)monomorphic inline caches, fewer deopts
avoid megamorphic hot functionsnormalise inputs at the boundary into one class
avoid repeated string concatenation in loopsbuild arrays and join, or write to streams
JSON: smaller payloads, schema serialisersstringify cost is linear in size; schemas skip type checks
reuse buffers and objects in hot loopsless allocation means less GC
move CPU work to workersthe loop stays responsive, latency for others stays low

When native code or WASM is justified: only when a profile shows a pure computation (not IO, not serialisation of data you could avoid sending) dominating CPU, the JavaScript is already reasonably optimised, and a benchmark of the native version shows a gain that survives the boundary-crossing cost. Write the benchmark first; it is usually where the idea dies.