Performance In Practice
Everything before this part was mechanism; this part is how to find out which mechanism is costing you, and in what order to fix it. Benchmarks that survive the optimiser, CPU profiles and their three shapes, allocation profiles and GC time, the hot-path checklist in order, one path fixed seven-fold step by step, production observation, and the course in one page.
Benchmarking without being fooled
"My benchmark says the new version is 40% faster. The deployed service is not. Which one is lying?"
Probably the benchmark, in one of five ways: it measured dead code the optimiser removed, it folded constant inputs, it timed an OSR'd top-level loop that production never runs, it blended tiers into one number, or it reported one noisy sample. The engine is a compiler that specialises on what it sees, and a benchmark is a tiny program the engine sees in full. Write the harness so the engine cannot see through it, and read the result knowing what tier it came from.
// a harness that survives the optimiser
import { performance } from 'node:perf_hooks'
let sink = 0 // an opaque sink: results escape here so nothing is dead
export function bench(name, fn, inputs, { warm = 2000, batches = 30, perBatch = 2000 } = {}) {
for (let i = 0; i < warm; i++) sink ^= fn(inputs[i % inputs.length]) | 0 // warm-up: reach the tier
const times = []
for (let b = 0; b < batches; b++) {
const t0 = performance.now()
for (let i = 0; i < perBatch; i++) sink ^= fn(inputs[i % inputs.length]) | 0 // varied inputs; consumed result
times.push((performance.now() - t0) / perBatch * 1e6) // ns per call
}
times.sort((a, b) => a - b)
const p = q => times[Math.floor(q * (times.length - 1))]
console.log(`${name}: median ${p(0.5).toFixed(1)} ns p10 ${p(0.1).toFixed(1)} p90 ${p(0.9).toFixed(1)}`)
}
// run with --allow-natives-syntax and check %GetOptimizationStatus(fn) after warm-up if the tier matters
// run twice with the order of cases swapped: order effects (one case's feedback polluting a shared helper) are real- Consume every result through something opaque (a global XOR sink, a library's
do_not_optimize). Unused values are dead; dead code is removed; the loop measures nothing. - Vary the inputs. Constants fold. Use an array of realistic inputs and index it with the loop counter.
- Measure a function called many times, not one top-level loop. Production calls functions; OSR code is a different compile.
- Warm up, then measure, then say which tier.
%GetOptimizationStatusunder--allow-natives-syntaxtells you. If production runs the code a hundred times total, measure the interpreter (--no-opt --no-maglevor simply no warm-up) because that is where it will live. - Many samples, a median and a spread. Thirty batches minimum. Compare p10 to p90 ranges; if they overlap, there is no difference.
- Order effects are real. Case A can pollute a shared helper's feedback so case B measures polymorphic code. Run each case in a fresh process when comparing shapes.
- Same machine state. Plugged in, no browser tab mining, no other benchmark. CPU frequency scaling alone is a 2× noise source on laptops.
- Can: which of two code shapes is cheaper for the engine, under stated conditions. "Map versus object for these keys", "for-of versus index in TurboFan", "this allocation eliminated or not".
- Cannot: whether it matters. A 10× win on a function that takes 0.1% of the request time is 0.09%. Only a profile of the real workload answers "does it matter", and it answers it before you write the benchmark.
- Also cannot: predict cache behaviour at scale (a benchmark's data fits in L2; production's does not), or GC behaviour under a real heap (the benchmark's young generation is empty; production's is full of other things).
node --trace-opt --trace-deopt bench.js while the benchmark runs. If the function under test is deoptimising, or never reaches the tier you expected, the number is not what you think it is.Reading a CPU profile
A sampled profile is the tool that answers "where does the time go" for a real workload. It costs little to record, it lies less than a benchmark, and it has three shapes that each point at a different kind of fix.
- Node:
node --cpu-prof app.jswrites a.cpuprofileon exit;--cpu-prof-intervalsets the sampling period (default 1 ms; 100 µs for short runs). Or the inspector:node --inspect, DevTools → Profiler. Orv8-profiler-next/ theinspectormodule to start and stop from code around a specific workload. - Browser: Performance panel record (includes rendering, GC, and the task structure) or the JavaScript Profiler panel (CPU only, finer sampling).
- Production: continuous profilers (Pyroscope, Datadog, Google Cloud Profiler) sample at low overhead across the fleet; the aggregate flame graph is the most honest performance data you will ever get.
- Source maps: apply them (DevTools does;
--enable-source-mapsin Node for stack traces) so bundled names resolve.
- Bottom-Up by self time first. The top ten functions by self time are where the CPU spends its instructions. If the top entry is your code, parts 3 to 6 apply to it. If it is
(garbage collector), part 5. If it is(program), the engine is working around your code: run--trace-deoptand--log-icnext. - Then the flame chart for structure. Find the top self-time function in the chart and look at the shape around it: a wide bar (a hot loop; optimise the loop), a comb (a per-item cost; restructure the iteration), a tower (deep stacks; flatten or inline).
- Total time for features. Sort by total to answer "what does rendering cost" or "what does validation cost" end to end, including everything beneath.
- Tier annotations. With the right settings (Node's
--prof+--prof-processoutput, or DevTools' JavaScript profiler details), frames are tagged interpreted, baseline or optimised. A hot function marked interpreted is a finding by itself. - Idle and waiting are gaps. A request that takes 200 ms with 5 ms of CPU is waiting on I/O; the profile cannot help; tracing (OpenTelemetry spans, async hooks) can.
Allocation profiles and GC time
CPU profiles show where instructions go; allocation profiles show where bytes come from. The two together explain (garbage collector) time: the functions that allocate the most are the ones feeding the scavenger, and the functions whose allocations survive are the ones feeding the marker.
- "How much GC time, and which kind?"
--trace-gcfor a log; aPerformanceObserveron'gc'entries for metrics; the Performance panel's GC bars. Scavenges every few hundred milliseconds with short pauses are normal; mark-compacts every few seconds are a sign the old generation is churning. - "Who allocates?" Allocation sampling (DevTools Memory → Allocation sampling;
node --heap-prof): a sampled profile by allocating function, low overhead, fine for production. The top entries are the allocation rate; part 5's cost model says which rate matters (scavenge frequency). - "Who allocates what survives?" Allocation instrumentation on timeline (DevTools): every allocation with its stack, and whether it was still alive at the end (the blue bars). Expensive; for a focused run.
- "What is retained?" Heap snapshots (part 5): shallow, retained, retainers, comparison.
- "Was this allocation eliminated?"
--trace-turbo-escapefor a specific function.
- High rate, all dies young, scavenges frequent but short: often acceptable. If the scavenge pauses matter (a frame budget), reduce the rate: reuse buffers, hoist allocations out of loops, avoid closures per item, return values instead of result objects, let escape analysis work by keeping callees inlinable.
- High rate, much survives, scavenges long: a big structure being built across many scavenges. Pretenuring helps automatically; build in fewer, larger steps; or build in a worker.
- Old generation growing, mark-compacts frequent: medium-lived objects (caches, per-request state that lives for the request). Bound the caches; consider scoping; check for leaks with snapshots.
- External memory growing: ArrayBuffers and Buffers not released; streams not consumed.
process.memoryUsage().external; not visible in the JS heap.
a frame budget, in allocation terms (60 fps, 16.7 ms): a scavenge of a mostly-dead 8 MB nursery: ~0.5 ms allocating 1 MB per frame → a scavenge every 8 frames → ~0.06 ms per frame on average, 0.5 ms on the frame it lands allocating 8 MB per frame → a scavenge every frame → 0.5 ms every frame, plus whatever survived the frame that gets the scavenge is the one that drops. per-frame allocation is a jank budget, not just a throughput one.
The hot-path checklist
// the hot-path checklist, as questions to ask of a profile's top function // 1. structure: is there a per-item cost that could be per-batch? a walker that could be compiled? an O(n²) hiding as includes/indexOf in a loop? // 2. tier: is it optimised? (%GetOptimizationStatus, or the profile's tier annotation). if not: --trace-opt for why (too big? deopt loop? eval?) // 3. shapes: --log-ic for polymorphic/megamorphic sites; --trace-deopt for wrong map / elements kind. fix at the construction site // 4. numbers: Smi vs double vs boxed; typed arrays for bulk; no NumberOrOddball (default the field); no BigInt in loops // 5. allocation: --trace-gc frequency; allocation profile top functions; temporaries that cross non-inlined calls; closures in loops // 6. strings: cons trees flattened repeatedly? dynamic keys internalised per access? regex rebuilt per call? // 7. calls: megamorphic call sites in the loop? getters that are not inlined? apply with big arrays? arguments in sloppy mode? // 8. async: an await per item where a batch would do? a microtask storm? promises for CPU work? // 9. data layout: array of objects vs struct of arrays; Map vs object; holey arrays; mixed element kinds // 10. measure again. one change at a time. keep the harness in the repo.
- Structure (any part, mostly 9 and 13): per-item costs that could be per-batch; generic walkers that could be compiled; O(n²) hiding in
includesorindexOfinside a loop; work done every call that could be done once. The largest wins are here and they need no engine knowledge beyond "the comb shape means per-item". - Tier (parts 0, 4): is the function optimised? If it never tiers up: too large (split it), a deopt loop (fix the feedback),
evalorwithnearby, or simply not hot enough to matter. - Shapes (part 3): megamorphic loads and stores; wrong-map deopts; holey or mixed arrays. Fix at the construction site. The most common silent regression and the one with the clearest tooling.
- Numbers (part 6): doubles stored in Smi fields; NumberOrOddball from undefined in arithmetic; boxing at non-inlined boundaries; BigInt in loops; parseFloat per element.
- Allocation (part 5): rate and survival; temporaries across non-inlined calls; closures and contexts per iteration; arrays rebuilt per call.
- Strings (part 6): cons trees read repeatedly; dynamic keys internalised; regexes constructed per call; two-byte strings where one-byte would do.
- Calls (part 8): megamorphic sites in the loop; non-inlined getters;
applywith large arrays; sloppyarguments. - Async (part 10): an await per item; promise churn for CPU-bound work; microtask storms starving the host.
- Data layout (part 12): array of objects versus struct of arrays; Map versus object; typed arrays for bulk numerics.
- Measure again. One change at a time, the harness in the repository, the profile re-recorded. A fix that cannot be measured did not happen.
A hot path, fixed in order
One worked example from profile to seven-fold improvement, with each step chosen by what a tool showed and verified by the next measurement. The order (structure, shapes, allocation, representation, micro) is the order of expected payoff and the order the tools suggest.
- A request handler that parses a JSON payload of records, validates each against a schema, maps them to internal objects, and serialises a response. p50 2.1 ms, p99 6 ms, at 2,000 requests per second across the fleet: 4.2 CPU-seconds per second of validation and mapping.
- The profile: validate 40% with a comb shape (per field per record through a generic schema walker), map 25%,
JSON.stringify15%,(garbage collector)12%,(program)8%. - The traces:
--trace-deoptshows map deoptimising with "wrong map" every few seconds;--log-icshows the field loads in validate at megamorphic;--trace-gcshows a scavenge every 400 ms with 30% survival.
- Structure: compile the validator. One generated function per schema with direct field access replaces the walker. The comb disappears; validate drops from 40% to 8%. 2.1 → 1.3 ms. No engine knowledge was needed to see this; the chart shape was enough.
- Shapes: one literal per record. The record factory had added optional fields conditionally; three maps resulted; map's loads were polymorphic and deopting on the fourth. Every field in the literal, null when absent. Deopts stop; ICs go monomorphic;
(program)drops to 2%. 1.3 → 0.9 ms. - Allocation: no per-field result objects. map returned
{ok, value}per field through a non-inlined call, so escape analysis could not remove it. Return the value, collect errors in a preallocated array. Scavenges go from every 400 ms to every 2 s; GC share from 12% to 3%. 0.9 → 0.6 ms. - Representation: plain response objects. The response records had a
toJSONmethod and Date fields;JSON.stringifytook its generic path. Build plain objects with pre-formatted strings once. The fast serialiser engages. 0.6 → 0.4 ms. - Micro: keys and lookups. A string key built per lookup becomes a nested Map; an
includeson a small array becomes a Set. 0.4 → 0.3 ms. Last because smallest. - Verification: the benchmark harness in the repository, the profile re-recorded after each step, and the fleet's CPU graph the day after deploy.
Production observation
- Event loop lag (Node:
perf_hooks.monitorEventLoopDelay; the browser: long tasks via PerformanceObserver): the single best health signal for a JavaScript process. Lag means the main thread is busy; the profile says with what. - GC time share and frequency (
PerformanceObserveron'gc'entries;--trace-gcin staging): scavenge and mark-compact counts and durations per minute. - Heap used and RSS over time (
process.memoryUsage): slope is the leak signal; external and arrayBuffers separately. - Continuous CPU profiling at low sampling rates across the fleet: aggregate flame graphs over hours, by version. Regressions show as a new function appearing in the top self-time list after a deploy.
- Request-level tracing (OpenTelemetry): where the wall time goes, including the waiting the CPU profile cannot see.
- In the browser: Core Web Vitals in the field (the browser course, part 12) with attribution; Long Animation Frames with script attribution for INP.
- CPU high: a CPU profile on one instance (
--cpu-profon a replica, or the inspector on a canary). Bottom-Up by self time. - Memory climbing: two heap snapshots an hour apart from the same instance (
--heapsnapshot-signal). Comparison view. - Latency spikes: GC logs around the spikes; event loop lag histogram; a trace of a slow request.
- Slow after a deploy: the diff, read for shape changes (a new optional property, a new conditional branch in a constructor, a prototype patch, a dependency upgrade that polyfills something), then the profile.
- Fast locally, slow in production: the tier (production functions may be cold), the data (production shapes may be polymorphic where test data was uniform), the heap (production has a full old generation), the CPU (production has no turbo boost). Reproduce with production-shaped data under
--cpu-prof.
- A benchmark harness in the repository, run in CI for the hot paths, with thresholds that fail on regression.
- A profile before and after any change that claims to be a performance change. Attached to the pull request.
- A budget for the critical paths (p50 and p99 latency, allocation per request, bundle size for the browser), reviewed when it moves.
- Shape discipline in code review: constructors that define every field; no conditional property additions; no prototype patches after startup; Maps for dynamic keys; typed arrays for bulk numerics. The rules from part 3 as review comments.
The course, in one page
- The engine is a compiler that speculates on observed types; fast JavaScript keeps the speculation true. (P0, P4)
- Parsing is paid before anything runs; lazy parsing halves it; the code cache removes it on repeat visits. (P1)
- Bytecode is the canonical form; feedback vectors are the engine's memory of your code; optimised code is a cache built from them. (P2)
- Objects have hidden classes; the same construction order gives the same class; inline caches are fast while the classes they see are few. (P3)
- Deoptimisation is the cost of a wrong guess; the trace names the guess; the fix is at the construction site. (P4)
- Most objects die young and cost nothing; medium-lived ones cost the most; the write barrier makes mutating old objects cost something. (P5)
- Small integers live in the word; doubles and strings live on the heap; concatenation is a tree until read. (P6)
- The prototype chain is walked once and guarded; define methods at definition time and leave prototypes alone. (P7)
- A closure is a pointer to a shared context; what it retains is the scope's captured variables, not its own. (P8)
- for...of is a protocol that is free over arrays in optimised code and a call per element otherwise; generators are resumable frames. (P9)
- A promise is a list of reactions; await is a suspend with one reaction; microtasks drain fully after every task. (P10)
- Modules are constructed, linked and evaluated as a graph; imports are references; the resolved key is the identity. (P11)
- Map is an ordered hash table; WeakMap is an ephemeron; a typed array is the one collection that never changes representation. (P12)
- A pattern is a shape over an engine feature and inherits its costs; know which feature before choosing. (P13)
- Benchmarks lie in five ways; profiles have three shapes; fix structure, then shapes, then allocation, then the rest; measure every step. (P14)
--cpu-prof, --trace-deopt, --log-ic, --trace-gc, in that order. Each flag will show you something in this course happening to your program. That is the whole point of having read it.