Part 4 · 7 chapters · ~55 min

Optimisation

The optimiser compiles your function for the types it has seen and leaves an exit for everything else. This part is speculation and guards and what they buy, the three compilers and why there are three, on-stack replacement for loops that never return, deoptimisation as the trace shows it, every reason you will meet with its fix, what the optimiser does and does not do, and the flags that expose all of it.

31

Speculation: the idea that makes dynamic languages fast

the question

"The feedback says o.x is always a Smi field at offset 12. What does the optimiser do with that, and what happens when it is wrong?"

It compiles code that assumes it, protected by a guard, and if the guard fails, it throws the code away and goes back to the interpreter at the exact point of failure. That is speculation, and it is the whole trick. A dynamically typed language cannot be compiled to fast code in general; it can be compiled to fast code for the cases that actually occur, as long as there is a way to detect the other cases and handle them slowly.

the pieces
  1. Feedback (parts 2 and 3): what was observed. Maps, representations, type kinds, call targets, elements kinds.
  2. Assumptions: the compiler turns observations into assumptions: "the receiver has map M", "this field is a Smi", "this add does not overflow", "this array is packed doubles", "Array.prototype has not been modified".
  3. Guards: each assumption that cannot be proven is checked at the cheapest possible point: a map compare, a tag test, the overflow flag, a prototype validity cell, a "stable map" dependency that is invalidated from outside rather than checked inline.
  4. Deoptimisation: when a guard fails, the optimised frame is translated back into an interpreter frame and execution resumes in Ignition. Feedback is updated. The optimised code may be discarded.
  5. Dependencies: some assumptions are not guarded inline but registered: "this code depends on map M being stable" or "on prototype P not changing". When the thing changes, every dependent code object is invalidated at once ("lazy deopt": the code bails at its next safepoint).
what speculation buys, concretely
  1. Property access becomes a fixed-offset load after one map check, instead of a hash lookup.
  2. Arithmetic becomes machine integer or float ops with hardware overflow detection, instead of a type dispatch.
  3. Calls become direct calls or inlined bodies, instead of indirect calls through a function object.
  4. Allocation can be eliminated when the compiler can see an object never escapes (escape analysis), because it knows the object's shape statically.
  5. Checks can be merged, hoisted out of loops, or proven redundant, because the compiler is reasoning about concrete maps rather than "any object".
the mental model
Optimised code is a specialisation of your function for the types it has seen, plus an exit for everything else. It is fast exactly when the exits are not taken.
SPECULATION
what the optimiser generates from feedback, and the guard that protects it
swipe the figure sideways, or tap expand for full screen
1/6
feedback
The feedback, after a warm-up: GetNamedProperty a → (map M, in-object offset 12, representation Smi). GetNamedProperty b → (M, offset 16, Smi). Add → SignedSmall. The optimiser reads all three before generating anything.
32

The three compilers

Sparkplug, Maglev and TurboFan are three different answers to "how much compile time is worth spending on this function". They share the bytecode as input and the deoptimiser as a safety net, and differ in what they read, what they build, and how long they take.

Sparkplug (2021)
  1. Reads: bytecode only. No feedback, no graph.
  2. Does: a single pass emitting machine code per bytecode: either a call to the baseline builtin for that instruction or a short inline sequence. Interpreter registers become fixed frame slots; the frame layout is identical to Ignition's, which makes OSR between them and deopt from Maglev to Sparkplug trivial.
  3. Produces: code that still consults every IC and does every check, but without decoding and dispatch. Roughly 1.5 to 2× Ignition.
  4. Costs: microseconds. Can compile in batches on a background thread.
  5. Never deopts: it made no assumptions. It is the floor that a function with chaotic feedback ends up at.
Maglev (2023)
  1. Reads: bytecode and feedback.
  2. Does: builds a control-flow graph with one node per bytecode operation, in SSA form, in a single pass; specialises each node on its feedback (CheckMap + LoadField, Int32Add with overflow check, Float64Mul); inlines small monomorphic callees; does a few local optimisations (redundant check elimination within a block, constant folding); allocates registers linearly.
  3. Produces: code that is most of the way to TurboFan for typical code. Several times Sparkplug.
  4. Costs: on the order of a millisecond for a mid-sized function, concurrent.
  5. Why it exists: the gap between Sparkplug and TurboFan was too big; many functions were warm enough to deserve better than baseline but not hot enough to justify TurboFan. Maglev also serves as a fast re-optimisation tier after a deopt.
TurboFan (2015, backend rewritten as Turboshaft from 2023)
  1. Reads: bytecode and feedback, plus dependencies on maps, prototypes, and global cells.
  2. Does: builds a sea-of-nodes graph (data and control dependencies as edges, not a fixed instruction order); runs typing, then a long pipeline: inlining (including of builtins like Array.prototype.map and the iterator protocol), escape analysis (removing allocations that never escape), load elimination, loop peeling and unrolling, redundant check elimination across the whole function, representation selection (which values stay as raw int32 or float64 in registers), then scheduling, instruction selection, global register allocation, code generation.
  3. Produces: the fastest code the engine can make. Tens of times Ignition on numeric code; less on code dominated by calls into the runtime or by memory traffic.
  4. Costs: milliseconds to tens of milliseconds, concurrent. Size limits: very large functions (by bytecode length) are not optimised at all; the limit is generous but reachable by generated code and by giant switch statements.
run it
node --trace-opt app.js | grep -o "target [A-Z]*" | sort | uniq -c: how many functions reached each tier. A healthy long-running service has a small TurboFan count, a larger Maglev count, and most functions never listed at all.
SPARKPLUG, MAGLEV, TURBOFAN
what each compiler does with the same bytecode
swipe the figure sideways, or tap expand for full screen
1/6
the bytecode
The input: bytecode for a loop that sums an array of points: for (const p of pts) t += p.x * p.y. About 30 bytecodes with six feedback slots: the iterator protocol, two property loads, a multiply, an add.
33

On-stack replacement

Tiering by call count has a hole: a function called once that loops forever. On-stack replacement fills it by letting a loop tier up while it is running, by swapping the frame underneath it.

how it works
  1. Budget on back-edges. The function's interrupt budget (the same counter that drives call-based tiering) is also decremented on every JumpLoop, by the size of the loop body jumped over. A hot loop exhausts it after some thousands of iterations.
  2. The runtime decides. On budget exhaustion the interpreter calls into the runtime, which notes the loop is hot and requests a compile with an OSR entry point for that loop header.
  3. The compile is concurrent. The interpreter keeps looping. The compiler produces code whose entry expects the interpreter's register state for that loop (which variables are live, in which registers) and starts at the loop header.
  4. The swap. At the next back-edge check that finds the OSR code ready, the interpreter frame's live registers are copied into the optimised frame's layout and control jumps into the optimised loop. The remaining iterations run optimised.
  5. Tiers: OSR can go from Ignition to Sparkplug (trivial, same frame layout), to Maglev, or to TurboFan. A loop can OSR twice: first to Maglev, later to TurboFan.
  6. Exit. OSR code can deopt like any other. A later normal call of the function uses the normally compiled version, not the OSR entry.
what it means in practice
  1. Top-level loops get optimised. A script that is one big loop still runs fast, after a warm-up of a few thousand iterations.
  2. Benchmark distortion. A microbenchmark as one long loop measures OSR'd code, which is compiled with knowledge of exactly one loop and may be optimised differently from the same code compiled through the normal path (different inlining decisions, since the OSR compile sees the whole function with a different entry). Part 14's harness avoids it by putting the measured code in a function called many times.
  3. Long-running loops with changing types can OSR, deopt, re-OSR. Same pathology as the deopt loop, same fix.
run it
node --trace-opt -e "let s=0; for (let i=0;i<1e7;i++) s+=i; console.log(s)" prints a line with reason: OSR or an OSR marker in the compile message. That is a loop tiering up with a call count of one.
ON-STACK REPLACEMENT
tiering up a loop that never returns
swipe the figure sideways, or tap expand for full screen
1/6
called once
main() is called once. Inside, a for loop runs a million iterations. Call-count tiering would never trigger: the call count is 1. The loop is where the time is.
34

Deoptimisation: reading the trace

A deopt is the engine saying: a guess I compiled on turned out wrong, here is which guess, here is where. It is not an error and not always a problem: a function that deopts once and re-optimises with better feedback is fine. A function that deopts repeatedly is losing time on every cycle and will eventually be abandoned.

code
$ node --trace-opt --trace-deopt app.js

[marking 0x3a2b <JSFunction process (sfi = 0x...)> for optimization to MAGLEV, reason: hot and stable]
[compiling method 0x3a2b <JSFunction process> (target MAGLEV), mode: concurrent]
[completed compiling 0x3a2b <JSFunction process> (target MAGLEV) - took 0.412, 1.103, 0.058 ms]
[marking 0x3a2b <JSFunction process> for optimization to TURBOFAN, reason: hot and stable]
[compiling method 0x3a2b <JSFunction process> (target TURBOFAN), mode: concurrent]
[completed compiling 0x3a2b <JSFunction process> (target TURBOFAN) - took 2.901, 9.775, 0.211 ms]
[bailout (kind: deopt-eager, reason: not a Smi): begin. deoptimizing 0x3a2b <JSFunction process>,
   opt id 7, bytecode offset 21, deopt exit 2, FP to SP delta 48, caller SP 0x7ffd..., pc 0x1f3a...]
[marking 0x3a2b <JSFunction process> for optimization to MAGLEV, reason: hot and stable]
...

; the three numbers after "took": prepare (main thread), execute (background), finalize (main thread), in ms
; "deopt-eager": the guard failed before the operation. "deopt-lazy": something invalidated the code
;   from outside (a prototype changed, a map was deprecated) and it bails at the next safepoint.
; "bytecode offset 21": line it up with --print-bytecode to find the exact operation.
the kinds
  1. Eager: a guard inside the code failed at the point of use. The frame is translated immediately. Most "wrong map", "not a Smi", "overflow" deopts are eager.
  2. Lazy: a dependency the code registered was invalidated from outside (a prototype was modified, a map was deprecated because a field generalised, a global constant was reassigned). The code is marked for deopt and bails at its next safepoint (a call return or loop back-edge). Many functions can lazily deopt at once from one change.
  3. Soft: the code reached a point where it has no feedback at all (a branch never taken during warm-up) and chose to deopt rather than compile a generic path. Expected on first visits to cold branches; a problem only if the branch is actually hot.
the mechanics of a bailout
  1. A deopt id at each exit point indexes a translation table that records, for that point, how to rebuild the interpreter frame: which optimised register or stack slot holds each interpreter register, the accumulator, the bytecode offset, the context, and for inlined functions, the whole chain of frames that must be materialised.
  2. Materialisation: objects that escape analysis had eliminated (never allocated) are allocated now, from their recorded field values, so the interpreter sees real objects.
  3. Resume: the interpreter frame is written onto the stack, and Ignition continues at the bytecode offset as if nothing had happened. The feedback for the failed site is updated.
  4. Cost: microseconds for the bailout itself, plus the lost optimised execution until re-optimisation, plus the compile.
the counter
Each function counts its deopts. Past a threshold (a handful to a dozen, by kind), the function is marked never-optimise and stays in Sparkplug silently. %GetOptimizationStatus(fn) with natives syntax returns a bit field; bit "never optimize" set is the thing to look for in a hot function that is unexpectedly slow.
A DEOPT LOOP
optimise, bail out, re-optimise, bail out
swipe the figure sideways, or tap expand for full screen
1/6
optimised
process(item) is hot. Its feedback says item.price is a Smi (integer cents). TurboFan compiles an integer path. --trace-opt: "marking process for optimization, reason: hot and stable".
35

Every deopt reason you will meet, and its fix

Reason (as traced)What happenedFix
wrong mapAn object with a different hidden class reached a monomorphic siteConsistent construction (same properties, same order, same prototype); see part 3
wrong call target / wrong functionA call site compiled for one callee saw anotherExpected for polymorphic callbacks; if hot and unintended, avoid passing different functions through the same site
not a Smi / not a heap number / not a numberA value's representation differed from the feedback (an int field got a double; a number got a string)Keep field and variable types stable; initialise numbers with the kind they will hold
overflowAn int32 operation overflowed; the result needs a doubleUsually fine once (it re-optimises with a float path); if frequent, use doubles from the start or Math.imul for intended wraparound
lost precision / minus zeroA float-to-int conversion did not round-trip; or a result was -0 where int was assumedRare; use | 0 or Math.trunc where an int is meant
wrong elements kind / holey elementsAn array's elements kind differed from feedbackConsistent array construction; no holes; no mixed types; see part 3
out of boundsAn index past the array length in code compiled assuming in-boundsCheck the loop bound; do not read past length in hot loops
insufficient type feedbackA site had no feedback when compiled (a branch not yet taken) and was hitExpected during warm-up; a problem only if persistent, which means the branch is hot and the function keeps re-optimising before it settles
prototype chain changed / wrong prototype / stable map dependencyLazy deopt: something modified a prototype or deprecated a map the code depended onNever modify prototypes after startup; never change an object's prototype; find who did it with --trace-maps
field type generalisation / deprecated mapA field's representation widened (Smi → Double → Tagged); every object with the old map migrates; code depending on the old map lazily deoptsOne-time cost; avoid by initialising fields with their eventual type
not a string / wrong instance typeA string operation got a non-string, or an instanceof check failedType consistency at the call site
division by zero / not an int32 in Math.imul etcInteger paths hit a case needing the double pathUsually fine; if hot, use float math deliberately
too many arguments / arguments adaptorA call with a different argument count than compiled for (mostly historical; cheap now)Call with consistent arity in hot paths
unexpected deopt in builtin / Array.prototype.* fast pathA builtin's fast path (map, forEach, indexOf) found a modified prototype, a holey array, or a callback that threwKeep arrays packed and prototypes untouched; the fast paths depend on both
Smi check on receiver / not a JSObjectA primitive (number, string) reached a property load compiled for an objectDo not mix primitives and objects through the same site
the procedure
--trace-deopt → the reason and bytecode offset → --print-bytecode --print-bytecode-filter=fn → the exact operation → the source line → the type or shape that varies → the construction site upstream that varies it. The fix is almost always at the last step, not in the function that deopted.
36

What the optimiser can and cannot do

Knowing the optimisations that exist tells you which code shapes are free and which are not. These are TurboFan's main passes, with the code pattern each one rewards.

what it does
  1. Inlining. Small, monomorphic callees are inlined into the caller, including builtins (Array.prototype.map, forEach, reduce, Math.*, string methods) and the array iterator protocol under for...of. Budgeted by cumulative bytecode size. Rewards: small functions, stable call targets, callbacks that are the same function each time.
  2. Escape analysis. An object allocated in the function (or an inlined callee) that never escapes (is not stored anywhere, not returned, not passed to a non-inlined call) is not allocated at all; its fields become local values. Rewards: small temporary objects and arrays ({x, y} returned from an inlined helper, destructured immediately).
  3. Load elimination. A property loaded twice with no intervening store to that object is loaded once. Rewards: nothing you need to do; it means o.x twice is not slower than caching it in a local.
  4. Check elimination. A map check or bounds check that an earlier check on the same value already covers is removed. Rewards: operating on the same object several times in sequence.
  5. Representation selection. Numbers that stay numeric are kept as raw int32 or float64 in registers, never boxed, across the whole function. Rewards: numeric code that does not mix types; arrays with consistent elements kinds.
  6. Loop optimisations. Loop-invariant code motion (a map check or a load hoisted out), loop peeling (the first iteration separated so the rest can assume warmed state), and bounds-check hoisting. Rewards: loops over arrays with a constant length and no writes to the array inside.
  7. Dead code and constant folding. Branches on constants, unreachable code after a deopt point, arithmetic on literals.
what it does not do
  1. Change algorithmic complexity. A nested loop is a nested loop.
  2. Optimise across megamorphic sites. A megamorphic load is a call; nothing is known after it.
  3. Eliminate allocations that escape. Objects pushed into an array, stored on this, or captured by a closure are allocated.
  4. Inline very large or polymorphic callees, or through apply with unknown arguments, or across try/catch in some older versions (fine now), or where the callee's feedback is megamorphic.
  5. Speculate on code it has not seen run. Cold branches are deopt points; the first time through is slow.
  6. Help functions that use eval, with, sloppy arguments aliasing, or that are too large. Those are not optimised at all.
  7. Recover from a deopt loop. After the limit, the function stays unoptimised; nothing you do at runtime brings it back short of a reload.
the pointer
Part 5 is the allocator and collector that escape analysis is trying to avoid. An allocation the optimiser could not eliminate is a young-generation object the scavenger will have to copy or discard.
37

Reading the optimiser: tools

ToolShowsUse it for
--trace-optMarking, compiling, completed, with tier, reason and timesDid it tier up; which tier; how long the compile took
--trace-deopt (and --trace-deopt-verbose)Each bailout: kind, reason, function, bytecode offset, and (verbose) the frame translationWhich guess failed, where
%GetOptimizationStatus(fn)A bit field: is function, never optimise, always optimise, maybe deopted, optimised, Maglev'd, TurboFan'd, interpreted, baseline, marked for optimisation...Asserting a tier in a benchmark; checking for "never optimize"
%OptimizeFunctionOnNextCall(fn) / %PrepareFunctionForOptimization(fn)Forces TurboFan on the next call (after one call to collect feedback)Testing optimised behaviour without a warm-up loop; reproducing a deopt deterministically
%OptimizeMaglevOnNextCall(fn)Same for MaglevComparing tiers
--print-opt-code --print-opt-code-filter=fn (needs a build with the disassembler, which d8 release builds and some Node builds have)The generated machine code with comments mapping to bytecodeSeeing exactly what speculation produced
--trace-turbo --trace-turbo-filter=fn + Turbolizer (tools/turbolizer in the V8 repo)The graph after every pass, visualisedWhy a value was boxed; why an inline did not happen; deep debugging
--trace-turbo-inliningEach inlining decision with the reason it was made or refused"Why is this callback not inlined"
--max-inlined-bytecode-size, --max-optimized-bytecode-sizeThe budgets (flags to read with --v8-options, rarely to change)Understanding why a big function was never optimised
--no-opt, --no-maglev, --no-sparkplugDisable tiersMeasuring how much each tier contributes; isolating a bug to a tier
Chrome Performance → Bottom-Up → enable "(optimized)/(interpreted)" annotations via the JavaScript profiler settings, or the older --prof + --prof-process output's "ticks" sectionsTime per function split by tierA hot function spending its time interpreted is the first thing to look at
run it, the deterministic way
node --allow-natives-syntax --trace-deopt -e "function f(o){return o.x+1} %PrepareFunctionForOptimization(f); f({x:1}); %OptimizeFunctionOnNextCall(f); f({x:1}); f({x:1,y:2})". The last call prints the wrong-map bailout. Change the third call's object to {x:1.5} and read the different reason.