Optimisation
The optimiser compiles your function for the types it has seen and leaves an exit for everything else. This part is speculation and guards and what they buy, the three compilers and why there are three, on-stack replacement for loops that never return, deoptimisation as the trace shows it, every reason you will meet with its fix, what the optimiser does and does not do, and the flags that expose all of it.
Speculation: the idea that makes dynamic languages fast
"The feedback says o.x is always a Smi field at offset 12. What does the optimiser do with that, and what happens when it is wrong?"
It compiles code that assumes it, protected by a guard, and if the guard fails, it throws the code away and goes back to the interpreter at the exact point of failure. That is speculation, and it is the whole trick. A dynamically typed language cannot be compiled to fast code in general; it can be compiled to fast code for the cases that actually occur, as long as there is a way to detect the other cases and handle them slowly.
- Feedback (parts 2 and 3): what was observed. Maps, representations, type kinds, call targets, elements kinds.
- Assumptions: the compiler turns observations into assumptions: "the receiver has map M", "this field is a Smi", "this add does not overflow", "this array is packed doubles", "Array.prototype has not been modified".
- Guards: each assumption that cannot be proven is checked at the cheapest possible point: a map compare, a tag test, the overflow flag, a prototype validity cell, a "stable map" dependency that is invalidated from outside rather than checked inline.
- Deoptimisation: when a guard fails, the optimised frame is translated back into an interpreter frame and execution resumes in Ignition. Feedback is updated. The optimised code may be discarded.
- Dependencies: some assumptions are not guarded inline but registered: "this code depends on map M being stable" or "on prototype P not changing". When the thing changes, every dependent code object is invalidated at once ("lazy deopt": the code bails at its next safepoint).
- Property access becomes a fixed-offset load after one map check, instead of a hash lookup.
- Arithmetic becomes machine integer or float ops with hardware overflow detection, instead of a type dispatch.
- Calls become direct calls or inlined bodies, instead of indirect calls through a function object.
- Allocation can be eliminated when the compiler can see an object never escapes (escape analysis), because it knows the object's shape statically.
- Checks can be merged, hoisted out of loops, or proven redundant, because the compiler is reasoning about concrete maps rather than "any object".
The three compilers
Sparkplug, Maglev and TurboFan are three different answers to "how much compile time is worth spending on this function". They share the bytecode as input and the deoptimiser as a safety net, and differ in what they read, what they build, and how long they take.
- Reads: bytecode only. No feedback, no graph.
- Does: a single pass emitting machine code per bytecode: either a call to the baseline builtin for that instruction or a short inline sequence. Interpreter registers become fixed frame slots; the frame layout is identical to Ignition's, which makes OSR between them and deopt from Maglev to Sparkplug trivial.
- Produces: code that still consults every IC and does every check, but without decoding and dispatch. Roughly 1.5 to 2× Ignition.
- Costs: microseconds. Can compile in batches on a background thread.
- Never deopts: it made no assumptions. It is the floor that a function with chaotic feedback ends up at.
- Reads: bytecode and feedback.
- Does: builds a control-flow graph with one node per bytecode operation, in SSA form, in a single pass; specialises each node on its feedback (CheckMap + LoadField, Int32Add with overflow check, Float64Mul); inlines small monomorphic callees; does a few local optimisations (redundant check elimination within a block, constant folding); allocates registers linearly.
- Produces: code that is most of the way to TurboFan for typical code. Several times Sparkplug.
- Costs: on the order of a millisecond for a mid-sized function, concurrent.
- Why it exists: the gap between Sparkplug and TurboFan was too big; many functions were warm enough to deserve better than baseline but not hot enough to justify TurboFan. Maglev also serves as a fast re-optimisation tier after a deopt.
- Reads: bytecode and feedback, plus dependencies on maps, prototypes, and global cells.
- Does: builds a sea-of-nodes graph (data and control dependencies as edges, not a fixed instruction order); runs typing, then a long pipeline: inlining (including of builtins like Array.prototype.map and the iterator protocol), escape analysis (removing allocations that never escape), load elimination, loop peeling and unrolling, redundant check elimination across the whole function, representation selection (which values stay as raw int32 or float64 in registers), then scheduling, instruction selection, global register allocation, code generation.
- Produces: the fastest code the engine can make. Tens of times Ignition on numeric code; less on code dominated by calls into the runtime or by memory traffic.
- Costs: milliseconds to tens of milliseconds, concurrent. Size limits: very large functions (by bytecode length) are not optimised at all; the limit is generous but reachable by generated code and by giant switch statements.
node --trace-opt app.js | grep -o "target [A-Z]*" | sort | uniq -c: how many functions reached each tier. A healthy long-running service has a small TurboFan count, a larger Maglev count, and most functions never listed at all.On-stack replacement
Tiering by call count has a hole: a function called once that loops forever. On-stack replacement fills it by letting a loop tier up while it is running, by swapping the frame underneath it.
- Budget on back-edges. The function's interrupt budget (the same counter that drives call-based tiering) is also decremented on every
JumpLoop, by the size of the loop body jumped over. A hot loop exhausts it after some thousands of iterations. - The runtime decides. On budget exhaustion the interpreter calls into the runtime, which notes the loop is hot and requests a compile with an OSR entry point for that loop header.
- The compile is concurrent. The interpreter keeps looping. The compiler produces code whose entry expects the interpreter's register state for that loop (which variables are live, in which registers) and starts at the loop header.
- The swap. At the next back-edge check that finds the OSR code ready, the interpreter frame's live registers are copied into the optimised frame's layout and control jumps into the optimised loop. The remaining iterations run optimised.
- Tiers: OSR can go from Ignition to Sparkplug (trivial, same frame layout), to Maglev, or to TurboFan. A loop can OSR twice: first to Maglev, later to TurboFan.
- Exit. OSR code can deopt like any other. A later normal call of the function uses the normally compiled version, not the OSR entry.
- Top-level loops get optimised. A script that is one big loop still runs fast, after a warm-up of a few thousand iterations.
- Benchmark distortion. A microbenchmark as one long loop measures OSR'd code, which is compiled with knowledge of exactly one loop and may be optimised differently from the same code compiled through the normal path (different inlining decisions, since the OSR compile sees the whole function with a different entry). Part 14's harness avoids it by putting the measured code in a function called many times.
- Long-running loops with changing types can OSR, deopt, re-OSR. Same pathology as the deopt loop, same fix.
node --trace-opt -e "let s=0; for (let i=0;i<1e7;i++) s+=i; console.log(s)" prints a line with reason: OSR or an OSR marker in the compile message. That is a loop tiering up with a call count of one.Deoptimisation: reading the trace
A deopt is the engine saying: a guess I compiled on turned out wrong, here is which guess, here is where. It is not an error and not always a problem: a function that deopts once and re-optimises with better feedback is fine. A function that deopts repeatedly is losing time on every cycle and will eventually be abandoned.
$ node --trace-opt --trace-deopt app.js [marking 0x3a2b <JSFunction process (sfi = 0x...)> for optimization to MAGLEV, reason: hot and stable] [compiling method 0x3a2b <JSFunction process> (target MAGLEV), mode: concurrent] [completed compiling 0x3a2b <JSFunction process> (target MAGLEV) - took 0.412, 1.103, 0.058 ms] [marking 0x3a2b <JSFunction process> for optimization to TURBOFAN, reason: hot and stable] [compiling method 0x3a2b <JSFunction process> (target TURBOFAN), mode: concurrent] [completed compiling 0x3a2b <JSFunction process> (target TURBOFAN) - took 2.901, 9.775, 0.211 ms] [bailout (kind: deopt-eager, reason: not a Smi): begin. deoptimizing 0x3a2b <JSFunction process>, opt id 7, bytecode offset 21, deopt exit 2, FP to SP delta 48, caller SP 0x7ffd..., pc 0x1f3a...] [marking 0x3a2b <JSFunction process> for optimization to MAGLEV, reason: hot and stable] ... ; the three numbers after "took": prepare (main thread), execute (background), finalize (main thread), in ms ; "deopt-eager": the guard failed before the operation. "deopt-lazy": something invalidated the code ; from outside (a prototype changed, a map was deprecated) and it bails at the next safepoint. ; "bytecode offset 21": line it up with --print-bytecode to find the exact operation.
- Eager: a guard inside the code failed at the point of use. The frame is translated immediately. Most "wrong map", "not a Smi", "overflow" deopts are eager.
- Lazy: a dependency the code registered was invalidated from outside (a prototype was modified, a map was deprecated because a field generalised, a global constant was reassigned). The code is marked for deopt and bails at its next safepoint (a call return or loop back-edge). Many functions can lazily deopt at once from one change.
- Soft: the code reached a point where it has no feedback at all (a branch never taken during warm-up) and chose to deopt rather than compile a generic path. Expected on first visits to cold branches; a problem only if the branch is actually hot.
- A deopt id at each exit point indexes a translation table that records, for that point, how to rebuild the interpreter frame: which optimised register or stack slot holds each interpreter register, the accumulator, the bytecode offset, the context, and for inlined functions, the whole chain of frames that must be materialised.
- Materialisation: objects that escape analysis had eliminated (never allocated) are allocated now, from their recorded field values, so the interpreter sees real objects.
- Resume: the interpreter frame is written onto the stack, and Ignition continues at the bytecode offset as if nothing had happened. The feedback for the failed site is updated.
- Cost: microseconds for the bailout itself, plus the lost optimised execution until re-optimisation, plus the compile.
%GetOptimizationStatus(fn) with natives syntax returns a bit field; bit "never optimize" set is the thing to look for in a hot function that is unexpectedly slow.Every deopt reason you will meet, and its fix
| Reason (as traced) | What happened | Fix |
|---|---|---|
| wrong map | An object with a different hidden class reached a monomorphic site | Consistent construction (same properties, same order, same prototype); see part 3 |
| wrong call target / wrong function | A call site compiled for one callee saw another | Expected for polymorphic callbacks; if hot and unintended, avoid passing different functions through the same site |
| not a Smi / not a heap number / not a number | A value's representation differed from the feedback (an int field got a double; a number got a string) | Keep field and variable types stable; initialise numbers with the kind they will hold |
| overflow | An int32 operation overflowed; the result needs a double | Usually fine once (it re-optimises with a float path); if frequent, use doubles from the start or Math.imul for intended wraparound |
| lost precision / minus zero | A float-to-int conversion did not round-trip; or a result was -0 where int was assumed | Rare; use | 0 or Math.trunc where an int is meant |
| wrong elements kind / holey elements | An array's elements kind differed from feedback | Consistent array construction; no holes; no mixed types; see part 3 |
| out of bounds | An index past the array length in code compiled assuming in-bounds | Check the loop bound; do not read past length in hot loops |
| insufficient type feedback | A site had no feedback when compiled (a branch not yet taken) and was hit | Expected during warm-up; a problem only if persistent, which means the branch is hot and the function keeps re-optimising before it settles |
| prototype chain changed / wrong prototype / stable map dependency | Lazy deopt: something modified a prototype or deprecated a map the code depended on | Never modify prototypes after startup; never change an object's prototype; find who did it with --trace-maps |
| field type generalisation / deprecated map | A field's representation widened (Smi → Double → Tagged); every object with the old map migrates; code depending on the old map lazily deopts | One-time cost; avoid by initialising fields with their eventual type |
| not a string / wrong instance type | A string operation got a non-string, or an instanceof check failed | Type consistency at the call site |
| division by zero / not an int32 in Math.imul etc | Integer paths hit a case needing the double path | Usually fine; if hot, use float math deliberately |
| too many arguments / arguments adaptor | A call with a different argument count than compiled for (mostly historical; cheap now) | Call with consistent arity in hot paths |
| unexpected deopt in builtin / Array.prototype.* fast path | A builtin's fast path (map, forEach, indexOf) found a modified prototype, a holey array, or a callback that threw | Keep arrays packed and prototypes untouched; the fast paths depend on both |
| Smi check on receiver / not a JSObject | A primitive (number, string) reached a property load compiled for an object | Do not mix primitives and objects through the same site |
--trace-deopt → the reason and bytecode offset → --print-bytecode --print-bytecode-filter=fn → the exact operation → the source line → the type or shape that varies → the construction site upstream that varies it. The fix is almost always at the last step, not in the function that deopted.What the optimiser can and cannot do
Knowing the optimisations that exist tells you which code shapes are free and which are not. These are TurboFan's main passes, with the code pattern each one rewards.
- Inlining. Small, monomorphic callees are inlined into the caller, including builtins (
Array.prototype.map,forEach,reduce,Math.*, string methods) and the array iterator protocol underfor...of. Budgeted by cumulative bytecode size. Rewards: small functions, stable call targets, callbacks that are the same function each time. - Escape analysis. An object allocated in the function (or an inlined callee) that never escapes (is not stored anywhere, not returned, not passed to a non-inlined call) is not allocated at all; its fields become local values. Rewards: small temporary objects and arrays (
{x, y}returned from an inlined helper, destructured immediately). - Load elimination. A property loaded twice with no intervening store to that object is loaded once. Rewards: nothing you need to do; it means
o.xtwice is not slower than caching it in a local. - Check elimination. A map check or bounds check that an earlier check on the same value already covers is removed. Rewards: operating on the same object several times in sequence.
- Representation selection. Numbers that stay numeric are kept as raw int32 or float64 in registers, never boxed, across the whole function. Rewards: numeric code that does not mix types; arrays with consistent elements kinds.
- Loop optimisations. Loop-invariant code motion (a map check or a load hoisted out), loop peeling (the first iteration separated so the rest can assume warmed state), and bounds-check hoisting. Rewards: loops over arrays with a constant length and no writes to the array inside.
- Dead code and constant folding. Branches on constants, unreachable code after a deopt point, arithmetic on literals.
- Change algorithmic complexity. A nested loop is a nested loop.
- Optimise across megamorphic sites. A megamorphic load is a call; nothing is known after it.
- Eliminate allocations that escape. Objects pushed into an array, stored on
this, or captured by a closure are allocated. - Inline very large or polymorphic callees, or through
applywith unknown arguments, or acrosstry/catchin some older versions (fine now), or where the callee's feedback is megamorphic. - Speculate on code it has not seen run. Cold branches are deopt points; the first time through is slow.
- Help functions that use
eval,with, sloppyargumentsaliasing, or that are too large. Those are not optimised at all. - Recover from a deopt loop. After the limit, the function stays unoptimised; nothing you do at runtime brings it back short of a reload.
Reading the optimiser: tools
| Tool | Shows | Use it for |
|---|---|---|
--trace-opt | Marking, compiling, completed, with tier, reason and times | Did it tier up; which tier; how long the compile took |
--trace-deopt (and --trace-deopt-verbose) | Each bailout: kind, reason, function, bytecode offset, and (verbose) the frame translation | Which guess failed, where |
%GetOptimizationStatus(fn) | A bit field: is function, never optimise, always optimise, maybe deopted, optimised, Maglev'd, TurboFan'd, interpreted, baseline, marked for optimisation... | Asserting a tier in a benchmark; checking for "never optimize" |
%OptimizeFunctionOnNextCall(fn) / %PrepareFunctionForOptimization(fn) | Forces TurboFan on the next call (after one call to collect feedback) | Testing optimised behaviour without a warm-up loop; reproducing a deopt deterministically |
%OptimizeMaglevOnNextCall(fn) | Same for Maglev | Comparing tiers |
--print-opt-code --print-opt-code-filter=fn (needs a build with the disassembler, which d8 release builds and some Node builds have) | The generated machine code with comments mapping to bytecode | Seeing exactly what speculation produced |
--trace-turbo --trace-turbo-filter=fn + Turbolizer (tools/turbolizer in the V8 repo) | The graph after every pass, visualised | Why a value was boxed; why an inline did not happen; deep debugging |
--trace-turbo-inlining | Each inlining decision with the reason it was made or refused | "Why is this callback not inlined" |
--max-inlined-bytecode-size, --max-optimized-bytecode-size | The budgets (flags to read with --v8-options, rarely to change) | Understanding why a big function was never optimised |
--no-opt, --no-maglev, --no-sparkplug | Disable tiers | Measuring how much each tier contributes; isolating a bug to a tier |
Chrome Performance → Bottom-Up → enable "(optimized)/(interpreted)" annotations via the JavaScript profiler settings, or the older --prof + --prof-process output's "ticks" sections | Time per function split by tier | A hot function spending its time interpreted is the first thing to look at |
node --allow-natives-syntax --trace-deopt -e "function f(o){return o.x+1} %PrepareFunctionForOptimization(f); f({x:1}); %OptimizeFunctionOnNextCall(f); f({x:1}); f({x:1,y:2})". The last call prints the wrong-map bailout. Change the third call's object to {x:1.5} and read the different reason.