Part 2 · 7 chapters · ~45 min

Ignition

Every function starts life as bytecode in a register machine with one accumulator, and most functions never leave. This part is why bytecode, how to read --print-bytecode line by line, the feedback vector and its slot kinds and lifecycle, what interpretation costs per instruction and what each tier removes, contexts and closures as they appear in bytecode, the constructs people ask about, and the tools.

17

Why bytecode, and what kind

the question

"If the engine compiles to machine code anyway, why have an interpreter at all?"

Because most code runs once, and compiling it to machine code would cost more than running it. Ignition compiles the AST to a compact bytecode in a single pass and interprets it; the bytecode is small (a few bytes per operation), quick to produce, and carries the feedback slots that every later tier depends on. It is also the one representation that survives: optimised code is a cache that can be discarded; bytecode is the function.

the design
  1. A register machine with an accumulator. Each frame has registers for parameters (a0, a1...) and locals and temporaries (r0, r1...), plus one implicit accumulator. Most instructions take their result from or leave it in the accumulator, which removes an operand from nearly every instruction and shrinks the bytecode.
  2. One byte per opcode, operands following, with prefix bytes (Wide, ExtraWide) for operands that need 16 or 32 bits. Typical instruction: one to four bytes. A function's bytecode is often smaller than its source.
  3. Feedback slot operands. Property loads, stores, calls, arithmetic, comparisons: each carries an index into the function's feedback vector, where the interpreter records what it saw. The slot is allocated by the bytecode generator; the vector itself is allocated lazily after a few calls.
  4. Constant pool. Strings, large numbers, nested function templates and object literal boilerplates referenced by index.
  5. Handlers in machine code. Each opcode's handler is written in CodeStubAssembler or Torque and compiled by TurboFan's backend when V8 is built, so the interpreter is not a C++ loop; it is a table of optimised machine-code snippets that tail-jump to each other.
what bytecode costs and saves
  1. Memory. V8 switched to Ignition (2016) partly to cut memory: the old baseline compiler produced machine code ten times larger than the bytecode. On a phone with many tabs, that mattered more than speed.
  2. Compile time. Bytecode generation is a single AST walk. Machine code generation, even unoptimised, is more.
  3. Deoptimisation target. When optimised code bails out, it reconstructs an interpreter frame and resumes at a bytecode offset. Having a precise, compact representation to fall back to is what makes speculation safe.
  4. Flushing. Bytecode for functions that have not run in a while can be discarded ("bytecode flushing") and regenerated from source on the next call. Another memory lever.
THE REGISTER MACHINE
Ignition's frame: registers, the accumulator, and the constant pool
swipe the figure sideways, or tap expand for full screen
1/6
the frame
function area(w, h) { const pad = 2; return (w + pad) * (h + pad) }. The frame: a0 = w, a1 = h (parameters), r0 = pad (one local register), and the accumulator, which holds the value in flight.
18

Reading --print-bytecode

The fastest way to understand what the engine does with a construct is to look at its bytecode. One flag, one filter, and the output is short enough to read in full.

code
$ node --print-bytecode --print-bytecode-filter=area area.js

[generated bytecode for function: area (0x... <SharedFunctionInfo area>)]
Bytecode length: 20
Parameter count 3            ; receiver + w + h
Register count 2             ; r0 (pad), r1 (temp)
Frame size 16
         0x... @    0 : 0d 02             LdaSmi [2]
         0x... @    2 : c4                Star0                  ; r0 = pad
         0x... @    3 : 0b 03             Ldar a0                ; w
         0x... @    5 : 39 fa 00          Add r0, [0]            ; feedback slot 0
         0x... @    8 : c3                Star1                  ; r1
         0x... @    9 : 0b 02             Ldar a1                ; h
         0x... @   11 : 39 fa 01          Add r0, [1]
         0x... @   14 : 3a f9 02          Mul r1, [2]
         0x... @   17 : ab                Return
Constant pool (size = 0)
Handler Table (size = 0)
Source Position Table (size = 7)
the header
  1. Parameter count includes the receiver (this), so a two-parameter function says 3. Parameters are registers a0, a1... in the frame.
  2. Register count is the number of local registers: named locals that were not context-allocated, plus temporaries the generator needed for subexpressions.
  3. Frame size in bytes: registers times pointer size.
the instructions
  1. Lda*, Sta*, Ldar, Star: loads into and stores from the accumulator. Star0, Star1 are short forms for the first few registers. LdaSmi, LdaZero, LdaUndefined, LdaConstant [i].
  2. Arithmetic and logic: Add r, [slot], Sub, Mul, Inc, BitwiseAnd... accumulator op register, result in accumulator, feedback in the slot.
  3. Property access: GetNamedProperty r, [name], [slot], SetNamedProperty, GetKeyedProperty, DefineNamedOwnProperty (for object literal initialisation, which skips the prototype chain checks a normal store needs).
  4. Calls: CallProperty1 r_fn, r_recv, r_arg, [slot], CallUndefinedReceiver, Construct. The suffix encodes the argument count for the common small cases.
  5. Control flow: Jump, JumpIfFalse, JumpIfToBooleanFalse, JumpLoop (the back edge, which is where on-stack replacement can happen), with offsets.
  6. Contexts and closures: CreateFunctionContext, PushContext, LdaContextSlot, StaCurrentContextSlot, CreateClosure. These are the register-versus-context decision from part 1 made visible.
code
function getX(o) { return o.x }
//   GetNamedProperty a0, [0], [1]     ; receiver a0, name constant [0] = "x", feedback slot [1]
//   Return

function setX(o, v) { o.x = v }
//   Ldar a1
//   SetNamedProperty a0, [0], [1]     ; same shape: receiver, name, slot

function callIt(f, x) { return f(x) }
//   CallUndefinedReceiver1 a0, a1, [0] ; one-arg call, receiver undefined, feedback slot 0 records the target

function make(x) { return { x, y: 2 } }
//   CreateObjectLiteral [0], [1], #41  ; boilerplate constant [0], allocation-site slot [1], flags
//   Star r0
//   Ldar a0
//   DefineNamedOwnProperty r0, [1], [3] ; set x on the fresh object (y came from the boilerplate)
run it
Put any function in a file, call it once, and run the flag with its name as the filter. Then change one thing (add a closure, use arguments, destructure a parameter) and diff the output; the engine tells you exactly what that syntax costs.
19

Feedback vectors: what the interpreter learns

The feedback vector is the most consequential data structure in V8 for a programmer, because it is the engine's memory of your code's behaviour and the sole input to every optimisation decision. Each function has one; each operation that could benefit from specialisation has a slot in it.

slot kinds
  1. Load and store slots (named and keyed property access): a list of (map, handler) pairs. The handler encodes where the property lives: in-object at offset N, out-of-object at index N, on the prototype at depth D, a getter, a dictionary lookup. Monomorphic = one pair; polymorphic = two to four; megamorphic = a sentinel meaning "use the global stub cache".
  2. Binary operation slots (Add, Sub, Mul, comparison): a lattice of type kinds. SignedSmall, Number, NumberOrOddball (numbers plus undefined/null/booleans), String, BigInt, Any. Only ever moves up.
  3. Call slots: the target (a JSFunction) when monomorphic, a count when polymorphic, megamorphic beyond. Plus a call count used by the optimiser's inlining heuristic.
  4. Construct slots and literal slots: an AllocationSite that tracks whether objects allocated here tend to survive (pretenuring) and, for array literals, the elements kind they end up with.
  5. Global load slots: cache a property cell for a global variable so Math or a module-level constant is a direct load after the first time.
  6. Instanceof, typeof, for-in, clone-object slots, each with its own small cache.
the lifecycle
  1. Uninitialised on allocation. The vector is allocated only after the function has been called a few times (the "feedback allocation budget"), so run-once functions never pay for it.
  2. Monomorphic after the first observation. The interpreter's inline cache (IC) now has a fast path: compare the map, use the handler.
  3. Polymorphic after up to four distinct maps. The IC does a linear scan. Still cheap.
  4. Megamorphic after more. The slot holds a sentinel; lookups go to the global stub cache (a hash keyed by map and name) and, on a miss, to a full runtime lookup. Order of magnitude slower. Permanent for that vector.
  5. Reset only when the vector is thrown away: on bytecode flushing of a cold function, or on some deopt sequences that clear feedback to get a fresh start.
run it
d8 (or node --trace-ic in recent versions, with --log-ic producing a log) prints transitions: LoadIC (0->1) at getX:1:21 x for uninitialised to monomorphic, (1->P) to polymorphic, (P->N) to megamorphic. A hot function whose loads go to N is the thing to fix first in any profile; part 3 is how.
THE FEEDBACK VECTOR
what each slot records, and how it changes
swipe the figure sideways, or tap expand for full screen
1/7
uninitialised
function getX(o) { return o.x }. One bytecode, GetNamedProperty r0 [0], one feedback slot. Before any call, slot 0 is UNINITIALIZED: the engine knows nothing.
20

Dispatch, and the cost of interpretation

Interpreting is slow for a specific reason: for every bytecode, the interpreter does the same bookkeeping around a small amount of real work. Measuring that overhead explains why baseline compilation exists and why it gives only a modest speedup.

what runs per instruction
  1. Dispatch. Load the next opcode byte; index the dispatch table; indirect jump. Ignition uses threaded dispatch (each handler jumps directly to the next) rather than a central switch, which gives the branch predictor one indirect jump per handler rather than one shared one.
  2. Operand decoding. Read register indices, slot indices and immediates from the bytecode stream; compute frame addresses.
  3. Feedback handling. Load the slot; branch on its state; after the operation, possibly update it.
  4. Type checks. Is the accumulator a Smi? Is the register a Smi? For a property load: is the receiver a heap object, does its map match?
  5. The work. One add and an overflow check; one load from an offset; one compare.
what each tier removes
  1. Sparkplug removes dispatch and operand decoding by compiling the bytecode to a straight-line sequence of handler calls (or inlined handler bodies) with registers mapped to fixed stack slots. Feedback and type checks remain. Speedup ~1.5 to 2×. Compile cost near zero.
  2. Maglev removes most type checks by trusting the feedback: if the slot says SignedSmall, emit an integer add guarded by a single check; if the load slot says map M1 at offset 0, emit a map check and a load. Feedback updates are gone (optimised code does not record; it only verifies). Speedup several times over Sparkplug.
  3. TurboFan removes checks it can prove redundant (the same map checked twice in a function), inlines callees so their checks merge with the caller's, hoists loads out of loops, allocates registers properly. The last few times faster.
worked numbers
rough costs for acc = acc + r0 with Smi operands:

  Ignition    ~30 to 40 machine instructions   (dispatch, decode, feedback, checks, add)
  Sparkplug   ~15 to 20                  (feedback, checks, add)
  Maglev      ~3                                (Smi check on the untrusted operand, add, overflow branch)
  TurboFan    1 to 2                           (the check proven away by an earlier one, or the value known to be an int)

the gap between the first and last line is why "hot and stable" code is 30× faster than code that stays in the interpreter.
DISPATCH
how the interpreter runs one instruction, and what Sparkplug removes
swipe the figure sideways, or tap expand for full screen
1/6
bytecode array
The bytecode array: [LdaSmi, 2, Star, r0, Ldar, a0, Add, r0, 0, ...]. A bytecode offset points at the current instruction. Each opcode is one byte; operands follow, with a prefix byte for wide operands.
21

Contexts and closures in bytecode

Part 1 decided which variables are captured. Here is what that decision looks like in bytecode: a context object created on entry, slots read and written through it, and a closure that is a function template plus a pointer to that context.

code
function counter() { let n = 0; return () => ++n }
// counter:
//   CreateFunctionContext [0], [1]     ; a Context with 1 slot, for n (captured by the arrow)
//   PushContext r0
//   LdaZero
//   StaCurrentContextSlot [2]          ; n lives in the context, not a register
//   CreateClosure [1], [0], #2         ; the arrow, bound to the current context
//   Return
// the arrow:
//   LdaCurrentContextSlot [2]          ; read n through the context pointer
//   Inc [0]
//   StaCurrentContextSlot [2]
//   Return
what the instructions do
  1. CreateFunctionContext allocates a Context on the heap with one slot per captured variable of this scope. It happens on every call to counter: one allocation per call, before any of the function's own work.
  2. PushContext makes it the current context; StaCurrentContextSlot [2] stores into slot 2 (slots 0 and 1 are the context's own header fields: scope info and previous context).
  3. CreateClosure allocates a JSFunction from a template (the SharedFunctionInfo in the constant pool) and the current context. The closure is two pointers: code and context. That is all a closure is.
  4. In the arrow, LdaCurrentContextSlot [2] reads n through the context pointer. If n were in an outer-outer scope, it would be LdaContextSlot with a depth operand: walk the context chain that many links first.
the costs this makes visible
  1. One context per scope activation that has captured variables. A loop with let and a closure in its body allocates a context per iteration (the per-iteration binding). Usually fine; in a hot loop creating a million closures, it is the allocation that dominates.
  2. Depth. A deeply nested closure reading a variable five scopes up walks five pointers. The optimiser can hoist that in a loop; the interpreter cannot.
  3. Sharing. All closures from one scope share its context. The context holds every captured variable of that scope. A tiny callback that captures nothing it needs still retains the large array a sibling closure captured.
  4. Materialised arguments (CreateMappedArguments in sloppy mode) and direct eval (CallRuntime [DeclareEvalVar] paths) are the expensive variants; strict mode and no eval keep the fast paths.
run it
--print-bytecode on a function with and without a closure capturing one of its locals. The appearance of CreateFunctionContext is the allocation you added; the switch from Ldar to LdaCurrentContextSlot is the indirection.
22

Bytecode for the constructs people ask about

SourceBytecode shapeWhat it tells you
for (let i = 0; i < n; i++)LdaZero; Star r0; [loop:] Ldar r0; TestLessThan r1 [s]; JumpIfFalse; ...body...; Ldar r0; Inc [s]; Star r0; JumpLoopIndex loops are a compare, a branch, an increment. The cheapest loop there is.
for (const x of arr)GetIterator; [loop:] CallProperty0 next; GetNamedProperty done; JumpIfToBooleanTrue; GetNamedProperty value; ...A call and two property loads per iteration in the interpreter. TurboFan recognises the array iterator protocol and reduces it to an index loop; Ignition does not.
arr.map(f)GetNamedProperty map; CallProperty1 ... [s]One call into a builtin written in Torque; the builtin loops. The callback call inside is a real call per element unless TurboFan inlines map and f together.
const {a, b} = objGetNamedProperty obj "a" [s1]; Star; GetNamedProperty obj "b" [s2]; StarDestructuring is plain property loads with their own feedback slots. Free.
[...arr], f(...args)CreateArrayFromIterable / CallWithSpreadBuiltins with fast paths for real arrays with unmodified iterators; generic iteration otherwise.
a?.bLdar a; JumpIfUndefinedOrNull [skip]; GetNamedProperty b; ...A branch and a load. No overhead beyond the check you asked for.
try { } catch { }Body as usual; a handler table entry mapping a bytecode range to a catch offsetZero cost on the non-throwing path. The old "try blocks are slow" was a Crankshaft limitation, gone since 2017.
class A { m() {} }CreateClosure per method; DefineClass builtin sets up the prototype and constructorMethods are closures on the prototype, created once at class definition.
async function f() { await x }...; Await / SuspendGenerator [s]; ResumeGenerator; ...An async function is a generator-like object with suspend and resume points; part 10.
typeof x === 'string'TestTypeOf [string]Specialised to a single instruction; no string comparison happens.
obj[key] with a string keyGetKeyedProperty r, [s]Keyed loads have ICs too; a constant key is as fast as a named load; a varying key over the same map is a hash lookup in the descriptor array or dictionary.
run it, pick one
The for...of row is the one to see: print the bytecode for an index loop and a for-of loop over the same array and count instructions per iteration. Then run both at 10 million iterations under --trace-opt and watch TurboFan close the gap. Part 9 has the iterator protocol in full; part 14 the honest measurement.
23

Reading Ignition: tools

ToolShowsUse it for
--print-bytecode --print-bytecode-filter=nameHeader, instructions with operands and slots, constant pool, handler tableWhat a construct compiles to; register versus context; slot counts
--trace-ic (d8) / --log-ic --logfile=- (node)Every inline cache state transition with location and property nameFinding polymorphic and megamorphic sites in a hot function
--allow-natives-syntax + %DebugPrint(fn)The SharedFunctionInfo and, if present, the feedback vector with each slot's stateInspecting the feedback of one function after a run
--trace-opt"marking for optimization", with the reason (hot and stable, small function, OSR)Confirming the interpreter's feedback was good enough to tier up
--no-opt --no-sparkplugRuns everything in IgnitionMeasuring interpreter-only cost; isolating optimiser effects in a benchmark
--no-lazy-feedback-allocationAllocate feedback vectors immediatelySeeing feedback for functions that would not otherwise get it
Chrome Performance → a function frame → "(interpreted)" / "(baseline)" / "(optimized)" suffixes in the Bottom-Up view with the right settingsWhich tier a sampled frame was inDiscovering that a hot function is still interpreted
--print-bytecode on the Node REPL inputBytecode for an expression typed liveQuick experiments without a file
the pointer
Part 3 is the maps that the load and store slots record: where they come from, how they transition, and the rules that keep a hot function's slots monomorphic. Everything in this part was "the slot records a map"; the next part is the map.