Ignition
Every function starts life as bytecode in a register machine with one accumulator, and most functions never leave. This part is why bytecode, how to read --print-bytecode line by line, the feedback vector and its slot kinds and lifecycle, what interpretation costs per instruction and what each tier removes, contexts and closures as they appear in bytecode, the constructs people ask about, and the tools.
Why bytecode, and what kind
"If the engine compiles to machine code anyway, why have an interpreter at all?"
Because most code runs once, and compiling it to machine code would cost more than running it. Ignition compiles the AST to a compact bytecode in a single pass and interprets it; the bytecode is small (a few bytes per operation), quick to produce, and carries the feedback slots that every later tier depends on. It is also the one representation that survives: optimised code is a cache that can be discarded; bytecode is the function.
- A register machine with an accumulator. Each frame has registers for parameters (a0, a1...) and locals and temporaries (r0, r1...), plus one implicit accumulator. Most instructions take their result from or leave it in the accumulator, which removes an operand from nearly every instruction and shrinks the bytecode.
- One byte per opcode, operands following, with prefix bytes (Wide, ExtraWide) for operands that need 16 or 32 bits. Typical instruction: one to four bytes. A function's bytecode is often smaller than its source.
- Feedback slot operands. Property loads, stores, calls, arithmetic, comparisons: each carries an index into the function's feedback vector, where the interpreter records what it saw. The slot is allocated by the bytecode generator; the vector itself is allocated lazily after a few calls.
- Constant pool. Strings, large numbers, nested function templates and object literal boilerplates referenced by index.
- Handlers in machine code. Each opcode's handler is written in CodeStubAssembler or Torque and compiled by TurboFan's backend when V8 is built, so the interpreter is not a C++ loop; it is a table of optimised machine-code snippets that tail-jump to each other.
- Memory. V8 switched to Ignition (2016) partly to cut memory: the old baseline compiler produced machine code ten times larger than the bytecode. On a phone with many tabs, that mattered more than speed.
- Compile time. Bytecode generation is a single AST walk. Machine code generation, even unoptimised, is more.
- Deoptimisation target. When optimised code bails out, it reconstructs an interpreter frame and resumes at a bytecode offset. Having a precise, compact representation to fall back to is what makes speculation safe.
- Flushing. Bytecode for functions that have not run in a while can be discarded ("bytecode flushing") and regenerated from source on the next call. Another memory lever.
Reading --print-bytecode
The fastest way to understand what the engine does with a construct is to look at its bytecode. One flag, one filter, and the output is short enough to read in full.
$ node --print-bytecode --print-bytecode-filter=area area.js
[generated bytecode for function: area (0x... <SharedFunctionInfo area>)]
Bytecode length: 20
Parameter count 3 ; receiver + w + h
Register count 2 ; r0 (pad), r1 (temp)
Frame size 16
0x... @ 0 : 0d 02 LdaSmi [2]
0x... @ 2 : c4 Star0 ; r0 = pad
0x... @ 3 : 0b 03 Ldar a0 ; w
0x... @ 5 : 39 fa 00 Add r0, [0] ; feedback slot 0
0x... @ 8 : c3 Star1 ; r1
0x... @ 9 : 0b 02 Ldar a1 ; h
0x... @ 11 : 39 fa 01 Add r0, [1]
0x... @ 14 : 3a f9 02 Mul r1, [2]
0x... @ 17 : ab Return
Constant pool (size = 0)
Handler Table (size = 0)
Source Position Table (size = 7)- Parameter count includes the receiver (
this), so a two-parameter function says 3. Parameters are registers a0, a1... in the frame. - Register count is the number of local registers: named locals that were not context-allocated, plus temporaries the generator needed for subexpressions.
- Frame size in bytes: registers times pointer size.
- Lda*, Sta*, Ldar, Star: loads into and stores from the accumulator.
Star0,Star1are short forms for the first few registers.LdaSmi,LdaZero,LdaUndefined,LdaConstant [i]. - Arithmetic and logic:
Add r, [slot],Sub,Mul,Inc,BitwiseAnd... accumulator op register, result in accumulator, feedback in the slot. - Property access:
GetNamedProperty r, [name], [slot],SetNamedProperty,GetKeyedProperty,DefineNamedOwnProperty(for object literal initialisation, which skips the prototype chain checks a normal store needs). - Calls:
CallProperty1 r_fn, r_recv, r_arg, [slot],CallUndefinedReceiver,Construct. The suffix encodes the argument count for the common small cases. - Control flow:
Jump,JumpIfFalse,JumpIfToBooleanFalse,JumpLoop(the back edge, which is where on-stack replacement can happen), with offsets. - Contexts and closures:
CreateFunctionContext,PushContext,LdaContextSlot,StaCurrentContextSlot,CreateClosure. These are the register-versus-context decision from part 1 made visible.
function getX(o) { return o.x }
// GetNamedProperty a0, [0], [1] ; receiver a0, name constant [0] = "x", feedback slot [1]
// Return
function setX(o, v) { o.x = v }
// Ldar a1
// SetNamedProperty a0, [0], [1] ; same shape: receiver, name, slot
function callIt(f, x) { return f(x) }
// CallUndefinedReceiver1 a0, a1, [0] ; one-arg call, receiver undefined, feedback slot 0 records the target
function make(x) { return { x, y: 2 } }
// CreateObjectLiteral [0], [1], #41 ; boilerplate constant [0], allocation-site slot [1], flags
// Star r0
// Ldar a0
// DefineNamedOwnProperty r0, [1], [3] ; set x on the fresh object (y came from the boilerplate)arguments, destructure a parameter) and diff the output; the engine tells you exactly what that syntax costs.Feedback vectors: what the interpreter learns
The feedback vector is the most consequential data structure in V8 for a programmer, because it is the engine's memory of your code's behaviour and the sole input to every optimisation decision. Each function has one; each operation that could benefit from specialisation has a slot in it.
- Load and store slots (named and keyed property access): a list of (map, handler) pairs. The handler encodes where the property lives: in-object at offset N, out-of-object at index N, on the prototype at depth D, a getter, a dictionary lookup. Monomorphic = one pair; polymorphic = two to four; megamorphic = a sentinel meaning "use the global stub cache".
- Binary operation slots (Add, Sub, Mul, comparison): a lattice of type kinds. SignedSmall, Number, NumberOrOddball (numbers plus undefined/null/booleans), String, BigInt, Any. Only ever moves up.
- Call slots: the target (a JSFunction) when monomorphic, a count when polymorphic, megamorphic beyond. Plus a call count used by the optimiser's inlining heuristic.
- Construct slots and literal slots: an AllocationSite that tracks whether objects allocated here tend to survive (pretenuring) and, for array literals, the elements kind they end up with.
- Global load slots: cache a property cell for a global variable so
Mathor a module-level constant is a direct load after the first time. - Instanceof, typeof, for-in, clone-object slots, each with its own small cache.
- Uninitialised on allocation. The vector is allocated only after the function has been called a few times (the "feedback allocation budget"), so run-once functions never pay for it.
- Monomorphic after the first observation. The interpreter's inline cache (IC) now has a fast path: compare the map, use the handler.
- Polymorphic after up to four distinct maps. The IC does a linear scan. Still cheap.
- Megamorphic after more. The slot holds a sentinel; lookups go to the global stub cache (a hash keyed by map and name) and, on a miss, to a full runtime lookup. Order of magnitude slower. Permanent for that vector.
- Reset only when the vector is thrown away: on bytecode flushing of a cold function, or on some deopt sequences that clear feedback to get a fresh start.
node --trace-ic in recent versions, with --log-ic producing a log) prints transitions: LoadIC (0->1) at getX:1:21 x for uninitialised to monomorphic, (1->P) to polymorphic, (P->N) to megamorphic. A hot function whose loads go to N is the thing to fix first in any profile; part 3 is how.Dispatch, and the cost of interpretation
Interpreting is slow for a specific reason: for every bytecode, the interpreter does the same bookkeeping around a small amount of real work. Measuring that overhead explains why baseline compilation exists and why it gives only a modest speedup.
- Dispatch. Load the next opcode byte; index the dispatch table; indirect jump. Ignition uses threaded dispatch (each handler jumps directly to the next) rather than a central switch, which gives the branch predictor one indirect jump per handler rather than one shared one.
- Operand decoding. Read register indices, slot indices and immediates from the bytecode stream; compute frame addresses.
- Feedback handling. Load the slot; branch on its state; after the operation, possibly update it.
- Type checks. Is the accumulator a Smi? Is the register a Smi? For a property load: is the receiver a heap object, does its map match?
- The work. One add and an overflow check; one load from an offset; one compare.
- Sparkplug removes dispatch and operand decoding by compiling the bytecode to a straight-line sequence of handler calls (or inlined handler bodies) with registers mapped to fixed stack slots. Feedback and type checks remain. Speedup ~1.5 to 2×. Compile cost near zero.
- Maglev removes most type checks by trusting the feedback: if the slot says SignedSmall, emit an integer add guarded by a single check; if the load slot says map M1 at offset 0, emit a map check and a load. Feedback updates are gone (optimised code does not record; it only verifies). Speedup several times over Sparkplug.
- TurboFan removes checks it can prove redundant (the same map checked twice in a function), inlines callees so their checks merge with the caller's, hoists loads out of loops, allocates registers properly. The last few times faster.
rough costs for acc = acc + r0 with Smi operands: Ignition ~30 to 40 machine instructions (dispatch, decode, feedback, checks, add) Sparkplug ~15 to 20 (feedback, checks, add) Maglev ~3 (Smi check on the untrusted operand, add, overflow branch) TurboFan 1 to 2 (the check proven away by an earlier one, or the value known to be an int) the gap between the first and last line is why "hot and stable" code is 30× faster than code that stays in the interpreter.
Contexts and closures in bytecode
Part 1 decided which variables are captured. Here is what that decision looks like in bytecode: a context object created on entry, slots read and written through it, and a closure that is a function template plus a pointer to that context.
function counter() { let n = 0; return () => ++n }
// counter:
// CreateFunctionContext [0], [1] ; a Context with 1 slot, for n (captured by the arrow)
// PushContext r0
// LdaZero
// StaCurrentContextSlot [2] ; n lives in the context, not a register
// CreateClosure [1], [0], #2 ; the arrow, bound to the current context
// Return
// the arrow:
// LdaCurrentContextSlot [2] ; read n through the context pointer
// Inc [0]
// StaCurrentContextSlot [2]
// ReturnCreateFunctionContextallocates a Context on the heap with one slot per captured variable of this scope. It happens on every call tocounter: one allocation per call, before any of the function's own work.PushContextmakes it the current context;StaCurrentContextSlot [2]stores into slot 2 (slots 0 and 1 are the context's own header fields: scope info and previous context).CreateClosureallocates a JSFunction from a template (the SharedFunctionInfo in the constant pool) and the current context. The closure is two pointers: code and context. That is all a closure is.- In the arrow,
LdaCurrentContextSlot [2]readsnthrough the context pointer. Ifnwere in an outer-outer scope, it would beLdaContextSlotwith a depth operand: walk the context chain that many links first.
- One context per scope activation that has captured variables. A loop with
letand a closure in its body allocates a context per iteration (the per-iteration binding). Usually fine; in a hot loop creating a million closures, it is the allocation that dominates. - Depth. A deeply nested closure reading a variable five scopes up walks five pointers. The optimiser can hoist that in a loop; the interpreter cannot.
- Sharing. All closures from one scope share its context. The context holds every captured variable of that scope. A tiny callback that captures nothing it needs still retains the large array a sibling closure captured.
- Materialised
arguments(CreateMappedArgumentsin sloppy mode) and directeval(CallRuntime [DeclareEvalVar]paths) are the expensive variants; strict mode and no eval keep the fast paths.
--print-bytecode on a function with and without a closure capturing one of its locals. The appearance of CreateFunctionContext is the allocation you added; the switch from Ldar to LdaCurrentContextSlot is the indirection.Bytecode for the constructs people ask about
| Source | Bytecode shape | What it tells you |
|---|---|---|
for (let i = 0; i < n; i++) | LdaZero; Star r0; [loop:] Ldar r0; TestLessThan r1 [s]; JumpIfFalse; ...body...; Ldar r0; Inc [s]; Star r0; JumpLoop | Index loops are a compare, a branch, an increment. The cheapest loop there is. |
for (const x of arr) | GetIterator; [loop:] CallProperty0 next; GetNamedProperty done; JumpIfToBooleanTrue; GetNamedProperty value; ... | A call and two property loads per iteration in the interpreter. TurboFan recognises the array iterator protocol and reduces it to an index loop; Ignition does not. |
arr.map(f) | GetNamedProperty map; CallProperty1 ... [s] | One call into a builtin written in Torque; the builtin loops. The callback call inside is a real call per element unless TurboFan inlines map and f together. |
const {a, b} = obj | GetNamedProperty obj "a" [s1]; Star; GetNamedProperty obj "b" [s2]; Star | Destructuring is plain property loads with their own feedback slots. Free. |
[...arr], f(...args) | CreateArrayFromIterable / CallWithSpread | Builtins with fast paths for real arrays with unmodified iterators; generic iteration otherwise. |
a?.b | Ldar a; JumpIfUndefinedOrNull [skip]; GetNamedProperty b; ... | A branch and a load. No overhead beyond the check you asked for. |
try { } catch { } | Body as usual; a handler table entry mapping a bytecode range to a catch offset | Zero cost on the non-throwing path. The old "try blocks are slow" was a Crankshaft limitation, gone since 2017. |
class A { m() {} } | CreateClosure per method; DefineClass builtin sets up the prototype and constructor | Methods are closures on the prototype, created once at class definition. |
async function f() { await x } | ...; Await / SuspendGenerator [s]; ResumeGenerator; ... | An async function is a generator-like object with suspend and resume points; part 10. |
typeof x === 'string' | TestTypeOf [string] | Specialised to a single instruction; no string comparison happens. |
obj[key] with a string key | GetKeyedProperty r, [s] | Keyed loads have ICs too; a constant key is as fast as a named load; a varying key over the same map is a hash lookup in the descriptor array or dictionary. |
for...of row is the one to see: print the bytecode for an index loop and a for-of loop over the same array and count instructions per iteration. Then run both at 10 million iterations under --trace-opt and watch TurboFan close the gap. Part 9 has the iterator protocol in full; part 14 the honest measurement.Reading Ignition: tools
| Tool | Shows | Use it for |
|---|---|---|
--print-bytecode --print-bytecode-filter=name | Header, instructions with operands and slots, constant pool, handler table | What a construct compiles to; register versus context; slot counts |
--trace-ic (d8) / --log-ic --logfile=- (node) | Every inline cache state transition with location and property name | Finding polymorphic and megamorphic sites in a hot function |
--allow-natives-syntax + %DebugPrint(fn) | The SharedFunctionInfo and, if present, the feedback vector with each slot's state | Inspecting the feedback of one function after a run |
--trace-opt | "marking for optimization", with the reason (hot and stable, small function, OSR) | Confirming the interpreter's feedback was good enough to tier up |
--no-opt --no-sparkplug | Runs everything in Ignition | Measuring interpreter-only cost; isolating optimiser effects in a benchmark |
--no-lazy-feedback-allocation | Allocate feedback vectors immediately | Seeing feedback for functions that would not otherwise get it |
| Chrome Performance → a function frame → "(interpreted)" / "(baseline)" / "(optimized)" suffixes in the Bottom-Up view with the right settings | Which tier a sampled frame was in | Discovering that a hot function is still interpreted |
--print-bytecode on the Node REPL input | Bytecode for an expression typed live | Quick experiments without a file |