Part 1 · 7 chapters · ~45 min

Parsing

Before a line of your code runs, every byte of it has been scanned, and the parts that run soon have been scanned twice. This part is the scanner and its three hard cases, lazy parsing and the pre-parser, the recursive descent parser and the AST it builds, scope analysis and the register-or-context decision, where startup time goes and the levers that move it, how function size and placement change the parser's work, and the flags that show all of it.

10

From bytes to tokens: the scanner

the question

"The bundle is 1 MB. How much of the startup time is the parser, and what in my code makes it worse?"

Parsing is the first thing the engine does to your code and the one cost paid before a single line runs. V8's parser is a hand-written recursive descent parser fed by a scanner that handles the language's lexical oddities; together they produce an AST and a scope tree that Ignition turns into bytecode. Lazy parsing is the trick that keeps a 1 MB bundle from costing a second of parse time.

the scanner's job
  1. Code units in, tokens out. Source is UTF-16 internally (one-byte where possible, part 6). The scanner produces identifiers, keywords, punctuators, numeric and string literals, template parts, regular expression literals, and tracks line terminators.
  2. Identifiers are internalised. Every identifier becomes a pointer into the string table, so obj.foo and obj["foo"] and a later foo in another function all refer to the same string object. Property names are compared by pointer from here on.
  3. Literals are converted once. Numbers become Smis or doubles at scan time; strings get their representation (one-byte or two-byte) decided; both go into the function's constant pool.
  4. Three context-dependent cases need the parser's help: whether / begins a regex or is division; where a template literal's ${ hands control to the expression parser and back; and whether a newline before a token should become a semicolon.
ASI, precisely
  1. A semicolon is inserted when the parser meets a token that cannot continue the current statement and a line terminator preceded it, or the token is }, or at end of input.
  2. Restricted productions: after return, break, continue, throw, a postfix ++/--, yield, async (before a function), and before =>, a newline forces a semicolon regardless. return\nvalue returns undefined.
  3. The hazards: a line starting with (, [, `, +, -, / continues the previous line's expression. a = b\n(c)() calls b. Minifiers and formatters that omit semicolons insert a leading ; on such lines for this reason.
run it
node -e "function f(){ return\n42 }; console.log(f())" prints undefined. The scanner flagged the newline; the parser inserted the semicolon after return. No warning, by design: ASI is a feature of the grammar, not an error recovery.
THE SCANNER
characters to tokens, with the hard cases
swipe the figure sideways, or tap expand for full screen
1/6
the slash
Input: a = b / c; The scanner emits Identifier(a), Punctuator(=), Identifier(b), and then sees "/". Is it division or the start of a regular expression? The scanner cannot know; it depends on what came before. The parser tells it: after an expression, division; after an operator or "(", a regex.
11

Lazy parsing: the pre-parser and the double-scan tax

Full parsing builds an AST and compiles to bytecode. Pre-parsing checks syntax, records function boundaries and captured variables, and builds nothing. V8 pre-parses every function it does not need yet, and fully parses it on first call. This halves startup parse time for typical bundles and introduces one cost: a function that is called soon after load is scanned twice.

what the pre-parser does
  1. Syntax check. A SyntaxError inside a never-called function is still thrown at load, because the pre-parser saw it. The grammar is the same; the output is not.
  2. Positions. Start and end offsets of each function, so the full parser can jump straight to it later.
  3. Scope data. Which variables inner functions reference from outer scopes. Needed so the outer function's context layout can be decided even when the inner functions are not compiled yet. Stored as "preparse data" on the function and reused by the full parse.
  4. Speed. Roughly twice the throughput of the full parser, because there is no tree to allocate and no bytecode to emit.
the eager heuristics
  1. Parenthesised functions. (function(){...}) and (()=>{...}) are compiled eagerly: the parentheses are read as "this will be called immediately" (the PIFE heuristic: Possibly-Invoked Function Expression). Webpack, Rollup and esbuild emit module wrappers this way on purpose.
  2. Immediately invoked in the same statement, with or without parentheses on some forms.
  3. Functions in a function being compiled eagerly may inherit eagerness one level down in some heuristics (the "eager inner" rule), so a module wrapper's direct children are often compiled at load too.
  4. The anti-pattern: a minifier that wraps every function in parentheses, or a bundler that emits one giant eager wrapper around everything, defeats laziness and compiles the whole bundle at load. Check with --trace-parse or the Performance panel's "Compile Code" events.
the double-scan tax, and the code cache
  1. A function called at load is pre-parsed (because the parser did not know) and then fully parsed on the call. Two scans of the same bytes. For the third of a bundle that typically runs at load, this is tens of milliseconds on a phone.
  2. Chrome's code cache: after a script runs, V8 serialises the bytecode of compiled functions; the browser stores it with the HTTP cache entry (after the second load for most scripts; immediately for service-worker-cached ones with the right flags). The next load deserialises instead of parsing: both scans gone for cached functions.
  3. Node: module.enableCompileCache() (v22+) does the same for CommonJS and ESM on disk; v8.startupSnapshot goes further by serialising the heap after initialisation.
run it
node --trace-parse script.js 2>&1 | head lists each function as it is pre-parsed or parsed and how long it took. A function that appears twice (preparse, then parse) within the first milliseconds was called at load and is a candidate for the eager hint; a function that appears once and never again was never called.
LAZY PARSING
pre-parse now, full parse on first call
swipe the figure sideways, or tap expand for full screen
1/7
the decision
The script arrives. The parser handles the top level eagerly: it must know every declaration to run the script. It reaches "function a() {". Decision: full parse or pre-parse?
12

The parser: recursive descent, and what the AST holds

V8's parser is a hand-written recursive descent parser with a few precedence-climbing loops for binary operators. Each grammar production is a C++ method; the AST is a tree of typed nodes; the parser also builds the scope tree and resolves every variable reference as it goes. The AST lives only long enough for Ignition to walk it.

the shape of the parser
  1. ParseProgram → ParseStatementList → ParseStatement → ... down to ParsePrimaryExpression. Each level consumes tokens from the scanner and returns a node. Lookahead is one token, with a few places that peek two.
  2. Expressions by precedence climbing: ParseBinaryExpression(prec) parses operands and loops on operators of at least the given precedence, which is how a + b * c gets the right tree without a method per precedence level.
  3. Arrow functions and destructuring are the hard part. (a, b) => a starts like a parenthesised expression; the parser parses it as one, then on seeing => reinterprets the expression as a parameter list ("cover grammar"). Same for ({a, b} = obj). This is why some syntax errors are reported at the arrow rather than inside the parens.
  4. Templates, regex, ASI: the scanner interactions from chapter 1.
  5. Error recovery: none. The first error aborts the script with a SyntaxError that includes the position. Unlike the HTML parser, the JavaScript grammar has no recovery rules.
what is in the AST
  1. Nodes for every construct: FunctionLiteral, Block, VariableDeclaration, Assignment, Call, Property, BinaryOperation, Conditional, ForStatement, ObjectLiteral with its properties, ClassLiteral with its members, and so on. Each carries source positions for stack traces and the debugger.
  2. Literal values already converted: a NumberLiteral holds a double or Smi; a StringLiteral holds the internalised string.
  3. Object literal "boilerplates": for {x: 1, y: 2}, the parser pre-computes the shape (hidden class) and the constant properties, so creating the object at runtime is a copy of a template rather than two property additions. Part 3 is why that matters.
  4. Not in the AST: types (there are none), and anything about how the code will run. The AST is syntax plus scope.
run it
There is no flag that prints V8's AST, but every popular JavaScript parser produces the standard ESTree shape: npx acorn --ecma2024 file.js or the AST Explorer site. The node types are the same ones the engine uses, which is why bundlers, linters and the engine agree on what the code means.
13

Scope analysis: registers or context

As the parser builds the AST it also builds a scope tree and resolves every identifier to the scope that declares it. One outcome of that resolution decides more about performance than most people expect: whether a variable can live in a register or must live in a heap-allocated context object because a closure captures it.

the scope kinds
  1. Script or module scope: top-level declarations. Module scope is a real lexical scope; classic script top-level var and function declarations become properties of the global object (slow: a global lookup is a property lookup on a dictionary-mode object unless the engine has cached it).
  2. Function scope: parameters, var, function declarations, arguments, this.
  3. Block scope: let, const, class, and the per-iteration bindings of for (let ...). The parser marks loop variables as per-iteration so closures created in the loop body see a fresh binding each time, which costs a context per iteration if a closure captures the variable.
  4. Catch scope, with scope (dynamic; disables resolution for everything inside), eval scope (a direct eval can declare variables at runtime, so every variable in enclosing scopes must be context-allocated and looked up dynamically).
resolution and allocation
  1. Each reference is resolved at parse time to a declaration in some enclosing scope, or to "global" if none. Dynamic lookups exist only under with and sloppy direct eval.
  2. A variable referenced only from its own function is stack-allocated: a register in Ignition's register file, a machine register or stack slot in optimised code. Free to read and write.
  3. A variable referenced from an inner function is context-allocated: a slot in a Context object created when the declaring scope is entered. Reads and writes go through the context pointer. The closure holds the context; the context holds all captured variables of that scope, not just the ones this closure uses.
  4. The sharing consequence: two closures from the same scope share one context. A small long-lived closure keeps alive every captured variable of its scope, including a large one captured only by a different, short-lived closure. Part 8 and the memory part return to this.
worked numbers
a function with and without capture:

  function sum(arr) { let t = 0; for (const x of arr) t += x; return t }     // t: register. fast.
  function sum(arr) { let t = 0; arr.forEach(x => { t += x }); return t }   // t: context slot. the arrow captures it.

  --print-bytecode shows the difference: Ldar r1 / Add vs LdaCurrentContextSlot [2] / Add / StaCurrentContextSlot [2]

neither is wrong. the second allocates a context per call and does a memory load per iteration; the optimiser often removes the difference, and sometimes cannot.
AST AND SCOPE RESOLUTION
what the parser builds for a small function
swipe the figure sideways, or tap expand for full screen
1/6
the AST
function outer(n) { let total = 0; const add = x => { total += x }; for (let i = 0; i < n; i++) add(i); return total }. The parser produces an AST: FunctionLiteral(outer) with a body of statements, each an expression tree.
14

Startup cost: measuring and reducing it

For a browser page, the engine's part of startup is scan, parse, compile, and the execution of top-level code. The host adds download, the preload scanner's ordering, and the main thread's other work. This chapter is the accounting and the levers, in order of effect.

the levers, largest first
  1. Ship less. Every engine cost is linear or worse in bytes. Route-level code splitting, removing unused dependencies, and not shipping polyfills to browsers that do not need them. Nothing else comes close.
  2. Run less at load. Top-level execution is often the biggest bar and is entirely yours: framework initialisation, module side effects, eager data fetching, synchronous hydration of the whole page. Defer whatever is below the fold or behind an interaction.
  3. Let the code cache work. Stable URLs for code that rarely changes (vendor chunks), so the cached bytecode survives deploys. A content hash that changes on every release invalidates the cache for every user on every release.
  4. Streaming parse. External scripts parse on a background thread as they download. Inline scripts and eval cannot. Large inline scripts are a startup anti-pattern for this reason as well as caching.
  5. Eager where it counts, lazy elsewhere. Module wrappers parenthesised (bundlers do this); everything else lazy. Verify with a trace that the bundle is not being compiled wholesale.
  6. Avoid the parser's slow paths. Direct eval and with force dynamic scope for everything around them. arguments in sloppy mode aliases parameters and forces a materialised object. Each is a per-function cost, but in a hot path it compounds.
measuring
  1. Chrome Performance panel: under each script's Evaluate Script, "Compile Script" (top-level) and "Compile Code" (lazy functions compiled on first call). The sum across the load is the engine's compile cost. Long "Compile Code" events after load mean big functions compiled late; many small ones in a burst mean a hot path being compiled piecemeal.
  2. Lighthouse: "Reduce JavaScript execution time" and "Minimize main-thread work" attribute time per script to parse, compile and execute.
  3. Node: --cpu-prof and look at the first frames; process.hrtime around require calls for a quick per-module cost; NODE_DEBUG=module for load order.
  4. Script streaming in the trace: "v8.parseOnBackground" events show background parsing; if they are absent for a large script, it was inline or loaded in a way that prevented streaming.
run it
Build the app twice, once as is and once with the largest route split out behind import(). Record both with 4× CPU throttling. The difference in the Compile and Evaluate bars is the number to put in the pull request.
WHERE PARSE TIME GOES
the same 1 MB bundle, three ways
swipe the figure sideways, or tap expand for full screen
1/5
cold
Cold, as shipped. Scan and pre-parse all 1 MB: ~150 ms. Full parse and compile the functions called at load (about a third of them, for a typical framework app): ~200 ms. Execute top level: ~250 ms. Total before the app is interactive: ~600 ms of engine work, plus the host's.
15

Function size, placement, and the parser

Three things about how code is laid out change what the parser does with it: how big functions are, where they are, and whether they are inside something that forces eagerness. None of these changes behaviour; all of them change startup.

size
  1. Big functions compile slowly and all at once. A 2,000-line function is compiled as a unit on first call: one long "Compile Code" event on the main thread. Splitting it into smaller functions lets the uncalled parts stay pre-parsed.
  2. Tiny functions have a floor. Each function has a SharedFunctionInfo, bytecode, a feedback vector, possibly a closure object. Thousands of one-line functions cost memory and per-call overhead until inlined. This is why bundlers inline trivial helpers.
  3. The optimiser has a size limit for inlining (bytecode length) and for optimisation at all. A huge function may never be optimised by TurboFan; part 4 has the thresholds.
placement
  1. Inner functions are re-parsed with their outer function. When outer is compiled lazily, every function inside it is pre-parsed again (the parser must walk through them). A deeply nested structure pays the pre-parse of inner functions once per enclosing level that gets compiled. Flat modules parse faster than deeply nested closures.
  2. Top-level code in modules is always eager. A module's body is compiled at link time whether or not anything is exported. Side-effect-free modules with heavy top-level computation should defer it into a function.
  3. Class bodies are compiled with the class: method bodies are lazy, but the class definition (field initialisers, static blocks) runs at definition time.
forced eagerness
  1. Parentheses (the PIFE hint), immediate invocation, and in some versions functions assigned to properties in object literals that are immediately used.
  2. A direct eval in a function makes every enclosing function's variables dynamically scoped and disables lazy parsing optimisations around it.
  3. The new Function and eval paths create a fresh parse each time, with no code cache and no streaming; a template engine that compiles templates with new Function at runtime pays full parse cost per template per load.
the pointer
Part 2 shows the bytecode the parser's AST becomes, and the register file that scope analysis sized. The scope decision made here (register or context) is visible as the difference between Ldar and LdaContextSlot there.
16

Reading the parser: flags and traces

ToolShowsUse it for
node --trace-parse script.jsEach parse and pre-parse with function name, size and timeWhich functions were compiled at load; the double-scan candidates
node --trace-lazy script.jsLazy compilation events: which function, whenConfirming a function stayed lazy until its first call
node --log-function-events --logfile=v8.logA log of parse, preparse, compile, first execution per function with timestampsThe full startup timeline per function; load it into a spreadsheet
Chrome Performance → Evaluate Script → Compile Script / Compile CodeMain-thread compile events per script and per lazy functionWhere compile time lands on the critical path
Chrome Performance → "Streaming compile" / v8.parseOnBackgroundBackground parsing of external scriptsWhether a script streamed; inline scripts never do
Lighthouse → Reduce JavaScript execution timePer-script parse, compile, execute totalsThe ranked list of scripts to split or defer
node --allow-natives-syntax -e "%GetOptimizationStatus(f)" (bit 3)Whether a function has been compiled at all, and to which tier"Is this function still lazy" without a trace
npx acorn / AST ExplorerThe ESTree AST for any sourceSeeing the tree the engine builds, in the same shape bundlers and linters use
run it, all of it
Take any real bundle. node --log-function-events --logfile=v8.log bundle.js, then sort the log by timestamp. The first few hundred lines are the startup story: parse, preparse, compile, first-execute, interleaved. Count how many functions were compiled in the first 100 ms and compare with how many were executed; the gap is eagerness you did not need.