Parsing
Before a line of your code runs, every byte of it has been scanned, and the parts that run soon have been scanned twice. This part is the scanner and its three hard cases, lazy parsing and the pre-parser, the recursive descent parser and the AST it builds, scope analysis and the register-or-context decision, where startup time goes and the levers that move it, how function size and placement change the parser's work, and the flags that show all of it.
From bytes to tokens: the scanner
"The bundle is 1 MB. How much of the startup time is the parser, and what in my code makes it worse?"
Parsing is the first thing the engine does to your code and the one cost paid before a single line runs. V8's parser is a hand-written recursive descent parser fed by a scanner that handles the language's lexical oddities; together they produce an AST and a scope tree that Ignition turns into bytecode. Lazy parsing is the trick that keeps a 1 MB bundle from costing a second of parse time.
- Code units in, tokens out. Source is UTF-16 internally (one-byte where possible, part 6). The scanner produces identifiers, keywords, punctuators, numeric and string literals, template parts, regular expression literals, and tracks line terminators.
- Identifiers are internalised. Every identifier becomes a pointer into the string table, so
obj.fooandobj["foo"]and a laterfooin another function all refer to the same string object. Property names are compared by pointer from here on. - Literals are converted once. Numbers become Smis or doubles at scan time; strings get their representation (one-byte or two-byte) decided; both go into the function's constant pool.
- Three context-dependent cases need the parser's help: whether
/begins a regex or is division; where a template literal's${hands control to the expression parser and back; and whether a newline before a token should become a semicolon.
- A semicolon is inserted when the parser meets a token that cannot continue the current statement and a line terminator preceded it, or the token is
}, or at end of input. - Restricted productions: after
return,break,continue,throw, a postfix++/--,yield,async(before a function), and before=>, a newline forces a semicolon regardless.return\nvaluereturns undefined. - The hazards: a line starting with
(,[,`,+,-,/continues the previous line's expression.a = b\n(c)()calls b. Minifiers and formatters that omit semicolons insert a leading;on such lines for this reason.
node -e "function f(){ return\n42 }; console.log(f())" prints undefined. The scanner flagged the newline; the parser inserted the semicolon after return. No warning, by design: ASI is a feature of the grammar, not an error recovery.Lazy parsing: the pre-parser and the double-scan tax
Full parsing builds an AST and compiles to bytecode. Pre-parsing checks syntax, records function boundaries and captured variables, and builds nothing. V8 pre-parses every function it does not need yet, and fully parses it on first call. This halves startup parse time for typical bundles and introduces one cost: a function that is called soon after load is scanned twice.
- Syntax check. A SyntaxError inside a never-called function is still thrown at load, because the pre-parser saw it. The grammar is the same; the output is not.
- Positions. Start and end offsets of each function, so the full parser can jump straight to it later.
- Scope data. Which variables inner functions reference from outer scopes. Needed so the outer function's context layout can be decided even when the inner functions are not compiled yet. Stored as "preparse data" on the function and reused by the full parse.
- Speed. Roughly twice the throughput of the full parser, because there is no tree to allocate and no bytecode to emit.
- Parenthesised functions.
(function(){...})and(()=>{...})are compiled eagerly: the parentheses are read as "this will be called immediately" (the PIFE heuristic: Possibly-Invoked Function Expression). Webpack, Rollup and esbuild emit module wrappers this way on purpose. - Immediately invoked in the same statement, with or without parentheses on some forms.
- Functions in a function being compiled eagerly may inherit eagerness one level down in some heuristics (the "eager inner" rule), so a module wrapper's direct children are often compiled at load too.
- The anti-pattern: a minifier that wraps every function in parentheses, or a bundler that emits one giant eager wrapper around everything, defeats laziness and compiles the whole bundle at load. Check with
--trace-parseor the Performance panel's "Compile Code" events.
- A function called at load is pre-parsed (because the parser did not know) and then fully parsed on the call. Two scans of the same bytes. For the third of a bundle that typically runs at load, this is tens of milliseconds on a phone.
- Chrome's code cache: after a script runs, V8 serialises the bytecode of compiled functions; the browser stores it with the HTTP cache entry (after the second load for most scripts; immediately for service-worker-cached ones with the right flags). The next load deserialises instead of parsing: both scans gone for cached functions.
- Node:
module.enableCompileCache()(v22+) does the same for CommonJS and ESM on disk;v8.startupSnapshotgoes further by serialising the heap after initialisation.
node --trace-parse script.js 2>&1 | head lists each function as it is pre-parsed or parsed and how long it took. A function that appears twice (preparse, then parse) within the first milliseconds was called at load and is a candidate for the eager hint; a function that appears once and never again was never called.The parser: recursive descent, and what the AST holds
V8's parser is a hand-written recursive descent parser with a few precedence-climbing loops for binary operators. Each grammar production is a C++ method; the AST is a tree of typed nodes; the parser also builds the scope tree and resolves every variable reference as it goes. The AST lives only long enough for Ignition to walk it.
- ParseProgram → ParseStatementList → ParseStatement → ... down to ParsePrimaryExpression. Each level consumes tokens from the scanner and returns a node. Lookahead is one token, with a few places that peek two.
- Expressions by precedence climbing:
ParseBinaryExpression(prec)parses operands and loops on operators of at least the given precedence, which is howa + b * cgets the right tree without a method per precedence level. - Arrow functions and destructuring are the hard part.
(a, b) => astarts like a parenthesised expression; the parser parses it as one, then on seeing=>reinterprets the expression as a parameter list ("cover grammar"). Same for({a, b} = obj). This is why some syntax errors are reported at the arrow rather than inside the parens. - Templates, regex, ASI: the scanner interactions from chapter 1.
- Error recovery: none. The first error aborts the script with a SyntaxError that includes the position. Unlike the HTML parser, the JavaScript grammar has no recovery rules.
- Nodes for every construct: FunctionLiteral, Block, VariableDeclaration, Assignment, Call, Property, BinaryOperation, Conditional, ForStatement, ObjectLiteral with its properties, ClassLiteral with its members, and so on. Each carries source positions for stack traces and the debugger.
- Literal values already converted: a NumberLiteral holds a double or Smi; a StringLiteral holds the internalised string.
- Object literal "boilerplates": for
{x: 1, y: 2}, the parser pre-computes the shape (hidden class) and the constant properties, so creating the object at runtime is a copy of a template rather than two property additions. Part 3 is why that matters. - Not in the AST: types (there are none), and anything about how the code will run. The AST is syntax plus scope.
npx acorn --ecma2024 file.js or the AST Explorer site. The node types are the same ones the engine uses, which is why bundlers, linters and the engine agree on what the code means.Scope analysis: registers or context
As the parser builds the AST it also builds a scope tree and resolves every identifier to the scope that declares it. One outcome of that resolution decides more about performance than most people expect: whether a variable can live in a register or must live in a heap-allocated context object because a closure captures it.
- Script or module scope: top-level declarations. Module scope is a real lexical scope; classic script top-level
varand function declarations become properties of the global object (slow: a global lookup is a property lookup on a dictionary-mode object unless the engine has cached it). - Function scope: parameters,
var, function declarations,arguments,this. - Block scope:
let,const,class, and the per-iteration bindings offor (let ...). The parser marks loop variables as per-iteration so closures created in the loop body see a fresh binding each time, which costs a context per iteration if a closure captures the variable. - Catch scope, with scope (dynamic; disables resolution for everything inside), eval scope (a direct
evalcan declare variables at runtime, so every variable in enclosing scopes must be context-allocated and looked up dynamically).
- Each reference is resolved at parse time to a declaration in some enclosing scope, or to "global" if none. Dynamic lookups exist only under
withand sloppy directeval. - A variable referenced only from its own function is stack-allocated: a register in Ignition's register file, a machine register or stack slot in optimised code. Free to read and write.
- A variable referenced from an inner function is context-allocated: a slot in a Context object created when the declaring scope is entered. Reads and writes go through the context pointer. The closure holds the context; the context holds all captured variables of that scope, not just the ones this closure uses.
- The sharing consequence: two closures from the same scope share one context. A small long-lived closure keeps alive every captured variable of its scope, including a large one captured only by a different, short-lived closure. Part 8 and the memory part return to this.
a function with and without capture:
function sum(arr) { let t = 0; for (const x of arr) t += x; return t } // t: register. fast.
function sum(arr) { let t = 0; arr.forEach(x => { t += x }); return t } // t: context slot. the arrow captures it.
--print-bytecode shows the difference: Ldar r1 / Add vs LdaCurrentContextSlot [2] / Add / StaCurrentContextSlot [2]
neither is wrong. the second allocates a context per call and does a memory load per iteration; the optimiser often removes the difference, and sometimes cannot.Startup cost: measuring and reducing it
For a browser page, the engine's part of startup is scan, parse, compile, and the execution of top-level code. The host adds download, the preload scanner's ordering, and the main thread's other work. This chapter is the accounting and the levers, in order of effect.
- Ship less. Every engine cost is linear or worse in bytes. Route-level code splitting, removing unused dependencies, and not shipping polyfills to browsers that do not need them. Nothing else comes close.
- Run less at load. Top-level execution is often the biggest bar and is entirely yours: framework initialisation, module side effects, eager data fetching, synchronous hydration of the whole page. Defer whatever is below the fold or behind an interaction.
- Let the code cache work. Stable URLs for code that rarely changes (vendor chunks), so the cached bytecode survives deploys. A content hash that changes on every release invalidates the cache for every user on every release.
- Streaming parse. External scripts parse on a background thread as they download. Inline scripts and
evalcannot. Large inline scripts are a startup anti-pattern for this reason as well as caching. - Eager where it counts, lazy elsewhere. Module wrappers parenthesised (bundlers do this); everything else lazy. Verify with a trace that the bundle is not being compiled wholesale.
- Avoid the parser's slow paths. Direct
evalandwithforce dynamic scope for everything around them.argumentsin sloppy mode aliases parameters and forces a materialised object. Each is a per-function cost, but in a hot path it compounds.
- Chrome Performance panel: under each script's Evaluate Script, "Compile Script" (top-level) and "Compile Code" (lazy functions compiled on first call). The sum across the load is the engine's compile cost. Long "Compile Code" events after load mean big functions compiled late; many small ones in a burst mean a hot path being compiled piecemeal.
- Lighthouse: "Reduce JavaScript execution time" and "Minimize main-thread work" attribute time per script to parse, compile and execute.
- Node:
--cpu-profand look at the first frames;process.hrtimearoundrequirecalls for a quick per-module cost;NODE_DEBUG=modulefor load order. - Script streaming in the trace: "v8.parseOnBackground" events show background parsing; if they are absent for a large script, it was inline or loaded in a way that prevented streaming.
import(). Record both with 4× CPU throttling. The difference in the Compile and Evaluate bars is the number to put in the pull request.Function size, placement, and the parser
Three things about how code is laid out change what the parser does with it: how big functions are, where they are, and whether they are inside something that forces eagerness. None of these changes behaviour; all of them change startup.
- Big functions compile slowly and all at once. A 2,000-line function is compiled as a unit on first call: one long "Compile Code" event on the main thread. Splitting it into smaller functions lets the uncalled parts stay pre-parsed.
- Tiny functions have a floor. Each function has a SharedFunctionInfo, bytecode, a feedback vector, possibly a closure object. Thousands of one-line functions cost memory and per-call overhead until inlined. This is why bundlers inline trivial helpers.
- The optimiser has a size limit for inlining (bytecode length) and for optimisation at all. A huge function may never be optimised by TurboFan; part 4 has the thresholds.
- Inner functions are re-parsed with their outer function. When
outeris compiled lazily, every function inside it is pre-parsed again (the parser must walk through them). A deeply nested structure pays the pre-parse of inner functions once per enclosing level that gets compiled. Flat modules parse faster than deeply nested closures. - Top-level code in modules is always eager. A module's body is compiled at link time whether or not anything is exported. Side-effect-free modules with heavy top-level computation should defer it into a function.
- Class bodies are compiled with the class: method bodies are lazy, but the class definition (field initialisers, static blocks) runs at definition time.
- Parentheses (the PIFE hint), immediate invocation, and in some versions functions assigned to properties in object literals that are immediately used.
- A direct
evalin a function makes every enclosing function's variables dynamically scoped and disables lazy parsing optimisations around it. - The
new Functionandevalpaths create a fresh parse each time, with no code cache and no streaming; a template engine that compiles templates withnew Functionat runtime pays full parse cost per template per load.
Ldar and LdaContextSlot there.Reading the parser: flags and traces
| Tool | Shows | Use it for |
|---|---|---|
node --trace-parse script.js | Each parse and pre-parse with function name, size and time | Which functions were compiled at load; the double-scan candidates |
node --trace-lazy script.js | Lazy compilation events: which function, when | Confirming a function stayed lazy until its first call |
node --log-function-events --logfile=v8.log | A log of parse, preparse, compile, first execution per function with timestamps | The full startup timeline per function; load it into a spreadsheet |
| Chrome Performance → Evaluate Script → Compile Script / Compile Code | Main-thread compile events per script and per lazy function | Where compile time lands on the critical path |
| Chrome Performance → "Streaming compile" / v8.parseOnBackground | Background parsing of external scripts | Whether a script streamed; inline scripts never do |
| Lighthouse → Reduce JavaScript execution time | Per-script parse, compile, execute totals | The ranked list of scripts to split or defer |
node --allow-natives-syntax -e "%GetOptimizationStatus(f)" (bit 3) | Whether a function has been compiled at all, and to which tier | "Is this function still lazy" without a trace |
npx acorn / AST Explorer | The ESTree AST for any source | Seeing the tree the engine builds, in the same shape bundlers and linters use |
node --log-function-events --logfile=v8.log bundle.js, then sort the log by timestamp. The first few hundred lines are the startup story: parse, preparse, compile, first-execute, interleaved. Count how many functions were compiled in the first 100 ms and compare with how many were executed; the gap is eagerness you did not need.