Part 2 · 11 chapters · ~80 min

Parsing

Bytes become characters, characters become tokens, tokens become a tree, and none of it is allowed to fail, which is why HTML is forgiving and why the repairs it makes are the source of a whole class of bugs. This part is the HTML tokenizer and tree builder in detail, scripts and the four ways to load them, the DOM as a data structure, shadow trees, CSS parsing into the CSSOM, what blocks rendering, and how a streamed response lets the top of a page paint before the bottom exists.

23

From bytes to characters

the question

"The response is bytes. How does the parser know what characters they are?"

It guesses, in a defined order, and a wrong guess shows up as mojibake or, worse, a re-parse. The HTML specification calls this encoding sniffing, and the order is fixed.

the order of precedence
  1. A byte order mark. EF BB BF at the start means UTF-8. FF FE or FE FF mean UTF-16. The BOM wins over everything else.
  2. The transport. Content-Type: text/html; charset=utf-8 on the response. This is the one you should be setting, and it is authoritative if present.
  3. Prescan of the first 1024 bytes. The browser looks for <meta charset> or <meta http-equiv="Content-Type"> in the first kilobyte before tokenising anything. Put it first in the head and it is found here.
  4. A meta tag found later. If the declaration appears after 1024 bytes, the parser may have already committed to a guess. On a mismatch it throws away the DOM and reparses from the start with the new encoding. Everything already fetched stays cached, but the time is gone.
  5. Heuristics and locale. With nothing declared, the browser guesses from byte patterns and the user's locale. Windows-1252 is the historical default in many locales. This is where mojibake comes from.
worked numbers
the two lines that make this a non-issue:

  Content-Type: text/html; charset=utf-8   ← on the response
  <meta charset="utf-8">                   ← first thing in <head>

utf-8 everywhere. the decoder is streaming, so bytes become characters as they arrive.
a detail that matters for streaming
The decoder is incremental: it emits characters as bytes arrive and holds back partial multibyte sequences at a chunk boundary. The tokenizer therefore starts on the first network chunk, which is why a server that flushes the head early gets the preload scanner working during its own think time.
24

The tokenizer: a state machine

The HTML tokenizer is a state machine with about eighty states, defined character by character in the specification. It never fails: every possible byte sequence has a defined next state. That is the single most important fact about it, and the reason HTML is forgiving in ways that cause their own problems.

the states you need by name
  1. Data. The default. Characters become character tokens. < transitions to tag open; & starts a character reference.
  2. Tag open, tag name, end tag open. Building a start or end tag token. Letters accumulate; whitespace moves to attributes; > emits.
  3. Attribute name, before/after attribute value, attribute value (double-quoted, single-quoted, unquoted). Three value states because each terminates differently. An unquoted value ends at whitespace or >, which is why class=a b gives you a stray attribute named b.
  4. RAWTEXT, RCDATA, script data. Entered after <style>, <textarea> and <title>, and <script> respectively. In these, < is literal. Only the exact matching end tag exits. RCDATA still decodes entities; RAWTEXT and script data do not.
  5. Markup declaration open. After <!: comment, doctype, or CDATA (only in foreign content).
  6. Character reference. &amp;, &#x27;, and the named references, with their own sub-machine and recovery for missing semicolons.
why the script data state matters to you
Inside <script>, the tokenizer is looking for one thing only: the characters </script. A JavaScript string containing that sequence ends the script element, regardless of quotes. Escape it as <\/script in any inline script, including ones you generate, including JSON embedded in a page. The previous part of this course hit exactly that bug while being built.
THE HTML TOKENIZER
a state machine over characters
swipe the figure sideways, or tap expand for full screen
1/7
data
The tokenizer starts in the data state, consuming characters one at a time. Ordinary characters become character tokens that will be merged into a text node.
25

Tree construction: the stack, the modes, the repairs

The tree builder consumes tokens and produces the DOM. It keeps two pieces of state: a stack of open elements and an insertion mode. The mode says where in the document we are; the stack says what is currently open. Every token is interpreted against both.

the insertion modes, roughly in order
  1. initial, before html, before head, in head, after head. The document skeleton. Missing html, head or body tags are implied; the builder creates them.
  2. in body. Where almost everything happens. Start tags create and push; end tags pop to the match; text appends to the current node.
  3. in table, in table body, in row, in cell. Tables have their own modes because their content model is strict: text directly inside a table is "foster parented" to before the table, which is why stray text in a table appears above it.
  4. in select, in template, in frameset. Special content models.
  5. text. Entered for RAWTEXT and script data, so the content becomes one text node.
  6. after body, after after body. Anything here is an error that gets appended to body anyway.
the repair rules you will actually meet
  1. Implied end tags. <p> is closed by any block-level start tag, <li> by the next <li>, <td> by the next <td> or <tr>. Omitting them is legal, and the parser is faster for it.
  2. The adoption agency algorithm. For mis-nested formatting elements (b, i, a, span is not one). It closes, reopens and clones so that formatting continues across the mis-nest. The clones are real DOM nodes you never wrote.
  3. Foster parenting. Content that is illegal inside a table is moved out to before the table.
  4. Stray end tags. Ignored, or in the case of </p> and </br>, turned into an empty element.
  5. Nested forms, nested anchors, nested buttons. Each has a rule that closes the outer one first.
why React and hydration care
Server-rendered HTML passes through this builder; client-rendered trees do not. If your JSX produces <p><div> or <a><a>, the server-rendered DOM will be repaired and the client's expected tree will not match. That is the "hydration mismatch" warning in its most common form: the parser fixed your markup and React noticed.
TREE CONSTRUCTION
tokens become a DOM, with error recovery
swipe the figure sideways, or tap expand for full screen
1/7
stack + mode
The tree builder holds a stack of open elements and a current insertion mode. It starts in the "initial" mode with an empty stack and a Document node.
26

Scripts: how they stop the parser, and the four ways to load them

the question

"Why is a script in the head so much worse than the same script at the end of body?"

Because the parser must pause for a synchronous script, and a pause in the head happens before anything has been rendered. The script could call document.write and insert arbitrary markup at the current position, so nothing after it can be built until it has run.

AttributeFetchExecuteOrderBlocks parsingUse for
(none)When the parser reaches itImmediately, parser pausedDocument orderYes, fullyAlmost nothing; inline bootstraps only
asyncImmediately, in parallelAs soon as fetched, parser interruptedArrival orderOnly while executingIndependent scripts: analytics, ads
deferImmediately, in parallelAfter parsing, before DOMContentLoadedDocument orderNoApplication code
type=moduleImmediately, with its import graphAfter parsing (deferred)Document orderNoModern application code
type=module asyncWith its graphAs soon as the graph is readyReadiness orderOnly while executingIndependent modules
inline, no srcNoneImmediatelyDocument orderYes, while runningTiny bootstraps, JSON data
the rules around them
  1. A sync script waits for all stylesheets above it. It might read getComputedStyle, so the engine blocks it until the CSSOM is complete. A slow stylesheet plus a sync script stalls the document.
  2. defer scripts run in order even if they arrive out of order. async scripts do not. Mixing them on interdependent code is a race condition.
  3. Inline scripts cannot be async or defer unless they are modules. An inline module is deferred.
  4. document.write after parsing has finished implicitly calls document.open, which wipes the document. From an async script it is ignored with a warning. Chrome blocks document.write of cross-origin scripts on slow connections outright.
  5. DOMContentLoaded fires after parsing and after deferred scripts; load fires after all subresources including images. Neither waits for async scripts.
the default you should write
<script type="module" src="app.js"> in the head. Deferred, ordered, strict, scoped, and the preload scanner sees it. The "scripts at the bottom of body" rule was a workaround from before defer existed.
SYNC, ASYNC, DEFER, MODULE
when a script fetches, when it runs, what it blocks
swipe the figure sideways, or tap expand for full screen
1/6
sync
A synchronous