Parsing
Bytes become characters, characters become tokens, tokens become a tree, and none of it is allowed to fail, which is why HTML is forgiving and why the repairs it makes are the source of a whole class of bugs. This part is the HTML tokenizer and tree builder in detail, scripts and the four ways to load them, the DOM as a data structure, shadow trees, CSS parsing into the CSSOM, what blocks rendering, and how a streamed response lets the top of a page paint before the bottom exists.
From bytes to characters
"The response is bytes. How does the parser know what characters they are?"
It guesses, in a defined order, and a wrong guess shows up as mojibake or, worse, a re-parse. The HTML specification calls this encoding sniffing, and the order is fixed.
- A byte order mark. EF BB BF at the start means UTF-8. FF FE or FE FF mean UTF-16. The BOM wins over everything else.
- The transport.
Content-Type: text/html; charset=utf-8on the response. This is the one you should be setting, and it is authoritative if present. - Prescan of the first 1024 bytes. The browser looks for
<meta charset>or<meta http-equiv="Content-Type">in the first kilobyte before tokenising anything. Put it first in the head and it is found here. - A meta tag found later. If the declaration appears after 1024 bytes, the parser may have already committed to a guess. On a mismatch it throws away the DOM and reparses from the start with the new encoding. Everything already fetched stays cached, but the time is gone.
- Heuristics and locale. With nothing declared, the browser guesses from byte patterns and the user's locale. Windows-1252 is the historical default in many locales. This is where mojibake comes from.
the two lines that make this a non-issue: Content-Type: text/html; charset=utf-8 ← on the response <meta charset="utf-8"> ← first thing in <head> utf-8 everywhere. the decoder is streaming, so bytes become characters as they arrive.
The tokenizer: a state machine
The HTML tokenizer is a state machine with about eighty states, defined character by character in the specification. It never fails: every possible byte sequence has a defined next state. That is the single most important fact about it, and the reason HTML is forgiving in ways that cause their own problems.
- Data. The default. Characters become character tokens.
<transitions to tag open;&starts a character reference. - Tag open, tag name, end tag open. Building a start or end tag token. Letters accumulate; whitespace moves to attributes;
>emits. - Attribute name, before/after attribute value, attribute value (double-quoted, single-quoted, unquoted). Three value states because each terminates differently. An unquoted value ends at whitespace or
>, which is whyclass=a bgives you a stray attribute namedb. - RAWTEXT, RCDATA, script data. Entered after
<style>,<textarea>and<title>, and<script>respectively. In these,<is literal. Only the exact matching end tag exits. RCDATA still decodes entities; RAWTEXT and script data do not. - Markup declaration open. After
<!: comment, doctype, or CDATA (only in foreign content). - Character reference.
&,', and the named references, with their own sub-machine and recovery for missing semicolons.
<script>, the tokenizer is looking for one thing only: the characters </script. A JavaScript string containing that sequence ends the script element, regardless of quotes. Escape it as <\/script in any inline script, including ones you generate, including JSON embedded in a page. The previous part of this course hit exactly that bug while being built.Tree construction: the stack, the modes, the repairs
The tree builder consumes tokens and produces the DOM. It keeps two pieces of state: a stack of open elements and an insertion mode. The mode says where in the document we are; the stack says what is currently open. Every token is interpreted against both.
- initial, before html, before head, in head, after head. The document skeleton. Missing
html,headorbodytags are implied; the builder creates them. - in body. Where almost everything happens. Start tags create and push; end tags pop to the match; text appends to the current node.
- in table, in table body, in row, in cell. Tables have their own modes because their content model is strict: text directly inside a table is "foster parented" to before the table, which is why stray text in a table appears above it.
- in select, in template, in frameset. Special content models.
- text. Entered for RAWTEXT and script data, so the content becomes one text node.
- after body, after after body. Anything here is an error that gets appended to body anyway.
- Implied end tags.
<p>is closed by any block-level start tag,<li>by the next<li>,<td>by the next<td>or<tr>. Omitting them is legal, and the parser is faster for it. - The adoption agency algorithm. For mis-nested formatting elements (
b,i,a,spanis not one). It closes, reopens and clones so that formatting continues across the mis-nest. The clones are real DOM nodes you never wrote. - Foster parenting. Content that is illegal inside a table is moved out to before the table.
- Stray end tags. Ignored, or in the case of
</p>and</br>, turned into an empty element. - Nested forms, nested anchors, nested buttons. Each has a rule that closes the outer one first.
<p><div> or <a><a>, the server-rendered DOM will be repaired and the client's expected tree will not match. That is the "hydration mismatch" warning in its most common form: the parser fixed your markup and React noticed.Scripts: how they stop the parser, and the four ways to load them
"Why is a script in the head so much worse than the same script at the end of body?"
Because the parser must pause for a synchronous script, and a pause in the head happens before anything has been rendered. The script could call document.write and insert arbitrary markup at the current position, so nothing after it can be built until it has run.
| Attribute | Fetch | Execute | Order | Blocks parsing | Use for |
|---|---|---|---|---|---|
| (none) | When the parser reaches it | Immediately, parser paused | Document order | Yes, fully | Almost nothing; inline bootstraps only |
| async | Immediately, in parallel | As soon as fetched, parser interrupted | Arrival order | Only while executing | Independent scripts: analytics, ads |
| defer | Immediately, in parallel | After parsing, before DOMContentLoaded | Document order | No | Application code |
| type=module | Immediately, with its import graph | After parsing (deferred) | Document order | No | Modern application code |
| type=module async | With its graph | As soon as the graph is ready | Readiness order | Only while executing | Independent modules |
| inline, no src | None | Immediately | Document order | Yes, while running | Tiny bootstraps, JSON data |
- A sync script waits for all stylesheets above it. It might read
getComputedStyle, so the engine blocks it until the CSSOM is complete. A slow stylesheet plus a sync script stalls the document. - defer scripts run in order even if they arrive out of order. async scripts do not. Mixing them on interdependent code is a race condition.
- Inline scripts cannot be async or defer unless they are modules. An inline module is deferred.
- document.write after parsing has finished implicitly calls document.open, which wipes the document. From an async script it is ignored with a warning. Chrome blocks document.write of cross-origin scripts on slow connections outright.
- DOMContentLoaded fires after parsing and after deferred scripts; load fires after all subresources including images. Neither waits for async scripts.
<script type="module" src="app.js"> in the head. Deferred, ordered, strict, scoped, and the preload scanner sees it. The "scripts at the bottom of body" rule was a workaround from before defer existed.