Part 14 · 11 chapters · ~75 min

Build Your Own Browser

Thirteen parts of mechanism, then the mechanism in code. The companion app is a browser engine in about 1,400 lines of TypeScript, with a React shell that shows every stage side by side and lets you click any element to follow it through. This part walks the files in pipeline order: tokenizer, tree builder, DOM and CSSOM, cascade, block and inline layout, display list and raster, the event loop and its sandbox, the glue, and then the exercises that turn each earlier part into an afternoon of code.

133

What we are building, and what we are leaving out

the question

"Thirteen parts of theory. Can I see the pipeline run, in code I can read end to end?"

Yes. The companion app at tiny/dist/ is a browser engine in about 1,400 lines of TypeScript: an HTML tokenizer, a tree builder, a DOM, a CSS parser, the cascade, block and inline layout, a painter to canvas, and an event loop with a script sandbox. A React shell shows every stage's output side by side and lets you click any element to see how it was tokenised, where it sits in the tree, which rules matched and why they won, what box it became and where, and which paint commands drew it.

the files, in pipeline order
  1. engine/tokenizer.ts: the state machine. 22 states, a transition log, raw text for script and style, entity decoding.
  2. engine/treebuilder.ts: the stack of open elements, four insertion modes, implied html/head/body, void elements, the closes-p rule, misnested end tags.
  3. engine/dom.ts: Document, Element, Text, Comment, inline style, a serialiser.
  4. engine/css.ts: stylesheet parser, selectors with specificity, shorthand expansion, the user agent stylesheet.
  5. engine/cascade.ts: right-to-left matching, the cascade sort, inheritance, computed values, a record of every decision.
  6. engine/layout.ts: the box tree with anonymous blocks, block layout with the auto width equation and margin collapsing, inline layout with line boxes measured by canvas.
  7. engine/paint.ts: a display list in paint order, rasterisation, paint flashing, the inspector overlay.
  8. engine/eventloop.ts: tasks, microtasks, rAF, timers, render-when-dirty, the sandbox, forced layout.
  9. engine/pipeline.ts, App.tsx, samples.ts: the glue, the shell, five documents to try.
deliberately absent
  1. Network. The HTML comes from the editor. Part 1 is a course of its own and a tokenizer that streams is an exercise below.
  2. A JavaScript engine. Scripts run in the host's V8 through new Function against our DOM and our queues. The JS course builds the interpreter.
  3. Flex, grid, floats, positioning, tables. Block and inline formatting contexts only. Flex is the first exercise.
  4. Images, fonts, forms. No decode, no font loading (we use whatever the canvas has), no form controls.
  5. The adoption agency algorithm. <b>bold <i>both</b> italic?</i> pops both; a real browser reopens the <i>. The "Misnested tags" sample shows the difference.
  6. A compositor. One canvas, repainted in full. Layers and tiles are part 5's story and an exercise.
how to read this part
Open the app in another tab. Each chapter names a file and quotes the lines that matter; the app's panes show those lines running. Change the sample, break something, watch which stage notices.
134

tokenizer.ts: characters in, tokens out

The tokenizer is a loop over characters with a current state. Each state is a case in a switch; each case looks at one character and either accumulates it, changes state, or emits a token. The spec has about 80 states; ours has 22, which covers everything the samples need and keeps the shape honest.

code
case 'Data':
  if (ch === '<') { flushText(); go('TagOpen', ch) }
  else { if (!text) textStart = pos; text += ch }
  break
case 'TagOpen':
  if (ch === '!') go('MarkupDeclarationOpen', ch)
  else if (ch === '/') go('EndTagOpen', ch)
  else if (/[a-zA-Z]/.test(ch)) { tag = { name: ch.toLowerCase(), attrs: [], ... }; go('TagName', ch) }
  else { text += '<' + ch; go('Data', ch) }   // "<" followed by a non-letter was text after all
  break
the things worth noticing
  1. Text is buffered, not emitted per character. text accumulates until a < arrives, then flushes as one Character token. Real tokenizers do the same, and that is why a 1 MB text node is one token.
  2. "<" is not always a tag. TagOpen only becomes TagName on a letter. a < b in text is text. The recovery branch appends the < back to the text buffer and returns to Data. No error thrown, ever: the tokenizer has no failure mode, only recovery.
  3. Attributes are collected on the tag record and only attached when the tag is emitted, which is why an unclosed tag at EOF loses its attributes along with itself.
  4. The transition log (go() records from, to, the character, and any emitted token) costs nothing to keep and is what the app's "show state transitions" shows. Real engines have the same information in their tracing.
raw text: the one place the tokenizer looks back
code
const emitTag = (ch) => {
  tokens.push(tag.end ? { type: 'EndTag', name } : { type: 'StartTag', name, attrs, selfClosing })
  // the one place the tokenizer depends on the tag it just emitted
  if (!tag.end && RAW_TEXT_ELEMENTS.has(tag.name)) { rawTag = tag.name; go('RawText', ch) }
  else go('Data', ch)
}
// ...
case 'RawTextEndTagName': {
  // only the matching end tag ends raw text; "</div>" inside <script> is script text
  const maybe = input.slice(pos, pos + rawTag.length).toLowerCase()
  if (maybe === rawTag && /[>\s]/.test(input[pos + rawTag.length] ?? '')) { ...emit Character(rawBuf), EndTag(rawTag)... }
  else { rawBuf += '</' + ch; go('RawText', ch) }
}

After emitting a <script> or <style> start tag, the tokenizer enters RawText, where < does not open a tag. Only the exact matching end tag leaves it. This is why "</script>" inside a JavaScript string ends the script block, why this course's own build tool escapes it as <\/script, and why a CSS file can contain <div> in a comment without the parser caring.

in the app
Tokens pane, "A page" sample. Find the Character token for the script body: one token, the whole script, because RawText swallowed it. Then edit the source: put </script> inside the console.log string and parse again. The script token ends early and the rest is HTML.
THE TOKENIZER AS A STATE MACHINE
tokenizer.ts on "&lt;p class=&quot;a&quot;&gt;Hi&lt;/p&gt;"
swipe the figure sideways, or tap expand for full screen
1/6
<
State Data. The character is "<". Flush any pending text as a Character token (none yet). Move to TagOpen.
135

treebuilder.ts: the stack of open elements

Tree construction takes the token stream and builds the DOM, using one data structure (a stack of elements that are open) and a set of rules about what each token does given the current insertion mode and the top of the stack. The rules are where "sloppy HTML always produces a tree" comes from.

the four modes we keep (of twenty-three)
  1. Initial: before anything. A doctype is noted; whitespace is ignored; anything else implies <html>.
  2. InHead: after <html>. Title, meta, link, style and script go into <head>, created on demand. Any other element, or text, closes head and opens body.
  3. InBody: where almost everything happens. Start tags push, end tags pop, and the special rules below apply.
  4. AfterBody: after </body>. Anything that follows is quietly put back into body, which is why content after </html> still renders.
code
let note = ''
if (CLOSES_P.has(name) && stack.some(e => e.tagName === 'p')) {
  popUntil('p')
  note = ' (open <p> closed first: block cannot nest in p)'
}
if (name === 'li' && current().tagName === 'li') { stack.pop(); note = ' (previous <li> closed first)' }
const el = insertElement(name, t.attrs)          // appends under current(), pushes
if (VOID_ELEMENTS.has(name) || t.selfClosing) { stack.pop(); note += ' (void: popped immediately)' }
record(t, `<${name}> inserted under <${parentName}>${note}`)
the rules in that excerpt
  1. CLOSES_P. A block-level start tag (div, p, h1, ul, section...) while a <p> is open pops the p first. A paragraph cannot contain a block, so the parser ends the paragraph. This is the single most common "why is my DOM not my markup" surprise.
  2. Implied li/dt/dd closing. A new <li> while one is open closes the open one. Lists without end tags are valid HTML.
  3. Void elements (br, img, input, meta...) are pushed and popped in one step. They can have no children, so an </img> later has nothing to match and is ignored.
  4. Self-closing on a non-void element (<div/>) is honoured here for clarity. The real parser ignores the slash on non-void elements, which is a common gotcha with inline SVG-style markup.
code
if (current()?.tagName === name) { stack.pop(); record(t, `</${name}> matches current node: popped`) }
else if (popUntil(name)) record(t, `</${name}> did not match: popped until <${name}> closed (misnested)`)
else record(t, `</${name}> with nothing to close: ignored`)

End tags have three outcomes: match the current node and pop; match something deeper and pop everything above it (misnesting, like <b><i></b>); or match nothing and be ignored. The spec's adoption agency algorithm handles the formatting-element case (b, i, a, em, strong...) more cleverly, by reopening the inner element after the pop so italic? in the sample stays italic. That is the one named algorithm this builder omits.

in the app
Tree build pane, "Misnested tags" sample. Read each step's action and the stack after it. Then the DOM pane: every element is where a real browser would put it, except the trailing "italic?", which the adoption agency would have wrapped in a second <i>.
THE STACK OF OPEN ELEMENTS
treebuilder.ts on "&lt;p&gt;One&lt;p&gt;Two&lt;div&gt;x&lt;/div&gt;&lt;/p&gt;"
swipe the figure sideways, or tap expand for full screen
1/6
implied html/head/body
Doctype and whitespace first: no nodes. The first

start tag, in mode Initial, implies , , and then , each pushed. Then

is inserted under body and pushed. Stack: html, body, p.

136

dom.ts and css.ts: the two trees the cascade joins

the DOM (dom.ts)
  1. Four node kinds: Document, Element, Text, Comment, each with parent, children, and a generator over descendants. Enough for selector matching (which needs parentElement and attributes) and for layout (which needs children in order).
  2. textContent as a getter and setter, because the sandbox exposes it to scripts and a setter that replaces children is the cheapest mutation to demonstrate.
  3. Inline style is parsed from the style attribute at setAttribute time into a Map, and written back when a script sets el.style.x. The cascade reads the Map; the serialiser writes the attribute.
  4. Constants: VOID_ELEMENTS and CLOSES_P live here because both the tree builder and anyone reading the DOM need them.
the CSSOM (css.ts)
  1. Rules are found by scanning for { and } after stripping comments. At-rules (@media, @font-face) are skipped: a real parser nests them; a tiny one notes they exist.
  2. Selectors are tokenised into compounds and combinators. Each compound can have a type, an id, classes and attribute selectors. Descendant (space) and child (>) combinators only; a pseudo-class drops the whole selector rather than matching wrongly.
  3. Specificity is computed at parse time as the (id, class, type) triple and stored on the selector. The app shows it next to each matched rule.
  4. Shorthands (margin, padding, border, background, font) are expanded into longhands at parse time, so the cascade only ever sees longhands. That is what real engines do, and it is why margin: 0 followed by margin-top: 1em works: four longhands, then one of them overridden.
  5. The UA stylesheet is a string at the bottom of the file: display: block for div and p, display: none for head and script, default margins on p and headings, bold on b and strong. Remove a line and parse again; watch what breaks. Every browser ships a version of this file, and it is the reason a bare <h1> is big.
in the app
Styles pane, any element. Rules marked ua come from that string. The struck-through declarations are the ones the cascade rejected; the next chapter is why.
137

cascade.ts: matching, sorting, inheriting, computing

For every element, in tree order: test every selector, collect the declarations of the rules that matched, sort them, pick a winner per property, fill in what nothing declared from the parent or the initial value, and turn relative lengths into pixels. That is the whole of style resolution, and it is the stage with the most interesting numbers.

code
// Right to left: the rightmost compound must match the element itself; then
// walk the ancestors for each combinator. Most selectors fail here and the walk never happens.
export function matchesSelector(el, sel) {
  const parts = sel.parts
  let idx = parts.length - 1
  if (!matchesCompound(el, parts[idx].compound)) return false
  let node = el
  while (idx > 0) {
    const comb = parts[idx].combinator
    idx--
    const target = parts[idx].compound
    if (comb === '>') { node = node.parentElement; if (!node || !matchesCompound(node, target)) return false }
    else { node = node.parentElement; while (node && !matchesCompound(node, target)) node = node.parentElement; if (!node) return false }
  }
  return true
}
right to left
  1. The rightmost compound is tested first, against the element itself. For .card p on a <div>, the test p fails immediately and the ancestor walk never runs. For the 976 selector tests the "A page" sample performs, most end here.
  2. Then each combinator walks up: child means exactly the parent; descendant means any ancestor, searched upward. The app's stats line reports the count: tests versus matches.
  3. Real engines add a bloom filter of ancestor ids, classes and tags so a descendant selector that cannot possibly match is rejected without the walk. Part 3 covered it; adding one here is an exercise.
code
// cascade sort: ascending by (origin+importance, specificity, order); the last one wins
const rank = (d) => {
  const originRank = d.important ? (d.source === 'ua' ? 5 : 4) : d.source === 'inline' ? 3 : d.source === 'author' ? 2 : 1
  return [originRank, d.specificity, d.order]
}
for (const prop of KNOWN) {
  const list = (byProp.get(prop) ?? []).sort(compareRank)
  const winner = list[list.length - 1]
  if (winner && winner.value !== 'inherit') cs.set(prop, { value: winner.value, from: 'declared', winner, candidates: list })
  else if ((INHERITED.has(prop) || winner?.value === 'inherit') && parentStyle?.has(prop)) cs.set(prop, { value: parentStyle.get(prop).value, from: 'inherited', candidates: list })
  else cs.set(prop, { value: INITIAL[prop], from: 'initial', candidates: list })
}
// computed values: font-size first (em on font-size refers to the parent), then lengths relative to it
const fs = toPx(cs.get('font-size').value, parentFs, parentFs)
for (const [prop, res] of cs) if (/^(margin|padding|border-.*-width|width|height)/.test(prop) && res.value !== 'auto' && !/%$/.test(res.value)) res.value = `${toPx(res.value, fs)}px`
the sort key, and what it means
  1. Origin and importance first. Normal declarations: UA, then author, then inline. Important declarations invert: author important beats inline, UA important beats everything (that is how display: none on <head> could be made unoverridable).
  2. Specificity second. The triple from the parser, as one number. Inline style gets a specificity above any selector.
  3. Source order last. Later rule wins. The sort is stable and ascending, and the last element of the list is the winner, which is the natural way to express "later beats earlier".
  4. Nothing matched? If the property is in INHERITED (color, font-*, line-height, text-align) or the winner said inherit, take the parent's computed value. Otherwise the initial value. The app marks each computed row as declared, inherited or initial.
  5. Computed values. font-size resolves first because em on it refers to the parent's size; then every other length resolves against this element's font-size. Percentages are left for layout, because they need the containing block's width, which does not exist yet.
in the app
"The cascade" sample. Click each paragraph in turn and read the Styles pane: the winner is uppermost, the losers are struck through, and the computed table says where each value came from. The one that says inherited for color is the span.
138

layout.ts, part one: boxes and blocks

Layout turns the styled DOM into a tree of boxes with geometry. It is two functions: buildBoxTree, which decides what boxes exist, and layoutBlock, which decides where they go. Inline content gets its own function in the next chapter.

the box tree
  1. One box per element with display block or inline; none for display: none and its subtree. Text nodes become text boxes.
  2. Anonymous block boxes. If a block container has both block and inline children, each run of inline children is wrapped in an anonymous block, so that every block container has either all-block or all-inline children. This is CSS 2.1 §9.2.1.1 and it is why the Boxes pane shows "anonymous block" rows that have no element.
  3. Whitespace-only text between blocks is dropped; inside an inline run it is kept and collapsed later.
code
// width: auto fills the containing block minus the horizontal edges (CSS 2.1 §10.3.3)
if (wv === 'auto') width = cb.width - horiz - box.margin.left - box.margin.right
else {
  width = px(wv, cb.width)
  const free = cb.width - width - horiz
  if (ml === 'auto' && mr === 'auto') box.margin.left = box.margin.right = free / 2   // centred
  else if (ml === 'auto') box.margin.left = free - box.margin.right
  else box.margin.right = free - box.margin.left                                     // over-constrained: margin-right gives
}
// vertical margin collapsing with the previous sibling: the larger wins
let collapsedTop = box.margin.top
if (prevMarginBottom !== null) collapsedTop = Math.max(box.margin.top, prevMarginBottom) - prevMarginBottom
box.content = { x: cb.x + box.margin.left + box.border.left + box.padding.left, y: cb.y + collapsedTop + box.border.top + box.padding.top, width, height: 0 }
block layout, line by line
  1. Width. auto means: the containing block's width, minus this box's horizontal margins, borders and padding. That is the equation that makes a div stretch and a nested div stretch inside it. An explicit width leaves free space, and the auto margins split it: both auto, centred; one auto, it takes all; neither, the right margin absorbs the difference (the over-constrained case).
  2. Position. x is the containing block's x plus the left margin, border and padding. y is the containing block's current y plus the top margin, collapsed with the previous sibling's bottom margin.
  3. Margin collapsing. Adjacent vertical margins between siblings collapse to the larger. The code tracks the previous sibling's bottom margin and adds only the excess of this box's top margin over it. Padding or a border on the parent stops collapsing with the parent, which is why the sample's first child sits 30 from the top of its padded parent and not 20.
  4. Children, then height. Block children are laid out top to bottom, each told the current y; inline children go through layoutInline. Height is the sum of what the children used, unless height was set.
  5. The return value is the total vertical space this box used in its parent's flow, including its collapsed top margin and its bottom margin, so the parent can advance its cursor.
in the app
"Layout: blocks and lines" sample, Boxes pane, "show layout log". Each line is one decision with its numbers. Click .fixed and read the box model diagram: 200 wide, margins computed to centre it, a 1px border counted in the border box.
FROM DOM TO BOXES TO LINES
layout.ts on a div with mixed children
swipe the figure sideways, or tap expand for full screen
1/6
DOM
DOM: div.card > ["Hello ", b > ["world"], div.inner > ["block"]]. Mixed children: inline content and a block.
139

layout.ts, part two: lines

Inline layout fills a block container's width with its inline content, one line box at a time. The shaping and measurement that part 11 described in detail is borrowed from the host canvas: measureText with the computed font gives the advance width of a run, and that is enough to break lines correctly.

code
const words = raw.trim().split(/\s+/)
if (leading && cursorX > 0) cursorX += space        // collapsed whitespace: a leading space only matters mid-line
words.forEach((w, idx) => {
  const text = idx > 0 ? ' ' + w : w
  let tw = ctx.measure(text, font)                 // canvas measureText: the shaper we borrow
  if (cursorX + tw > width && line.items.length) {   // does not fit, and the line is not empty
    flush()                                        // close the line box, start the next
    tw = ctx.measure(w, font)
  }
  line.items.push({ box, x: cursorX, width: tw, text, font, fontSize: fs, color, baseline: fs * 0.8 })
  cursorX += tw
  line.height = Math.max(line.height, lh)
})
if (trailing) cursorX += space
what the loop does
  1. Whitespace collapsing. A text node's internal whitespace becomes single spaces (the split). A leading space is honoured only mid-line; a trailing one is carried to the next run so was <b>tokenised</b>, keeps its space and loses nothing.
  2. Measure, then place. Each word (with its preceding space) is measured in the run's font. If it fits in the remaining width, it is appended at the cursor. If not, and the line is not empty, the line is flushed and the word starts the next line without its leading space.
  3. Line height. Each line box is as tall as the tallest run on it (the computed line-height of that run's style). A bold 24px word in a 16px paragraph makes that one line taller, which is the behaviour you see in a real page.
  4. Text align is applied at flush: the whole line's items are shifted by half or all of the remaining space.
  5. Inline element boxes (<b>, <em>) contribute their horizontal padding, border and margin to the cursor at their start and end, and get a fragment per line they span: the union of their items on that line. The painter draws their background per fragment, which is why an inline element that wraps gets two background rectangles with no vertical padding between them.
what a real engine adds
  1. Shaping. We measure words; HarfBuzz shapes runs with ligatures, kerning and contextual forms. Our line breaks are correct for Latin text and wrong for anything where the shaped width differs from the sum of words.
  2. Break opportunities. We break at spaces. UAX #14 breaks after hyphens, before CJK characters, never inside a non-breaking space, and so on.
  3. Bidi. We place left to right. A real line can have runs in both directions, reordered after breaking.
  4. Vertical alignment and baselines. We centre each run in its line box and put the baseline at 80% of the font size. Real engines align baselines from font metrics and honour vertical-align.
  5. A shaping cache. Blink caches shaped words; we call measureText every time. For a few paragraphs it does not matter; for a document it would.
in the app
"Layout: blocks and lines" sample. Click the <em> that wraps: two fragments in the Boxes pane, two rectangles on the canvas. Narrow the .narrow width in the source and parse again; watch the line count change in the layout log.
140

paint.ts: a display list, then pixels

Painting is two stages on purpose. First, walk the layout tree in paint order and record a list of drawing commands, with no canvas involved. Second, replay the list onto a canvas. The list is what the Display list pane shows, and separating it from rasterisation is what lets a real engine repaint without re-deciding what to paint, raster on another thread, and split the output into layers.

code
// CSS 2.1 appendix E, abbreviated: for each block, its background and border,
// then its block children in order, then its inline content.
const paintBlock = (box) => {
  const bb = borderBox(box)
  if (bg !== 'transparent') commands.push({ op: 'rect', rect: bb, color: bg, box, why: 'background' })
  if (hasBorder) commands.push({ op: 'border', rect: bb, widths, colors, box })
  if (box.lines) {
    for (const b of inlineBoxesIn(box)) for (const f of b.fragments) commands.push({ op: 'rect', rect: f, color: bgOf(b), box: b })
    for (const ln of box.lines) for (const it of ln.items) commands.push({ op: 'text', x: box.content.x + it.x, y: ln.y + it.baseline + (ln.height - it.fontSize) / 2, text: it.text, font: it.font, color: it.color, box: it.box })
  } else for (const c of box.children) paintBlock(c)
}
paint order
  1. A block's background, then its border, then its children. So a child's background paints over its parent's. CSS 2.1 Appendix E has the full order with floats, positioned elements and z-index; with only normal flow, this is all of it.
  2. Within an inline formatting context: inline backgrounds first (per fragment), then the text, line by line. That is why a span's background sits under its text and under nothing else.
  3. Text position. x is the container's content x plus the item's x on the line. y is the line's y plus the baseline offset, adjusted to centre the run's font size within the line height. The canvas is told textBaseline = 'alphabetic' so the y means the baseline.
rasterise, and the two overlays
  1. rasterise(list, ctx) clears to white and replays each command: fillRect for backgrounds, four rects for the border sides, fillText for text with the item's font and colour. The canvas is sized to the viewport width times devicePixelRatio and scaled, so text is sharp on a retina screen (the DPR lesson from part 11).
  2. Paint flashing. After each render, every box that was painted is tinted green for half a second: the same idea as DevTools' Rendering → Paint flashing. Because this engine repaints everything, everything flashes; a real engine would flash only the invalidated area, and making that true here is an exercise.
  3. The highlight overlay. The selected element's margin (orange), border box (green) and content (blue) are drawn on top, like the Elements panel's hover overlay. For inline boxes, each fragment is highlighted.
in the app
Display list pane. Count the commands; note that the first is the <html> background only if one was set, and that text commands carry their resolved font string. Click the painted page: the hit test walks the box tree and picks the last box whose border box or fragment contains the point, which is the same rule as the paint order in reverse.
141

eventloop.ts: tasks, microtasks, frames, and a sandbox

The event loop is what makes the engine a browser rather than a renderer. Scripts run as tasks; their promise reactions run as microtasks before the task ends; rendering happens at most once per frame and only if something changed; timers become tasks when due. The simulated clock makes the ordering visible without waiting.

code
step(): boolean {
  for (const t of dueTimers()) this.tasks.push({ label: t.label, run: t.run })   // 1. promote
  const hadTask = this.tasks.length > 0
  if (hadTask) {                                                                 // 2. one task
    const t = this.tasks.shift()!
    this.emit('task', t.label); t.run(); this.time += 2
    this.drainMicrotasks()                                                       // 3. all microtasks
  }
  const frameDue = this.time % this.frameInterval < 2 || !hadTask
  if (frameDue) {                                                                // 4. rendering opportunity
    for (const r of this.takeRafs()) { this.emit('raf', r.label); r.run(this.time); this.drainMicrotasks() }
    if (this._dirty) { this.hooks.render(this.dirtyReasons.join(', ')); this._dirty = false }
    if (!hadTask) this.time += this.frameInterval
  }
  return this.tasks.length + this.microtasks.length + this.rafs.length + this.timers.length > 0 || this._dirty
}
what the five steps enforce
  1. One task at a time. A task runs to completion; nothing interrupts it. A script that loops forever would hang the loop, exactly as it hangs a tab.
  2. Microtasks drain completely after every task and after every rAF callback. A microtask that queues a microtask runs in the same drain; the runaway guard at 10,000 is the only thing a real browser does not have.
  3. Rendering is an opportunity, not a task. It happens when a frame is due (every 16 simulated milliseconds) or when the loop is otherwise idle, and only if the dirty flag is set. Three mutations in one task produce one render. This is the single most important fact about browser performance and it falls out of six lines.
  4. rAF callbacks run before the render, inside the same frame, so a rAF that mutates the DOM is rendered in that frame, not the next one.
  5. Timers are promoted, not run. A due timer becomes a task at the front of the next step. setTimeout(fn, 0) is not immediate; it is "after the current task and its microtasks, and after any task already queued".
code
// the script sees only these; every mutation marks the loop dirty with a reason
const fn = new Function('document', 'console', 'setTimeout', 'clearTimeout', 'queueMicrotask', 'requestAnimationFrame', 'Promise', src)
fn(sandbox.document, sandbox.console, sandbox.setTimeout, ..., sandbox.Promise)

class ScriptElement {
  set textContent(v) { this.el.textContent = v; this.loop.markDirty(`textContent on <${this.el.tagName}>`) }
  get offsetHeight() {
    // a layout read: forces a synchronous layout if the loop is dirty. the host reports it on the timeline.
    return this.loop.dirty ? forcedLayoutHook(this.loop, this.el) : lastKnownHeight(this.el)
  }
}
the sandbox
  1. Globals by parameter. new Function with named parameters means the script's document is our object, not the host page's. It is not a security boundary (the script can still reach window by other means); it is a teaching boundary.
  2. Every mutation marks dirty with a reason. The Event loop pane shows the reasons when the render happens. In a real engine the reasons are invalidation sets and dirty bits on the layout tree; the idea is identical.
  3. TinyPromise exists because the host's Promise would schedule reactions on the host's microtask queue, invisible to our loop. Ours pushes reactions onto our queue so that Promise.resolve().then() and queueMicrotask() interleave correctly in the timeline.
  4. offsetHeight forces layout. If the loop is dirty when a script reads a layout property, the hook runs style and layout synchronously (no paint) and the timeline shows "FORCED synchronous layout". Write, read, write in one task: two layouts where one would do. Part 4's layout thrash, reproduced in the "Event loop" sample's timer.
in the app
"Event loop" sample. Press Step loop once: the script task, then its two microtasks in order. Again: the timeout-0 task, because the frame is not due yet. Again: rAF, then the single render for three textContent writes and one className change. Keep stepping to the 40 ms timer and watch the forced layout appear before the render.
THE TINY EVENT LOOP
eventloop.ts step(), and the sandbox that feeds it
swipe the figure sideways, or tap expand for full screen
1/6
script task
The page script is one task. It runs through new Function with our document, setTimeout, queueMicrotask, requestAnimationFrame, Promise and console passed as the only globals. Nothing it does touches the real page.
142

pipeline.ts and App.tsx: the glue, and the shell

pipeline.ts
  1. parseStage(html): tokenize, build the tree, collect the author CSS from every <style> element's text, and time each. Runs once per Parse + render.
  2. renderStages(document, css, measure, viewportWidth): cascade, box tree, layout, display list, timed. Runs on every render the loop decides to do, and on every forced layout, against the same Document object that scripts mutated.
  3. Timings are performance.now() deltas and appear on the stage strip. They are real: a longer sample costs more in the cascade (selector tests scale with elements times rules) and in layout (measureText per word).
App.tsx
  1. One state object, the PipelineOutput, replaced whole on each render so React re-renders the panes. The loop keeps its own reference (current) and hands React a copy after each render it performs.
  2. The measurer is a hidden canvas context; measure(text, font) sets the font and reads measureText().width. This is the only place the engine depends on the host browser's text stack.
  3. Selection crosses panes. The selected Element is state; the DOM pane, the Boxes pane, the Styles pane and the canvas overlay all read it. Clicking the canvas hit-tests the box tree.
  4. Paint flashing is a Set of boxes painted by the last render, drawn tinted for 500 ms and then cleared by repainting without the overlay.
  5. Samples are plain strings in samples.ts. Add your own; the select is generated from the list.
worked numbers
what the stage strip shows for the "A page" sample on a laptop:

  tokenizer   ~0.5 ms   (1.3 KB of HTML, 60 tokens, 300 transitions)
  tree builder ~0.4 ms
  cascade    ~1.2 ms   (16 elements × 61 selectors = 976 tests, 31 matches)
  layout     ~1.0 ms   (27 steps; measureText is most of it)
  paint      ~0.3 ms   (89 commands)

the proportions are the same as a real page's: style and layout dominate, parsing is cheap, painting a display list is cheap. raster, which we do on the host canvas, is the part a real engine moves off the main thread.
143

Exercises: the next thing to add, in order of payoff

parsing
  1. Stream the tokenizer. Make tokenize a generator that accepts chunks and yields tokens as they complete, with the state preserved between chunks. Then feed it the sample 100 bytes at a time with a delay and render after each chunk: progressive rendering, from part 2.
  2. The adoption agency algorithm. Implement the formatting-element list and the reconstruction step so <b><i></b>x</i> produces two <i> elements. The spec section is long but mechanical.
  3. Script blocking. Pause tree construction when a <script> end tag arrives, run the script as a task against the partial DOM, then resume. Then add defer. Then add a preload scanner that lists the resources it would have fetched.
style
  1. Pseudo-classes. :first-child, :hover (with a hovered element in state), :not(). Each one is a predicate in matchesCompound.
  2. An ancestor bloom filter. Before the descendant walk, check whether any ancestor could have the target's id, class or tag using a filter built on the way down. Count the tests saved.
  3. Incremental style. When a script changes a class, recompute style only for that element and its descendants, and only re-layout its containing block. Mark the others as reused in the layout log.
  4. Cascade layers (@layer) as a fourth component of the sort key, between origin and specificity.
layout
  1. Flexbox, single line. Measure each item's max-content width (lay it out with infinite width), sum, distribute free space by flex-grow, lay each item out at its final width, align cross-axis. This is the heart of the algorithm from part 4 and fits in 80 lines.
  2. Percentages and min/max-width. Resolve percentage widths against the containing block in layoutBlock; clamp.
  3. Positioning. position: relative as an offset after layout; absolute as a separate pass against the nearest positioned ancestor; a z-index-aware paint order.
  4. Images. An <img> that fetches its src, decodes with createImageBitmap, reserves its box from width and height attributes, and marks dirty on load. Then watch the layout shift when the attributes are missing.
paint and loop
  1. Damage rectangles. Track which boxes changed geometry or style and flash only those; then rasterise only the union of their rects.
  2. A compositor layer. Give any element with transform its own canvas, paint it once, and move it with a transform on the canvas element without re-running paint. Measure the difference on a 1,000-node page.
  3. Input. Pointer events from the canvas, hit-tested to our boxes, dispatched to listeners registered through the sandbox with bubbling. Then make :hover work.
  4. Long task breakup. A scheduler.yield() in the sandbox that splits a script across tasks, with a render in between. Show it with a script that appends 5,000 rows.
the point of the exercise
Every one of these is a chapter from parts 1 to 12 made concrete. The engine is small enough that each addition is an afternoon, and after a few of them the real engine's source, when you open it, is no longer foreign: the function names differ, the shape does not.