Part 11 · 8 chapters · ~65 min

Input, Text, Images, Media

The parts of the page a user actually touches, reads, and watches each have a pipeline the browser runs for them. This part is the event sequence behind a tap and the three controls over it, focus and the accessibility tree, what native forms do for free, the text pipeline from codepoints to shaped glyphs and why marks float, image selection, priority, decoding and the LCP checklist, the three drawing surfaces, and the media pipeline from range request to hardware overlay.

110

The input pipeline, revisited at the event level

the question

"A tap on the phone fires how many events, and why does my click handler run late?"

Part 6 traced an input from the OS to the compositor to the main thread. This chapter is what the main thread does with it: the sequence of events the browser synthesises from one physical action, the delay it inserts while deciding what the action meant, and the three mechanisms (pointer events, touch-action, passive listeners) that let the page take control of that.

the event families
  1. Pointer events (pointerdown, pointermove, pointerup, pointercancel, pointerenter/leave/over/out): one model for mouse, touch and pen. pointerType, pointerId for multi-touch, pressure, tiltX. The one to write new code against.
  2. Touch events (touchstart, touchmove, touchend, touchcancel): the older model, with touches lists. Still fired; still needed for a few multi-touch gesture details; the one that creates scroll-blocking problems.
  3. Mouse events: on a touch device, synthesised after the touch sequence for compatibility. mouseover, mouseenter, mousemove, mousedown, mouseup, then click.
  4. Click: device-independent, fired by keyboard (Enter and Space on buttons), by assistive technology, by touch, by mouse. The one to use for activation.
  5. Keyboard: keydown (repeats), keypress (deprecated), keyup. key for the meaning ("a", "Enter", "ArrowLeft"), code for the physical key ("KeyA"), isComposing during IME composition (East Asian input, where keydown does not mean a character).
the delays and how to remove them
  1. The 300 ms tap delay: the browser waits to distinguish a tap from a double-tap zoom. Removed by <meta name="viewport" content="width=device-width"> or touch-action: manipulation on the element. Every modern page has the former; the delay persists on pages that do not.
  2. Scroll-blocking listeners: a touchstart or touchmove or wheel listener that might preventDefault forces the compositor to wait for the main thread before scrolling. { passive: true } promises not to; Chrome defaults document-level touch and wheel listeners to passive and warns.
  3. touch-action in CSS declares which gestures the browser handles on an element: pan-y (you handle horizontal; browser scrolls vertical), none (you handle everything; a map or a slider), manipulation (pan and pinch; no double-tap zoom). The compositor reads it without asking the main thread.
  4. Pointer capture: el.setPointerCapture(pointerId) routes all subsequent pointer events for that pointer to the element even after it leaves. Drag handles without document-level listeners.
ONE TAP, EVERY EVENT
the sequence the browser synthesises from a finger
swipe the figure sideways, or tap expand for full screen
1/6
down
Finger down. pointerdown fires first (pointerType: "touch"), then touchstart. If the touchstart listener is not passive, the browser must wait for it to return before it can start scrolling, because the listener might call preventDefault.
111

Focus, keyboard, and the accessibility tree

The keyboard is the input that proves whether a UI is operable. Focus is the browser's record of where keyboard input goes, and the accessibility tree is what assistive technology reads instead of the DOM. Both are built by the browser from the same DOM; both are broken by the same shortcuts.

focus
  1. Focusable by default: links with href, buttons, inputs, selects, textareas, iframes, elements with tabindex, summary, audio and video with controls. A div with a click handler is not, which is why it is unreachable by keyboard.
  2. tabindex="0" adds to the tab order; -1 makes focusable by script only (for moving focus to a heading after route change); positive values reorder and are a maintenance trap.
  3. Sequential navigation follows DOM order, not visual order. A CSS order or grid placement that changes visual order leaves focus jumping around the screen.
  4. Focus visibility: :focus-visible matches when the browser would show a ring (keyboard, not mouse). Never outline: none without a replacement on :focus-visible.
  5. Focus and removal: removing the focused element sends focus to body. Dialogs must trap and restore; <dialog> with showModal() and the inert attribute on the rest of the page do this natively now.
  6. Focus events: focusin/focusout bubble; focus/blur do not. relatedTarget says where focus came from or is going.
the accessibility tree
  1. Built from the DOM plus ARIA, per element: role (button, heading, textbox), name (from content, label, aria-label, aria-labelledby, alt, title, in a defined order), state (checked, expanded, disabled), and properties. Hidden elements (display: none, aria-hidden) are pruned.
  2. Exposed through platform APIs (UIA, IAccessible2, AX on macOS, AT-SPI on Linux) to screen readers, switch access, voice control. The screen reader never sees the DOM.
  3. Native elements come with roles, names, states and keyboard behaviour. A <button> has role button, its text as name, Enter and Space activation, focusability. A <div role="button"> needs tabindex, a keydown handler, and the name, and will still miss something.
  4. Live regions (aria-live="polite", role="status", role="alert") are how a change away from focus is announced. Route changes, toasts, validation summaries.
  5. In DevTools: Elements → Accessibility pane shows the computed role, name and the name-computation chain. The full-page accessibility tree view replaces the DOM tree with what AT sees. The "Lighthouse accessibility" audit catches the mechanical issues; the keyboard catches the rest.
112

Forms: the browser's own state machine

what a native form does without JavaScript
  1. Submission: serialises named controls (urlencoded, multipart for files), navigates or fetches to the action with the method. formdata event lets script modify the payload; new FormData(form) reads it.
  2. Constraint validation: required, pattern, min, max, type=email, minlength. checkValidity(), reportValidity(), the :invalid and :user-invalid pseudo-classes (the latter only after interaction, which is what you want), setCustomValidity() for server-side errors.
  3. Autofill: keyed by autocomplete tokens (given-name, cc-number, one-time-code, new-password). Correct tokens are the difference between a one-tap checkout and a typed one. Password managers rely on them.
  4. Implicit submission: Enter in a text field submits if there is a submit button or only one field. Custom form components that forget this break a reflex.
  5. Restoration: the browser restores form values on back and reload (autocomplete=off disables for a field). SPAs that re-render lose this.
the input event model
  1. beforeinput fires before the DOM changes with inputType (insertText, deleteContentBackward, insertFromPaste, formatBold) and data. Cancelable. The hook for rich text editors and input masks.
  2. input fires after every change, including IME composition updates and autofill. change fires on commit (blur, Enter, selection).
  3. Composition: compositionstart, compositionupdate, compositionend for IME input. Between start and end, keydown values are not characters; filtering on keydown breaks Chinese, Japanese and Korean entry.
  4. Selection: selectionchange on the document; input.selectionStart; window.getSelection() for contenteditable.
the design rule
Use native inputs and extend them. A custom select built from divs loses: keyboard type-ahead, mobile pickers, autofill, form participation, validation, and screen reader semantics. The Customizable Select (appearance: base-select) and form-associated custom elements (ElementInternals) exist so you can style without rebuilding.
113

Text: from codepoints to glyphs

Text is most of most pages, and its rendering is a pipeline of its own inside layout. Knowing the stages explains the bugs: wrong font in one word, accents floating, Arabic letters disconnected, a line breaking inside a number.

the stages
  1. Codepoints. The DOM holds UTF-16; the pipeline works on codepoints; the user thinks in grapheme clusters. "é".length may be 1 or 2 depending on normalisation (NFC vs NFD). Intl.Segmenter gives graphemes, words and sentences correctly.
  2. Bidi. The Unicode Bidirectional Algorithm resolves direction per run. dir attribute, dir="auto", unicode-bidi: isolate (default with dir), plaintext. Mixed-direction UI (an Arabic name in an English table) needs isolation or the punctuation jumps.
  3. Itemisation. Runs of one script, direction, and font.
  4. Font selection and fallback. For each character, the first family in font-family that has it; else the system fallback. unicode-range scopes a web font to the characters it should claim; font-display decides what shows while it loads (part 3).
  5. Shaping. HarfBuzz maps codepoints to positioned glyphs using the font's tables: ligatures, contextual forms, mark positioning, kerning. Yoruba tone marks on dotted vowels (ọ́, ẹ̀) depend on the font having mark-to-base positioning; without it the marks float. Blink caches shaped words.
  6. Line breaking. UAX #14 break classes, dictionary segmentation for Thai, Lao, Khmer, Japanese, hyphenation with hyphens: auto plus a lang. text-wrap: balance for headings, pretty for paragraphs; overflow-wrap: anywhere for long URLs; word-break: keep-all for CJK.
  7. Paint. Glyph outlines rasterised at the size with subpixel antialiasing (grayscale on transformed or composited layers, which is why text on a will-change layer looks lighter), cached per glyph and size.
metrics and performance
  1. Font metrics (ascent, descent, line-gap) decide line height and baseline; two fonts at the same size have different line boxes, which is the layout shift when a web font replaces its fallback. size-adjust, ascent-override, descent-override on a fallback @font-face match them.
  2. Variable fonts carry axes (weight, width, optical size) in one file: one download instead of six, and animatable weight.
  3. Subsetting by unicode-range and by glyph set cuts a 300 KB font to 30 KB. WOFF2 with Brotli.
  4. Cost: shaping is per run; a 50,000-word document with one font and size shapes once and caches; changing font-size on a container re-shapes everything in it.
FROM CODEPOINTS TO GLYPHS
the text pipeline inside layout
swipe the figure sideways, or tap expand for full screen
1/6
codepoints
Input: a string of Unicode codepoints. "Ẹ̀kọ́ 123 العربية". Not characters: a visible letter may be several codepoints (base plus combining marks), and several codepoints may become one glyph.
114

Images: selection, loading, decoding, and the LCP image

Images are most of a page's bytes and usually its largest contentful paint. The browser gives you control over which file is chosen, when it is fetched, how urgently, and when it is decoded; using those controls is most of image performance.

choosing the file
  1. srcset with width descriptors + sizes: the browser picks the smallest source that covers the slot at the device pixel ratio. sizes must describe the layout width, in CSS terms, before layout exists; sizes="auto" with lazy images lets the browser use the actual width.
  2. <picture> with <source type>: format negotiation (AVIF, then WebP, then JPEG) and art direction (a different crop per breakpoint via media).
  3. Formats: AVIF smallest, slowest to decode; WebP broadly supported and fast; JPEG universal; PNG for sharp edges and alpha where WebP is not possible; SVG for vectors. Image CDNs negotiate via Accept and resize per URL parameter, which replaces most of the above.
loading and priority
  1. The preload scanner starts image fetches from HTML before layout. CSS background images and JS-inserted images start only when their rule or script runs, which is why the LCP image belongs in HTML.
  2. Priority: images are low until layout places them in the viewport, then medium. fetchpriority="high" on the LCP image moves it ahead of everything else from the start. One or two per page, not all.
  3. loading="lazy" for below-the-fold. Never for the first viewport. It also disables the preload scanner for that image.
  4. Preload (<link rel="preload" as="image" imagesrcset imagesizes>) for an LCP image the scanner cannot see (CSS background, injected by a framework).
decoding and memory
  1. Decode is CPU: tens of milliseconds for large photos. decoding="async" (default now) keeps it off the main thread; the frame appears one frame later. img.decode() to await before showing.
  2. Decoded size is layout size in Chrome; the download is the full file. Serving a 4000 px image for a 400 px slot wastes bandwidth, not decode.
  3. Decoded bitmaps are cached and evicted under memory pressure; a long gallery re-decodes on scroll back. content-visibility: auto on offscreen sections lets the browser skip their rendering work entirely.
  4. Reserve the box: width and height attributes, or aspect-ratio in CSS, on every image. No layout shift when bytes arrive.
the LCP image checklist
In the HTML, not CSS. fetchpriority="high". Not lazy. Width and height set. Modern format via picture or CDN negotiation. Sized to the slot via srcset and sizes. Served from the same origin or a preconnected one. That list is most of an LCP fix.
AN IMAGE'S LIFE
from src to a texture on the GPU
swipe the figure sideways, or tap expand for full screen
1/6
select + fetch
The preload scanner sees before the parser gets there and picks a candidate: it evaluates sizes against the viewport, picks the smallest source with a density at or above the device pixel ratio, and starts the fetch. Priority: low until layout proves it is in the viewport; fetchpriority="high" overrides for the LCP image.
115

Canvas, SVG and WebGL/WebGPU: the three drawing surfaces

SVGCanvas 2DWebGL / WebGPU
ModelRetained: a DOM of shapesImmediate: draw commands to a bitmapImmediate: GPU programs over buffers
ScalesResolution-independentFixed pixels; set width×DPR for sharpnessFixed pixels; same
InteractionDOM events per shape, CSS, accessibilityYou hit-test; no accessibility without extra workYou hit-test
Cost scales withNumber of nodes (thousands is slow)Pixels drawn per frameVertices, fragments, state changes
ThreadsMainMain, or OffscreenCanvas in a workerMain, or OffscreenCanvas in a worker
Best forIcons, charts with hundreds of elements, anything that needs CSS or events per shapeCharts with thousands of points, image editing, games with simple graphics3D, large datasets, shaders, anything GPU-shaped
what the browser does with each
  1. SVG is parsed into the DOM, styled by the cascade, laid out and painted like HTML. Inline SVG shares the document; <img src="x.svg"> is isolated (no scripts, no external resources, no CSS from the page). Filters and masks are expensive; animating a filter re-rasters.
  2. Canvas 2D records draw calls; Chrome rasterises them on the GPU where available (accelerated canvas) and composites the result as a texture. getImageData forces a GPU-to-CPU readback, which stalls; avoid it in a frame loop. willReadFrequently: true at context creation keeps the canvas on the CPU if you must read back.
  3. WebGL is OpenGL ES through ANGLE (translated to Direct3D, Metal, Vulkan). WebGPU is a modern API with compute shaders, explicit pipelines, and far less driver overhead; it is the forward path and is shipped in all major engines now.
  4. OffscreenCanvas moves any of them to a worker: canvas.transferControlToOffscreen(), then draw from the worker; the main thread stays free. The way to run a heavy visualisation without janking the UI.
  5. DPR: a canvas's bitmap is width × height pixels; its CSS size is separate. Set the bitmap to CSS size × devicePixelRatio and scale the context, or it is blurry on every phone.
116

Audio and video: the media pipeline

the element pipeline
  1. Fetch with range requests; find the index (moov); buffer ahead according to preload. The server must support ranges or seeking re-downloads.
  2. Demux the container into streams; decode each, in hardware where the codec and resolution are supported; render frames to a compositor surface or a hardware overlay, synced to the audio clock.
  3. navigator.mediaCapabilities.decodingInfo() says whether a codec and resolution will be smooth and powerEfficient on this device. Pick formats with that, not with canPlayType alone.
  4. Hardware overlays are the low-power path; anything that forces the video through the compositor (opacity, filters, transforms, overlapping translucent elements) costs battery.
Media Source Extensions and adaptive streaming
  1. MSE: JavaScript creates a MediaSource, attaches it as the element's src via an object URL, adds SourceBuffers per track, and appends fetched segments. The browser decodes what is appended.
  2. ABR: the player (hls.js, dash.js, Shaka) measures throughput per segment, picks the next quality, and switches at segment boundaries. Buffer-based and throughput-based algorithms; the trade is rebuffering against quality.
  3. Safari on iPhone has no MSE; it plays HLS natively from a .m3u8 URL. Ship both paths.
  4. Low latency: LL-HLS and LL-DASH with partial segments; WebRTC for sub-second.
the rest
  1. Autoplay: muted autoplay is allowed; sound needs user activation or a high engagement score. play() rejects when blocked; catch it and show a play button.
  2. Encrypted Media Extensions: Widevine (Chrome, Firefox), FairPlay (Safari), PlayReady (Edge). License exchange via MediaKeySession; decoded frames stay in a protected path.
  3. Web Audio: a graph of nodes (source, gain, filter, analyser) processed on the audio thread; AudioWorklet for custom DSP. Latency in the single-digit milliseconds.
  4. WebCodecs: direct access to encoders and decoders with VideoFrame and AudioData objects; for editors, screen recorders, and custom transports. Media Session API: lock screen and hardware key integration. Picture-in-Picture and the Document PiP variant.
  5. Capture: getUserMedia (camera, mic), getDisplayMedia (screen), MediaRecorder (to WebM/MP4), and WebRTC for real-time transport (its own course).
VIDEO PLAYBACK PIPELINE
demux, decode, render, and what MSE changes
swipe the figure sideways, or tap expand for full screen
1/6
fetch + index
src="movie.mp4". The browser fetches the first bytes to find the moov atom (the index); if it is at the end of the file, it fetches the end too, then seeks with HTTP range requests. preload="metadata" stops after the index; preload="auto" buffers ahead.
117

Reading input, text and media in DevTools

WhereShowsUse it for
Performance → Interactions trackEach interaction with input delay, processing, presentation delayINP; which handler was slow; whether the delay was before or after the handler
Console → monitorEvents(el, 'pointer')Every event of a family as it firesSeeing the real sequence for a tap or drag
Rendering → Scrolling performance issuesOverlays on elements with non-passive listeners or touch-action issuesFinding scroll-blocking handlers
Elements → Accessibility pane / full a11y treeRole, name, name computation, states; the tree AT readsWhy a control has no name; what a div-button is missing
Elements → Computed → Rendered fontsWhich font file actually rendered each glyph runFallback happened where; the mixed-font word
Network → Img filter + Priority columnEach image's chosen source, size, priority, and whether lazyWas the LCP image high priority; was a 4000 px file served to a 400 px slot
Performance → LCP marker → elementThe LCP element and its load timing breakdown (TTFB, load delay, load time, render delay)Which of the four LCP phases is the problem
Media panelEach media element: codec, decoder (hardware or software), buffered ranges, dropped frames, eventsWhy video stutters; whether HW decode is in use; MSE append timing
Rendering → Layer borders / Layers panelCanvas and video layers, their sizesGPU memory from large canvases; whether video got an overlay
go to the lab
  1. Route /input/tap-sequence: log every event from a tap and a drag on a touch device (or DevTools device mode). Then add touch-action and passive, and compare the sequence and the Interactions track.
  2. Route /input/div-button: a div with a click handler next to a real button. Tab through; use a screen reader; read the Accessibility pane for each.
  3. Route /text/fallback: Yoruba and Arabic text in a font without the glyphs. Rendered fonts pane shows the fallback; add unicode-range and a proper font; watch the marks sit correctly.
  4. Route /images/lcp: the same hero image as a CSS background, as a lazy img, and as an img with fetchpriority=high. Record each; compare LCP phases.
  5. Route /media/decode: the same clip in H.264, VP9, and AV1. Media panel shows which decoded in hardware; mediaCapabilities says why.