Part 1 · 2 chapters · ~20 min
Accessibility in Depth: Focus, Keyboard, Screen Readers
The tab sequence and roving focus inside composites, the keyboard model per widget, deliberate focus management on every change, intended and unintended traps, and the keyboard pass; then how screen readers read a page: the readers, reading versus focus mode, what is announced, landmarks and headings as the map, names that stand alone, and announcing dynamic content.
3
Focus and keyboard models
everything pointer-reachable is keyboard-reachable, in order, visibly
- The tab sequence: DOM order over focusable elements plus
tabindex="0";-1for script-only focus; positive values forbidden; the visual order must match the DOM order (CSS reordering that disagrees fails 1.3.2 and 2.4.3 at once); a skip link as the first stop (2.4.1). - Roving tabindex in composites (tabs, menus, radio groups, toolbars, listboxes, grids): one child at 0, the rest at -1, arrows move both, Home and End jump, typeahead selects; or
aria-activedescendantkeeping DOM focus on the container (the FSD course M6's virtualised list, where the active row may not exist). - The models per widget from the WAI-ARIA Authoring Practices: tabs, menus, listbox, combobox, dialog, grid, tree, disclosure, each with its keys. Implementing from the pattern beats inventing; borrowing a primitive (the Design course part 2) beats implementing.
- Focus management on change: a route change focuses the new heading or main; a dialog takes focus in and returns it to the trigger; a deleted item moves focus to the next item or the list; a submit error moves focus to the first invalid field; a toast never takes focus (a live region announces it). Never to body; never nowhere.
- Traps and inert: a modal is the only intended trap (Tab cycles, Escape leaves, the rest inert, which the dialog element and the inert attribute provide); a widget that catches Tab, an embed, or an offscreen element that is still focusable (focus vanishes) are bugs. Hidden means display: none or inert, not offscreen.
- Testing: unplug the mouse and do the task: reach every control, operate every widget by its model, see focus at all times (2.4.7; not obscured by sticky headers, 2.4.11), leave every trap, return after every dialog and route. Scripts per widget; a human pass per release.
FOCUS AND KEYBOARD MODELS
the tab sequence, roving focus inside composites, focus management on change, and the traps
swipe the figure sideways, or tap expand for full screen
1/6
the tab sequence
The tab sequence: DOM order over natively focusable elements plus tabindex="0"; tabindex="-1" makes an element focusable by script but not by Tab; positive tabindex values are forbidden in practice (they reorder unpredictably). The visual order must match the DOM order (CSS reordering with flex order or grid placement that disagrees with the DOM is criterion 1.3.2 and 2.4.3 failing at once). A skip link as the first focusable element jumps past repeated navigation (2.4.1).
4
How a screen reader reads a page
the tree, spoken, with ways to move that are not Tab
- The readers: VoiceOver (macOS and iOS, built in), NVDA (Windows, free, the most used), JAWS (Windows, enterprise), TalkBack (Android), Narrator. Test on VoiceOver and NVDA at least, with the browsers their users pair them with.
- Reading versus focus mode: in reading mode the arrows move through every node including static text and single keys are shortcuts (H heading, B button, F field, L list, T table, K link); focus mode sends keys to the widget and switches on automatically for inputs.
role="application"forces focus mode and removes the shortcuts: almost always wrong. - What is announced: role, name, state and context ("Send money, button"; "Amount, edit text, invalid, Must be a number"; "Transfers, tab, 2 of 4, selected"; "list, 3 items"). The information set is the tree's. A div with a click handler is announced as text alone: the user does not know it is interactive.
- Landmarks and headings as the map: banner, navigation (named when there are several), one main, complementary, contentinfo, named forms and regions; one h1 naming the page, h2 sections in order, no skipped levels, no styled divs as headings. The user jumps by landmark then heading; without them the page is a wall.
- Names that work read aloud and alone: a link says where it goes (never "click here"); a button says what it does; an icon button has the verb; alt describes what matters or is empty; a table has a caption; a field's label is the question. The rotor lists names out of context.
- Dynamic content: a change the user did not cause is announced only through a live region (part 2) or a focus move; loading through aria-busy and a live "Loaded 20 transactions"; a submit error through focus to the field; a route through the focused heading. Silence is the common failure: the sighted user saw it change and the screen-reader user heard nothing.
the exercise
Turn on VoiceOver (⌘F5), close your eyes, and complete your product's main task by keyboard and ear. Every moment you do not know where you are, what a thing is, or what just happened is a finding; write them in order.
HOW A SCREEN READER READS A PAGE
VoiceOver, NVDA, JAWS and TalkBack: the reading modes, the rotor, landmarks, headings, and what the user hears
swipe the figure sideways, or tap expand for full screen
1/6
the readers
The readers and where they are: VoiceOver (macOS and iOS, built in: ⌘F5; the VO keys; the rotor by two-finger twist on iOS), NVDA (Windows, free, the most used with Firefox and Chrome), JAWS (Windows, commercial, widespread in enterprises), TalkBack (Android), Narrator (Windows, built in). Each has its own verbosity and its own quirks; a product is tested on VoiceOver and NVDA at least, with the browser pairings their users have.