Part 9 · 2 chapters · ~20 min

Lighthouse and Recorder

What a Lighthouse score is made of (a lab run, five weighted metrics, variance, audits, the other categories, and field data that outranks it), and the Recorder that turns a user flow into a replay, a measurement and an exported test.

24

Lighthouse: what the score is made of

the run and the score
  1. A lab run: a fresh profile, a simulated device (a mid-range phone on slow 4G by default, or desktop), simulated throttling by default (measured unthrottled, modelled slow) or DevTools throttling (real, slower, closer), one load. Modes: Navigation (a load), Timespan (a flow you perform), Snapshot (the page as is, for accessibility and best-practice audits).
  2. The Performance score is a weighted curve over five metrics: LCP 25%, TBT 30% (Total Blocking Time, the sum of blocking over 50 ms per task: the lab stand-in for INP), CLS 25%, FCP 10%, Speed Index 10%. INP needs real interactions and appears only in the field section.
  3. Variance: the same page scores 72, 85, 68 across runs. One score is noise; the median of five is a measurement; under ten points between single runs is not evidence. The CLI or PageSpeed Insights (Google's machines, plus field data) for consistency.
the audits, and the other categories
  1. Opportunities (with estimated savings: render-blocking resources, unsized images, unused JavaScript, text compression, preconnect) and Diagnostics (long tasks, DOM size, third-party cost, the layout-shift culprit element). Each names a mechanism; the savings are a model; measure after applying.
  2. Accessibility (axe-core rules: contrast, names, roles, order: necessary, not sufficient; a screen reader still has to be used), Best Practices (HTTPS, console errors, deprecated APIs, aspect ratios), SEO (meta, crawlability, link text).
  3. Field versus lab: the report's top shows CrUX field data (28-day p75 LCP, INP, CLS from real Chrome users) when it exists. That is what users have. When field and lab disagree, the field is right and the lab is a hypothesis about why (the Browser course part 12).
go to the lab
  1. Run Lighthouse (mobile, Navigation) on /images/lcp three times for each variant; write the medians; compare the LCP audit's phases with what the Performance panel showed in part 5.
  2. Run a Snapshot on /input/div-button: the accessibility audit flags the unrepaired div (no role, no name). The repaired one passes: and still try it with a screen reader.
  3. Run a Timespan on /perf/inp-phases while clicking add-to-cart with all toggles on, then with all off: read TBT for each.
WHAT A LIGHTHOUSE SCORE IS MADE OF
a lab run on a simulated device, five metrics weighted into one number, and the audits beneath
swipe the figure sideways, or tap expand for full screen
1/6
the run
The run: Lighthouse navigates to the URL in a fresh profile with throttling (simulated by default: it measures unthrottled and models slow 4G; "DevTools throttling" applies real throttling and is slower but closer), collects a trace and the DOM, and runs the audits. Mode: Navigation (a load), Timespan (a user flow you perform), Snapshot (the page as it is now, for accessibility and best-practice audits without a load).
25

The Recorder: flows you can replay, measure and export

record, replay, measure, export
  1. Record: name the flow, perform the actions; each step captures its type, several selectors (CSS, ARIA, text, XPath, pierce) and its timing. Edit brittle selectors (generated class names) to ARIA or text, which survive a redesign.
  2. Replay at normal, slow or extra-slow, with network and CPU throttling, with a breakpoint on any step to inspect the page there. A failing step names its reason (selector not found, timeout, behind a modal), which is often the bug.
  3. Measure performance: replays with the Performance panel recording and opens the trace; the same flow after a change is the before-and-after without hand-clicking twice. Lighthouse Timespan on the flow gives TBT, CLS and interaction audits for what a user does rather than for a load.
  4. Export as a Puppeteer or Playwright script, replay JSON, or a Lighthouse flow; assert steps become test assertions. One artefact is a smoke test and a performance regression test in CI.
where it fits
Part 0's method had reproduce first and measure-again last. The Recorder makes the first a file and the last a replay. "Checkout is slow after three items" becomes a flow on every PR.
go to the lab
  1. Record a flow on /perf/inp-phases: navigate, click add-to-cart three times, assert "cart: 3". Replay it with 4× CPU; set a breakpoint on the second click and read the Interactions track so far.
  2. "Measure performance" on that flow with all toggles on; then turn the toggles off and measure again; compare the three interaction durations in each trace.
  3. Export it as a Playwright script and read the selectors it chose; change the recorded .btn.primary to the ARIA selector and export again.
  4. Record /nav/spa-router: three navigations with no fixes, replay with the fixes on, and assert the h2 text on each page. That is an accessibility regression test in four steps.
THE RECORDER
a user flow captured, replayed, measured and exported
swipe the figure sideways, or tap expand for full screen
1/6
record
Record: name the flow, click record, perform the actions in the page (DevTools captures clicks, typing, navigations, scrolls, and assertions you add), stop. Each step shows its type, its selectors (several: CSS, ARIA, text, XPath, pierce for shadow DOM) and its timing. Edit a step's selector when the recorded one is brittle (a generated class name).