Part 8 · 2 chapters · ~20 min

On-Call for Frontend

What pages a frontend engineer (new error signatures, flow success by version, vitals by version and tier, synthetics on the tail browsers, the blank-page beacon) and what must not, then the on-call itself: a rotation from OWNERS, two-minute triage, levers with time-to-fleet, automated communication, blameless reviews that produce mechanisms, and the on-call's own SLO.

17

What pages a frontend engineer

alerts built on client telemetry, by version and surface
  1. A new error signature: a fingerprint first seen on this version crossing N occurrences across M users in T minutes pages the surface's owner; thresholds per surface; seen-before signatures are tracked, not paged (the Architecture course part 3's error telemetry).
  2. Flow success rate by version: start and success events for the key flows; the ratio per version per ring against the previous version; a relative drop beyond tolerance pages. This is the frontend's availability SLO ("99.5% of started checkouts reach confirmation"), measured in the field, by version.
  3. Web Vitals by version and tier: p75 LCP and INP per surface per tier per version from RUM; a regression during the rings halts the train (part 6); a sustained one outside them pages the owner. Without the version dimension a slow network day is a regression.
  4. Synthetic monitors: a scripted flow (the DevTools course part 9's Recorder export) from several regions on several browsers including the tail WebViews (the Architecture course part 7), every few minutes. They catch what produces no error: a blank page, a regional CDN fault, an expired certificate.
  5. The blank-page beacon: the inline capture in the HTML, before the bundle, sends a beacon when the entry fails to parse or run; a count by version and user agent pages. The one alert that works when the application cannot.
  6. What must not page: a single user's error, a lab regression, a metric without a version, a degraded third party the app handles, anything not an agreed SLO. An unactionable page trains the on-call to ignore pages.
WHAT PAGES A FRONTEND ENGINEER
the alerts that wake someone for a client, and the ones that should not
swipe the figure sideways, or tap expand for full screen
1/6
new signature
New error signature: the error reporter groups by fingerprint (stack plus message); a fingerprint first seen on this version, crossing N occurrences across M users within T minutes, pages the surface's owner. The threshold is per surface (checkout lower than settings); a signature seen before on previous versions is a known issue, tracked but not paged.
18

The on-call: rotation, runbooks, levers, reviews

who, what first, what to pull, what after
  1. The rotation: a week, primary and secondary, from the surface's OWNERS so the person paged knows the code; release engineering as the escalation for fleet levers; follow-the-sun handovers with a written note (the Platformization course); on-call time budgeted and compensated.
  2. Triage in two minutes: which version (did it start with a deploy?), which surface, which population (market, tier, browser, locale), users per minute, still happening. Five questions, five dashboard panels; the answers choose the lever.
  3. The levers and their time-to-fleet: a kill switch (seconds, the flag poll); halt the train and serve the previous artefact (minutes for new loads; open tabs keep the new until they reload); a forced-reload beacon (hours, or a deadline); a CDN purge (minutes to an hour per edge); a hotfix through the lane (hours). The runbook names the lever per failure class and what it will not fix, written before the incident by the people who will be paged.
  4. Communication: the alert updates the status page component itself (a human updating a status page is a human not fixing it); an incident channel every fifteen minutes with version, blast radius, lever and time; a timeline kept as it happens; support reads the channel and tells users what engineers know.
  5. The review, within a week, blameless, with the timeline: detection source (a user report means an alert is missing), time-to-detect and time-to-mitigate against the SLO, what the lever did and did not fix, which guardrail, budget, synthetic or test would have caught it before promotion, and actions with owners and dates tracked to completion (part 7). The output is a mechanism, never "be more careful".
  6. The on-call's own SLO: time-to-detect p50 under five minutes, time-to-mitigate p50 under thirty, pages per week per person under a threshold, actionable pages above 80%. Frequent unactionable pages are a signal about the alerts, not the engineers. A boring on-call is the goal, like a boring train.
the transferable part
Put the version on every client metric and error; write the runbook before the incident with each lever's reach; automate the status page from the alert; review blamelessly for mechanisms. A team of thirty with one surface can do all four; the rotation and the escalation tiers are what scale adds.
THE ON-CALL: ROTATION, RUNBOOKS, LEVERS, REVIEWS
who is paged, what they do first, what they can pull, and what happens after
swipe the figure sideways, or tap expand for full screen
1/6
the rotation
The rotation: a week at a time, two people (primary and secondary), drawn from the surface's OWNERS (part 7) so the person paged knows the code; release engineering has its own rotation for fleet levers and is the escalation; a follow-the-sun handoff for global teams (the Platformization course) with a written handover. On-call time is budgeted (a week on-call is a week of less feature work) and compensated.