Part 7 · 4 chapters · ~30 min
What Breaks at 100M
Incident logs at a hundred million users: the one percent that was a million people and a blank page with no error, WebViews and hybrid shells as a platform with their own lifecycle and a versioned bridge, and shipping as release trains with rings, experiments with guardrails and correctness checks, and a client fleet you cannot recall. Then the course in one table.
24
The one percent that was a million people
the incident
- The distribution at 100M: Chrome 71%, Safari 14%, Samsung 5%, Firefox 3%, and a 7% tail across dozens of WebViews and in-app browsers, each under 1% and each a population of hundreds of thousands to a million. The field user-agent data had it; nobody had read it as a list of markets.
- The incident: the build target moved from ES2019 to ES2022 (9% smaller, part 1). An OEM WebView with 1.2% of users, 60% of them in one country, could not parse a class-field construct: a SyntaxError at module evaluation, before anything ran. A blank page for a million people, noticed the next day from a page-view drop.
- Why no signal: the error SDK was inside the script that failed to parse. The fix for the signal: a few hundred bytes of inline ES5 in the HTML that registers
window.onerrorfirst and beacons anything that happens before the SDK arrives (part 3's inline capture exists for this). - The fix for the code: a differential build (
type="module"for engines that pass,nomodulewith a lower target and polyfills for the rest); a browserslist derived from the field distribution with a coverage threshold (everything above 0.2% of sessions, reviewed quarterly); a CI check that parses the entry with the oldest supported engine. - The process fix: the user-agent distribution as a quarterly artefact with an owner; a published support list derived from it; the two largest tail WebViews in the device lab; a synthetic monitor loading the page from the affected market on that browser, so a blank page is minutes, not a day.
- The rule: the tail is a list of populations with sizes and markets; support is a threshold on it; telemetry must work when nothing else does; a blank page has no error and needs a synthetic signal. The number that broke was one percent.
INCIDENT: THE ONE PERCENT THAT WAS A MILLION PEOPLE
a browser version the team had never heard of, a syntax error, and a blank page for a country
swipe the figure sideways, or tap expand for full screen
1/6
the distribution
The distribution at 100M: Chrome 120+ 71%; Safari 16+ 14%; Samsung Internet 5%; Firefox 3%; a tail of 7% across dozens of WebViews (OEM shells, in-app browsers of social and messaging apps, older Android system WebViews), each under 1% and each a population of hundreds of thousands to a million. The field data by user agent (part 3) had the tail; nobody had read it as a list of markets.
25
WebViews and hybrid shells
the web app inside native apps
- Three embeddings: the company's hybrid shell (native navigation, web screens); a native app with a few web screens; other companies' in-app browsers (their user agent, cookie policy and script injections). At 100M, in-app browsers are a top-five "browser".
- The engine: on iOS every WebView is the OS's engine (an iOS 15 user has a 2021 engine); on Android the System WebView updates through the store or the app bundles its own. The user agent reports it inconsistently; a feature-detection beacon on load is the reliable census.
- The lifecycle: suspended (timers stop, sockets drop) or killed (state gone) whenever the native app is backgrounded; a return may be a resume mid-flow or a cold start. Save state on hidden; treat every show as a possible cold start; re-establish sockets (the FSD course M8). The Browser course part 10 with the native app's moods on top.
- The bridge: a message channel with a schema (
{ method, id, params }and{ id, result | error }), versioned on both sides. The web asks the shell its bridge version at start and adapts per method (fall back to the web API, or hide the feature), because app-store releases take weeks and users do not update. - Gotchas: storage partitioned per WebView or app; no third-party cookies; the back button is the shell's; safe-area insets; link handoff to the system browser; OAuth redirects and downloads behave differently; the user agent lies. Every one is a bug report that says "works in Safari".
- The rule: the WebView is a platform with its own distribution, lifecycle and cadence; the bridge is a versioned contract the web degrades across; the device lab holds the real app builds; and the web team owns a dashboard of shell versions in the field, because the web will be blamed for what the shell did.
WEBVIEWS AND HYBRID SHELLS
the same web app inside a native app: a different browser, a different lifecycle, a bridge, and a release you do not control
swipe the figure sideways, or tap expand for full screen
1/6
the shapes
The shapes: the company's own app embedding the web app for most screens (a hybrid shell: native navigation and tabs, web content); a native app embedding a few web screens (help, legal, promotions); the web app opened in another app's in-app browser (a social app's link preview, which is a WebView with that app's user agent, cookie policy and JavaScript injections).
26
Release trains, experiments at scale, and the client you cannot recall
shipping to a hundred million
- Trains: a cut at a fixed time from main; rings of internal (an hour), 1% (four hours), 10% (a day), 100%; a change not ready waits for the next train; hotfixes are their own train with a named approver. The schedule is the forcing function that keeps changes small.
- Ring health: per-version dashboards (part 6) compared against the previous version on the same ring: error rate, crash-free sessions, LCP and INP p75, conversion, by locale and device tier; promotion automatic within tolerance; a breach halts the train and pages the owner.
- Experiments: every change that could move a metric ships as a sticky randomised variant with a pre-registered primary metric, guardrails, a sample and duration from the effect size, and an analysis that reports a confidence interval. The platform owns assignment, exposure logging and analysis; teams own hypotheses.
- Experiment correctness: sample ratio mismatch (52/48 on a 50/50 split means biased exposure logging, often slow devices failing to log: the result is invalid); peeking (sequential tests or fixed horizons); interactions between concurrent experiments (layers or exclusion); novelty (a week is not a month). The platform enforces these or the organisation ships noise.
the client you cannot recall
- Where the bytes are: the HTTP cache (to max-age), a service worker's Cache Storage (until the worker updates), open tabs (until a reload: some have been open for weeks). A rollback stops serving the artefact; it does not remove it.
- The levers and their time-to-fleet: a kill switch (seconds: running clients poll flags); a service-worker update check on every load and hourly (minutes to hours); a version beacon that prompts old clients to reload, forced past a deadline for a security fix (hours); CDN purge plus cache lifetimes chosen knowing all this (up to a day for stubborn caches).
- Incident response: a runbook per failure class (blank page, broken checkout, a leaking feature, a security fix) naming the lever and its time-to-fleet; a status page driven by the same alerts; a review that adds a guardrail or a synthetic for what was not watched. The SRE course has the server side of the same incident.
RELEASE TRAINS, EXPERIMENTS AT SCALE, AND THE CLIENT YOU CANNOT RECALL
shipping to a hundred million people on a schedule, measuring every change, and responding when the bytes are already on the device
swipe the figure sideways, or tap expand for full screen
1/6
the train
The train: a cut at a fixed time (say 09:00 Tuesday and Thursday) from main; the artefact goes through rings: internal employees (ring 0, an hour), 1% (ring 1, four hours), 10% (ring 2, a day), 100%. A change that is not ready waits for the next train; nothing is rushed onto a train. Hotfixes are their own train with a shorter ring schedule and a named approver. The schedule is the forcing function that keeps changes small.
27
The course, in one table
| scale | what broke | the number | the architecture change | the course parts |
|---|---|---|---|---|
| 100k | The bundle | Entry bytes, growing 2 KB a day | A per-chunk budget in CI, ratcheted down | 1, 5 |
| 100k | The cache | Deploys per cache lifetime | Chunks by change frequency; a hash-churn check; a hit-rate alert | 1, 5 |
| 100k | State | Places an entity can live (ceiling: one) | One owner per kind, enforced by lint | 0, 5; React 7 |
| 100k | Tests | Flake rate | A flake budget, quarantine, a smoke set on the preview | 2, 5 |
| 1M | Distance | Far users over near | The edge assembles the page; cacheability on every response | 6; Cloud |
| 1M | Exposure | Users per minute of a bad release | Deploy is not release: flags, canaries, pointer rollbacks | 6 |
| 1M | Devices | Share of users on hardware the team lacks | Budgets per tier; a device lab; real-device CI | 6; FSD |
| 1M | Languages, support | Keys × locales; tickets per engineer-hour | Pipelines: translation CI, Intl, tickets with diagnoses attached | 6, 3 |
| 100M | The tail | One percent | Support as a threshold on a list; differential builds; inline capture; synthetics | 7, 3 |
| 100M | WebViews | Shell versions in the field | A versioned bridge; aggressive save, defensive resume; the shells in the lab | 7; Browser 10 |
| 100M | Shipping | The fleet's caches | Trains with rings; experiments with guardrails; kill switches and a runbook with time-to-fleet | 7; SRE |
what to keep
- Every architecture decision has a number that breaks it; write the number down when you make the decision, and watch it.
- Boundaries, pipelines, tests and telemetry are the decisions that are expensive to reverse; spend design time in proportion.
- At each order of magnitude the same number changes kind: a slow-growing scalar, then a distribution to cut by dimension, then a list of populations with markets.
- The fix is a measurement and a line: a budget, an alert, a lint rule, a flake rate, a ring guardrail. The cleanup is the symptom; the line is the architecture.
the exercise, for the course
Take the frontend you own and fill in the table for it: for each row, what is the number today, is anyone watching it, and what is the line? The rows with no watcher are the next incidents, in order of how fast their number is moving.