Part 7 · 4 chapters · ~30 min

What Breaks at 100M

Incident logs at a hundred million users: the one percent that was a million people and a blank page with no error, WebViews and hybrid shells as a platform with their own lifecycle and a versioned bridge, and shipping as release trains with rings, experiments with guardrails and correctness checks, and a client fleet you cannot recall. Then the course in one table.

24

The one percent that was a million people

the incident
  1. The distribution at 100M: Chrome 71%, Safari 14%, Samsung 5%, Firefox 3%, and a 7% tail across dozens of WebViews and in-app browsers, each under 1% and each a population of hundreds of thousands to a million. The field user-agent data had it; nobody had read it as a list of markets.
  2. The incident: the build target moved from ES2019 to ES2022 (9% smaller, part 1). An OEM WebView with 1.2% of users, 60% of them in one country, could not parse a class-field construct: a SyntaxError at module evaluation, before anything ran. A blank page for a million people, noticed the next day from a page-view drop.
  3. Why no signal: the error SDK was inside the script that failed to parse. The fix for the signal: a few hundred bytes of inline ES5 in the HTML that registers window.onerror first and beacons anything that happens before the SDK arrives (part 3's inline capture exists for this).
  4. The fix for the code: a differential build (type="module" for engines that pass, nomodule with a lower target and polyfills for the rest); a browserslist derived from the field distribution with a coverage threshold (everything above 0.2% of sessions, reviewed quarterly); a CI check that parses the entry with the oldest supported engine.
  5. The process fix: the user-agent distribution as a quarterly artefact with an owner; a published support list derived from it; the two largest tail WebViews in the device lab; a synthetic monitor loading the page from the affected market on that browser, so a blank page is minutes, not a day.
  6. The rule: the tail is a list of populations with sizes and markets; support is a threshold on it; telemetry must work when nothing else does; a blank page has no error and needs a synthetic signal. The number that broke was one percent.
INCIDENT: THE ONE PERCENT THAT WAS A MILLION PEOPLE
a browser version the team had never heard of, a syntax error, and a blank page for a country
swipe the figure sideways, or tap expand for full screen
1/6
the distribution
The distribution at 100M: Chrome 120+ 71%; Safari 16+ 14%; Samsung Internet 5%; Firefox 3%; a tail of 7% across dozens of WebViews (OEM shells, in-app browsers of social and messaging apps, older Android system WebViews), each under 1% and each a population of hundreds of thousands to a million. The field data by user agent (part 3) had the tail; nobody had read it as a list of markets.
25

WebViews and hybrid shells

the web app inside native apps
  1. Three embeddings: the company's hybrid shell (native navigation, web screens); a native app with a few web screens; other companies' in-app browsers (their user agent, cookie policy and script injections). At 100M, in-app browsers are a top-five "browser".
  2. The engine: on iOS every WebView is the OS's engine (an iOS 15 user has a 2021 engine); on Android the System WebView updates through the store or the app bundles its own. The user agent reports it inconsistently; a feature-detection beacon on load is the reliable census.
  3. The lifecycle: suspended (timers stop, sockets drop) or killed (state gone) whenever the native app is backgrounded; a return may be a resume mid-flow or a cold start. Save state on hidden; treat every show as a possible cold start; re-establish sockets (the FSD course M8). The Browser course part 10 with the native app's moods on top.
  4. The bridge: a message channel with a schema ({ method, id, params } and { id, result | error }), versioned on both sides. The web asks the shell its bridge version at start and adapts per method (fall back to the web API, or hide the feature), because app-store releases take weeks and users do not update.
  5. Gotchas: storage partitioned per WebView or app; no third-party cookies; the back button is the shell's; safe-area insets; link handoff to the system browser; OAuth redirects and downloads behave differently; the user agent lies. Every one is a bug report that says "works in Safari".
  6. The rule: the WebView is a platform with its own distribution, lifecycle and cadence; the bridge is a versioned contract the web degrades across; the device lab holds the real app builds; and the web team owns a dashboard of shell versions in the field, because the web will be blamed for what the shell did.
WEBVIEWS AND HYBRID SHELLS
the same web app inside a native app: a different browser, a different lifecycle, a bridge, and a release you do not control
swipe the figure sideways, or tap expand for full screen
1/6
the shapes
The shapes: the company's own app embedding the web app for most screens (a hybrid shell: native navigation and tabs, web content); a native app embedding a few web screens (help, legal, promotions); the web app opened in another app's in-app browser (a social app's link preview, which is a WebView with that app's user agent, cookie policy and JavaScript injections).
26

Release trains, experiments at scale, and the client you cannot recall

shipping to a hundred million
  1. Trains: a cut at a fixed time from main; rings of internal (an hour), 1% (four hours), 10% (a day), 100%; a change not ready waits for the next train; hotfixes are their own train with a named approver. The schedule is the forcing function that keeps changes small.
  2. Ring health: per-version dashboards (part 6) compared against the previous version on the same ring: error rate, crash-free sessions, LCP and INP p75, conversion, by locale and device tier; promotion automatic within tolerance; a breach halts the train and pages the owner.
  3. Experiments: every change that could move a metric ships as a sticky randomised variant with a pre-registered primary metric, guardrails, a sample and duration from the effect size, and an analysis that reports a confidence interval. The platform owns assignment, exposure logging and analysis; teams own hypotheses.
  4. Experiment correctness: sample ratio mismatch (52/48 on a 50/50 split means biased exposure logging, often slow devices failing to log: the result is invalid); peeking (sequential tests or fixed horizons); interactions between concurrent experiments (layers or exclusion); novelty (a week is not a month). The platform enforces these or the organisation ships noise.
the client you cannot recall
  1. Where the bytes are: the HTTP cache (to max-age), a service worker's Cache Storage (until the worker updates), open tabs (until a reload: some have been open for weeks). A rollback stops serving the artefact; it does not remove it.
  2. The levers and their time-to-fleet: a kill switch (seconds: running clients poll flags); a service-worker update check on every load and hourly (minutes to hours); a version beacon that prompts old clients to reload, forced past a deadline for a security fix (hours); CDN purge plus cache lifetimes chosen knowing all this (up to a day for stubborn caches).
  3. Incident response: a runbook per failure class (blank page, broken checkout, a leaking feature, a security fix) naming the lever and its time-to-fleet; a status page driven by the same alerts; a review that adds a guardrail or a synthetic for what was not watched. The SRE course has the server side of the same incident.
RELEASE TRAINS, EXPERIMENTS AT SCALE, AND THE CLIENT YOU CANNOT RECALL
shipping to a hundred million people on a schedule, measuring every change, and responding when the bytes are already on the device
swipe the figure sideways, or tap expand for full screen
1/6
the train
The train: a cut at a fixed time (say 09:00 Tuesday and Thursday) from main; the artefact goes through rings: internal employees (ring 0, an hour), 1% (ring 1, four hours), 10% (ring 2, a day), 100%. A change that is not ready waits for the next train; nothing is rushed onto a train. Hotfixes are their own train with a shorter ring schedule and a named approver. The schedule is the forcing function that keeps changes small.
27

The course, in one table

scalewhat brokethe numberthe architecture changethe course parts
100kThe bundleEntry bytes, growing 2 KB a dayA per-chunk budget in CI, ratcheted down1, 5
100kThe cacheDeploys per cache lifetimeChunks by change frequency; a hash-churn check; a hit-rate alert1, 5
100kStatePlaces an entity can live (ceiling: one)One owner per kind, enforced by lint0, 5; React 7
100kTestsFlake rateA flake budget, quarantine, a smoke set on the preview2, 5
1MDistanceFar users over nearThe edge assembles the page; cacheability on every response6; Cloud
1MExposureUsers per minute of a bad releaseDeploy is not release: flags, canaries, pointer rollbacks6
1MDevicesShare of users on hardware the team lacksBudgets per tier; a device lab; real-device CI6; FSD
1MLanguages, supportKeys × locales; tickets per engineer-hourPipelines: translation CI, Intl, tickets with diagnoses attached6, 3
100MThe tailOne percentSupport as a threshold on a list; differential builds; inline capture; synthetics7, 3
100MWebViewsShell versions in the fieldA versioned bridge; aggressive save, defensive resume; the shells in the lab7; Browser 10
100MShippingThe fleet's cachesTrains with rings; experiments with guardrails; kill switches and a runbook with time-to-fleet7; SRE
what to keep
  1. Every architecture decision has a number that breaks it; write the number down when you make the decision, and watch it.
  2. Boundaries, pipelines, tests and telemetry are the decisions that are expensive to reverse; spend design time in proportion.
  3. At each order of magnitude the same number changes kind: a slow-growing scalar, then a distribution to cut by dimension, then a list of populations with markets.
  4. The fix is a measurement and a line: a budget, an alert, a lint rule, a flake rate, a ring guardrail. The cleanup is the symptom; the line is the architecture.
the exercise, for the course
Take the frontend you own and fill in the table for it: for each row, what is the number today, is anyone watching it, and what is the line? The rows with no watcher are the next incidents, in order of how fast their number is moving.