Part 6 · 4 chapters · ~25 min
What Breaks at 1M
Incident logs at a million users: an origin that served every page to forty countries, a deploy that was a release and reached everyone in minutes, a laptop that was the test device while a third of users were on 2 GB phones, and i18n and support becoming pipelines. Each ends in a cut by dimension and a control per cut.
20
The origin that served every page
the incident
- The map: one origin region; a CDN caching the chunks but passing the HTML and every API call through. A far user paid a 250 ms TLS handshake to the origin, 350 ms for the HTML, then 350 ms per API call. LCP p75 far 4.1 s, near 1.6 s.
- The tell: LCP cut by country was bimodal and TTFB by country explained it entirely; the chunks were fine everywhere. A HAR from a far user showed the handshake and the waiting, twice.
- HTML at the edge: the shell cached with a short TTL and stale-while-revalidate; personalisation moved to a small API call or an edge function reading a cookie; TLS terminated at the edge with a warm connection to the origin. HTML 20 ms everywhere.
- API at the edge: every response declares its cacheability; public responses cached by URL and purged on change; per-user slow-changing responses cached with the session in the key and a short TTL; dynamic calls still reach the origin but over the warm connection, without the handshake.
- A second region with read replicas only after the first two steps: most of the far TTFB was the handshake and the HTML, not the database (the Cloud course for what multi-region costs).
- The result: far LCP 4.1 to 2.0 s; origin requests down 70%; cost shifted from origin compute to cheaper edge requests. The rule: the edge assembles the page, the origin does what only it can, every response declares its cacheability. The number that broke was the ratio of far users to near, and it crossed one when the product went global.
INCIDENT: THE ORIGIN THAT SERVED EVERY PAGE
a million users in forty countries, one region, and a TTFB that was the distance
swipe the figure sideways, or tap expand for full screen
1/6
the map
The map: one origin in one region (compute, the API, the database); a CDN with 200 points of presence serving the hashed chunks (cached, fast everywhere) but passing the HTML and every API call through to the origin (uncached, slow in proportion to distance). A user in the far region: TLS to the origin 250 ms, HTML 350 ms, chunks 40 ms (edge), API 350 ms, API 350 ms. LCP p75 far: 4.1 s; near: 1.6 s.
21
The deploy that was a release
the incident
- Before: merge, build, deploy, and every user had the new bundle on their next load. A regression's blast radius was 100% from minute one; the only control was a rollback with the same latency as a deploy.
- The incident: a currency formatting change rounded 0.005 wrongly in de-DE and fr-FR; the payment provider rejected mismatched totals; 14,000 failed checkouts in eleven minutes. The error rate by locale spiked in sixty seconds (part 3); the rollback took eleven because it was manual with a purge. The incident was the eleven minutes.
- Flags: a remote boolean or variant evaluated per user; the code ships with both paths; kill switches default to off with one job. The flag service is critical-path: bootstrapped into the HTML or cached locally, and failing safe (a missing flag is off).
- Canaries: the new artefact served to a sticky slice (a cookie or a hash of the user id) with dashboards per version (error rate, LCP, INP, conversion, by locale). Promotion is a decision with numbers; a regression stops at the slice.
- The rollout after: deploy dark; 1% for an hour; 10% for a day; 100%. Each step a pointer change at the edge, seconds to apply and reverse. A regression at 1% is ten thousand users for a minute, not a million for eleven.
- The rule: deploy is not release; artefacts are addressable by hash at the edge; users are sticky to a version; the rollback is a pointer; risky changes ship behind a flag with a kill switch; dashboards cut by version and flag. The number that broke was users per minute of exposure; at a million users, one minute is the incident.
INCIDENT: THE DEPLOY THAT WAS A RELEASE
one artefact, every user, and a checkout bug that reached a million people in four minutes
swipe the figure sideways, or tap expand for full screen
1/6
before
Before: merge → build → deploy = release. The new bundle is live for everyone on the next load. A regression's blast radius is 100% of traffic from minute one; the only control is the rollback, which is a deploy in reverse with the same latency.
22
The laptop that was the test device
the incident
- The field: INP p75 380 ms overall; by device class 90 ms on 8 GB-plus, 160 on 4 GB, 620 on 2 GB. The 2 GB tier was 30% of sessions and 42% of new signups: the growth markets. Every engineer had a 16 GB laptop and a flagship phone.
- The lab: Lighthouse's simulated throttling models one mid-tier phone; its TBT said 300 ms. A real 2 GB device with background apps, slow storage and thermal throttling said 620. The lab measured a device nobody had.
- The real device: a $120 Android, remote debugging over USB, the Performance panel on the main flow: JSON.parse of a 900 KB response 180 ms, an unwindowed 3,000-row render 210 ms, a chat widget's script evaluation 300 ms on every route change, GC pauses under memory pressure.
- Code fixes, each a course part: paginate (FSD M1), window the list (M6), defer the widget to idle and interaction (Browser 6), cut the entry for this tier (part 5). 2 GB INP 620 to 210 ms; 8 GB 90 to 70. Nobody with a laptop would have found them.
- Process fixes: a device lab spanning the field distribution; real-device runs of the smoke flows in CI with budgets per tier; per-tier dashboards by default; every performance PR states the tier it measured on. "Mobile" became three budgets.
- The rule: the device distribution is measured, not assumed; budgets are per tier; CI runs on the worst tier that is a real market; engineers have access to one. The number that broke was the share of users on devices the team did not have; at a million users that share is a market.
INCIDENT: THE LAPTOP THAT WAS THE TEST DEVICE
a million users, a device distribution nobody had measured, and an INP that was fine on every engineer's machine
swipe the figure sideways, or tap expand for full screen
1/6
the field
The field: INP p75 380 ms overall; cut by device class (from deviceMemory and cores, part 3): 90 ms on 8 GB-plus devices, 160 ms on 4 GB, 620 ms on 2 GB. The 2 GB tier was 30% of sessions and 42% of new signups (the markets the product was growing in). Every engineer had a 16 GB laptop and a flagship phone.
23
i18n, support load, and the shape of a million
two more, briefly
- i18n: fourteen locales added one at a time, each as a hand-written JSON file in the entry (part 5's 220 KB). At a million users: catalogues per locale loaded lazily by route; a translation pipeline (extraction from source, a translation service, PRs from the service, a CI check that no key is missing in any locale); pluralisation and formatting through
Intl(dates, numbers, currencies, lists, relative time) rather than string templates; RTL as a layout mode tested in the visual regression suite; the locale as a dimension on every metric (the formatting incident above was found by it). The number that broke was keys times locales times engineers editing JSON by hand. - Support load: at a million users the support queue is a metric. Tickets by category, tagged to a route and a release, on the same dashboard as errors; a "report a problem" affordance that attaches the session id, the route, the browser and the last twenty breadcrumbs (part 3) so a ticket opens to a trace; feature flags that support can read per user ("which version is this customer on?"); a status page driven by the same alerts as the on-call. The number that broke was tickets per engineer-hour, and the fix was making each ticket arrive with its diagnosis attached.
what a million users breaks
- Distance, because the users are everywhere and the origin is somewhere. The edge assembles the page.
- Exposure, because a deploy reaches a million people in minutes. Deploy is not release.
- Devices, because the distribution is wider than the team's. Budgets per tier, measured on real hardware.
- Languages and support, because both are now pipelines, not tasks.
the pattern
At a hundred thousand the numbers grew slowly; at a million they are distributions. Every fix in this part replaced one number with a cut by dimension (country, version, device tier, locale) and a control per cut. Part 7 is what happens when the tail of each distribution becomes its own population.