Part 5 · 4 chapters · ~25 min

What Breaks at 100k

Four incident logs: a bundle that grew two kilobytes a day until the phone cohort stopped converting, a cache hit rate of sixty percent from a deploy a day and hashes that changed without content, state that lived in four places and disagreed, and a test suite nobody believed. Each ends in the measurement and the line that would have caught it.

16

The bundle that grew two kilobytes a day

the incident
  1. Month 0: entry 280 KB gzip, LCP p75 on phones 2.1 s, no size budget, the analyser run when someone remembers.
  2. Month 6: 520 KB. moment for one label (68 KB); an icon barrel with side effects (190 KB for 4 KB used); a chart library imported at the top of a module every route imports (110 KB for the admin page). LCP 2.9 s. Desktop fine; nobody notices.
  3. Month 12: 880 KB. IE11 polyfills two years after support ended (140 KB); three versions of one utility across chunks; all fourteen locales in the entry (220 KB). LCP 3.8 s.
  4. Month 18: 1,380 KB. Phone conversion down 23% against desktop; a PM raises it; a week lands on Evaluate Script at 1.4 s on a mid phone. "The entry chunk is five times its size from eighteen months ago and nobody set a line."
  5. The cleanup, by bytes: route-level splitting (380 KB), lazy locales (200 KB), drop the polyfills and raise the target (140 KB), per-icon imports (186 KB), date-fns for moment (60 KB), one utility version (40 KB). Entry 370 KB; LCP 2.3 s within a release.
  6. The architecture change: a per-chunk budget in CI set at the new size and ratcheted down quarterly; the delta on every PR; the treemap on every build; a dependency checklist (size, ESM, sideEffects, alternatives) in the PR template; a quarterly bundle review. The budget is the fix; the cleanup was the symptom.
INCIDENT: THE BUNDLE THAT GREW 2 KB A DAY
eighteen months of PRs, no budget, and the day the phone cohort stopped converting
swipe the figure sideways, or tap expand for full screen
1/6
month 0
Month 0: entry 280 KB (gzip), vendor 140 KB, LCP p75 on phones 2.1 s. No size budget in CI; the analyser is run by one engineer when they remember. Each PR is reviewed for correctness; nobody reviews bytes because nobody can see them.
17

The sixty percent cache hit rate

the incident
  1. The setup: one entry and one vendor chunk, but module ids assigned by build order, so adding any module renamed every chunk. A deploy a day; 40% returning users; 520 KB re-downloaded by each of them daily (21 GB a day that should have been zero); CDN hit rate 60% because every deploy invalidated the chunks at every edge.
  2. The tell: repeat-visit LCP p75 equal to first-visit LCP p75. Caching was doing nothing. The bandwidth bill was the first symptom anyone saw.
  3. The diagnosis: the Network panel on a repeat visit after a deploy: every chunk a full 200 with correct immutable headers and new filenames. The bundler's stats: module ids changing between builds. The headers were right; the names were wrong.
  4. The fix: deterministic module ids; chunks split by change frequency (vendor quarterly, ui-kit monthly, app daily, per-route on its PRs); a runtime chunk so the manifest change does not ripple. A deploy now changes the app chunk and a route or two.
  5. The result: ~60 KB per deploy for a returning user; hit rate 94%; repeat-visit LCP 1.4 s against first-visit 2.3 s; a third off the bandwidth line. The HTML stays no-cache and names chunks mostly already on the device.
  6. The rule: chunking by change frequency written in the build config with its reason; a CI check comparing chunk hashes between main and the PR ("N chunks changed"; a vendor hash change without a dependency change fails); the hit rate on a dashboard with an alert under 85%. The number that broke was deploys per cache lifetime; the fix made them independent.
INCIDENT: THE 60% CACHE HIT RATE
a deploy a day, a hash per deploy, and a CDN that never got warm
swipe the figure sideways, or tap expand for full screen
1/6
the setup
The setup: one entry chunk and one vendor chunk, but the vendor chunk's hash was derived from the whole build (module ids assigned by order; adding a module shifted every id). Every deploy: new hashes for both. Repeat visitors: 100% of bytes re-downloaded. CDN: a cold cache per deploy per edge.
18

State sprawl

the incident
  1. The map, drawn during the incident and never before: a global store with a transactions array from year one; a query cache with pages from year two; form state for drafts; URL state for filters; a selected id in three places. The same entity, four copies, four update paths.
  2. The bug: "mark reviewed" in the detail panel updated the query cache; the list read the global store; the list showed unreviewed until a reload re-fetched the store. Two engineers a year apart, each right about their path. "A reload fixes it" is the signature of two sources of truth.
  3. The rule (the React course part 7): server state in the query cache and nowhere else; UI state in the component or a small UI store; URL state in the URL; derived state computed, never stored. One owner per kind.
  4. The migration: the dependency graph lists 31 reads of the store's transactions; a codemod for the mechanical ones, hand edits for the rest; delete the slice; the type-check catches the misses; a flag switches the list's data source so it ships incrementally. Two weeks.
  5. Enforcement: a lint rule forbidding server-shaped data in the UI store's types; a review checklist question ("which kind of state is this, and where does that kind live"); the state map kept in the docs. Sprawl returns the moment the rule is unwritten.
  6. The result: one copy of each entity; list and panel agree by construction (the FSD course M1 v3); the global store shrank from 4,000 lines to a 300-line UI store. The number that broke was the count of places an entity can live, and it broke at two.
INCIDENT: STATE SPRAWL
four stores, two caches, one truth nobody could name, and the bug that lived between them
swipe the figure sideways, or tap expand for full screen
1/6
the map
The map, drawn during the incident: the global store held a transactions array (fetched on app load, updated by some mutations); the query cache held pages of transactions (fetched per view, invalidated by other mutations); form state held the detail panel's draft; URL state held the filters; and a "selected" id lived in three places. The same entity, four copies, updated by different code paths.
19

Flaky tests, and the shape of 100k

the fourth incident, briefly
  1. The suite: 140 E2E tests, 11% flake rate, 40-minute runs, retried three times silently. Engineers re-ran on red by reflex; a real failure (a broken checkout) sat red for two days as "probably flaky". The incident was the checkout.
  2. The fix was part 2's discipline: flake rate per test measured from the runner's history; the 22 tests over 5% quarantined with owners and dates; the suite split into a 6-flow smoke set that gates merges from the preview deploy and the rest on main; retries reduced to one and reported; the five causes worked through test by test (timing fixed by retrying assertions, selectors moved to roles, shared state replaced by API seeding with run ids, a payment sandbox stubbed at the edge, and two real races in the app fixed). Flake rate 1.2% within a quarter; the smoke set under five minutes; a red merge gate is believed again.
what a hundred thousand users breaks
  1. The bundle, because growth is gradual and nobody set a line. The fix is a budget in CI.
  2. The cache, because deploys became daily and the chunking ignored change frequency. The fix is chunks by frequency and a churn check.
  3. State, because a second mechanism was added beside the first and nothing said which owned what. The fix is one owner per kind, enforced.
  4. Tests, because the suite grew faster than its discipline. The fix is a flake budget and a smoke set.
the pattern
Every 100k incident is a number that grew slowly with nothing watching it. The architecture change in each case was a measurement and a line: a budget, a hit-rate alert, a lint rule, a flake rate. Parts 6 and 7 are what happens when the numbers grow by orders of magnitude instead.