Part 5 · 4 chapters · ~25 min
What Breaks at 100k
Four incident logs: a bundle that grew two kilobytes a day until the phone cohort stopped converting, a cache hit rate of sixty percent from a deploy a day and hashes that changed without content, state that lived in four places and disagreed, and a test suite nobody believed. Each ends in the measurement and the line that would have caught it.
16
The bundle that grew two kilobytes a day
the incident
- Month 0: entry 280 KB gzip, LCP p75 on phones 2.1 s, no size budget, the analyser run when someone remembers.
- Month 6: 520 KB. moment for one label (68 KB); an icon barrel with side effects (190 KB for 4 KB used); a chart library imported at the top of a module every route imports (110 KB for the admin page). LCP 2.9 s. Desktop fine; nobody notices.
- Month 12: 880 KB. IE11 polyfills two years after support ended (140 KB); three versions of one utility across chunks; all fourteen locales in the entry (220 KB). LCP 3.8 s.
- Month 18: 1,380 KB. Phone conversion down 23% against desktop; a PM raises it; a week lands on Evaluate Script at 1.4 s on a mid phone. "The entry chunk is five times its size from eighteen months ago and nobody set a line."
- The cleanup, by bytes: route-level splitting (380 KB), lazy locales (200 KB), drop the polyfills and raise the target (140 KB), per-icon imports (186 KB), date-fns for moment (60 KB), one utility version (40 KB). Entry 370 KB; LCP 2.3 s within a release.
- The architecture change: a per-chunk budget in CI set at the new size and ratcheted down quarterly; the delta on every PR; the treemap on every build; a dependency checklist (size, ESM, sideEffects, alternatives) in the PR template; a quarterly bundle review. The budget is the fix; the cleanup was the symptom.
INCIDENT: THE BUNDLE THAT GREW 2 KB A DAY
eighteen months of PRs, no budget, and the day the phone cohort stopped converting
swipe the figure sideways, or tap expand for full screen
1/6
month 0
Month 0: entry 280 KB (gzip), vendor 140 KB, LCP p75 on phones 2.1 s. No size budget in CI; the analyser is run by one engineer when they remember. Each PR is reviewed for correctness; nobody reviews bytes because nobody can see them.
17
The sixty percent cache hit rate
the incident
- The setup: one entry and one vendor chunk, but module ids assigned by build order, so adding any module renamed every chunk. A deploy a day; 40% returning users; 520 KB re-downloaded by each of them daily (21 GB a day that should have been zero); CDN hit rate 60% because every deploy invalidated the chunks at every edge.
- The tell: repeat-visit LCP p75 equal to first-visit LCP p75. Caching was doing nothing. The bandwidth bill was the first symptom anyone saw.
- The diagnosis: the Network panel on a repeat visit after a deploy: every chunk a full 200 with correct immutable headers and new filenames. The bundler's stats: module ids changing between builds. The headers were right; the names were wrong.
- The fix: deterministic module ids; chunks split by change frequency (vendor quarterly, ui-kit monthly, app daily, per-route on its PRs); a runtime chunk so the manifest change does not ripple. A deploy now changes the app chunk and a route or two.
- The result: ~60 KB per deploy for a returning user; hit rate 94%; repeat-visit LCP 1.4 s against first-visit 2.3 s; a third off the bandwidth line. The HTML stays no-cache and names chunks mostly already on the device.
- The rule: chunking by change frequency written in the build config with its reason; a CI check comparing chunk hashes between main and the PR ("N chunks changed"; a vendor hash change without a dependency change fails); the hit rate on a dashboard with an alert under 85%. The number that broke was deploys per cache lifetime; the fix made them independent.
INCIDENT: THE 60% CACHE HIT RATE
a deploy a day, a hash per deploy, and a CDN that never got warm
swipe the figure sideways, or tap expand for full screen
1/6
the setup
The setup: one entry chunk and one vendor chunk, but the vendor chunk's hash was derived from the whole build (module ids assigned by order; adding a module shifted every id). Every deploy: new hashes for both. Repeat visitors: 100% of bytes re-downloaded. CDN: a cold cache per deploy per edge.
18
State sprawl
the incident
- The map, drawn during the incident and never before: a global store with a transactions array from year one; a query cache with pages from year two; form state for drafts; URL state for filters; a selected id in three places. The same entity, four copies, four update paths.
- The bug: "mark reviewed" in the detail panel updated the query cache; the list read the global store; the list showed unreviewed until a reload re-fetched the store. Two engineers a year apart, each right about their path. "A reload fixes it" is the signature of two sources of truth.
- The rule (the React course part 7): server state in the query cache and nowhere else; UI state in the component or a small UI store; URL state in the URL; derived state computed, never stored. One owner per kind.
- The migration: the dependency graph lists 31 reads of the store's transactions; a codemod for the mechanical ones, hand edits for the rest; delete the slice; the type-check catches the misses; a flag switches the list's data source so it ships incrementally. Two weeks.
- Enforcement: a lint rule forbidding server-shaped data in the UI store's types; a review checklist question ("which kind of state is this, and where does that kind live"); the state map kept in the docs. Sprawl returns the moment the rule is unwritten.
- The result: one copy of each entity; list and panel agree by construction (the FSD course M1 v3); the global store shrank from 4,000 lines to a 300-line UI store. The number that broke was the count of places an entity can live, and it broke at two.
INCIDENT: STATE SPRAWL
four stores, two caches, one truth nobody could name, and the bug that lived between them
swipe the figure sideways, or tap expand for full screen
1/6
the map
The map, drawn during the incident: the global store held a transactions array (fetched on app load, updated by some mutations); the query cache held pages of transactions (fetched per view, invalidated by other mutations); form state held the detail panel's draft; URL state held the filters; and a "selected" id lived in three places. The same entity, four copies, updated by different code paths.
19
Flaky tests, and the shape of 100k
the fourth incident, briefly
- The suite: 140 E2E tests, 11% flake rate, 40-minute runs, retried three times silently. Engineers re-ran on red by reflex; a real failure (a broken checkout) sat red for two days as "probably flaky". The incident was the checkout.
- The fix was part 2's discipline: flake rate per test measured from the runner's history; the 22 tests over 5% quarantined with owners and dates; the suite split into a 6-flow smoke set that gates merges from the preview deploy and the rest on main; retries reduced to one and reported; the five causes worked through test by test (timing fixed by retrying assertions, selectors moved to roles, shared state replaced by API seeding with run ids, a payment sandbox stubbed at the edge, and two real races in the app fixed). Flake rate 1.2% within a quarter; the smoke set under five minutes; a red merge gate is believed again.
what a hundred thousand users breaks
- The bundle, because growth is gradual and nobody set a line. The fix is a budget in CI.
- The cache, because deploys became daily and the chunking ignored change frequency. The fix is chunks by frequency and a churn check.
- State, because a second mechanism was added beside the first and nothing said which owned what. The fix is one owner per kind, enforced.
- Tests, because the suite grew faster than its discipline. The fix is a flake budget and a smoke set.
the pattern
Every 100k incident is a number that grew slowly with nothing watching it. The architecture change in each case was a measurement and a line: a budget, a hit-rate alert, a lint rule, a flake rate. Parts 6 and 7 are what happens when the numbers grow by orders of magnitude instead.