Part 2 · 3 chapters · ~25 min

Testing Strategy

The pyramid and where it lies for a frontend, the diamond with a thick middle of component tests, the other kinds by what they catch, contract tests that make the seam between frontend and API tested once from both sides, and E2E that stays green through five fixes and a flake budget.

8

The pyramid, and where it lies

code
// the thick middle: a feature rendered with its real parts, the network mocked from the contract, interacted with like a user
import { render, screen } from '@testing-library/react'
import userEvent from '@testing-library/user-event'
import { http, HttpResponse } from 'msw'
import { server } from '@/test/msw'                       // handlers generated from the OpenAPI contract; this test overrides one
import { TransferMoney } from '@features/transfer-money'   // through the public API (part 0); real form, real store, real validation

test('a transfer over the limit shows the reason and never submits', async () => {
  const posted: unknown[] = []
  server.use(http.post('/api/transfers', async ({ request }) => { posted.push(await request.json()); return HttpResponse.json({ id: 't_1' }) }))
  render(<TransferMoney account={{ id: 'a_1', balance: 500_00 }} />, { wrapper: AppProviders })   // real providers: query cache, router, i18n

  const user = userEvent.setup()
  await user.type(screen.getByRole('textbox', { name: /amount/i }), '600')
  await user.click(screen.getByRole('button', { name: /send/i }))

  expect(await screen.findByText(/exceeds your balance by 100\.00/i)).toBeVisible()   // what a user sees; retries until it appears
  expect(posted).toHaveLength(0)                                                    // and what the server never received
})
// what this covers: the form wiring, the money formatting, the validation rule, the store update, the button's disabled state, the error's accessible text
// what it does not: the real API (the contract test does), the real browser rendering (a smoke E2E does), visual regressions (a screenshot test can)
the diamond
  1. The pyramid's cost reasoning is right; its proportions assume unit tests find most bugs. For UI code, most bugs are wiring (the wrong prop, the undefined selector, the hidden button, the double submit) and wiring is exactly what mocked-children unit tests delete.
  2. A thin base of pure-logic unit tests: formatters, parsers, reducers, selectors, money and date arithmetic, algorithms. Fast, exhaustive, property-based where the input space is large. Thin because the frontend's pure logic is a small fraction of its code; thick in cases because those functions have many.
  3. A thick middle of component and integration tests: a feature rendered with its real children and store, the network mocked (MSW) from the contract, interacted with like a user (role and name), asserted on what a user sees. One test covers the wiring of twenty modules in 50 ms and fails for reasons users would notice. Testing Library on jsdom for most; a real browser (Vitest browser mode, Playwright component tests) for layout-dependent behaviour.
  4. A thin top of E2E flows whose failure is an incident: login, the main page, the main action, payment, logout. Five to twenty, against the preview deploy; they catch the real API, CDN and browser; they cost minutes per run and hours per flake.
  5. Numbers for a mid-sized app: ~400 unit (2 s), ~800 middle (90 s; 20 s affected-only), ~15 E2E (6 min). Coverage of the wiring is the goal, not a percentage.
the other kinds, by what they catch
  1. Visual regression (screenshot diffs per component state, in a real browser): catches CSS and layout changes no assertion would; costs a review of diffs per PR and a tolerance setting. Worth it for a design system; noisy for feature pages.
  2. Accessibility (axe in component tests; a screen-reader pass manually): catches missing names, roles and contrast; necessary, not sufficient.
  3. Performance (Lighthouse CI budgets, size-limit): part 1's gates; they are tests with numbers.
  4. Type-check is a test too: the cheapest one, and the one that catches the renamed field when types are generated from the contract (next chapter).
THE PYRAMID, AND WHERE IT LIES
test kinds by what they catch, what they cost, and how often they lie
swipe the figure sideways, or tap expand for full screen
1/6
the pyramid
The textbook pyramid: unit (fast, isolated, many), integration (slower, fewer), E2E (slow, brittle, few). The reasoning is cost per run and cost per failure to diagnose. The reasoning is right; the proportions assume the unit tests catch most bugs, which for UI code they do not.
9

Contract tests at the seam

the untested seam
  1. The frontend mocks the API; the API mocks the frontend; both suites are green; the API renames customer_name to customerName; the page renders "undefined". The mock was true once and nothing kept it true.
the contract
  1. A shared, versioned description: an OpenAPI document or a GraphQL schema in a repo both teams own, with examples. From it, generated: TypeScript types (openapi-typescript; a renamed field becomes a compile error), MSW handlers from the examples (mocks correct by construction), request validation in the API, and a diff tool that fails CI on a breaking change (a removed field, a changed type) without a version bump.
  2. Consumer-driven (Pact): the frontend's tests record what they use (request shape, the response fields actually read, with matchers like "a string here"); the pact is published; the API's CI replays every pact against the real service. The API cannot break a consumer it does not know about, because the pacts are the list of consumers.
  3. Assert what you use: fields, types, nullability, enums, status codes per case, pagination shape. Not every field: over-pinning turns the contract into a veto on harmless additions.
  4. In CI: the frontend regenerates types and mocks on every build; the API verifies against pacts or examples; a contract change is a PR both teams review with a changelog entry; breaking changes ship behind a version or alongside the old field until consumers migrate.
  5. The payoff: the thick middle's mocks are true; E2E is no longer the only thing that catches renames; "can we deploy the API" is answered by a test. The Big-company FE course's BFF puts the contract in a layer the frontend team owns.
CONTRACT TESTS AT THE SEAM
the frontend and the API agree on a shape, and both are tested against it
swipe the figure sideways, or tap expand for full screen
1/6
the failure
The failure: the frontend's MSW handlers return { customer_name } because that is what the API returned when the test was written; the API now returns { customerName }; the API's own tests pass (they test the new shape); the frontend's tests pass (they test the mock); the page renders "undefined". Both suites green; the seam untested.
10

E2E that stays green

five causes, five fixes
  1. Timing: asserting before the app is ready. Retrying assertions (expect(locator).toBeVisible()), waiting for the specific thing not a duration, and the app exposing readiness. Never sleep(2000): it is a guess about a machine you do not control.
  2. Selectors: coupled to markup. Roles and accessible names (getByRole("button", { name: "Submit" })), which double as an accessibility check; data-testid only where no accessible handle exists; the Recorder's selector alternatives (DevTools part 9).
  3. Shared state: order-dependent tests, leftover data. Seed through the API with a unique run id; ephemeral environments; any order, in parallel, or they are not independent.
  4. Environment: real third parties, slow runners. Stub at the network edge (route interception); run on the preview deploy with its own API; timeouts from the p99 of the real environment.
  5. Real nondeterminism: a race, a date boundary, a random key. The test is right; fix the app; it earned its keep that day. Never "fix" it with a wait.
the discipline
  1. Every failure triaged within a day; a flake quarantined with an owner and a date in the skip reason, not deleted; the flake rate per test tracked (the runner's history) and budgeted (under 1% per test, 5% per suite); one retry allowed and reported, never silent.
  2. The smoke set (five flows) gates merges from the preview; the full suite runs on main and nightly and pages the owner. A suite nobody trusts costs more than no suite.
the budget, as a sentence
"Our tests cost N minutes per PR and catch M% of the bugs that reach staging; the E2E flake rate is under 2%; a contract change cannot merge without both teams." If you cannot fill in the numbers, the strategy is a hope.
E2E THAT STAYS GREEN
the five causes of flakiness, and the discipline that keeps a suite trusted
swipe the figure sideways, or tap expand for full screen
1/6
timing
Timing: the test clicks and asserts before the data arrives. Fix: assertions that wait (Playwright's expect(locator).toBeVisible() retries until timeout; never sleep(2000)); wait for the specific thing (the row with this text), not for a duration; and the app exposes readiness (a data-ready attribute, or the absence of a loading state) rather than the test guessing.