Part 12 · 5 chapters · ~30 min

Testing React

A test that resembles how the software is used survives refactors and catches regressions; one that knows the implementation does neither. This part is the test levels and what each sees and survives, component tests by role and user event, mocking at the wire with MSW, the causes of flakiness and the fix for each, Suspense and transitions in tests, and visual regression with stories.

59

What to test, at which level

the question

"The suite is 2,000 unit tests that break on every refactor and 40 end-to-end tests that fail on Tuesdays. What should it be instead?"

Mostly component tests that interact with the rendered UI the way a user does and assert what a user would see, written so they do not know the implementation. Some integration tests per feature with the network mocked at the wire. A few end-to-end tests on the paths that lose money if they break. Visual tests on the design system. Unit tests where logic is dense enough to deserve them. The principle (Testing Library's): the more a test resembles how the software is used, the more confidence it gives, and the more refactors it survives.

code
// what to test, by component kind
// a presentational component (a Badge): a story + visual test; a component test only if it has logic (a variant with a11y implications)
// a form: component test of the full flow (fill, validate, submit, errors, pending, success); the schema has its own unit tests
// a list with selection, sorting, filtering: component test of the behaviours the user sees; a unit test of the sort/filter functions
// a data-driven page: integration test with MSW: loading → content; error → retry; mutation → updated list
// a hook with branching logic (useUndo, usePagination): renderHook tests of the transitions
// a custom input with a mask: component test of typing behaviour (what appears for what is typed; cursor position if you manage it)
// a provider / context: tested through consumers; a direct test only for its reducer
// a route guard / redirect: integration test with the router: unauthenticated → login; authenticated → page
// accessibility: every component test queries by role, so a missing label fails. add vitest-axe / jest-axe for a scan per rendered state
// what not to test: that React renders props (it does); the exact DOM structure; class names; that a mock was called with the right args when the UI result is assertable instead
the levels and their contracts
  1. Hook unit tests (renderHook): state transitions and derived values. For reducers and hooks with branching. They know nothing of markup, which is their strength and their limit.
  2. Component tests (render, screen, userEvent): the user's view. They query by role and accessible name, interact through real event sequences, and assert visible outcomes. They survive swaps of state libraries, form libraries, and markup that preserves semantics. The centre of the suite.
  3. Integration tests: a route with its providers and a mocked API, driving a flow across components: loading, content, error, mutation, cache update. The seams between components are where the bugs that component tests miss live.
  4. End-to-end: the built app, a real browser, real or seeded backend. Few, critical, run on merge. Flakiness is environmental and is managed by keeping the count small.
  5. Visual regression: screenshots per story or page, diffed against baselines. For the design system and layout-critical pages. Reviewed like code.
the measure
A test is good if it fails when a user-visible behaviour breaks and passes when the implementation changes. A test that asserts on state shape, class names, or that a mock was called is testing the implementation; it will fail on the next refactor and pass on the next bug.
WHAT EACH TEST LEVEL SEES
and what each one survives
swipe the figure sideways, or tap expand for full screen
1/6
hook unit
Hook unit test: renderHook(() => useCart()); act(() => result.current.add(item)); expect(result.current.total).toBe(…). Sees: state logic, derived values, edge cases of the reducer. Survives: any markup change. Misses: everything about rendering and interaction. Right for: complex hooks and reducers; wrong for: hooks whose only job is to feed a component.
60

Component tests: queries, events, assertions

code
// a component test in the Testing Library style: queries by what the user perceives; events as the user does them
import { render, screen } from '@testing-library/react'
import userEvent from '@testing-library/user-event'

test('transfers money after confirmation', async () => {
  const user = userEvent.setup()
  render(<TransferForm balance={500_00} />, { wrapper: Providers })            // Providers: QueryClient, router, store; the test's own instances
  await user.type(screen.getByRole('textbox', { name: /amount/i }), '150.50')  // by role + accessible name, not by test id or class
  await user.type(screen.getByRole('textbox', { name: /account/i }), '0123456789')
  await user.click(screen.getByRole('button', { name: /review/i }))
  expect(screen.getByRole('heading', { name: /confirm/i })).toBeInTheDocument()
  expect(screen.getByText('₦150.50')).toBeInTheDocument()                        // the parsed, formatted amount: the thing that must be right
  await user.click(screen.getByRole('button', { name: /send/i }))
  expect(await screen.findByRole('status')).toHaveTextContent(/sent/i)           // findBy: waits for the async result (MSW answered)
  expect(screen.getByRole('button', { name: /send/i })).toBeDisabled()           // no double submit
})
// what this test does not know: the state library, the form library, whether the amount is controlled, the component tree, class names.
// what it would catch: a float parse (₦150.5), a missing disabled state, a confirmation that shows the typed string, a broken label (the query fails: an a11y bug IS a test failure).
// the queries, in priority: getByRole > getByLabelText > getByPlaceholderText > getByText > getByDisplayValue > getByAltText > getByTitle > getByTestId (last resort)
queries
  1. By role and name first: getByRole('button', { name: /send/i }) finds what a screen reader would announce. It fails if the element has no accessible name, which is an accessibility bug surfacing as a test failure: the right incentive.
  2. Then by label, placeholder, text, display value. Test ids last, for things with no semantic handle (a decorative container you must assert exists).
  3. getBy throws if absent (use for "it is there now"); queryBy returns null (use for "it is not there"); findBy waits and polls (use for anything after an await). getAllBy and friends for lists.
  4. within(container) scopes queries to a region (a row, a dialog) when the page has several similar elements.
events
  1. userEvent simulates the full sequence a device produces: type fires keydown, keypress, input, keyup per character and respects maxlength and disabled; click fires pointer and mouse events and focus; keyboard for shortcuts; upload for files; selectOptions; tab for focus order. Always awaited; set up with userEvent.setup() per test.
  2. fireEvent dispatches one synthetic event with no sequence; it can pass tests that fail for users (a disabled button "clicked"; an input changed without keystrokes). Reserve it for events userEvent cannot produce (a custom event, a resize).
assertions
  1. On what the user sees: text content, roles present, attributes that matter (toBeDisabled, toBeChecked, toHaveAccessibleName, toHaveValue, toBeVisible) via jest-dom matchers.
  2. Not on: component state, hook return values, the DOM structure, class names (unless the class is the product, as in a design system), or mock call arguments when the UI result can be asserted instead.
  3. Negative assertions need queryBy and care: "the error is not shown" after an async operation should wait for the success state first, or it passes trivially before the error would have appeared.
61

Mocking the right boundary

code
// the network boundary: Mock Service Worker intercepts fetch at the request level; components and query caches run for real
import { http, HttpResponse } from 'msw'
import { setupServer } from 'msw/node'
export const server = setupServer(
  http.get('/api/posts/:id', ({ params }) => HttpResponse.json({ id: params.id, title: 'Hello', likes: 41 })),
  http.post('/api/transfers', async ({ request }) => {
    const body = await request.json()
    if (body.amount > 500_00) return HttpResponse.json({ error: 'insufficient' }, { status: 422 })   // the error path is a handler too
    return HttpResponse.json({ id: 'tx_1', amount: body.amount })
  }),
)
beforeAll(() => server.listen({ onUnhandledRequest: 'error' }))   // an unmocked request fails the test: no accidental real calls
afterEach(() => server.resetHandlers())
afterAll(() => server.close())
// per-test overrides: server.use(http.get('/api/posts/:id', () => HttpResponse.error()))   // network failure for this test only
// why not mock fetch / axios / the query hook: you would be testing your mocks. MSW leaves the real code path intact to the wire.
// the same handlers run in the browser (service worker) for Storybook and manual testing: one definition of "the API" for dev and test
where to mock
  1. The network, at the wire (MSW): intercept requests; return responses. Everything inside the app runs for real: fetch calls, the query cache, retries, invalidation, serialisation. The handlers double as a contract with the API and as a dev server for Storybook.
  2. Not the data layer: mocking useQuery or the API module tests the mock, skips the cache behaviour, and breaks when the data layer changes.
  3. Not fetch itself (vi.fn() on global fetch): fragile to how the code builds requests; MSW matches on URL and method and lets the request code be real.
  4. Time: fake timers for debounces, intervals and timeouts; advance explicitly. Never sleep.
  5. Randomness and ids: seed or stub crypto.randomUUID and Math.random where output must be deterministic (an idempotency key you assert was sent).
  6. Browser APIs the environment lacks: jsdom has no layout (getBoundingClientRect returns zeros), no IntersectionObserver, no matchMedia; stub them per test or use a browser-mode runner (Vitest browser mode, Playwright component testing) for components that depend on layout.
  7. Modules with side effects (analytics, error reporting): mock at the module boundary; assert they were called only when the behaviour is the call itself.
test setup that scales
  1. A render wrapper that provides the app's providers with test instances: a fresh QueryClient (retries off, gcTime: Infinity to avoid cross-test garbage collection timing), a memory router at a route, a store with initial state. One helper; every test uses it.
  2. Factories for data (a makePost() with overrides) so tests say what matters and defaults cover the rest.
  3. MSW handlers per domain, composed into the server; per-test overrides with server.use for error and edge cases.
  4. onUnhandledRequest: 'error' so a request without a handler fails loudly rather than hanging a findBy until timeout.
62

Async, timing, and flakiness

code
// async and timing: the rules that remove flakiness
// 1. await what you query for: findBy* (waits up to 1 s, polling) for anything that appears after an await; getBy* only for what is there now
expect(await screen.findByText(/sent/i)).toBeInTheDocument()
// 2. userEvent, not fireEvent: userEvent.type() fires the sequence a real keyboard does (keydown, keypress, input, keyup per character), goes through
//    React's event system, and respects disabled and pointer-events. fireEvent.change() skips all that and passes tests that fail for users
// 3. act() is implicit in RTL's render / userEvent / findBy; explicit act() is for state updates you trigger outside those (a store.dispatch, a timer)
// 4. fake timers: vi.useFakeTimers(); userEvent.setup({ advanceTimers: vi.advanceTimersByTime }); advance for debounces, not sleep
// 5. waitFor(() => expect(…)) for conditions that are not a single element appearing (a count reaching 3; a mock called twice)
// 6. never: setTimeout in a test to "let things settle"; snapshotting the whole DOM (brittle, reviews nothing); asserting on state (test outputs, not internals)
// 7. Suspense and transitions: findBy handles the fallback → content change; for isPending UI, assert the pending state then the final
// 8. cleanup: RTL unmounts after each test; the QueryClient must be per test (a shared cache leaks data between tests); MSW handlers reset per test
why tests flake, and the fix for each
  1. Asserting before the async result: getBy right after a click that triggers a fetch. Use findBy, which polls until found or 1 s.
  2. Sleeping: await new Promise(r => setTimeout(r, 500)) passes on a fast machine and fails on CI. Replace with a findBy or waitFor on the condition.
  3. Shared state between tests: a module-level QueryClient or store; data from one test visible in the next. Per-test instances in the wrapper.
  4. Real timers with debounces: a 300 ms debounce makes the test wait 300 ms, or race. Fake timers advanced explicitly.
  5. Unawaited userEvent: every userEvent call returns a promise; a missing await lets the assertion run before the event sequence finishes.
  6. act warnings: "an update was not wrapped in act": a state update happened outside RTL's awareness, usually from a resolved promise after the test ended, or a timer. Await the thing that resolves; clean up timers; or wrap the external trigger in act.
  7. Order dependence: tests that pass alone and fail together: shared mocks not reset, handlers not reset, a leaked subscription. afterEach resets; MSW resetHandlers; RTL auto-cleanup.
  8. Layout-dependent behaviour in jsdom: virtualised lists render nothing because the container has zero height. Stub the measurements, or test in a browser runner.
Suspense, transitions, and concurrent features
  1. Suspense fallbacks: findBy the content; optionally assert the fallback first with getBy if the data is held (an MSW handler with a delay or a deferred promise).
  2. Transitions: assert the pending affordance (a dimmed list, an aria-busy) then the final state; findBy covers the lane timing.
  3. Strict Mode in tests: enable it in the wrapper so double-invoked effects are caught in CI, not in production.
  4. Server Components: as of React 19 the component-testing story is the client tree with the RSC payload as fixture data, or end-to-end; unit-test server components as async functions returning elements, and their data functions directly.
63

Visual regression, Storybook, and the tools

stories as the unit of visual testing
  1. A story per state: default, loading, error, empty, long content, RTL, dark theme, each breakpoint that matters. Stories document the component, drive visual tests, and serve as the manual QA surface.
  2. Snapshots: Chromatic (cloud, per-story, with review), Playwright toHaveScreenshot (local or CI, deterministic fonts and viewport required), Storybook's test runner with jest-image-snapshot. Thresholds for antialiasing; baselines committed or stored by the service.
  3. Interaction tests in stories (play functions with Testing Library) let one story be both the visual fixture and a behaviour test.
  4. Accessibility scans per story (axe via the a11y addon) catch contrast, missing names, and ARIA misuse in every state, which component tests miss unless they query the exact element.
NeedToolNotes
Component and hook testsVitest (or Jest) + Testing Library + user-event + jest-domjsdom or happy-dom for DOM; Vitest browser mode when layout matters
Network mockingMSWOne set of handlers for tests, Storybook, and local dev
End-to-endPlaywrightTrace viewer for failures; component testing mode for layout-dependent components
Visual regressionChromatic, Playwright screenshots, PercyDeterministic fonts, animations disabled, fixed dates
Accessibilityvitest-axe / jest-axe in component tests; Storybook a11y addon; Playwright axeRole-based queries are the first line; axe the second
Coveragev8 or istanbul via the runnerCoverage of user flows matters more than line coverage; do not chase the number
Performance regressionReact Profiler API in tests (onRender callback asserting render counts); Lighthouse CI for pages"This interaction renders at most N components" is a testable claim
Type-level teststsc; expectTypeOf (Vitest); tsdFor polymorphic component types and public library APIs
the pointer
Part 13 is React on the wire: hydration and its mismatches, islands and partial hydration in other frameworks, and the hybrid and WebView cases. The tests above run in jsdom or a browser against a client tree; the next part is what happens when the first render was somewhere else.