Testing React
A test that resembles how the software is used survives refactors and catches regressions; one that knows the implementation does neither. This part is the test levels and what each sees and survives, component tests by role and user event, mocking at the wire with MSW, the causes of flakiness and the fix for each, Suspense and transitions in tests, and visual regression with stories.
What to test, at which level
"The suite is 2,000 unit tests that break on every refactor and 40 end-to-end tests that fail on Tuesdays. What should it be instead?"
Mostly component tests that interact with the rendered UI the way a user does and assert what a user would see, written so they do not know the implementation. Some integration tests per feature with the network mocked at the wire. A few end-to-end tests on the paths that lose money if they break. Visual tests on the design system. Unit tests where logic is dense enough to deserve them. The principle (Testing Library's): the more a test resembles how the software is used, the more confidence it gives, and the more refactors it survives.
// what to test, by component kind // a presentational component (a Badge): a story + visual test; a component test only if it has logic (a variant with a11y implications) // a form: component test of the full flow (fill, validate, submit, errors, pending, success); the schema has its own unit tests // a list with selection, sorting, filtering: component test of the behaviours the user sees; a unit test of the sort/filter functions // a data-driven page: integration test with MSW: loading → content; error → retry; mutation → updated list // a hook with branching logic (useUndo, usePagination): renderHook tests of the transitions // a custom input with a mask: component test of typing behaviour (what appears for what is typed; cursor position if you manage it) // a provider / context: tested through consumers; a direct test only for its reducer // a route guard / redirect: integration test with the router: unauthenticated → login; authenticated → page // accessibility: every component test queries by role, so a missing label fails. add vitest-axe / jest-axe for a scan per rendered state // what not to test: that React renders props (it does); the exact DOM structure; class names; that a mock was called with the right args when the UI result is assertable instead
- Hook unit tests (
renderHook): state transitions and derived values. For reducers and hooks with branching. They know nothing of markup, which is their strength and their limit. - Component tests (
render,screen,userEvent): the user's view. They query by role and accessible name, interact through real event sequences, and assert visible outcomes. They survive swaps of state libraries, form libraries, and markup that preserves semantics. The centre of the suite. - Integration tests: a route with its providers and a mocked API, driving a flow across components: loading, content, error, mutation, cache update. The seams between components are where the bugs that component tests miss live.
- End-to-end: the built app, a real browser, real or seeded backend. Few, critical, run on merge. Flakiness is environmental and is managed by keeping the count small.
- Visual regression: screenshots per story or page, diffed against baselines. For the design system and layout-critical pages. Reviewed like code.
Component tests: queries, events, assertions
// a component test in the Testing Library style: queries by what the user perceives; events as the user does them
import { render, screen } from '@testing-library/react'
import userEvent from '@testing-library/user-event'
test('transfers money after confirmation', async () => {
const user = userEvent.setup()
render(<TransferForm balance={500_00} />, { wrapper: Providers }) // Providers: QueryClient, router, store; the test's own instances
await user.type(screen.getByRole('textbox', { name: /amount/i }), '150.50') // by role + accessible name, not by test id or class
await user.type(screen.getByRole('textbox', { name: /account/i }), '0123456789')
await user.click(screen.getByRole('button', { name: /review/i }))
expect(screen.getByRole('heading', { name: /confirm/i })).toBeInTheDocument()
expect(screen.getByText('₦150.50')).toBeInTheDocument() // the parsed, formatted amount: the thing that must be right
await user.click(screen.getByRole('button', { name: /send/i }))
expect(await screen.findByRole('status')).toHaveTextContent(/sent/i) // findBy: waits for the async result (MSW answered)
expect(screen.getByRole('button', { name: /send/i })).toBeDisabled() // no double submit
})
// what this test does not know: the state library, the form library, whether the amount is controlled, the component tree, class names.
// what it would catch: a float parse (₦150.5), a missing disabled state, a confirmation that shows the typed string, a broken label (the query fails: an a11y bug IS a test failure).
// the queries, in priority: getByRole > getByLabelText > getByPlaceholderText > getByText > getByDisplayValue > getByAltText > getByTitle > getByTestId (last resort)- By role and name first:
getByRole('button', { name: /send/i })finds what a screen reader would announce. It fails if the element has no accessible name, which is an accessibility bug surfacing as a test failure: the right incentive. - Then by label, placeholder, text, display value. Test ids last, for things with no semantic handle (a decorative container you must assert exists).
getBythrows if absent (use for "it is there now");queryByreturns null (use for "it is not there");findBywaits and polls (use for anything after an await).getAllByand friends for lists.within(container)scopes queries to a region (a row, a dialog) when the page has several similar elements.
userEventsimulates the full sequence a device produces:typefires keydown, keypress, input, keyup per character and respects maxlength and disabled;clickfires pointer and mouse events and focus;keyboardfor shortcuts;uploadfor files;selectOptions;tabfor focus order. Alwaysawaited; set up withuserEvent.setup()per test.fireEventdispatches one synthetic event with no sequence; it can pass tests that fail for users (a disabled button "clicked"; an input changed without keystrokes). Reserve it for events userEvent cannot produce (a custom event, a resize).
- On what the user sees: text content, roles present, attributes that matter (
toBeDisabled,toBeChecked,toHaveAccessibleName,toHaveValue,toBeVisible) via jest-dom matchers. - Not on: component state, hook return values, the DOM structure, class names (unless the class is the product, as in a design system), or mock call arguments when the UI result can be asserted instead.
- Negative assertions need
queryByand care: "the error is not shown" after an async operation should wait for the success state first, or it passes trivially before the error would have appeared.
Mocking the right boundary
// the network boundary: Mock Service Worker intercepts fetch at the request level; components and query caches run for real
import { http, HttpResponse } from 'msw'
import { setupServer } from 'msw/node'
export const server = setupServer(
http.get('/api/posts/:id', ({ params }) => HttpResponse.json({ id: params.id, title: 'Hello', likes: 41 })),
http.post('/api/transfers', async ({ request }) => {
const body = await request.json()
if (body.amount > 500_00) return HttpResponse.json({ error: 'insufficient' }, { status: 422 }) // the error path is a handler too
return HttpResponse.json({ id: 'tx_1', amount: body.amount })
}),
)
beforeAll(() => server.listen({ onUnhandledRequest: 'error' })) // an unmocked request fails the test: no accidental real calls
afterEach(() => server.resetHandlers())
afterAll(() => server.close())
// per-test overrides: server.use(http.get('/api/posts/:id', () => HttpResponse.error())) // network failure for this test only
// why not mock fetch / axios / the query hook: you would be testing your mocks. MSW leaves the real code path intact to the wire.
// the same handlers run in the browser (service worker) for Storybook and manual testing: one definition of "the API" for dev and test- The network, at the wire (MSW): intercept requests; return responses. Everything inside the app runs for real: fetch calls, the query cache, retries, invalidation, serialisation. The handlers double as a contract with the API and as a dev server for Storybook.
- Not the data layer: mocking
useQueryor the API module tests the mock, skips the cache behaviour, and breaks when the data layer changes. - Not
fetchitself (vi.fn()on global fetch): fragile to how the code builds requests; MSW matches on URL and method and lets the request code be real. - Time: fake timers for debounces, intervals and timeouts; advance explicitly. Never sleep.
- Randomness and ids: seed or stub
crypto.randomUUIDandMath.randomwhere output must be deterministic (an idempotency key you assert was sent). - Browser APIs the environment lacks: jsdom has no layout (
getBoundingClientRectreturns zeros), noIntersectionObserver, nomatchMedia; stub them per test or use a browser-mode runner (Vitest browser mode, Playwright component testing) for components that depend on layout. - Modules with side effects (analytics, error reporting): mock at the module boundary; assert they were called only when the behaviour is the call itself.
- A
renderwrapper that provides the app's providers with test instances: a freshQueryClient(retries off,gcTime: Infinityto avoid cross-test garbage collection timing), a memory router at a route, a store with initial state. One helper; every test uses it. - Factories for data (a
makePost()with overrides) so tests say what matters and defaults cover the rest. - MSW handlers per domain, composed into the server; per-test overrides with
server.usefor error and edge cases. onUnhandledRequest: 'error'so a request without a handler fails loudly rather than hanging afindByuntil timeout.
Async, timing, and flakiness
// async and timing: the rules that remove flakiness
// 1. await what you query for: findBy* (waits up to 1 s, polling) for anything that appears after an await; getBy* only for what is there now
expect(await screen.findByText(/sent/i)).toBeInTheDocument()
// 2. userEvent, not fireEvent: userEvent.type() fires the sequence a real keyboard does (keydown, keypress, input, keyup per character), goes through
// React's event system, and respects disabled and pointer-events. fireEvent.change() skips all that and passes tests that fail for users
// 3. act() is implicit in RTL's render / userEvent / findBy; explicit act() is for state updates you trigger outside those (a store.dispatch, a timer)
// 4. fake timers: vi.useFakeTimers(); userEvent.setup({ advanceTimers: vi.advanceTimersByTime }); advance for debounces, not sleep
// 5. waitFor(() => expect(…)) for conditions that are not a single element appearing (a count reaching 3; a mock called twice)
// 6. never: setTimeout in a test to "let things settle"; snapshotting the whole DOM (brittle, reviews nothing); asserting on state (test outputs, not internals)
// 7. Suspense and transitions: findBy handles the fallback → content change; for isPending UI, assert the pending state then the final
// 8. cleanup: RTL unmounts after each test; the QueryClient must be per test (a shared cache leaks data between tests); MSW handlers reset per test- Asserting before the async result:
getByright after a click that triggers a fetch. UsefindBy, which polls until found or 1 s. - Sleeping:
await new Promise(r => setTimeout(r, 500))passes on a fast machine and fails on CI. Replace with afindByorwaitForon the condition. - Shared state between tests: a module-level QueryClient or store; data from one test visible in the next. Per-test instances in the wrapper.
- Real timers with debounces: a 300 ms debounce makes the test wait 300 ms, or race. Fake timers advanced explicitly.
- Unawaited userEvent: every userEvent call returns a promise; a missing
awaitlets the assertion run before the event sequence finishes. - act warnings: "an update was not wrapped in act": a state update happened outside RTL's awareness, usually from a resolved promise after the test ended, or a timer. Await the thing that resolves; clean up timers; or wrap the external trigger in
act. - Order dependence: tests that pass alone and fail together: shared mocks not reset, handlers not reset, a leaked subscription.
afterEachresets; MSWresetHandlers; RTL auto-cleanup. - Layout-dependent behaviour in jsdom: virtualised lists render nothing because the container has zero height. Stub the measurements, or test in a browser runner.
- Suspense fallbacks:
findBythe content; optionally assert the fallback first withgetByif the data is held (an MSW handler with a delay or a deferred promise). - Transitions: assert the pending affordance (a dimmed list, an
aria-busy) then the final state;findBycovers the lane timing. - Strict Mode in tests: enable it in the wrapper so double-invoked effects are caught in CI, not in production.
- Server Components: as of React 19 the component-testing story is the client tree with the RSC payload as fixture data, or end-to-end; unit-test server components as async functions returning elements, and their data functions directly.
Visual regression, Storybook, and the tools
- A story per state: default, loading, error, empty, long content, RTL, dark theme, each breakpoint that matters. Stories document the component, drive visual tests, and serve as the manual QA surface.
- Snapshots: Chromatic (cloud, per-story, with review), Playwright
toHaveScreenshot(local or CI, deterministic fonts and viewport required), Storybook's test runner withjest-image-snapshot. Thresholds for antialiasing; baselines committed or stored by the service. - Interaction tests in stories (
playfunctions with Testing Library) let one story be both the visual fixture and a behaviour test. - Accessibility scans per story (axe via the a11y addon) catch contrast, missing names, and ARIA misuse in every state, which component tests miss unless they query the exact element.
| Need | Tool | Notes |
|---|---|---|
| Component and hook tests | Vitest (or Jest) + Testing Library + user-event + jest-dom | jsdom or happy-dom for DOM; Vitest browser mode when layout matters |
| Network mocking | MSW | One set of handlers for tests, Storybook, and local dev |
| End-to-end | Playwright | Trace viewer for failures; component testing mode for layout-dependent components |
| Visual regression | Chromatic, Playwright screenshots, Percy | Deterministic fonts, animations disabled, fixed dates |
| Accessibility | vitest-axe / jest-axe in component tests; Storybook a11y addon; Playwright axe | Role-based queries are the first line; axe the second |
| Coverage | v8 or istanbul via the runner | Coverage of user flows matters more than line coverage; do not chase the number |
| Performance regression | React Profiler API in tests (onRender callback asserting render counts); Lighthouse CI for pages | "This interaction renders at most N components" is a testable claim |
| Type-level tests | tsc; expectTypeOf (Vitest); tsd | For polymorphic component types and public library APIs |