Part 7 · 2 chapters · ~15 min
The Frontend's Share
Client SLOs: what server SLIs miss, journey success including designed failures, interaction latency per tier, crash-free sessions per release, budgets that gate rings, and server and client side by side; then the short list of client conditions that should page a frontend engineer, and the tickets that should not.
11
Client SLOs and the crash-free rate
reliability where users experience it
- What the server misses: delivery, client exceptions, crashes, third parties.
- Journey success: the user reached either the success state or a designed failure state.
- Interaction latency per device tier.
- Crash-free sessions per release: the client's availability SLI.
- Client budgets gate the release rings.
- Show server and client SLOs together. When they diverge, look at delivery or the client.
code
// RUM: journey events that become the client SLI
rum.journey('transfer', 'started', { route: '/transfer' });
try {
const res = await api.createTransfer(form, { idempotencyKey });
rum.journey('transfer', 'succeeded', { status: res.status }); // good
} catch (e) {
if (isDesignedFailure(e)) rum.journey('transfer', 'designed_failure', { code: e.code }); // good: e.g. insufficient funds
else rum.journey('transfer', 'failed', { kind: classify(e) }); // bad
}
// SLI = (succeeded + designed_failure) / started, per release, per device tier
// sessions that end with 'started' and nothing else (tab closed on a blank screen) count as badCLIENT SLOS
measuring reliability where users experience it: in the browser and the app, journey by journey
swipe the figure sideways, or tap expand for full screen
1/6
what the server misses
What the server misses: DNS and CDN failures, a bundle that fails to load on a flaky network, a JavaScript exception after a successful API response, an app crash on low memory, a third-party script blocking the main thread, a release that works on Chrome and breaks on an older Safari. Every one of these is a user-visible failure with a perfect server SLI.
12
What pages a frontend engineer
a short list that passes the paging test
- Page: the app does not load, as seen by multi-region synthetic checks and RUM.
- Page: a journey fails on the client while the server looks fine.
- Page: crash-free sessions collapse after a release.
- Page: a critical third-party script is broken.
- Ticket, not page: budgets, regressions and warnings.
- To make it work: client SLOs, runbooks, shared dashboards, one incident process.
the whole stack, one number
The courses meet here. The Browser and Disciplines courses give the client measurements. Cloud, Infra and Distributed Systems give the server. This course turns both into SLOs for each journey, with a budget, alerts, an incident process and postmortems. A staff engineer can follow a transfer from the tap on Pay to the ledger commit and say how reliable it is, what it costs, and who gets paged when it breaks.
WHAT PAGES A FRONTEND ENGINEER
the short list of client-side conditions that deserve a page, and everything that should be a ticket instead
swipe the figure sideways, or tap expand for full screen
1/6
app not loading
Page: the app does not load. Synthetic probes from several regions (including Lagos and Nairobi) fail to render the shell, or RUM page-load success drops sharply: a CDN misconfiguration, an expired certificate on the asset domain, a deploy that shipped a broken index.html, a CSP that now blocks the main bundle.