Part 5 · 2 chapters · ~15 min

Client Observability

Observing from an unknown device in a tab that can close mid-report: the collect, send, sample, ingest and store pipeline with one envelope on every event; then errors from capture to symbolication, fingerprinting, replay, the backend trace join, ownership and the crash-free rate a release gate can use.

11

The client observability pipeline

The client is a hostile place to observe from: an unknown device, a tab that closes mid-report, and a network that is often the thing failing. The pipeline is built for that.

collect, send, sample, ingest, store
  1. Collect vitals with attribution, errors from every source, Resource Timing and Server-Timing for the calls you care about, long animation frames, and custom events, all under one envelope: release, route template (never the raw URL), tier, connection, session id, hashed user.
  2. Send batched, flushed on visibilitychange to hidden with sendBeacon or keepalive fetch; final INP and CLS are only known then.
  3. Sample by session with the rate in the envelope; errors and vitals near 100%, traces and replays 1 to 10%.
  4. Ingest: symbolicate with private source maps, fingerprint, drop bots, enforce a PII allowlist, join traceparent.
  5. Store and ask: histograms (p75 from the distribution, never averaged), issues per release, event rows for funnels.
  6. Its own budget: the SDK costs bytes and main-thread time; load it after first paint and measure it like any dependency.
code
// rum.ts: the minimum viable client pipeline
import { onLCP, onINP, onCLS } from 'web-vitals/attribution';

const env = {
  release: __RELEASE__, route: currentRouteTemplate(), tier: deviceTier(),
  conn: (navigator as any).connection?.effectiveType ?? 'unknown',
  session: sessionId(), rate: SAMPLE_RATE,
};
const queue: object[] = [];
const push = (type: string, data: object) => queue.push({ type, ...env, ...data, t: performance.now() });

onLCP(m => push('lcp', { v: m.value, el: m.attribution.target, phase: m.attribution }));
onINP(m => push('inp', { v: m.value, target: m.attribution.interactionTarget,
  input: m.attribution.inputDelay, proc: m.attribution.processingDuration,
  pres: m.attribution.presentationDelay }));
onCLS(m => push('cls', { v: m.value, src: m.attribution.largestShiftTarget }));

addEventListener('error', e => push('error', { msg: e.message, stack: e.error?.stack }), true);
addEventListener('unhandledrejection', e => push('error', { msg: String(e.reason), stack: e.reason?.stack }));

addEventListener('visibilitychange', () => {
  if (document.visibilityState !== 'hidden' || !queue.length) return;
  navigator.sendBeacon('/rum', JSON.stringify(queue.splice(0)));
});
the trap
Logging the full URL as a dimension. Query strings carry tokens, emails and search terms (PII), and every distinct URL is a new series (cardinality). Send the route template, and allowlist the parameters you need.
CLIENT OBSERVABILITY: THE PIPELINE
what the client sends, how it gets out of a closing tab, and what the backend does with it before anyone can ask a question
swipe the figure sideways, or tap expand for full screen
1/6
collect
Collect: web-vitals with attribution (LCP element and phase, INP target and phase, CLS source), errors (window.onerror, unhandledrejection, React error boundaries, failed resources), network (Resource Timing for the API calls you care about, status and duration; the Server-Timing header to join the backend), long tasks and long animation frames, and custom events (step reached, flag variant). One envelope on everything: release, route template (not the URL: cardinality and PII), device tier, connection, session id, hashed user id.
12

Errors, replay, crash-free and the trace

from a minified stack to an owned issue
  1. Capture from window error, unhandledrejection, error boundaries and capture-phase resource errors; load cross-origin scripts with crossorigin and CORS or get only "Script error."
  2. Symbolicate and fingerprint by release and debug id; tune the fingerprint so one bug is one issue.
  3. Context: breadcrumbs, the envelope, and a masked, sampled replay for sessions with an error.
  4. Join the backend with traceparent so the client error links to the server span.
  5. Route by the top frame through CODEOWNERS; page on new-in-release above a threshold.
  6. Crash-free sessions per release and route, beside the vitals in the ring gate.
signalcostprivacy risksampleanswers
vitalstinylow~100%is it slower, where, for whom
errorssmallmedium (messages can carry data)100%, deduplicatedwhat broke, how often, since when
breadcrumbssmallmediumwith errorswhat led to it
tracesmediumlow1 to 10%which hop was slow
replaylargehigh (mask by default)1 to 5% + error sessionswhat the user saw
the exercise
Take your top client error this week. Can you name the release it started in, the team that owns it, the server span it touched and the share of sessions it hit, each in under a minute? Each "no" is a missing link in the chain above.
ERRORS, REPLAY, CRASH-FREE AND THE TRACE
turning a stack trace from a phone you will never see into an issue with an owner, a cause and a count
swipe the figure sideways, or tap expand for full screen
1/6
capture
Capture from every source: window error (sync throws, with the Error when the script is same-origin or loaded with crossorigin and CORS headers; otherwise "Script error." with nothing: the Browser course's CORS), unhandledrejection (async throws), React error boundaries (render throws; they do not catch handlers or async), resource errors (a capture-phase error listener on window for failed scripts, images, CSS), and console.error as a breadcrumb. Filter extension frames (chrome-extension://, moz-extension://) at the SDK.