Part 3 · 2 chapters · ~20 min

Observability

Errors with the context that makes them bugs, metrics with attribution and dimensions read as percentiles, traces that cross from the client into the service that was slow, the ids that join the three, and the costs in bytes, requests and events with sampling per signal and the privacy constraints on all of it.

11

Three signals, shared ids

code
// telemetry.ts: three signals, shared ids, one beacon, sampled by session, configured remotely
const session = { id: crypto.randomUUID(), release: __RELEASE__, sampled: hash(sessionId) < config.metricsRate }   // 10% of sessions send metrics
const q: unknown[] = []
const emit = (e: object) => { q.push({ ...e, session: session.id, release: session.release, route: currentRoute(), t: Date.now() }) }

// errors: always captured; deduped by fingerprint at the backend
window.addEventListener('error', e => emit({ kind: 'error', msg: e.message, stack: e.error?.stack, crumbs: crumbs.slice(-20) }))
window.addEventListener('unhandledrejection', e => emit({ kind: 'error', msg: String(e.reason), stack: e.reason?.stack, crumbs: crumbs.slice(-20) }))
// the React error boundary calls emit({ kind: 'error', componentStack }) too

// metrics: web-vitals with attribution, only for sampled sessions; custom marks for your own phases
if (session.sampled) {
  onLCP(m => emit({ kind: 'metric', name: 'LCP', value: m.value, el: m.attribution.element, nav: m.navigationType }), { reportAllChanges: false })
  onINP(m => emit({ kind: 'metric', name: 'INP', value: m.value, target: m.attribution.interactionTarget, type: m.attribution.interactionType }))
  onCLS(m => emit({ kind: 'metric', name: 'CLS', value: m.value, src: m.attribution.largestShiftTarget }))
  performance.getEntriesByType('measure').forEach(e => emit({ kind: 'metric', name: e.name, value: e.duration }))   // route→painted, search→results
}

// traces: one id per page load, propagated on every request; the API joins the rest
export const traceparent = () => `00-${session.traceId}-${spanId()}-${session.traceSampled ? '01' : '00'}`   // W3C; traceSampled: 1% head
api.use(req => req.headers.set('traceparent', traceparent()))

// one beacon on hide (survives the page closing); keepalive fetch as fallback; flush at most every 10 s while visible
document.addEventListener('visibilitychange', () => { if (document.hidden && q.length) { navigator.sendBeacon('/telemetry', JSON.stringify(q.splice(0))) } })
errors
  1. Capture at three edges: window.onerror, unhandledrejection, and the React error boundary. The SDK batches and sends; stacks are symbolicated server-side from the hidden maps (part 1).
  2. Context makes it a bug: release (did this deploy start it?), route (one page or all?), browser and device (one engine or all?), breadcrumbs (the last twenty clicks, navigations, requests and console lines: the reproduction steps nobody wrote), and the fingerprint count by users (one user retrying versus ten thousand users once).
  3. The dashboard's first two views: new errors by release; the top ten by users affected.
metrics
  1. web-vitals with attribution: LCP with its element, INP with its target and type, CLS with its largest shift source, TTFB, FCP, and the navigation type (a bfcache restore is a different population). Custom phases with performance.mark and measure: route change to content painted, search typed to results shown, each checkout step.
  2. Dimensions on every metric: route, release, device class (from deviceMemory and cores: coarse but cuttable), connection type, country, logged-in state.
  3. Percentiles by dimension, side by side across releases: LCP p75 by route and release finds the deploy that regressed one page; INP p95 by device class finds the phone tier that suffers. Averages hide both; one global number hides which thing to fix.
traces
  1. One trace id per page load, a span per request, carried in the W3C traceparent header; the API logs and propagates it; the backend joins. A slow page in the field becomes client span (3.2 s) containing API span (2.1 s) containing DB span (1.9 s): the frontend engineer points at the query without a meeting. OpenTelemetry's browser SDK does the headers and spans.
  2. The join: the same session id and release on all three signals, so an error opens to its trace, a slow metric to its trace, a trace to its errors. One tool or three, the ids make it one system.
THE THREE SIGNALS FROM A CLIENT
errors with context, metrics with attribution, traces that cross the API
swipe the figure sideways, or tap expand for full screen
1/6
errors
Errors: window.onerror and unhandledrejection catch what escapes; an error boundary catches what React throws in render; the SDK (Sentry, Bugsnag, a home-grown one) batches and sends. The payload: message, stack (symbolicated server-side from the hidden source maps), release and environment, route, user id (or a session id), breadcrumbs (the last twenty clicks, navigations, requests, console messages), browser, OS, device, viewport.
12

What it costs, sampling, and privacy

three costs
  1. Bytes: a 40 KB SDK in a 180 KB entry is 22% of the budget. Inline a few hundred bytes that queue early errors; load the full SDK after load and flush the queue. web-vitals is 2 KB; keep it early.
  2. Requests: five metrics is five requests per view unless batched: one sendBeacon on visibilitychange to hidden (survives the page closing; keepalive fetch as fallback). CLS and INP finalise late and must report on hide, not on load (the Browser course part 12).
  3. Events: a million views a day times five metrics times twenty dimensions, plus half a million stacks: backends price per event or per gigabyte, and the sampling policy is the bill.
sampling, per signal
  1. Errors: 100%, deduplicated by fingerprint (a storm is one row with a count), with a per-session cap against a loop that throws every frame.
  2. Metrics: by session, not by event: 10% of sessions (a hash of the session id under a threshold) send everything; percentiles from a random 10% are unbiased, while per-event sampling that drops slow events is not. The rate is remote config: raise it during an incident, lower it after, without a deploy.
  3. Traces: 1% head-sampled at the client (the sampled flag in traceparent, respected by every service) for a flat baseline, plus 100% of slow or errored traces kept by tail sampling at the backend.
privacy
  1. Telemetry is data about people: mask input values in breadcrumbs (the SDKs do by default), never log full URLs with tokens, truncate IPs at the backend, gate by consent where the law requires (metrics without identifiers are usually fine; session replay is not), and set a retention policy. The Trust course owns the rules; the architecture leaves room for them.
the test of an observability setup
Pick last week's worst field report. Can you open it to the error, the error to the trace, the trace to the query, and the metric to the release that caused it, without asking anyone? If any link is missing, that is the next thing to build.
WHAT IT COSTS, AND SAMPLING
bytes on the client, events on the wire, dollars at the backend, and the rates that keep it honest
swipe the figure sideways, or tap expand for full screen
1/6
bytes
Bytes: the SDK is in the critical path if it loads early (it must, to catch early errors). 40 KB of SDK on a 180 KB entry is 22%. Load the error capture inline and tiny (a few hundred bytes that queue errors), and the full SDK after load; the queued errors flush when it arrives. web-vitals is 2 KB; keep it.