Part 4 · 6 chapters · ~45 min

M4: Visualisations and Dashboards

An operations dashboard in four rounds: SVG charts against the pixel budget and canvas for the scatter, reduction of millions of samples to the thousand pixels that can show them, live updates coalesced into frames with per-series subscriptions, and the forty-panel dashboard as a query plan with a shared clock, lazy panels and declared budgets.

26

The brief and the questions

the brief

"An operations dashboard: request rates, error rates, latency percentiles, a map of traffic by region, and a scatter of slow requests. Live, so on-call can watch an incident. Teams want to build their own dashboards from the same panels. It should feel instant."

the questions, and the answers
  1. How many series, at what rate? ~2,000 series in the store; a panel shows 1 to 20; samples every 1 s; the live stream is 50 updates a second across a typical dashboard.
  2. How far back, at what resolution? Default 1 hour (3,600 samples per series); up to 30 days (2.6M samples per series). A chart is ~1,000 px wide.
  3. How many panels per dashboard; how many dashboards? 8 typical, 40 at the extreme; 300 dashboards after a year of team building.
  4. Marks on screen? A line of 1k points per series; the scatter is up to 200k slow requests over a day; the map is 5k regions.
  5. Devices? Laptops; a wall screen (a 4K TV running Chrome for days); occasionally a phone during an incident.
  6. What is "instant"? Load under 2 s; a time-range change under 200 ms to first panel; live updates smooth (no dropped frames visible).
  7. Interactions? Hover tooltips, zoom and pan on time, click to drill down, shared time across panels, variables (region, service).
  8. Backend? A time-series store (Prometheus-like) with range queries at a resolution; a WebSocket for live samples.
the requirements, with numbers
  1. FR: line, bar, percentile, map and scatter panels; live mode; time range and refresh; variables; hover, zoom, drill-down; user-built dashboards with saved layouts.
  2. NFR: load under 2 s at 8 panels (under 4 s at 40); time change under 200 ms; 60 fps during live updates at 50 per second; 30-day ranges without more than ~4k points per series on the client; a wall screen running for a week without memory growth; the scatter at 200k points interactive.
the two budgets
Sixty frames a second is a pixel budget: marks on screen × updates per second, against what SVG, canvas or the GPU can draw in 16 ms. The data is a bytes budget: samples × series against the pixels that can show them. Both are computed from the panel layout, and neither depends on how many samples the store holds.
27

v1: SVG charts, raw series, and the pixel budget

code
// v1: a React + D3 line chart; the API returns the raw series; every point is an SVG element or a path segment
function Line({ series }: { series: Sample[] }) {              // 3,600 samples for an hour: fine
  const x = scaleTime().domain(extent(series, d => d.t)).range([0, W]), y = scaleLinear().domain([0, max(series, d => d.v)]).range([H, 0])
  const d = line<Sample>().x(s => x(s.t)).y(s => y(s.v))(series)
  return <svg width={W} height={H}><path d={d} fill="none" stroke="currentColor" />{series.map(s => <circle key={s.t} cx={x(s.t)} cy={y(s.v)} r={2} />)}</svg>
}
// 1 hour: 3,600 circles: ok. 30 days: 2.6M samples requested (40 MB), 2.6M circles attempted: the tab dies before the first frame
what v1 is right about
  1. SVG for small charts is correct and should stay: tooltips, CSS hover, crisp zoom, screen-reader text, all from the DOM. Eight panels of 1k-point lines render in a few ms.
  2. The break is two numbers at once: the 30-day range sends 2.6M samples (bytes) and SVG tries to make 2.6M nodes (pixels). The scatter at 200k points breaks SVG on its own (~20k nodes is the practical ceiling: an update costs ~120 ms).
the renderer rule
  1. Under ~2k marks: SVG. 2k to ~100k: Canvas 2D, batched into one path, with a quadtree or picking canvas for hit testing and an offscreen summary for accessibility. Over 100k, or animating 100k: WebGL through regl, deck.gl or three.js.
  2. A dashboard mixes renderers: SVG for the lines, canvas for the scatter, WebGL for the map. The choice is per panel, by mark count after reduction.
code
// v2: a canvas scatter in React: React owns the element and the data; a draw function owns the pixels; a quadtree owns hit testing
function Scatter({ points, width, height, onHover }: Props) {
  const ref = useRef<HTMLCanvasElement>(null)
  const tree = useMemo(() => quadtree(points, p => p.x, p => p.y), [points])          // d3-quadtree: O(n) build, O(log n) find
  useEffect(() => {
    const c = ref.current!; const dpr = devicePixelRatio
    c.width = width * dpr; c.height = height * dpr                                      // backing store at DPR; CSS size stays width×height
    const ctx = c.getContext('2d')!; ctx.scale(dpr, dpr); ctx.clearRect(0, 0, width, height)
    ctx.fillStyle = 'rgba(190,24,93,.6)'; ctx.beginPath()
    for (const p of points) { ctx.moveTo(p.x + 2, p.y); ctx.arc(p.x, p.y, 2, 0, Math.PI * 2) }   // one path for all marks
    ctx.fill()                                                                           // one fill: 20k marks in ~1 ms
  }, [points, width, height])
  return <canvas ref={ref} style={{ width, height }}
    onMouseMove={e => { const r = e.currentTarget.getBoundingClientRect(); onHover(tree.find(e.clientX - r.left, e.clientY - r.top, 6)) }}
    role="img" aria-label={`Scatter of ${points.length} points`} />                      // plus an offscreen table or summary for AT
}
// hit testing alternatives: a picking canvas (each mark drawn in a unique colour; read the pixel under the cursor: O(1), exact shapes);
// the quadtree is simpler and fine for points. for the renderer decision: 20k here; past ~100k, regl/deck.gl
the sentence
  1. v1.5 (canvas for the scatter) buys 200k interactive points and pays: hand-written hit testing (complexity), an accessibility fallback (complexity), the loss of CSS and DOM events on marks (capability). The bytes problem remains: the next round.
SIXTY FRAMES IS A PIXEL BUDGET
SVG, Canvas 2D and WebGL against the number of marks on screen
swipe the figure sideways, or tap expand for full screen
1/6
1k SVG
1,000 points as SVG circles: 1,000 DOM nodes; an update sets 1,000 cx/cy attributes; style recalc + layout + paint ~6 ms. Fine. 60 fps with room. Tooltips, CSS hover, accessibility, text, all free because it is the DOM.
28

Round two: reduce before you render

The 30-day range asks for 2.6M samples per series to draw a 1,000 px line. The number that broke is samples against pixels; a chart 1,000 px wide can show 1,000 x positions, and four points per column (min, max, first, last) reproduce the line exactly. v2 reduces at the source.

v2
  1. Server-side aggregation at a resolution: GET /series?from&to&buckets=1000 returns per-bucket min/max/avg/count from the store's precomputed aggregates (continuous aggregates, recording rules). 32 KB per series at any range; the client never holds a raw series.
  2. Downsampling where raw is needed (a worker over a full series once, or the server): M4 (min/max/first/last per column) or LTTB (one shape-preserving point per bucket). 1M to 4k points, pixel-identical.
  3. Scatters bin: 200k slow requests into a 200 × 100 histogram drawn as a heatmap; individual points appear when a zoomed region holds under ~2k. The eye reads density past a few thousand marks anyway.
  4. Levels of detail: pre-aggregated 1 min, 1 h, 1 day; the client picks the level whose bucket count for the viewport is nearest 1,000; zoom crosses levels and refetches; a cache keyed by (series, level, range) makes panning hit neighbours.
  5. The rule: never more points than pixels × 4; aggregate at the source; bin scatters; pick levels by viewport. Choose the renderer after reduction, on the reduced count (often SVG is then enough).
the sentence
  1. v2 buys 30-day ranges at 32 KB per series and a scatter that stays interactive; pays: aggregation in the store (money, and a backend dependency), a resolution parameter every query must carry (complexity), level boundaries that visibly re-render on zoom (consistency; crossfade or hold the coarser level), and the loss of exact values at coarse levels (avg and min/max, not every sample: capability; drill down for the raw).
REDUCE BEFORE YOU RENDER
a thousand pixels cannot show a million points, so do not send them
swipe the figure sideways, or tap expand for full screen
1/6
1M into 1k px
1M samples over 30 days, a 1,000 px chart: 1,000 samples per pixel column. Drawn naively, each column is a vertical smear of 1,000 overlapping segments; the shape is right but the cost is 1M points of bytes (~16 MB), parse, and draw.
29

Round three: live mode at fifty updates a second

code
// v3: live updates become frames. a ring per series, one rAF, per-series subscriptions
const buffers = new Map<string, Ring>()                 // latest N samples per series
const subs = new Map<string, Set<() => void>>()
let scheduled = false
socket.onmessage = e => {
  for (const { series, t, v } of decode(e.data)) buffers.get(series)?.push(t, v)      // no render here
  if (!scheduled) { scheduled = true; requestAnimationFrame(flush) }
}
function flush() {
  scheduled = false
  const touched = new Set<string>()
  for (const [s, ring] of buffers) if (ring.dirty) { ring.commit(); touched.add(s) }   // swap into the readable snapshot
  for (const s of touched) subs.get(s)?.forEach(fn => fn())                           // only charts on touched series re-render
}
document.addEventListener('visibilitychange', () => { if (document.hidden) buffers.forEach(r => r.keepLatestOnly()) })   // no rAF while hidden; do not hoard
// a chart: const snap = useSyncExternalStore(cb => subscribe(series, cb), () => buffers.get(series)!.snapshot)
// streaming canvas line: ctx.drawImage(canvas, -1, 0); draw the new column at the right edge; axes on a second canvas layer that redraws on scale change only
the break
  1. Live mode renders on every WebSocket message: 50 commits a second across twelve charts, ~600 ms of work per second. Frames drop; the wall screen's fan runs; on-call sees a stuttering incident. The number is messages per second × render cost; the unit must become the frame.
v3
  1. Coalesce into a ring per series; one requestAnimationFrame drains all rings and commits a snapshot; at most 60 applies a second, fewer when nothing arrived. A late frame applies everything pending: no loss, no queue growth.
  2. Per-series subscriptions: each chart subscribes to its series through useSyncExternalStore; a frame that touched three series re-renders three charts. One dashboard state tree would re-render all twelve.
  3. Partial canvas redraws: a streaming line shifts by one column: drawImage the canvas onto itself shifted, draw the new column; axes on a separate layer that redraws only on a scale change. A column of work per frame.
  4. Backpressure on the render side: rings keep the latest N; if applies fall behind, intermediates are dropped (a chart is not an audit log). A hidden tab has no rAF: keep only the latest value per series and apply on visibility; never hoard an hour of messages. The protocol side (acks, resubscribe, ordering) is M8.
the sentence
  1. v3 buys 60 fps at 50 updates a second and a wall screen that runs for a week; pays: a scheduler and ring buffers (complexity), per-series stores instead of one tree (complexity), canvas-specific redraw code (complexity), and "frames, not messages": a chart shows the latest frame's truth, which is right everywhere except a trade tape (M9).
LIVE UPDATES AND THE FRAME
coalesce, schedule, redraw only what changed
swipe the figure sideways, or tap expand for full screen
1/6
naive
Naive: a WebSocket message → setState → React render of the dashboard → every chart re-renders → canvas redraws. At 50 messages a second: 50 commits, each ~12 ms with twelve charts: 600 ms of work per second. Frames drop; the browser course part 12's long-task log fills; the fan spins.
30

Round four: forty panels and three hundred dashboards

Teams built their own dashboards: forty panels, duplicated queries, panels below the fold refreshing every five seconds, each panel owning its own time range. Load is 40 requests and 1.3 MB; a time change is a second of jank; the backend sees 40 × users × refresh queries. The numbers that broke are panels × queries × refresh, and the dashboard became an application.

v4
  1. A query layer: panels declare queries; the layer deduplicates identical ones (one request, many subscribers), caches by (query, range, level), and batches ten queries per HTTP call where the backend allows. Forty panels become ~15 requests. It is a query cache (the React course part 8) with a domain key and a batcher.
  2. Shared time and variables as a store: range, refresh interval, "now" anchor, region, service. Every panel's query key derives from it; a change updates keys; results arrive and panels render. Panels never own time. The URL carries the store so a shared link reproduces the view.
  3. Lazy panels: IntersectionObserver with a 200 px margin; off-screen panels keep their last result and skip refreshes; content-visibility: auto skips their layout and paint. Eight visible of forty do eight queries per tick.
  4. Per-panel budgets, declared in the editor: max points (feeds reduction), refresh interval (not everything needs 5 s), renderer (by mark count). A panel over budget shows a warning in edit mode instead of slowing everyone.
  5. Drill-down and layout: click a point → navigate with variables in the URL to a filtered dashboard. The saved layout model is versioned: 300 dashboards outlive the code that made them; migrations are part of the system.
the sentence, and the stop
  1. v4 buys forty-panel dashboards loading under two seconds and changing time under 200 ms, with a backend cost that follows distinct queries rather than panels; pays: a query layer (complexity), a global time store (complexity), lazy discipline (complexity), budgets that authors must set (friction), and a versioned layout model with migrations (complexity).
  2. Stop: every number in the brief is met. Alerting, annotations and collaborative editing of dashboards are features; the collaborative part is M5.
V4: THE DASHBOARD AS A SYSTEM
a query layer, a layout of panels, shared time, and the cost of each panel
swipe the figure sideways, or tap expand for full screen
1/6
forty panels
Forty panels each with a query: 40 requests on load, 40 on every time-range change, 40 on every refresh tick. Several panels ask for the same series with different charts. Panels below the fold query and render anyway. The load is 40 × 32 KB = 1.3 MB and 40 renders; a time change is a second of jank.
31

The whole board, and the exercise

RoundThe numberThe breakThe designPaid in
v18 panels, 1-hour ranges200k-point scatter in SVG; 30-day ranges as raw samplesSVG for small charts; canvas with a quadtree for the scatter; WebGL for the mapHand hit-testing; an a11y fallback; no DOM on marks
v22.6M samples for 1,000 pxSamples against pixelsAggregate at a resolution; M4/LTTB; bin scatters; levels of detail with a cacheStore aggregation; a resolution parameter; level-boundary re-renders; averages not samples
v350 updates/s × 12 chartsA render per messageRings + one rAF; per-series subscriptions; partial canvas redraws; render-side backpressureA scheduler; per-series stores; canvas code; frames not messages
v440 panels; 300 dashboardsPanels × queries × refresh; panel-owned timeQuery layer (dedupe, cache, batch); time store; lazy panels; declared budgets; versioned layoutsA query layer; a global store; discipline; friction; migrations
what the sequence teaches
  1. Count marks, multiply by rate, pick the renderer; SVG, canvas and WebGL each have a ceiling in marks per frame.
  2. Reduce to the pixels at the source; the chart cannot show what it was sent, so do not send it.
  3. Frames are the unit of live rendering; messages are inputs to frames, and a ring plus one rAF turns one into the other.
  4. A dashboard is a query plan with a shared clock; the panel is the leaf, and the system is dedupe, cache, laziness and budgets.
the exercise
Take one chart you own. Count the marks it draws at the largest range users pick, and the points the API returns for it. If points exceed pixels × 4, reduction is missing; if marks exceed 2k in SVG, the renderer is wrong; if it renders on every message, the frame is missing. Fix the one that is furthest over.