M4: Visualisations and Dashboards
An operations dashboard in four rounds: SVG charts against the pixel budget and canvas for the scatter, reduction of millions of samples to the thousand pixels that can show them, live updates coalesced into frames with per-series subscriptions, and the forty-panel dashboard as a query plan with a shared clock, lazy panels and declared budgets.
The brief and the questions
"An operations dashboard: request rates, error rates, latency percentiles, a map of traffic by region, and a scatter of slow requests. Live, so on-call can watch an incident. Teams want to build their own dashboards from the same panels. It should feel instant."
- How many series, at what rate? ~2,000 series in the store; a panel shows 1 to 20; samples every 1 s; the live stream is 50 updates a second across a typical dashboard.
- How far back, at what resolution? Default 1 hour (3,600 samples per series); up to 30 days (2.6M samples per series). A chart is ~1,000 px wide.
- How many panels per dashboard; how many dashboards? 8 typical, 40 at the extreme; 300 dashboards after a year of team building.
- Marks on screen? A line of 1k points per series; the scatter is up to 200k slow requests over a day; the map is 5k regions.
- Devices? Laptops; a wall screen (a 4K TV running Chrome for days); occasionally a phone during an incident.
- What is "instant"? Load under 2 s; a time-range change under 200 ms to first panel; live updates smooth (no dropped frames visible).
- Interactions? Hover tooltips, zoom and pan on time, click to drill down, shared time across panels, variables (region, service).
- Backend? A time-series store (Prometheus-like) with range queries at a resolution; a WebSocket for live samples.
- FR: line, bar, percentile, map and scatter panels; live mode; time range and refresh; variables; hover, zoom, drill-down; user-built dashboards with saved layouts.
- NFR: load under 2 s at 8 panels (under 4 s at 40); time change under 200 ms; 60 fps during live updates at 50 per second; 30-day ranges without more than ~4k points per series on the client; a wall screen running for a week without memory growth; the scatter at 200k points interactive.
v1: SVG charts, raw series, and the pixel budget
// v1: a React + D3 line chart; the API returns the raw series; every point is an SVG element or a path segment
function Line({ series }: { series: Sample[] }) { // 3,600 samples for an hour: fine
const x = scaleTime().domain(extent(series, d => d.t)).range([0, W]), y = scaleLinear().domain([0, max(series, d => d.v)]).range([H, 0])
const d = line<Sample>().x(s => x(s.t)).y(s => y(s.v))(series)
return <svg width={W} height={H}><path d={d} fill="none" stroke="currentColor" />{series.map(s => <circle key={s.t} cx={x(s.t)} cy={y(s.v)} r={2} />)}</svg>
}
// 1 hour: 3,600 circles: ok. 30 days: 2.6M samples requested (40 MB), 2.6M circles attempted: the tab dies before the first frame- SVG for small charts is correct and should stay: tooltips, CSS hover, crisp zoom, screen-reader text, all from the DOM. Eight panels of 1k-point lines render in a few ms.
- The break is two numbers at once: the 30-day range sends 2.6M samples (bytes) and SVG tries to make 2.6M nodes (pixels). The scatter at 200k points breaks SVG on its own (~20k nodes is the practical ceiling: an update costs ~120 ms).
- Under ~2k marks: SVG. 2k to ~100k: Canvas 2D, batched into one path, with a quadtree or picking canvas for hit testing and an offscreen summary for accessibility. Over 100k, or animating 100k: WebGL through regl, deck.gl or three.js.
- A dashboard mixes renderers: SVG for the lines, canvas for the scatter, WebGL for the map. The choice is per panel, by mark count after reduction.
// v2: a canvas scatter in React: React owns the element and the data; a draw function owns the pixels; a quadtree owns hit testing
function Scatter({ points, width, height, onHover }: Props) {
const ref = useRef<HTMLCanvasElement>(null)
const tree = useMemo(() => quadtree(points, p => p.x, p => p.y), [points]) // d3-quadtree: O(n) build, O(log n) find
useEffect(() => {
const c = ref.current!; const dpr = devicePixelRatio
c.width = width * dpr; c.height = height * dpr // backing store at DPR; CSS size stays width×height
const ctx = c.getContext('2d')!; ctx.scale(dpr, dpr); ctx.clearRect(0, 0, width, height)
ctx.fillStyle = 'rgba(190,24,93,.6)'; ctx.beginPath()
for (const p of points) { ctx.moveTo(p.x + 2, p.y); ctx.arc(p.x, p.y, 2, 0, Math.PI * 2) } // one path for all marks
ctx.fill() // one fill: 20k marks in ~1 ms
}, [points, width, height])
return <canvas ref={ref} style={{ width, height }}
onMouseMove={e => { const r = e.currentTarget.getBoundingClientRect(); onHover(tree.find(e.clientX - r.left, e.clientY - r.top, 6)) }}
role="img" aria-label={`Scatter of ${points.length} points`} /> // plus an offscreen table or summary for AT
}
// hit testing alternatives: a picking canvas (each mark drawn in a unique colour; read the pixel under the cursor: O(1), exact shapes);
// the quadtree is simpler and fine for points. for the renderer decision: 20k here; past ~100k, regl/deck.gl- v1.5 (canvas for the scatter) buys 200k interactive points and pays: hand-written hit testing (complexity), an accessibility fallback (complexity), the loss of CSS and DOM events on marks (capability). The bytes problem remains: the next round.
Round two: reduce before you render
The 30-day range asks for 2.6M samples per series to draw a 1,000 px line. The number that broke is samples against pixels; a chart 1,000 px wide can show 1,000 x positions, and four points per column (min, max, first, last) reproduce the line exactly. v2 reduces at the source.
- Server-side aggregation at a resolution:
GET /series?from&to&buckets=1000returns per-bucket min/max/avg/count from the store's precomputed aggregates (continuous aggregates, recording rules). 32 KB per series at any range; the client never holds a raw series. - Downsampling where raw is needed (a worker over a full series once, or the server): M4 (min/max/first/last per column) or LTTB (one shape-preserving point per bucket). 1M to 4k points, pixel-identical.
- Scatters bin: 200k slow requests into a 200 × 100 histogram drawn as a heatmap; individual points appear when a zoomed region holds under ~2k. The eye reads density past a few thousand marks anyway.
- Levels of detail: pre-aggregated 1 min, 1 h, 1 day; the client picks the level whose bucket count for the viewport is nearest 1,000; zoom crosses levels and refetches; a cache keyed by (series, level, range) makes panning hit neighbours.
- The rule: never more points than pixels × 4; aggregate at the source; bin scatters; pick levels by viewport. Choose the renderer after reduction, on the reduced count (often SVG is then enough).
- v2 buys 30-day ranges at 32 KB per series and a scatter that stays interactive; pays: aggregation in the store (money, and a backend dependency), a resolution parameter every query must carry (complexity), level boundaries that visibly re-render on zoom (consistency; crossfade or hold the coarser level), and the loss of exact values at coarse levels (avg and min/max, not every sample: capability; drill down for the raw).
Round three: live mode at fifty updates a second
// v3: live updates become frames. a ring per series, one rAF, per-series subscriptions
const buffers = new Map<string, Ring>() // latest N samples per series
const subs = new Map<string, Set<() => void>>()
let scheduled = false
socket.onmessage = e => {
for (const { series, t, v } of decode(e.data)) buffers.get(series)?.push(t, v) // no render here
if (!scheduled) { scheduled = true; requestAnimationFrame(flush) }
}
function flush() {
scheduled = false
const touched = new Set<string>()
for (const [s, ring] of buffers) if (ring.dirty) { ring.commit(); touched.add(s) } // swap into the readable snapshot
for (const s of touched) subs.get(s)?.forEach(fn => fn()) // only charts on touched series re-render
}
document.addEventListener('visibilitychange', () => { if (document.hidden) buffers.forEach(r => r.keepLatestOnly()) }) // no rAF while hidden; do not hoard
// a chart: const snap = useSyncExternalStore(cb => subscribe(series, cb), () => buffers.get(series)!.snapshot)
// streaming canvas line: ctx.drawImage(canvas, -1, 0); draw the new column at the right edge; axes on a second canvas layer that redraws on scale change only- Live mode renders on every WebSocket message: 50 commits a second across twelve charts, ~600 ms of work per second. Frames drop; the wall screen's fan runs; on-call sees a stuttering incident. The number is messages per second × render cost; the unit must become the frame.
- Coalesce into a ring per series; one
requestAnimationFramedrains all rings and commits a snapshot; at most 60 applies a second, fewer when nothing arrived. A late frame applies everything pending: no loss, no queue growth. - Per-series subscriptions: each chart subscribes to its series through
useSyncExternalStore; a frame that touched three series re-renders three charts. One dashboard state tree would re-render all twelve. - Partial canvas redraws: a streaming line shifts by one column:
drawImagethe canvas onto itself shifted, draw the new column; axes on a separate layer that redraws only on a scale change. A column of work per frame. - Backpressure on the render side: rings keep the latest N; if applies fall behind, intermediates are dropped (a chart is not an audit log). A hidden tab has no rAF: keep only the latest value per series and apply on visibility; never hoard an hour of messages. The protocol side (acks, resubscribe, ordering) is M8.
- v3 buys 60 fps at 50 updates a second and a wall screen that runs for a week; pays: a scheduler and ring buffers (complexity), per-series stores instead of one tree (complexity), canvas-specific redraw code (complexity), and "frames, not messages": a chart shows the latest frame's truth, which is right everywhere except a trade tape (M9).
Round four: forty panels and three hundred dashboards
Teams built their own dashboards: forty panels, duplicated queries, panels below the fold refreshing every five seconds, each panel owning its own time range. Load is 40 requests and 1.3 MB; a time change is a second of jank; the backend sees 40 × users × refresh queries. The numbers that broke are panels × queries × refresh, and the dashboard became an application.
- A query layer: panels declare queries; the layer deduplicates identical ones (one request, many subscribers), caches by (query, range, level), and batches ten queries per HTTP call where the backend allows. Forty panels become ~15 requests. It is a query cache (the React course part 8) with a domain key and a batcher.
- Shared time and variables as a store: range, refresh interval, "now" anchor, region, service. Every panel's query key derives from it; a change updates keys; results arrive and panels render. Panels never own time. The URL carries the store so a shared link reproduces the view.
- Lazy panels: IntersectionObserver with a 200 px margin; off-screen panels keep their last result and skip refreshes;
content-visibility: autoskips their layout and paint. Eight visible of forty do eight queries per tick. - Per-panel budgets, declared in the editor: max points (feeds reduction), refresh interval (not everything needs 5 s), renderer (by mark count). A panel over budget shows a warning in edit mode instead of slowing everyone.
- Drill-down and layout: click a point → navigate with variables in the URL to a filtered dashboard. The saved layout model is versioned: 300 dashboards outlive the code that made them; migrations are part of the system.
- v4 buys forty-panel dashboards loading under two seconds and changing time under 200 ms, with a backend cost that follows distinct queries rather than panels; pays: a query layer (complexity), a global time store (complexity), lazy discipline (complexity), budgets that authors must set (friction), and a versioned layout model with migrations (complexity).
- Stop: every number in the brief is met. Alerting, annotations and collaborative editing of dashboards are features; the collaborative part is M5.
The whole board, and the exercise
| Round | The number | The break | The design | Paid in |
|---|---|---|---|---|
| v1 | 8 panels, 1-hour ranges | 200k-point scatter in SVG; 30-day ranges as raw samples | SVG for small charts; canvas with a quadtree for the scatter; WebGL for the map | Hand hit-testing; an a11y fallback; no DOM on marks |
| v2 | 2.6M samples for 1,000 px | Samples against pixels | Aggregate at a resolution; M4/LTTB; bin scatters; levels of detail with a cache | Store aggregation; a resolution parameter; level-boundary re-renders; averages not samples |
| v3 | 50 updates/s × 12 charts | A render per message | Rings + one rAF; per-series subscriptions; partial canvas redraws; render-side backpressure | A scheduler; per-series stores; canvas code; frames not messages |
| v4 | 40 panels; 300 dashboards | Panels × queries × refresh; panel-owned time | Query layer (dedupe, cache, batch); time store; lazy panels; declared budgets; versioned layouts | A query layer; a global store; discipline; friction; migrations |
- Count marks, multiply by rate, pick the renderer; SVG, canvas and WebGL each have a ceiling in marks per frame.
- Reduce to the pixels at the source; the chart cannot show what it was sent, so do not send it.
- Frames are the unit of live rendering; messages are inputs to frames, and a ring plus one rAF turns one into the other.
- A dashboard is a query plan with a shared clock; the panel is the leaf, and the system is dedupe, cache, laziness and budgets.