Part 3 · 6 chapters · ~45 min

M3: A Media-Intensive Web App

A property-listings product from two thousand listings to a million visitors, in four rounds: originals everywhere and the decode budget that killed the phone, an upload pipeline with client resize, signed direct-to-storage and resumable parts, video as a bitrate ladder with an adaptive player, and delivery through immutable URLs on a CDN with a first-screen order.

20

The brief and the questions

the brief

"Agents upload photos of a property from their phones, we show them in a listing grid and a detail page with a gallery, and buyers browse on their phones. We want video tours next quarter. It must look good and load fast. We have about two thousand listings."

the questions, and the answers
  1. How many photos per listing, at what size? 20 to 40; phone originals, 12 megapixels, 3 MB each. Two years: 10k listings, 300k photos.
  2. Who uploads, on what? 300 agents, phones in the field, often 3G. Who views? A million monthly visitors, 70% on phones, median 4G.
  3. Where are photos shown? A grid (400 px slots), a detail gallery (full width), a lightbox (full screen), thumbnails in search results (120 px). Video: a tour per listing, up to 2 minutes.
  4. What is "fast"? The listing page LCP under 2.5 s on 4G; the grid scrolls at 60 fps; an upload of 40 photos finishes on 3G without the agent babysitting it.
  5. Can we process server-side? Yes; a CDN and object storage are available.
  6. Privacy? EXIF GPS must not leak.
  7. Cost sensitivity? Media delivery is expected to be the biggest infrastructure line; they want to see the number.
the requirements, with numbers
  1. FR: multi-photo upload from phones with progress and resume; a grid, a gallery, a lightbox, thumbnails; video tours with a player; posters; EXIF stripped.
  2. NFR: LCP under 2.5 s on 4G at p75; first-screen bytes under 1 MB; grid at 60 fps on a mid phone; decoded-pixel memory under 100 MB on a phone; 40-photo upload on 3G completes with resume; 1M visitors a month; a visible cost per GB served.
the two numbers nobody asks for
File size is the wire cost; pixel count × 4 is the decode and memory cost, and it does not depend on the file size. A grid of 40 originals is 120 MB on the wire and 1.9 GB decoded. The second number killed the field test; the first passed the audit.
21

v1: originals everywhere, and the decode budget

code
// v1: upload the original through the API; show the original in every slot with CSS sizing. looks right in the office on fibre
<img src={photo.url} style={{ width: 400, height: 300, objectFit: 'cover' }} />
// the grid: 40 × 3 MB = 120 MB on the wire; 40 × 4000×3000×4 B = 1.9 GB decoded; the phone tab has ~300 MB
// what the user sees on a phone: a long blank grid, then images that flicker in and out as the browser evicts and re-decodes bitmaps on scroll, then a crash
what v1 got right
  1. The product: the grid, the gallery, the lightbox, the upload. On office laptops it works and the design is approved.
  2. The lesson it teaches by failing: every image is a decode budget. 12 MP × 4 bytes = 48 MB per photo in memory regardless of the 3 MB JPEG; a 400 px slot uses 1% of it.
the fix is the slot
  1. A variant ladder (400, 800, 1600, original) built server-side on upload; srcset and sizes so the browser picks by slot width × DPR; the original only in the lightbox.
  2. Formats: WebP by default (~25% smaller than JPEG), AVIF where decode time is affordable (~50% smaller, ~2× decode), JPEG fallback. <picture> or Accept negotiation at the CDN.
  3. Attributes: width and height (no layout shift), loading="lazy" below the fold, decoding="async", fetchpriority="high" on the hero only, a blurhash placeholder.
  4. The rule: sum over visible slots of (w × DPR) × (h × DPR) × 4 under ~100 MB on a phone with one screen of margin; first-screen wire bytes under 1 MB on 4G. Both computed from the layout before a photo exists.
code
// v1 → v2 markup: the slot decides the pixels; the browser picks the variant; the LCP image is explicit
<img
  src={img.url(800)}                                   // fallback
  srcSet={`${img.url(400)} 400w, ${img.url(800)} 800w, ${img.url(1600)} 1600w`}
  sizes="(max-width: 640px) 100vw, (max-width: 1024px) 50vw, 400px"
  width={img.w} height={img.h}                         // reserves the slot: no layout shift
  loading={isHero ? 'eager' : 'lazy'} fetchPriority={isHero ? 'high' : 'auto'} decoding="async"
  style={{ background: `url(${blurhashToDataUrl(img.blurhash)})`, backgroundSize: 'cover' }}   // 20 bytes of placeholder
  alt={img.alt}
/>
// <picture> with AVIF then WebP sources when the ladder has both; or let the CDN negotiate on Accept and keep one URL
// the hero: also <link rel="preload" as="image" imagesrcset=… imagesizes=…> in the HTML head so it starts before the CSS is parsed
the sentence
  1. v1.5 buys a grid that fits a phone (77 MB decoded, 2 MB wire instead of 1.9 GB and 120 MB) and pays: a variant ladder on the server (money, a job per upload), a markup discipline every image must follow (complexity), and a placeholder scheme (complexity, small).
THE DECODE BUDGET
every image is pixels × 4 bytes, paid on decode, regardless of the file size
swipe the figure sideways, or tap expand for full screen
1/6
one photo
One listing photo: 4000 × 3000 JPEG, 3 MB on the wire. Decoded: 4000 × 3000 × 4 = 48 MB. Decode time on a laptop ~80 ms (off the main thread in Chrome, but the bitmap must exist before paint); on a mid phone ~300 ms. Displayed in a 400 × 300 slot: 47.5 MB of the 48 was decoded to be thrown away by the downscale.
22

Round two: the upload pipeline

code
// v2: client-side resize in a worker. the phone already decoded the photo to preview it; resize where the pixels are
// worker
self.onmessage = async ({ data: { file, maxW } }) => {
  const bmp = await createImageBitmap(file, { imageOrientation: 'from-image' })   // honours EXIF rotation; decode happens here, off the main thread
  const scale = Math.min(1, maxW / bmp.width)
  const w = Math.round(bmp.width * scale), h = Math.round(bmp.height * scale)
  const canvas = new OffscreenCanvas(w, h); canvas.getContext('2d')!.drawImage(bmp, 0, 0, w, h); bmp.close()   // release the 48 MB bitmap now
  const blob = await canvas.convertToBlob({ type: 'image/webp', quality: 0.85 })      // ~400 KB from 3 MB; EXIF (GPS, device) is gone
  self.postMessage({ blob, w, h })
}
// main: one worker, a queue of files, three concurrent resizes (each holds a decoded bitmap: 48 MB × 3 is the memory line on a phone)
// upload: GET /uploads/sign?type=image/webp&size=… → { url, key } → fetch(url, { method: 'PUT', body: blob }) with progress via XHR or a ReadableStream body
// resumable (videos, flaky links): multipart create → parts of 5 MB, 3 in flight → record each acked part {key, partNo, etag} in IndexedDB → complete; resume = skip acked parts
the break
  1. Agents on 3G upload 40 × 3 MB through the API: 20 minutes, no progress, a failure at minute 19 restarts the file. The API buffers 120 MB per listing in flight; at 50 concurrent agents it is 6 GB on the server. The number that broke is bytes through the API × concurrent uploaders.
v2
  1. Resize on the client, in a worker: createImageBitmap with imageOrientation: 'from-image', draw to an OffscreenCanvas at 2000 px wide, convertToBlob as WebP at 0.85. 3 MB becomes ~400 KB; the re-encode strips EXIF (GPS gone by default). Three resizes in flight at most: each holds a 48 MB bitmap and the phone has ~300 MB.
  2. Direct to storage: the API signs a PUT URL (scoped by content type, size range, key, short expiry); the client PUTs the blob to the bucket; the API sees a completion callback. Bucket CORS allows PUT from the origin. The server never holds bytes.
  3. Resumable: multipart with 5 MB parts, three in flight, acked parts recorded in IndexedDB (key, part number, etag); resume after a drop or a reload by skipping acked parts; progress is acked bytes. tus for a generic server; S3 multipart natively. Photos rarely need it; videos always do.
  4. Processing job: on completion, build the ladder (WebP and AVIF), a blurhash, a moderation pass. The listing shows the client-resized image optimistically and swaps to CDN variants when the job completes (a subscription from M8, or a poll).
the sentence
  1. v2 buys uploads that survive 3G and a server with no bytes in its path; pays: a resize worker and EXIF handling (complexity), a signed-URL flow that is a write grant and needs a security review (complexity), a resumable protocol with client-side part tracking (complexity), a processing pipeline (money), and a window where the listing shows an unprocessed image (consistency).
V2: THE UPLOAD PIPELINE
client resize, direct-to-storage, resumable chunks, and a processing job
swipe the figure sideways, or tap expand for full screen
1/6
v1: through the API
v1: → FormData → POST /api/listings/:id/photos → the server reads the body into memory, writes to storage, returns. 40 × 3 MB on 3G is 20 minutes with no progress and one failure restarts the file. The API server buffers 120 MB per listing in flight: at 50 agents it is the server's memory, not the client's, that breaks first.
23

Round three: video is a bitrate ladder

Video tours arrive: a 60-second phone clip is 150 MB at 50 Mbps. v1 for video (upload the file, <video src> it) streams 150 MB to every viewer and plays nothing for ten seconds on 4G. The number that broke is bitrate against the viewer's connection, and it is a different number for every viewer.

v3
  1. Upload on the resumable multipart path; no client transcode (WebCodecs can encode, but on a phone it is minutes and a hot battery). The listing shows "processing" with a poster from the first frame.
  2. Transcode job: a ladder of renditions (240p/400 kbps to 1080p/5 Mbps), H.264 everywhere and AV1 or HEVC where supported for ~40% fewer bytes, segmented at 4 s, with an HLS (or DASH) manifest. ~2 minutes for a 60 s clip with renditions in parallel; ~70 MB per minute stored.
  3. Playback: native HLS on Safari; hls.js via Media Source Extensions elsewhere. The player starts low for a fast first frame (~1 s on 4G from a 200 KB segment), measures throughput per segment, buffers 10 to 30 s ahead, switches with hysteresis. Targets: rebuffer ratio under 1%, few switches per minute.
  4. Grids: posters, never autoplaying players. A phone has ~2 hardware decoders; a grid of 12 autoplays falls to software decode, drops frames and drains the battery. The manifest loads on hover or tap. A sprite sheet of thumbnails (one image, 100 frames) gives scrubbing previews for one request.
the sentence
  1. v3 buys video that starts in a second and adapts; pays: transcode latency and compute (money; a "processing" state to design), storage per minute (money), a ~100 KB player library (bytes), and a decoder budget that forbids autoplay grids (capability).
V3: VIDEO IS A BITRATE LADDER
transcoding, adaptive streaming, and the player's decisions
swipe the figure sideways, or tap expand for full screen
1/6
upload
The upload: the same resumable multipart path as photos, with no client-side transcode (the browser can decode video to canvas but encoding in JS is impractical; WebCodecs can encode but at phone-draining cost; the server does it). 150 MB on 4G is ~5 minutes with progress and resume; the listing shows "processing" with the first frame as a poster.
24

Round four: delivery at a million visitors

Ten thousand listings and a million monthly visitors: media is now the cost line and the LCP line. Each listing page is 1.5 MB of images at p50; the origin serves them all; LCP is 3.4 s at p75 on 4G because the hero image queues behind thumbnails and a font. The numbers that broke are bytes × visitors (money) and request order (LCP).

v4
  1. Content-hashed, immutable URLs: /img/{hash}/{w}x{h}.{fmt}, Cache-Control: public, max-age=31536000, immutable. A new upload is a new hash; nothing is ever updated in place; the CDN and the browser cache forever and invalidation does not exist.
  2. On-demand variants at the edge: a width not in the ladder is generated from the 1600 variant, cached, served; an allowlist of widths (or a signed URL) so an attacker cannot mint a million sizes.
  3. The first screen, in order: HTML → CSS → the hero (preloaded in the head with imagesrcset, fetchpriority=high) → thumbnails lazy beyond the fold → video manifests only on intent. ~300 KB for the first screen; LCP ~1.8 s on 4G. Fonts: font-display: swap and a preload for the hero's font.
  4. The cost line: 1M visitors × 1.5 MB = 1.5 TB a month: tens of dollars through a CDN, hundreds from an origin plus its CPU and bandwidth. AVIF and right-sizing halve the bytes; the CDN takes the origin to near zero.
  5. Measurement in the field: LCP p75 by page with the LCP element named (web-vitals to analytics), bytes per session by type (Resource Timing transferSize, sampled), CDN hit ratio over 95%, p95 image load by variant. A regression is one page that forgot the ladder.
the sentence, and the stop
  1. v4 buys LCP under 2 s at a million visitors and a cost line that follows the CDN; pays: a URL scheme that can never change (complexity), an edge function with a security boundary (complexity), a priority discipline every page must follow (complexity), field measurement (money).
  2. Stop: every number in the brief is met with margin. 360° tours and floor-plan rendering are M4's territory (canvas and WebGL); offline agent uploads are M5's.
V4: DELIVERY AT SCALE
CDN, cache keys, priorities, and the first screen on 4G
swipe the figure sideways, or tap expand for full screen
1/6
immutable URLs
URLs: /img/{hash}/{w}x{h}.{fmt} where hash is of the original's content; a new upload is a new hash, so every variant can be cached forever (Cache-Control: public, max-age=31536000, immutable) at the CDN and in the browser. No cache invalidation problem exists because nothing is ever updated in place.
25

The whole board, and the exercise

RoundThe numberThe breakThe designPaid in
v12k listings, office fibre40 originals = 1.9 GB decoded on a 300 MB phoneVariant ladder; srcset/sizes; WebP/AVIF; lazy, async, priority, placeholdersA ladder job; a markup discipline
v2300 agents on 3G; 50 concurrentBytes through the API × uploaders; no resumeClient resize in a worker; signed direct-to-storage; resumable multipart; processing jobA worker; a write-grant flow; a protocol; a pipeline; an optimistic window
v3150 MB clips on 4G viewersOne bitrate cannot fit every connectionTranscode ladder; HLS/DASH; adaptive player; posters and sprites; no autoplay gridsTranscode latency and compute; storage; a player; a decoder cap
v41M visitors; 1.5 MB per pageCost × visitors; LCP behind thumbnailsImmutable hashed URLs on a CDN; edge variants with an allowlist; first-screen order; field metricsA fixed URL scheme; an edge function; a priority discipline; measurement
what the sequence teaches
  1. Pixels, not bytes, are the memory and decode budget; compute it from the layout before any upload.
  2. Bytes never go through the API; the API signs and the storage carries. Resumability is parts plus a record of acked parts.
  3. Video is a ladder and a control loop; the pipeline builds the ladder, the player runs the loop, and grids show posters.
  4. Delivery is a URL scheme, a CDN and an order; immutable hashes remove invalidation, and the first screen's request order is the LCP.
the exercise
Open a media-heavy page you own on a phone with DevTools remote debugging; read the decoded image memory in the Memory panel and the transfer sizes in Network. Compute the slot budget from the layout and compare. If the decoded number is over 100 MB, the ladder is missing somewhere; find the image that is 4000 px wide in a 400 px slot.