Part 9 · 1 chapters · ~10 min

The GPU and the Compositor

What a GPU is built for, lockstep lanes and divergence, the cost of moving data and unified memory on Apple Silicon, the vertex, raster and fragment pipeline, how browsers paint into layers and composite on the GPU, and why transform and opacity animate smoothly when width and top do not.

15

From shader cores to smooth animation

code
/* compositor-only: smooth even while the main thread is busy */
.toast        { transform: translateY(16px); opacity: 0; transition: transform .2s, opacity .2s; }
.toast.show   { transform: translateY(0);    opacity: 1; }

/* layout every frame: drops frames on a cheap phone */
.toast-bad      { top: 16px; transition: top .2s; }
.toast-bad.show { top: 0; }

/* layer memory: a full-screen layer at 1170 × 2532 device pixels */
/* 1170 × 2532 × 4 bytes ≈ 11.8 MB of GPU memory, for one layer */
property animatedwork per frameruns on
transform, opacitycompositecompositor thread and GPU
color, background, box-shadowpaint + compositemain thread, then GPU
width, height, top, marginlayout + paint + compositemain thread, then GPU
where this connects
The Browser course follows the same pipeline from the renderer's side, and the Design course uses these rules for motion. Here the point is why: compositing reuses textures already on the GPU, and anything that changes the textures has to go back through the main thread.
THE GPU AND THE COMPOSITOR
thousands of simple cores, memory far from the CPU, and why transform and opacity animate smoothly
swipe the figure sideways, or tap expand for full screen
1/6
CPU vs GPU
CPU against GPU: a CPU has a few large cores tuned for fast single-thread work with branches and big caches; a GPU has thousands of small cores grouped into warps or wavefronts (typically 32 or 64 lanes) that execute the same instruction on different data. Divergent branches inside a group run one after the other.