Part 8 · 1 chapters · ~11 min
Concurrency in Hardware
What happens when cores share memory: cache coherence over 128-byte lines, the measured cost of atomics, memory ordering on ARM and why atomics act as barriers, locks built from compare-and-swap and wait/notify, and false sharing measured at about 4.7 times slower on this machine.
14
Sharing memory between cores
code
// the false-sharing experiment, as run on this machine (node, Apple M3 Pro)
import { Worker } from 'node:worker_threads';
const N = 20_000_000;
const code = `const { workerData: { sab, idx, n }, parentPort } = require('node:worker_threads');
const a = new Int32Array(sab); for (let i = 0; i < n; i++) Atomics.add(a, idx, 1); parentPort.postMessage(1);`;
async function run(stride: number) {
const sab = new SharedArrayBuffer(1024); const t = performance.now();
await Promise.all([0, 1].map(k => new Promise<void>(r => {
const w = new Worker(code, { eval: true, workerData: { sab, idx: k * stride, n: N } });
w.on('message', () => { w.terminate(); r(); });
})));
return performance.now() - t;
}
await run(1); // slots 0 and 1: same 128-byte line → about 690 ms
await run(32); // slots 0 and 32: 128 bytes apart → about 146 ms| measurement (this machine) | result |
|---|---|
cache line size (sysctl hw.cachelinesize) | 128 bytes |
20M increments, one thread: plain vs Atomics.add | 23 ms vs 127 ms |
2 workers × 20M Atomics.add: adjacent vs 128 B apart | about 690 ms vs about 146 ms (three runs each) |
why a frontend engineer cares
Workers, WebAssembly threads and audio worklets all share memory this way. The lesson carries upward: contention, not computation, is usually what makes concurrent code slow, at every scale from cache lines to database rows to distributed locks.
CONCURRENCY AT THE HARDWARE LEVEL
cache coherence, atomics, memory ordering, locks and false sharing, measured on this machine
swipe the figure sideways, or tap expand for full screen
1/6
coherence
Cache coherence: each core caches lines (128 bytes on this M3 Pro, as reported by hw.cachelinesize). A protocol such as MESI tracks each line's state: when one core writes, the other cores' copies are invalidated, and the line moves to the writer. Shared writes mean lines bouncing between cores.