Part 0 · 2 chapters · ~12 min
Disks and SSD Internals
Hard disks (seeks, rotation, sequential versus random), NAND flash (SLC to QLC, pages and erase blocks), the flash translation layer, garbage collection and write amplification, TRIM, wear and endurance ratings (TBW, DWPD), NVMe queues, volatile drive caches and power-loss protection, and measured durability costs.
1
Durability, measured
code
// the benchmark (fsync.c): 4 KB writes to one file, three modes
for (int i = 0; i < n; i++) {
write(fd, buf, 4096);
if (mode == 1) fsync(fd); // macOS: to the drive, not through its cache
if (mode == 2) fcntl(fd, F_FULLFSYNC); // macOS: flush the drive cache too
}
// Apple M3 Pro, two runs: write only 7.6-10.1 µs · write+fsync 50.3-50.9 µs · write+F_FULLFSYNC 2,842-3,166 µsOn Linux, fsync is specified to flush the device's volatile write cache as well (via a cache flush command), so Linux fsync on a consumer SSD costs closer to the F_FULLFSYNC number. Enterprise SSDs with power-loss protection have capacitors that let them treat their cache as durable, so flushes return quickly; this is one of the biggest differences between consumer and data-centre drives for databases.
WHAT DURABILITY COSTS, MEASURED
one 4 KB write per commit, Apple M3 Pro internal SSD, macOS, clang -O2
swipe the figure sideways, or tap expand for full screen
1/4
page cache
write() returned after copying into the page cache: 7.6 to 10.1 µs per 4 KB across two runs. Nothing is durable yet; a power cut loses it.
fast, not durable~100,000 per second
2
Flash and the FTL
| device | random 4 KB read (typical order) | notes |
|---|---|---|
| 7,200 rpm hard disk | ~5-10 ms (seek + half a rotation) | ~100-200 random IOPS; good sequential bandwidth |
| SATA SSD | ~100 µs | limited by the SATA interface and one queue |
| NVMe SSD | tens of µs to ~100 µs | many deep queues, hundreds of thousands of IOPS |
INSIDE AN SSD
flash cannot overwrite in place
swipe the figure sideways, or tap expand for full screen
1/4
no overwrite
NAND flash pages can be written only after their whole erase block is erased. So an update to LBA 42 goes to a fresh page somewhere else.
write elsewhere, never in placeerase is per block