Part 4 · 1 chapters · ~14 min

File Systems and Storage

From write() to the flash cell: blocks, inodes and atomic rename, the page cache, journaling, the SSD's translation layer and garbage collection, fsync measured at about 1,300 times a plain write, and group commit, with the durable-write recipe and a durability table.

9

File systems, SSDs and fsync

durability is a contract you ask for
  1. Blocks, inodes, directories. Rename is atomic.
  2. The page cache: write() returns before anything is durable.
  3. Journaling keeps the structure consistent through a crash.
  4. SSDs remap everything internally, with garbage collection and write amplification.
  5. fsync waits for all of it, and costs about 1,300× a plain write here.
  6. Group commit shares one fsync across many writes.
code
// the durable-write recipe: write a temp file, fsync it, rename over the original, fsync the directory
import { openSync, writeSync, fsyncSync, closeSync, renameSync } from 'node:fs';
import { dirname } from 'node:path';
function writeFileDurably(path: string, data: Buffer) {
  const tmp = path + '.tmp';
  const fd = openSync(tmp, 'w'); writeSync(fd, data); fsyncSync(fd); closeSync(fd);   // data on disk
  renameSync(tmp, path);                                                             // atomic swap
  const dfd = openSync(dirname(path), 'r'); fsyncSync(dfd); closeSync(dfd);         // the rename itself on disk
}
// skip any step and some crash leaves the old file, an empty file, or a half-written one
what you didsurvives a process crash?survives a power cut?
write() returnedyes (it is in the page cache)no
write() + fsync() on the fileyesyes, for the data; a new file's directory entry may not be
temp file + fsync + rename + fsync(dir)yesyes: old or new, never half
a database commit with synchronous_commit onyesyes (the WAL was fsynced)
Redis appendfsync everysecyesup to about one second of writes can be lost
measured
The numbers in the figure come from appending 2,000 records of 128 bytes in Node on this machine: 1.6 µs per write with no fsync, 2,137 µs with fsync on every write, and 29 µs with one fsync per 100 writes. The Build Your Own storage engine's WAL fsyncs every write for clarity. Group commit is how a real engine stays both durable and fast.
FROM write() TO THE FLASH CELL
blocks, inodes, the page cache, the journal, the SSD's translation layer, and what fsync actually waits for
swipe the figure sideways, or tap expand for full screen
1/6
blocks, inodes
Files are blocks plus metadata: a file system divides the device into blocks (usually 4 KB); an inode records a file's size, permissions, timestamps and which blocks hold its data; a directory maps names to inode numbers. Renaming a file changes a directory entry, not the data, which is why rename is atomic and cheap.