Part 4 · 1 chapters · ~14 min
File Systems and Storage
From write() to the flash cell: blocks, inodes and atomic rename, the page cache, journaling, the SSD's translation layer and garbage collection, fsync measured at about 1,300 times a plain write, and group commit, with the durable-write recipe and a durability table.
9
File systems, SSDs and fsync
durability is a contract you ask for
- Blocks, inodes, directories. Rename is atomic.
- The page cache: write() returns before anything is durable.
- Journaling keeps the structure consistent through a crash.
- SSDs remap everything internally, with garbage collection and write amplification.
- fsync waits for all of it, and costs about 1,300× a plain write here.
- Group commit shares one fsync across many writes.
code
// the durable-write recipe: write a temp file, fsync it, rename over the original, fsync the directory
import { openSync, writeSync, fsyncSync, closeSync, renameSync } from 'node:fs';
import { dirname } from 'node:path';
function writeFileDurably(path: string, data: Buffer) {
const tmp = path + '.tmp';
const fd = openSync(tmp, 'w'); writeSync(fd, data); fsyncSync(fd); closeSync(fd); // data on disk
renameSync(tmp, path); // atomic swap
const dfd = openSync(dirname(path), 'r'); fsyncSync(dfd); closeSync(dfd); // the rename itself on disk
}
// skip any step and some crash leaves the old file, an empty file, or a half-written one| what you did | survives a process crash? | survives a power cut? |
|---|---|---|
| write() returned | yes (it is in the page cache) | no |
| write() + fsync() on the file | yes | yes, for the data; a new file's directory entry may not be |
| temp file + fsync + rename + fsync(dir) | yes | yes: old or new, never half |
| a database commit with synchronous_commit on | yes | yes (the WAL was fsynced) |
| Redis appendfsync everysec | yes | up to about one second of writes can be lost |
measured
The numbers in the figure come from appending 2,000 records of 128 bytes in Node on this machine: 1.6 µs per write with no fsync, 2,137 µs with fsync on every write, and 29 µs with one fsync per 100 writes. The Build Your Own storage engine's WAL fsyncs every write for clarity. Group commit is how a real engine stays both durable and fast.
FROM write() TO THE FLASH CELL
blocks, inodes, the page cache, the journal, the SSD's translation layer, and what fsync actually waits for
swipe the figure sideways, or tap expand for full screen
1/6
blocks, inodes
Files are blocks plus metadata: a file system divides the device into blocks (usually 4 KB); an inode records a file's size, permissions, timestamps and which blocks hold its data; a directory maps names to inode numbers. Renaming a file changes a directory entry, not the data, which is why rename is atomic and cheap.