Part 3 · 1 chapters · ~8 min
Journaling and Copy-on-Write
The crash consistency problem, fsck as the old answer, write-ahead journaling (physical and logical, metadata-only versus full data), checkpointing, copy-on-write with atomic root updates, soft updates, log-structured file systems, the same choices in databases (WAL versus shadow paging), and what applications must still do.
5
Two answers to one problem
| journaling (ext4, XFS, NTFS) | copy-on-write (ZFS, btrfs, APFS) | |
|---|---|---|
| after a crash | replay the journal (seconds) | use the last committed root (instant) |
| extra writes | journaled blocks written twice | parents rewritten up to the root |
| snapshots | not native (use LVM) | native and cheap |
| random overwrites | in place, stays contiguous | fragment over time |
| database analogue | WAL + checkpoints (Postgres, InnoDB) | shadow paging (LMDB) |
What applications still owe: file systems keep their own structures consistent, not your data. Ordering and durability of application data still need fsync at the right points (the rename pattern in part 2, a WAL in a database).
CRASH CONSISTENCY: JOURNAL VERSUS COPY-ON-WRITE
appending a block to a file needs three updates
swipe the figure sideways, or tap expand for full screen
1/4
the problem
Appending one block touches the free-space bitmap, the inode and the data block. A crash between writes leaves them inconsistent: a leaked block or an inode pointing at garbage.
three writes, crash betweeninconsistent structures