Part 3 · 1 chapters · ~8 min

Journaling and Copy-on-Write

The crash consistency problem, fsck as the old answer, write-ahead journaling (physical and logical, metadata-only versus full data), checkpointing, copy-on-write with atomic root updates, soft updates, log-structured file systems, the same choices in databases (WAL versus shadow paging), and what applications must still do.

5

Two answers to one problem

journaling (ext4, XFS, NTFS)copy-on-write (ZFS, btrfs, APFS)
after a crashreplay the journal (seconds)use the last committed root (instant)
extra writesjournaled blocks written twiceparents rewritten up to the root
snapshotsnot native (use LVM)native and cheap
random overwritesin place, stays contiguousfragment over time
database analogueWAL + checkpoints (Postgres, InnoDB)shadow paging (LMDB)

What applications still owe: file systems keep their own structures consistent, not your data. Ordering and durability of application data still need fsync at the right points (the rename pattern in part 2, a WAL in a database).

CRASH CONSISTENCY: JOURNAL VERSUS COPY-ON-WRITE
appending a block to a file needs three updates
file systemjournalmain structures
swipe the figure sideways, or tap expand for full screen
1/4
the problem
Appending one block touches the free-space bitmap, the inode and the data block. A crash between writes leaves them inconsistent: a leaked block or an inode pointing at garbage.
three writes, crash betweeninconsistent structures