Part 1 · 1 chapters · ~8 min

ext4 and XFS

Inodes, directories and blocks, ext4 block groups and extents, journaling modes (ordered, writeback, journal), delayed allocation, XFS allocation groups and B+trees, scaling with threads, the atomic rename pattern and the 2009 ext4 zero-length file lesson, mount options, and checking with fsck and xfs_repair.

3

Layouts and safe writes

code
// atomic file replace (C): the pattern editors and databases use
int fd = open("state.json.tmp", O_WRONLY | O_CREAT | O_TRUNC, 0644);
write_all(fd, data, len);
fsync(fd); close(fd);                         // 1. data durable in the temp file
rename("state.json.tmp", "state.json");       // 2. atomic: readers see old or new, never half
int dir = open(".", O_RDONLY); fsync(dir);    // 3. make the rename itself durable
close(dir);

stat state.json                               # inode number, blocks, links
filefrag -v big.db                            # ext4: extents of a file
xfs_info /data                                # XFS: allocation groups (agcount), block size
tune2fs -l /dev/nvme0n1p1 | grep -i features  # ext4 features: extent, has_journal, dir_index …

The 2009 lesson: when ext4 introduced delayed allocation, applications that wrote a new file and renamed it without fsync could end up with zero-length files after a crash, because the rename was committed before the data was allocated. ext4 added a heuristic for this pattern, but the portable fix is the fsync in step 1.

EXT4 AND XFS
the two default Linux file systems
ext4 layoutBlock groups with inode tables andbitmaps; extents map files toblocks.ext4 journalDefault data=ordered: metadatajournaled, data written first.XFS layoutAllocation groups with B+trees forfree space and inodes: parallelallocation.XFS strengthsLarge files, many threads, bigfile systems; default on RHEL.delayed allocationBlocks chosen at writeback, notwrite(): better layout.the rename danceWrite temp, fsync, rename, fsyncdir: atomic file replace.
swipe the figure sideways, or tap expand for full screen
1/4
ext4
ext4 divides the disk into block groups and maps each file with extents (start block plus length) instead of block lists, so large files need little metadata.
block groups and extentsDebian and Ubuntu default