Part 3 · 1 chapters · ~8 min

File Systems

Why to measure file system latency rather than disk latency, page cache hit ratios, cache misses and readahead, synchronous writes and fsync latency, file system locks and metadata operations, ext4slower and fileslower, cachestat, and directory scaling.

4

Latency the application feels

code
cachestat-bpfcc 1                # page cache hits, misses, dirtied pages per second, hit ratio
ext4slower-bpfcc 10              # every ext4 operation slower than 10 ms: process, op, bytes, latency, file
fileslower-bpfcc 10              # same, for synchronous file reads and writes from any file system
opensnoop-bpfcc -T               # which files are opened, failing opens (missing config, ENOENT storms)
strace -f -e trace=fsync,fdatasync -T -p $(pgrep -o postgres)   # how long each fsync takes

Typical orders of magnitude, not measurements: a page cache hit is microseconds, a local NVMe read about 100 µs, a cloud network volume around a millisecond, and a spinning disk seek several milliseconds. A database whose working set stops fitting in the page cache can slow down sharply even though nothing else changed.

WHERE FILE SYSTEM LATENCY COMES FROM (TYPICAL ORDERS)
one 4 KB read, by where it is served
page cache hit~1 µs or lessNVMe SSD read~100 µsnetwork block storage (cloud volume)~0.5-2 msspinning disk seek~5-10 ms
swipe the figure sideways, or tap expand for full screen
1/4
cached
A read served from the page cache costs about a memory copy: well under a few microseconds.
memory speedthe common case when warm