Part 6 · 1 chapters · ~8 min

Distributed File Systems

The Google File System design (2003) and HDFS, metadata servers and large chunks, replication pipelines and rack awareness, the small-files problem, Ceph with RADOS and CRUSH, POSIX semantics versus object semantics, network file systems (NFS, EFS) and their latency, and choosing between block, file and object storage.

8

Metadata and data paths

needusewhy
a database's data directoryblock storage (local NVMe, EBS, Persistent Disk)low latency, fsync semantics, one writer
shared files across many servers (legacy apps, ML datasets)network file system (NFS, EFS, CephFS)POSIX semantics, shared, higher latency
backups, media, data lake, WAL archivesobject storage (S3, GCS, R2, MinIO, Ceph RGW)cheap, durable, HTTP API, immutable objects
very large batch analytics on-premisesHDFS or Cephdata locality, throughput
code
ceph osd tree                      # failure domains: hosts and racks
ceph osd map rbd my-object         # CRUSH computes placement: no lookup table
hdfs fsck /logs -files -blocks -locations   # HDFS: which DataNodes hold each block

POSIX is expensive to distribute: rename, locking, append and consistent directory listings across many clients need coordination. Object storage drops most of these guarantees, which is why it scales further and costs less.

A DISTRIBUTED FILE SYSTEM (GFS / HDFS SHAPE)
one metadata service, many data servers
clientopen /logs/day.logmetadata servernamespace, chunk mapdata server 1chunk 7 replicadata server 2chunk 7 replicadata server 3chunk 7 replica
swipe the figure sideways, or tap expand for full screen
1/4
metadata
A metadata server (GFS master, HDFS NameNode) holds the namespace and which servers store each large chunk (64-128 MB). Clients ask it where to go, then talk to data servers directly.
ask metadata, then read datacontrol and data paths split