Part 1 · 1 chapters · ~8 min
Brokers, Topics, Partitions and Segments
Brokers and clusters, topics and partitions as the unit of parallelism and order, keys and the default partitioner, choosing partition counts, segments with offset and time indexes, the page cache and zero-copy sendfile, retention by time and size, and tiered storage.
3
Partitions and keys
code
kafka-topics.sh --create --topic transfers --partitions 12 --replication-factor 3 --bootstrap-server b:9092 \ --config retention.ms=604800000 --config min.insync.replicas=2 kafka-topics.sh --describe --topic transfers --bootstrap-server b:9092 # Partition: 0 Leader: 1 Replicas: 1,2,3 Isr: 1,2,3 ... ls /var/lib/kafka/data/transfers-0/ 00000000000000000000.log 00000000000000000000.index 00000000000000000000.timeindex leader-epoch-checkpoint
| choosing partitions | consideration |
|---|---|
| max consumer parallelism | one partition is read by at most one consumer in a group: partitions ≥ peak consumers |
| throughput | target per-partition throughput (a few to tens of MB/s) × partitions ≥ peak |
| changing later | adding partitions changes key → partition mapping: per-key order breaks for in-flight keys |
| too many | more files, more replication work, slower leader elections; thousands per broker are fine with KRaft, millions are not |
TOPICS, PARTITIONS, SEGMENTS
how one topic spreads over brokers and disks
swipe the figure sideways, or tap expand for full screen
1/5
partitions
A topic is split into partitions: the unit of parallelism and ordering. Each partition is an ordered log on one leader broker (plus replicas). More partitions allow more parallel consumers.
partitions = parallelism and ordering unitsorder only within a partition