Part 1 · 1 chapters · ~8 min

Brokers, Topics, Partitions and Segments

Brokers and clusters, topics and partitions as the unit of parallelism and order, keys and the default partitioner, choosing partition counts, segments with offset and time indexes, the page cache and zero-copy sendfile, retention by time and size, and tiered storage.

3

Partitions and keys

code
kafka-topics.sh --create --topic transfers --partitions 12 --replication-factor 3 --bootstrap-server b:9092 \
  --config retention.ms=604800000 --config min.insync.replicas=2
kafka-topics.sh --describe --topic transfers --bootstrap-server b:9092
# Partition: 0  Leader: 1  Replicas: 1,2,3  Isr: 1,2,3   ...

ls /var/lib/kafka/data/transfers-0/
00000000000000000000.log  00000000000000000000.index  00000000000000000000.timeindex  leader-epoch-checkpoint
choosing partitionsconsideration
max consumer parallelismone partition is read by at most one consumer in a group: partitions ≥ peak consumers
throughputtarget per-partition throughput (a few to tens of MB/s) × partitions ≥ peak
changing lateradding partitions changes key → partition mapping: per-key order breaks for in-flight keys
too manymore files, more replication work, slower leader elections; thousands per broker are fine with KRaft, millions are not
TOPICS, PARTITIONS, SEGMENTS
how one topic spreads over brokers and disks
topic: transfers6 partitionspartition 0leader broker 1partition 1leader broker 2partition 2leader broker 3segments00000000.log + .index + .timeindexkey → partitionhash(account_id) mod 6retention7 days or forever
swipe the figure sideways, or tap expand for full screen
1/5
partitions
A topic is split into partitions: the unit of parallelism and ordering. Each partition is an ordered log on one leader broker (plus replicas). More partitions allow more parallel consumers.
partitions = parallelism and ordering unitsorder only within a partition