10 parts · 11 chapters

Kafka and Event Streaming

Kafka is a distributed, replicated, partitioned log: producers append, consumers read at their own pace, and data stays for as long as you choose. That simple abstraction underlies event-driven architectures, CDC pipelines, stream processing and audit trails at companies of every size.

Ten parts: the log as the abstraction; brokers, topics, partitions and segments on disk; replication, in-sync replicas, leader election and KRaft; producers with batching, acks and idempotence; consumers, groups, rebalancing and offsets; exactly-once and transactions; schemas and evolution; Kafka Connect and CDC; Kafka Streams and Flink basics; and operating Kafka, with Redpanda, Pulsar and NATS as alternatives.

the log · brokers, topics, partitions · replication and KRaft · producers · consumers and groups · exactly-once · schemas · Connect and CDC · Streams and Flink · operating Kafka and alternativesmid → staff · backend and data engineers building event-driven systems
the logAppend-only, ordered per partition, retained, replayable.
durabilityReplication factor, min.insync.replicas, acks=all, KRaft.
producersBatching, compression, keys and partitioning, idempotence.
consumersGroups, rebalancing, offsets, at-least-once processing.
guaranteesExactly-once within Kafka, idempotent consumers outside it.
ecosystemSchema Registry, Connect, Debezium, Kafka Streams, Flink.
Built on Distributed Systems and ScalingDistributed Systems part 7 introduced logs and queues; Scaling Databases part 7 used Kafka for CDC and the outbox. The BYO course built a mini Kafka.