10 parts · 11 chapters
Kafka and Event Streaming
Kafka is a distributed, replicated, partitioned log: producers append, consumers read at their own pace, and data stays for as long as you choose. That simple abstraction underlies event-driven architectures, CDC pipelines, stream processing and audit trails at companies of every size.
Ten parts: the log as the abstraction; brokers, topics, partitions and segments on disk; replication, in-sync replicas, leader election and KRaft; producers with batching, acks and idempotence; consumers, groups, rebalancing and offsets; exactly-once and transactions; schemas and evolution; Kafka Connect and CDC; Kafka Streams and Flink basics; and operating Kafka, with Redpanda, Pulsar and NATS as alternatives.
the logAppend-only, ordered per partition, retained, replayable.
durabilityReplication factor, min.insync.replicas, acks=all, KRaft.
producersBatching, compression, keys and partitioning, idempotence.
consumersGroups, rebalancing, offsets, at-least-once processing.
guaranteesExactly-once within Kafka, idempotent consumers outside it.
ecosystemSchema Registry, Connect, Debezium, Kafka Streams, Flink.
00
The Log as the Abstraction
A log, not a queue · Where a log fits
2 ch · ~12 min01Brokers, Topics, Partitions and Segments
Partitions and keys
1 ch · ~8 min02Replication, ISR and KRaft
Durability settings
1 ch · ~8 min03Producers
Batching, compression and idempotence
1 ch · ~8 min04Consumers and Consumer Groups
Groups, offsets and lag
1 ch · ~8 min05Exactly-Once and Transactions
Transactions inside Kafka
1 ch · ~8 min06Schemas and Evolution
Contracts for events
1 ch · ~8 min07Kafka Connect and CDC
Connectors instead of code
1 ch · ~8 min08Kafka Streams and Flink Basics
Processing streams
1 ch · ~8 min09Operating Kafka and Alternatives
Running it, and alternatives
1 ch · ~8 minBuilt on Distributed Systems and ScalingDistributed Systems part 7 introduced logs and queues; Scaling Databases part 7 used Kafka for CDC and the outbox. The BYO course built a mini Kafka.