Part 9 · 1 chapters · ~8 min

Operating Kafka and Alternatives

Sizing brokers, disks and partitions, monitoring (under-replicated partitions, ISR shrinks, consumer lag, request latency), rebalancing partitions across brokers, upgrades, security (TLS, SASL, ACLs), managed Kafka (MSK, Confluent Cloud), and alternatives: Redpanda, Apache Pulsar, NATS JetStream, and cloud queues.

11

Running it, and alternatives

metricalert when
UnderReplicatedPartitions> 0 for more than a few minutes
OfflinePartitionsCount> 0: data unavailable
ISR shrink / expand ratefrequent shrinks: slow or overloaded brokers
consumer lag (per group)growing steadily
request latency p99 (produce, fetch)rising with load
disk usageabove ~70%: retention or capacity change needed
alternativedifferencechoose when
RedpandaKafka API, C++ thread-per-core design, no JVMyou want Kafka compatibility with simpler operations
Apache Pulsarcompute and storage separated (BookKeeper), multi-tenancy, queues and streamsmany tenants, geo-replication, tiered storage at scale
NATS JetStreamlightweight, simple operations, subjects with persistencemoderate volumes, edge and IoT, Go-centric stacks
SQS/SNS, Pub/Sub, Event Hubsmanaged queues and pub/subyou need delivery, not a retained replayable log

Security basics: TLS everywhere, SASL (SCRAM or OAuth) for clients, ACLs per topic and group with least privilege, and quotas so one client cannot starve others.