Part 9 · 1 chapters · ~8 min
Operating Kafka and Alternatives
Sizing brokers, disks and partitions, monitoring (under-replicated partitions, ISR shrinks, consumer lag, request latency), rebalancing partitions across brokers, upgrades, security (TLS, SASL, ACLs), managed Kafka (MSK, Confluent Cloud), and alternatives: Redpanda, Apache Pulsar, NATS JetStream, and cloud queues.
11
Running it, and alternatives
| metric | alert when |
|---|---|
| UnderReplicatedPartitions | > 0 for more than a few minutes |
| OfflinePartitionsCount | > 0: data unavailable |
| ISR shrink / expand rate | frequent shrinks: slow or overloaded brokers |
| consumer lag (per group) | growing steadily |
| request latency p99 (produce, fetch) | rising with load |
| disk usage | above ~70%: retention or capacity change needed |
| alternative | difference | choose when |
|---|---|---|
| Redpanda | Kafka API, C++ thread-per-core design, no JVM | you want Kafka compatibility with simpler operations |
| Apache Pulsar | compute and storage separated (BookKeeper), multi-tenancy, queues and streams | many tenants, geo-replication, tiered storage at scale |
| NATS JetStream | lightweight, simple operations, subjects with persistence | moderate volumes, edge and IoT, Go-centric stacks |
| SQS/SNS, Pub/Sub, Event Hubs | managed queues and pub/sub | you need delivery, not a retained replayable log |
Security basics: TLS everywhere, SASL (SCRAM or OAuth) for clients, ACLs per topic and group with least privilege, and quotas so one client cannot starve others.