The event backbone: Kafka and its substitutes
Part 4 of the CBA module chose Kafka and partitioned by account id, because a credit and a debit on one account must arrive in order. Every managed alternative can satisfy that requirement, which means ordering does not decide this. Retention, replay precision and the surrounding toolchain do, and this part works out which of those you are actually trading away.
What Part 4 actually requires of a log
Same discipline as the ledger: write the requirement before opening the catalogue, because the requirement eliminates options that otherwise look interchangeable.
- Total ordering per account. Part 4 chose
account_idas the partition key because a credit and a debit on one account must arrive in order. - Durable retention with replay. A consumer down for six hours loses nothing; a new consumer rebuilds from the beginning.
- Independent consumer positions. Six consumers, six offsets, no coordination.
- At-least-once delivery. Duplicates are fine; loss is not.
- Exactly-once. Part 4 chapter 29 retired that phrase.
- Sub-millisecond latency. The lag budgets were 5 seconds to 15 minutes.
- Global ordering. Explicitly rejected, because it caps you at one partition.
the first requirement is the one that decides the platform: total ordering per key Kafka gives it through partitions. Pub/Sub does not have partitions. that single fact shapes the whole of this part.
MSK, MSK Serverless, and self-managed Kafka
| Self-managed | MSK | MSK Serverless | |
|---|---|---|---|
| Broker ops | Yours | AWS patches and replaces | Invisible |
| Partition control | Full | Full | Automatic, less control |
| Cost shape | Instances | Instances plus storage | Per GB and per partition-hour |
| Scaling | Manual rebalance | Manual, with tooling | Automatic |
| Fit | Only with real Kafka expertise | The default for a bank | Good for spiky or unknown load |
- Partition rebalancing is still a thought. Adding brokers does not move existing partitions without Cruise Control or manual reassignment.
- Consumer group health is still yours: lag, rebalancing storms and the Part 4 chapter 51 settings.
- Topic configuration is still yours, and
min.insync.replicas=2is still the setting everyone forgets. - Schema registry is separate: Glue Schema Registry or Confluent.
Kinesis Data Streams: the shard model
Kinesis: a stream is a set of shards, each with a hash range
per shard: 1 MB/s or 1,000 records/s in
2 MB/s out, shared across consumers
or 2 MB/s each with enhanced fan-out
ordering is per shard, and the partition key hashes to a shard
so per-account ordering works exactly as it does in Kafka| Kafka / MSK | Kinesis | |
|---|---|---|
| Ordering unit | Partition | Shard. Same idea |
| Retention | Configurable, effectively unlimited | Up to 365 days, billed |
| Consumer model | Consumer groups, broker-coordinated | KCL with a DynamoDB lease table |
| Resharding | Painful | Split and merge, online but fiddly |
| Throughput ceiling | Add partitions and brokers | Add shards, quota-limited |
| Ecosystem | Enormous | AWS-centric |
| Ops burden | Moderate even on MSK | Lower |
Pub/Sub: no partitions, and what that costs
Kafka: N partitions, each consumed by exactly one member ordering is a consequence of the assignment parallelism is bounded by partition count Pub/Sub: no partitions at all. a topic is a logical stream ordering is opt-in per message via an ordering key messages with the same key are delivered in order, one at a time parallelism is bounded by the number of distinct keys for us that is 20 million accounts, so parallelism is not the issue.
- No partition count to size. Part 4 spent a chapter deriving 48 partitions; on Pub/Sub that decision does not exist.
- Ordering costs throughput per key. An ordered key is delivered serially, and a slow message blocks its own key. For one account that is correct behaviour and it is worth knowing.
- No consumer group rebalancing, which removes the Part 4 chapter 51 rebalancing storm entirely. That is a genuine simplification.
- Replay is by seek to a timestamp or a snapshot, not by offset. Workable, and less precise.
- Retention is 7 days by default, up to 31. Shorter than Part 4 wanted, which pushes rebuild-from-source onto the warehouse.
Pub/Sub Lite, and Managed Kafka on GCP
| Pub/Sub | Pub/Sub Lite | Managed Service for Kafka | |
|---|---|---|---|
| Model | Serverless, no capacity | Partitioned, provisioned | Actual Kafka |
| Ordering | Per ordering key | Per partition | Per partition |
| Cost at steady high volume | Higher | Much lower | Instance-based |
| Ops | None | Capacity planning returns | Kafka ops, managed brokers |
| Ecosystem | GCP-native | GCP-native | Full Kafka ecosystem |
| Status | Mature | Mature, niche | Newer |
Ordering: the property that decides this
Everything in this part reduces to one question, so it deserves stating directly.
we need: total order of events per account. Kafka / MSK → partition by account_id Kinesis → partition key hashes to a shard Pub/Sub → ordering key = account_id Managed Kafka (GCP) → partition by account_id all four can do it. the differences are elsewhere: retention, replay precision, ecosystem, and ops burden.
- How long can a consumer be broken before it cannot catch up? Retention answers this, and Part 4 wanted 30 days.
- How precisely can you replay? Offsets are exact; timestamps are approximate. For a reconciliation rebuild, exactness matters.
- Does the surrounding toolchain assume Kafka? Debezium, Connect and the registry all do.
- What is the failure mode under load? Kafka pushes back through lag; Pub/Sub grows an invisible backlog until the subscription expires.
Retention, replay, and rebuilding a consumer
why Part 4 wanted 30 days: a consumer with a bug is fixed and must reprocess a new consumer must build its state from history an incident may need a replay of a window if retention < time-to-fix, the log cannot help you and rebuilding falls to the warehouse, which is hours behind and shaped differently.
Kinesis: up to 365 days, billed by GB-hour beyond 24 hours.
Managed Kafka: same as MSK, configurable.
The outbox relay on each platform
Part 4 chapter 49 built a relay that polls unpublished outbox rows with FOR UPDATE SKIP LOCKED and publishes before marking. That code is platform-independent; where it runs is not.
Not Lambda: the poll loop is continuous and connection churn is the Part 1 chapter 16 problem.
- Publish latency is dominated by the poll interval, not by the broker.
LISTEN/NOTIFYon Postgres wakes the relay immediately and cuts p99 publish latency substantially. - Batch size versus latency is a real tuning knob: 500 rows per poll at 50 ms gives good throughput and acceptable latency.
- The relay is a critical path for six consumers, so it gets its own SLO and its own alert on oldest-unpublished-age, per Part 13.
- Producer config still matters:
acks=all,enable.idempotence=true, andmin.insync.replicas=2on the topic.
CDC: DMS, Datastream, and Debezium
| DMS | Datastream | Debezium | |
|---|---|---|---|
| Runs on | AWS managed | GCP managed | You run it, on Connect or embedded |
| Sources | Very broad | Postgres, MySQL, Oracle | Broad |
| Targets | S3, Redshift, Kinesis, many | BigQuery, GCS | Kafka |
| Schema evolution | Limited | Reasonable | Good, with the registry |
| Replication slot risk | Yes | Yes | Yes |
| Fit here | Analytics load into S3 | Direct to BigQuery, very clean | When the target is Kafka |
The decision, with numbers
Roughly $2,400 to $3,200 per month at our volume, dominated by broker hours.
Pub/Sub at ~1,500 events/s is roughly $1,100 to $1,600 per month, and genuinely less to operate.