Part 2 · 10 chapters · ~55 min

The event backbone: Kafka and its substitutes

Part 4 of the CBA module chose Kafka and partitioned by account id, because a credit and a debit on one account must arrive in order. Every managed alternative can satisfy that requirement, which means ordering does not decide this. Retention, replay precision and the surrounding toolchain do, and this part works out which of those you are actually trading away.

19

What Part 4 actually requires of a log

Same discipline as the ledger: write the requirement before opening the catalogue, because the requirement eliminates options that otherwise look interchangeable.

non-negotiable
  1. Total ordering per account. Part 4 chose account_id as the partition key because a credit and a debit on one account must arrive in order.
  2. Durable retention with replay. A consumer down for six hours loses nothing; a new consumer rebuilds from the beginning.
  3. Independent consumer positions. Six consumers, six offsets, no coordination.
  4. At-least-once delivery. Duplicates are fine; loss is not.
nice, and not required
  1. Exactly-once. Part 4 chapter 29 retired that phrase.
  2. Sub-millisecond latency. The lag budgets were 5 seconds to 15 minutes.
  3. Global ordering. Explicitly rejected, because it caps you at one partition.
worked numbers
the first requirement is the one that decides the platform:

  total ordering per key

Kafka gives it through partitions. Pub/Sub does not have partitions.
that single fact shapes the whole of this part.
20

MSK, MSK Serverless, and self-managed Kafka

Self-managedMSKMSK Serverless
Broker opsYoursAWS patches and replacesInvisible
Partition controlFullFullAutomatic, less control
Cost shapeInstancesInstances plus storagePer GB and per partition-hour
ScalingManual rebalanceManual, with toolingAutomatic
FitOnly with real Kafka expertiseThe default for a bankGood for spiky or unknown load
what MSK does not do for you
  1. Partition rebalancing is still a thought. Adding brokers does not move existing partitions without Cruise Control or manual reassignment.
  2. Consumer group health is still yours: lag, rebalancing storms and the Part 4 chapter 51 settings.
  3. Topic configuration is still yours, and min.insync.replicas=2 is still the setting everyone forgets.
  4. Schema registry is separate: Glue Schema Registry or Confluent.
what you are actually buying
Broker lifecycle, not Kafka expertise. MSK removes patching, disk management and replacement, which is real and substantial. It does not remove the need to understand partitions, ISR and consumer groups. A team that does not understand Kafka will have the same incidents on MSK, just without the 3am disk-full page.
21

Kinesis Data Streams: the shard model

worked numbers
Kinesis: a stream is a set of shards, each with a hash range

  per shard: 1 MB/s or 1,000 records/s in
             2 MB/s out, shared across consumers
             or 2 MB/s each with enhanced fan-out

ordering is per shard, and the partition key hashes to a shard
so per-account ordering works exactly as it does in Kafka
Kafka / MSKKinesis
Ordering unitPartitionShard. Same idea
RetentionConfigurable, effectively unlimitedUp to 365 days, billed
Consumer modelConsumer groups, broker-coordinatedKCL with a DynamoDB lease table
ReshardingPainfulSplit and merge, online but fiddly
Throughput ceilingAdd partitions and brokersAdd shards, quota-limited
EcosystemEnormousAWS-centric
Ops burdenModerate even on MSKLower
the honest comparison
Kinesis is genuinely simpler to operate and genuinely more limited. For our event backbone the decisive issues are retention, since Part 4 wanted 30 days so a consumer can be rebuilt, and the ecosystem, since Kafka Connect, Debezium and the schema registry all assume Kafka. Kinesis is the right answer for a pure AWS shop with simple consumers and the wrong one when you want the Kafka toolchain.
22

Pub/Sub: no partitions, and what that costs

worked numbers
Kafka: N partitions, each consumed by exactly one member
  ordering is a consequence of the assignment
  parallelism is bounded by partition count

Pub/Sub: no partitions at all. a topic is a logical stream
  ordering is opt-in per message via an ordering key
  messages with the same key are delivered in order, one at a time
  parallelism is bounded by the number of distinct keys

for us that is 20 million accounts, so parallelism is not the issue.
what changes in practice
  1. No partition count to size. Part 4 spent a chapter deriving 48 partitions; on Pub/Sub that decision does not exist.
  2. Ordering costs throughput per key. An ordered key is delivered serially, and a slow message blocks its own key. For one account that is correct behaviour and it is worth knowing.
  3. No consumer group rebalancing, which removes the Part 4 chapter 51 rebalancing storm entirely. That is a genuine simplification.
  4. Replay is by seek to a timestamp or a snapshot, not by offset. Workable, and less precise.
  5. Retention is 7 days by default, up to 31. Shorter than Part 4 wanted, which pushes rebuild-from-source onto the warehouse.
the tradeoff, stated fairly
Pub/Sub removes two real operational problems, partition sizing and rebalancing storms, and introduces one real constraint, shorter retention with less precise replay. For most systems that is a good trade. For a ledger where rebuilding a consumer from the log is part of the recovery story, the retention limit is the thing to interrogate.
ordering
Kafka partitions versus Pub/Sub ordering keys
swipe the figure sideways, or tap expand for full screen
1/7
partitions
Kafka splits a topic into partitions. Part 4 derived 48 of them for our event volume.
23

Pub/Sub Lite, and Managed Kafka on GCP

Pub/SubPub/Sub LiteManaged Service for Kafka
ModelServerless, no capacityPartitioned, provisionedActual Kafka
OrderingPer ordering keyPer partitionPer partition
Cost at steady high volumeHigherMuch lowerInstance-based
OpsNoneCapacity planning returnsKafka ops, managed brokers
EcosystemGCP-nativeGCP-nativeFull Kafka ecosystem
StatusMatureMature, nicheNewer
the GCP answer for our system
Managed Service for Kafka, if the Kafka ecosystem matters, which for us it does: Part 4 uses the schema registry, and Part 11 uses Debezium for CDC. Pub/Sub if you are willing to give up that toolchain in exchange for having no capacity decisions at all. Pub/Sub Lite exists mostly as a cost answer at very high steady volume and is rarely the right first choice.
24

Ordering: the property that decides this

Everything in this part reduces to one question, so it deserves stating directly.

worked numbers
we need: total order of events per account.

Kafka / MSK           → partition by account_id
Kinesis                 → partition key hashes to a shard
Pub/Sub                → ordering key = account_id
Managed Kafka (GCP) → partition by account_id

all four can do it. the differences are elsewhere:
retention, replay precision, ecosystem, and ops burden.
the questions that actually separate them
  1. How long can a consumer be broken before it cannot catch up? Retention answers this, and Part 4 wanted 30 days.
  2. How precisely can you replay? Offsets are exact; timestamps are approximate. For a reconciliation rebuild, exactness matters.
  3. Does the surrounding toolchain assume Kafka? Debezium, Connect and the registry all do.
  4. What is the failure mode under load? Kafka pushes back through lag; Pub/Sub grows an invisible backlog until the subscription expires.
25

Retention, replay, and rebuilding a consumer

worked numbers
why Part 4 wanted 30 days:

  a consumer with a bug is fixed and must reprocess
  a new consumer must build its state from history
  an incident may need a replay of a window

if retention < time-to-fix, the log cannot help you
and rebuilding falls to the warehouse, which is
hours behind and shaped differently.
AWS
GCP
Tiered storage
MSK: retention is configurable per topic, limited only by storage. pushes older segments to cheaper storage, so 30 days is affordable.

Kinesis: up to 365 days, billed by GB-hour beyond 24 hours.
Pub/Sub: 7 days default, 31 maximum. Snapshots let you mark a point and replay from it for up to 7 days.

Managed Kafka: same as MSK, configurable.
the design consequence
If you choose Pub/Sub, write the rebuild path explicitly: a consumer rebuilding from beyond retention reads the bronze layer from Part 11 instead of the topic. That is a real, workable design and it must be built and drilled before it is needed, not discovered during an incident. Shorter retention is acceptable when the fallback is designed; it is dangerous when it is assumed.
26

The outbox relay on each platform

Part 4 chapter 49 built a relay that polls unpublished outbox rows with FOR UPDATE SKIP LOCKED and publishes before marking. That code is platform-independent; where it runs is not.

AWS
GCP
ECS Fargate task
, long-lived, one or more replicas. Polls with SKIP LOCKED so replicas do not collide.

Not Lambda: the poll loop is continuous and connection churn is the Part 1 chapter 16 problem.
Cloud Run service with minimum instances set above zero, or a GKE deployment. Same reasoning exactly.
the platform-specific details worth knowing
  1. Publish latency is dominated by the poll interval, not by the broker. LISTEN/NOTIFY on Postgres wakes the relay immediately and cuts p99 publish latency substantially.
  2. Batch size versus latency is a real tuning knob: 500 rows per poll at 50 ms gives good throughput and acceptable latency.
  3. The relay is a critical path for six consumers, so it gets its own SLO and its own alert on oldest-unpublished-age, per Part 13.
  4. Producer config still matters: acks=all, enable.idempotence=true, and min.insync.replicas=2 on the topic.
27

CDC: DMS, Datastream, and Debezium

DMSDatastreamDebezium
Runs onAWS managedGCP managedYou run it, on Connect or embedded
SourcesVery broadPostgres, MySQL, OracleBroad
TargetsS3, Redshift, Kinesis, manyBigQuery, GCSKafka
Schema evolutionLimitedReasonableGood, with the registry
Replication slot riskYesYesYes
Fit hereAnalytics load into S3Direct to BigQuery, very cleanWhen the target is Kafka
the risk all three share, restated
Part 4 chapter 50 warned it and it is worth repeating because it has taken down production databases: a stalled CDC consumer prevents WAL reclamation and can fill the primary’s disk. On a managed database you may not be able to simply delete the slot. Alert on replication slot lag at the same severity as disk space, because they are the same alert with a delay.
28

The decision, with numbers

AWS
GCP
MSK
, 3 brokers across 3 AZs, tiered storage, 30-day retention on the entries topic. Glue Schema Registry.

Roughly $2,400 to $3,200 per month at our volume, dominated by broker hours.
Managed Service for Kafka for the same architecture, or Pub/Sub accepting shorter retention and a designed rebuild path.

Pub/Sub at ~1,500 events/s is roughly $1,100 to $1,600 per month, and genuinely less to operate.
the answer
"MSK on AWS, Managed Kafka on GCP, and the deciding factors are retention and ecosystem rather than ordering, because all four candidates can do per-account ordering. Part 4 wanted 30 days of retention so a broken consumer can be fixed and replayed, and the surrounding tooling, Debezium and the schema registry, assumes Kafka. Pub/Sub is a defensible choice that removes partition sizing and rebalancing storms entirely, and I would take it if I were willing to build and drill the rebuild-from-warehouse path. What I would not do is choose it by accident and discover the retention limit during an incident."