Event-Driven and Streaming Data
CDC done properly (log-based against query-based and trigger-based), Debezium and schema evolution, the outbox pattern, Kafka as a database log with retention, compaction and exactly-once, stream processing with windows and state, Lambda against Kappa, the OLTP to OLAP split, and when to add a columnar store.
CDC, Debezium and the outbox
| CDC style | how | catches deletes? | load on DB |
|---|---|---|---|
| query-based | poll WHERE updated_at > last | no (unless soft deletes) | repeated scans; misses fast double updates |
| trigger-based | triggers write to a change table | yes | doubles write cost in the transaction |
| log-based | read the WAL or binlog (Debezium) | yes, with before images | minimal; needs slots or binlog retention |
Schema evolution is the hard part: a column rename in the database breaks every consumer of the raw change stream. Publish domain events through an outbox (a contract you control) rather than raw table changes, and register schemas (Avro or Protobuf with a schema registry and compatibility rules).
Kafka as a log, stream processing and the OLAP split
| Kafka feature | use as a database log |
|---|---|
| retention by time or size | a replayable history for days or forever (tiered storage) |
| log compaction | keep the latest value per key: a changelog that rebuilds a table |
| idempotent producers and transactions | exactly-once read-process-write within Kafka |
| partitions keyed by aggregate id | ordering per account or per order |
Stream processing (Kafka Streams, Flink) keeps state per key and computes over windows: tumbling (fixed, non-overlapping), hopping (overlapping), session (gaps close them). Lambda architecture runs a batch path and a streaming path and merges them; Kappa uses only the stream and reprocesses by replaying the log. Kappa wins when the log is retained and replay is fast.
The OLTP to OLAP split: analytical queries (scans of months of data, group-bys across all merchants) do not belong on the transactional primary. Stream changes into a columnar store (ClickHouse, BigQuery, Snowflake, or DuckDB over Parquet) where scans of a few columns over billions of rows take seconds. Add one when dashboards start appearing in your slow-query top ten.