Part 0 · 2 chapters · ~12 min

The Log as the Abstraction

Logs as the shared idea behind databases, replication and integration, logs versus queues, retention and replay, events as facts, Kafka's origin at LinkedIn, and where a log fits in an architecture (CDC, event-driven services, stream processing, audit).

1

A log, not a queue

Kafka was built at LinkedIn (2011) to move activity and database change data between hundreds of systems. Its design choice was to make a durable, replayable log the central integration point, instead of point-to-point connections between every pair of systems.

THE LOG
append at the end, read from anywhere, keep it for as long as you like
producersappendpartition logoffset 0 1 2 3 4 5 6 7 ...fraud serviceat offset 7 (live)analyticsat offset 3 (behind)new search indexreplaying from 0
swipe the figure sideways, or tap expand for full screen
1/4
append-only
A log is an ordered, append-only sequence of records, each with an offset. Producers only add to the end; nothing is modified in place.
ordered, append-only, offsetsnothing modified in place
2

Where a log fits

useexample
event-driven servicesTransferCompleted consumed by notifications, fraud, loyalty, analytics
change data captureevery row change in Postgres streamed to a search index and a warehouse (Scaling course P7)
stream processingreal-time fraud features, rolling balances, alerts
audit and replayan immutable history of business events, replayed to rebuild views or debug
buffering between systemsabsorbing bursts so slow consumers do not slow producers

Events are facts, named in the past tense (TransferCompleted, not CompleteTransfer). Commands ask; events tell. Mixing them in one topic is a common design mistake.