Part 1 · 1 chapters · ~12 min

Designing Data-Intensive Applications

Kleppmann's book reduced to the ideas that change how you build: precise reliability and scalability, data models as choices, LSM versus B-tree and row versus column, replication and partitioning anomalies, isolation levels and write skew, and derived data kept in sync by logs.

3

Designing Data-Intensive Applications

The best single book on how data systems work underneath. It is precise about definitions and honest about trade-offs. Read it when the Distributed Systems course raises a question it cannot fully answer.

chaptersread forapply to
1precise definitions; percentilesevery SLO conversation (the SRE course)
3LSM versus B-tree; row versus columnchoosing a store; the Build Your Own storage engine
5, 6replication anomalies; partitioningread-your-writes on balances; sharding by tenant
7isolation levels and their anomaliesevery place money moves: write skew is real
8, 9the trouble with distributed systems; consistency and consensusthe Distributed Systems course, deeper
11, 12streams, derived data, the future of data systemsevent-driven and CDC architectures
code
-- write skew under snapshot isolation (DDIA ch. 7), the fintech version
-- two withdrawals from a shared wallet, each checking the balance, both commit
BEGIN ISOLATION LEVEL REPEATABLE READ;                       -- snapshot isolation in Postgres
SELECT balance_kobo FROM wallets WHERE id = 'w_7';           -- both see 50,000
-- app: 50,000 ≥ 40,000, ok
INSERT INTO ledger(wallet_id, amount_kobo) VALUES ('w_7', -4000000);
COMMIT;                                                      -- both succeed → −30,000
-- fixes: SERIALIZABLE (one aborts), or SELECT … FOR UPDATE, or a single conditional UPDATE:
UPDATE wallets SET balance_kobo = balance_kobo - 4000000 WHERE id = 'w_7' AND balance_kobo >= 4000000;
where it has aged
The first edition predates widespread serverless databases, cloud-native storage (Aurora, AlloyDB) and the current generation of stream processors, so some product examples are dated. The ideas have not aged. A second edition is in progress.
DESIGNING DATA-INTENSIVE APPLICATIONS
Martin Kleppmann, 2017 (second edition in progress): the ideas that change how you build
swipe the figure sideways, or tap expand for full screen
1/6
Reliable, scalable, maintainable
Reliability, scalability, maintainability: Kleppmann insists on precise definitions. A fault is one component deviating; a failure is the system stopping service; reliability is tolerating faults. Load is described by parameters that matter to your system; performance by percentiles, because averages hide the users who suffer.