Distributed Systems
A distributed system is one in which the failure of a computer you did not know existed can make your own computer unusable. Leslie Lamport's line is still the best definition, because the defining property is partial failure: some parts work, some do not, and you often cannot tell which. Every hard problem in this course, from ordering events to agreeing on a leader to moving money between two services exactly once, comes from that one fact plus a network that loses, delays, duplicates and reorders messages.
Fourteen parts, in the order the ideas build on each other. Why distributed is different; time and ordering without a shared clock; consensus and Raft step by step; replication in its three shapes; partitioning; the consistency models and what each costs; failure detection and idempotency; logs, queues and streams; transactions across services with sagas and the outbox; and the papers that proved what is possible, read for what each one changed. Four further parts follow: CRDTs and client sync, caching across the system, membership with gossip and consistent hashing, and testing distributed systems.