Part 1 · 3 chapters · ~20 min

The Architect's View

Failure domains physical and logical, cells, shuffle sharding and static stability; availability math in series and parallel and the independence trap; cost models from provisioned to serverless; and the well-architected pillars as six questions with explicit trade-offs.

4

Failure domains and blast radius

know what fails together
  1. Physical domains nest: process, host, rack, zone, region, provider. Each level up costs more and protects against rarer events.
  2. Losing a zone with three zones is no outage, provided the two survivors can carry the load. Run at about 66% per zone or less.
  3. Logical domains (deploys, configuration, certificates, DNS, shared dependencies) cross every zone. Most big outages are logical.
  4. Cells are independent stacks, each serving a slice of customers and deployed one at a time.
  5. Shuffle sharding gives each customer a small random set of workers, so one bad tenant hurts almost no one else.
  6. Static stability: the data plane keeps serving when the control plane is down.
failure domainexample eventwhat survives it
processOOM killa supervisor restart; a second process
hosthardware faultreplicas on different hosts (spread placement)
zonepower or network loss in one data centrereplicas in three zones, sized to lose one
regioncontrol-plane or regional network eventa second region with data replicated and a tested failover
deploya bad releasecanaries, cells, rings, fast rollback
config / cert / DNSan expired certificate, a bad flagstaged config rollout, expiry alerts, last-known-good caching
dependencythe payment provider is downa second provider, a queue, a degraded mode
FAILURE DOMAINS AND BLAST RADIUS
every component lives inside nested boundaries that fail together; design so the boundary that fails takes as little as possible with it
swipe the figure sideways, or tap expand for full screen
1/6
nesting
Physical domains nest: process, host, rack, zone, region, provider. A replica on the same host as its primary protects against a process crash and nothing else; replicas in three zones survive a zone; a second region survives a region. Each level up costs more latency and more money, and protects against rarer events.
5

Availability math

nines compose
  1. Nines as downtime: 99.9% allows 43 minutes a month and 99.99% allows 4.3. Each nine costs more than ten times the one before.
  2. In series, availabilities multiply. Every synchronous dependency subtracts from the total.
  3. In parallel, failures multiply: the system is down only if every replica is down. Redundancy is the only way to beat the availability of your parts.
  4. The parallel formula assumes independence. Correlated failures break it.
  5. A dependency that can fail without failing the request leaves the chain.
  6. The target is capped by your dependencies and justified by your users.
code
// availability of a request path, the way you'd sketch it in a design review
const serial   = (...a) => a.reduce((x, y) => x * y, 1);
const parallel = (...a) => 1 - a.reduce((x, y) => x * (1 - y), 1);

const api      = parallel(0.99, 0.99, 0.99);        // 3 replicas in 3 zones: 0.999999
const db       = 0.9995;                            // managed, multi-AZ SLA
const psp      = parallel(0.999, 0.999);            // two payment providers: 0.999999
const checkout = serial(0.9999 /* CDN */, 0.9999 /* LB */, api, db, psp);
// ≈ 0.9993: the DB is now the weakest link; the next nine is there, not in more API replicas
the trap
Those replica numbers assume independent failures. Three replicas running the same bad deploy fail together, so the real figure for that path is the availability of your deploy process. That is why canaries and cells (above) matter as much as replica count.
AVAILABILITY MATH
nines, serial and parallel composition, and why a chain of good components makes a worse system
swipe the figure sideways, or tap expand for full screen
1/6
nines
Nines as downtime: 99% is 3.65 days a year (7.2 hours a month), 99.9% is 8.8 hours a year (43 minutes a month), 99.95% about 22 minutes a month, 99.99% about 52 minutes a year (4.3 minutes a month), 99.999% about 5 minutes a year. Each nine is ten times harder and usually more than ten times more expensive.
6

Cost models and the well-architected pillars

six lenses on every design
  1. Operational excellence: can you run it at 3 a.m.?
  2. Security: what can a stolen credential reach?
  3. Reliability: what takes it down, and for how long?
  4. Performance efficiency: where does the p99 go under peak load?
  5. Cost optimisation: what does one more customer cost, and who owns that number?
  6. Sustainability usually lines up with cost. The other trade-offs get decided and recorded in an ADR.
cost modelyou pay forfitswatch out for
provisioned (VMs, RDS)capacity, by the second, used or notsteady loadidle capacity; sizing for peak
committed (Savings Plans, CUDs)a spend or capacity commitment for 1 to 3 yearsthe baseline you are sure ofcommitting to what you later shrink
spot / preemptiblespare capacity at a deep discountbatch jobs, CI, stateless workersreclamation with little notice
serverless (Lambda, Cloud Run)requests and execution timespiky or low loadhigh steady load costing more than VMs
per-request managed (DynamoDB on-demand, S3)operations and storageunpredictable traffichot loops and chatty clients
unit economics
The number that matters is cost per unit of business: per active user, per transaction, per GB processed. A bill that grows with revenue is healthy. A bill that grows faster than revenue is an architecture problem. Part 8 shows how to tag spend so that number can be computed.
THE WELL-ARCHITECTED PILLARS
six lenses every design review uses, and the trade-offs between them
swipe the figure sideways, or tap expand for full screen
1/6
operations
Operational excellence: can you run it? Infrastructure as code, small reversible changes, runbooks, observability, post-incident learning. Question: when this breaks at 3 a.m., what does the on-call engineer see and do?