Part 1 · 3 chapters · ~20 min
The Architect's View
Failure domains physical and logical, cells, shuffle sharding and static stability; availability math in series and parallel and the independence trap; cost models from provisioned to serverless; and the well-architected pillars as six questions with explicit trade-offs.
4
Failure domains and blast radius
know what fails together
- Physical domains nest: process, host, rack, zone, region, provider. Each level up costs more and protects against rarer events.
- Losing a zone with three zones is no outage, provided the two survivors can carry the load. Run at about 66% per zone or less.
- Logical domains (deploys, configuration, certificates, DNS, shared dependencies) cross every zone. Most big outages are logical.
- Cells are independent stacks, each serving a slice of customers and deployed one at a time.
- Shuffle sharding gives each customer a small random set of workers, so one bad tenant hurts almost no one else.
- Static stability: the data plane keeps serving when the control plane is down.
| failure domain | example event | what survives it |
|---|---|---|
| process | OOM kill | a supervisor restart; a second process |
| host | hardware fault | replicas on different hosts (spread placement) |
| zone | power or network loss in one data centre | replicas in three zones, sized to lose one |
| region | control-plane or regional network event | a second region with data replicated and a tested failover |
| deploy | a bad release | canaries, cells, rings, fast rollback |
| config / cert / DNS | an expired certificate, a bad flag | staged config rollout, expiry alerts, last-known-good caching |
| dependency | the payment provider is down | a second provider, a queue, a degraded mode |
FAILURE DOMAINS AND BLAST RADIUS
every component lives inside nested boundaries that fail together; design so the boundary that fails takes as little as possible with it
swipe the figure sideways, or tap expand for full screen
1/6
nesting
Physical domains nest: process, host, rack, zone, region, provider. A replica on the same host as its primary protects against a process crash and nothing else; replicas in three zones survive a zone; a second region survives a region. Each level up costs more latency and more money, and protects against rarer events.
5
Availability math
nines compose
- Nines as downtime: 99.9% allows 43 minutes a month and 99.99% allows 4.3. Each nine costs more than ten times the one before.
- In series, availabilities multiply. Every synchronous dependency subtracts from the total.
- In parallel, failures multiply: the system is down only if every replica is down. Redundancy is the only way to beat the availability of your parts.
- The parallel formula assumes independence. Correlated failures break it.
- A dependency that can fail without failing the request leaves the chain.
- The target is capped by your dependencies and justified by your users.
code
// availability of a request path, the way you'd sketch it in a design review const serial = (...a) => a.reduce((x, y) => x * y, 1); const parallel = (...a) => 1 - a.reduce((x, y) => x * (1 - y), 1); const api = parallel(0.99, 0.99, 0.99); // 3 replicas in 3 zones: 0.999999 const db = 0.9995; // managed, multi-AZ SLA const psp = parallel(0.999, 0.999); // two payment providers: 0.999999 const checkout = serial(0.9999 /* CDN */, 0.9999 /* LB */, api, db, psp); // ≈ 0.9993: the DB is now the weakest link; the next nine is there, not in more API replicas
the trap
Those replica numbers assume independent failures. Three replicas running the same bad deploy fail together, so the real figure for that path is the availability of your deploy process. That is why canaries and cells (above) matter as much as replica count.
AVAILABILITY MATH
nines, serial and parallel composition, and why a chain of good components makes a worse system
swipe the figure sideways, or tap expand for full screen
1/6
nines
Nines as downtime: 99% is 3.65 days a year (7.2 hours a month), 99.9% is 8.8 hours a year (43 minutes a month), 99.95% about 22 minutes a month, 99.99% about 52 minutes a year (4.3 minutes a month), 99.999% about 5 minutes a year. Each nine is ten times harder and usually more than ten times more expensive.
6
Cost models and the well-architected pillars
six lenses on every design
- Operational excellence: can you run it at 3 a.m.?
- Security: what can a stolen credential reach?
- Reliability: what takes it down, and for how long?
- Performance efficiency: where does the p99 go under peak load?
- Cost optimisation: what does one more customer cost, and who owns that number?
- Sustainability usually lines up with cost. The other trade-offs get decided and recorded in an ADR.
| cost model | you pay for | fits | watch out for |
|---|---|---|---|
| provisioned (VMs, RDS) | capacity, by the second, used or not | steady load | idle capacity; sizing for peak |
| committed (Savings Plans, CUDs) | a spend or capacity commitment for 1 to 3 years | the baseline you are sure of | committing to what you later shrink |
| spot / preemptible | spare capacity at a deep discount | batch jobs, CI, stateless workers | reclamation with little notice |
| serverless (Lambda, Cloud Run) | requests and execution time | spiky or low load | high steady load costing more than VMs |
| per-request managed (DynamoDB on-demand, S3) | operations and storage | unpredictable traffic | hot loops and chatty clients |
unit economics
The number that matters is cost per unit of business: per active user, per transaction, per GB processed. A bill that grows with revenue is healthy. A bill that grows faster than revenue is an architecture problem. Part 8 shows how to tag spend so that number can be computed.
THE WELL-ARCHITECTED PILLARS
six lenses every design review uses, and the trade-offs between them
swipe the figure sideways, or tap expand for full screen
1/6
operations
Operational excellence: can you run it? Infrastructure as code, small reversible changes, runbooks, observability, post-incident learning. Question: when this breaks at 3 a.m., what does the on-call engineer see and do?