Part 1 · 9 chapters · ~55 min
Reliability: designing for the failure you will get
Reliability is not uptime and it is not redundancy. It is the set of decisions that determine what happens when a dependency is slow rather than dead, when every client retries at the same instant, and when the queue that was supposed to absorb a spike grows without bound instead. This part is those decisions, with the mechanisms named.
9
Availability arithmetic, and why nines mislead
worked numbers
99% = 3.65 days a year = 7.2 hours a month 99.9% = 8.77 hours a year = 43 minutes a month 99.99% = 52.6 minutes a year = 4.3 minutes a month 99.999% = 5.3 minutes a year = 26 seconds a month but availability composes badly: 6 dependencies at 99.9% in series = 0.999⁶ = 99.4% = 52 hours a year, which is Part 4 chapter 43 again
why the number misleads on its own
- It averages over time. Four minutes down at 03:00 on a Sunday and four minutes down on salary day are the same number and completely different events.
- It averages over users. 99.9% overall can be 100% for most customers and 0% for one shard, which is 20,000 people with a totally broken service.
- It hides partial failure. A system where transfers work and balances do not is "up" by most measurements.
- It says nothing about correctness. Part 13: the four golden signals can be green while the money is wrong.
what to use instead
Availability per critical user journey, measured as the fraction of attempts that succeeded, sliced by shard and region. "99.94% of transfer attempts succeeded, worst shard 99.81%" is actionable. A single organisation-wide uptime figure is a number for a slide.
10
Failure modes: crash, slow, wrong, and partitioned
| Mode | What happens | Difficulty | Defence |
|---|---|---|---|
| Crash | Process dies, connection refused | Easiest | Health checks, restart, redundancy |
| Slow | Responds, eventually | Much harder | Timeouts, breakers, bulkheads |
| Wrong | Responds fast, with bad data | Hardest | Validation, invariants, reconciliation |
| Partitioned | Some can reach it, some cannot | Hard | Quorum, fencing, designed CAP choice |
| Byzantine | Inconsistent answers to different callers | Hardest | Rarely defended against outside consensus systems |
worked numbers
why slow is worse than crash:
a crashed dependency → fails fast, breaker opens, you degrade
a slow dependency → holds your threads, fills your pool,
and takes down paths that never touched it
this is why every call has a timeout, and why the cloud module
capped autoscaling: slow dependencies amplify.the mode nobody designs for
Wrong. A service returning plausible bad data passes every health check and every latency alert. In a bank the defence is structural rather than operational: the zero-sum invariant, the conditional write, and Part 10 reconciliation exist to catch "wrong" specifically, because no amount of redundancy detects it.
11
Redundancy that actually helps
the questions that separate real redundancy from the appearance of it
- Is the failure independent? Three instances in one AZ share a power domain. Three replicas of the same buggy build share the bug.
- Does failover work, measured? Cloud module Part 1: the documented RTO is a fraction of the real one. Untested redundancy is a hypothesis.
- Is there capacity to absorb the loss? Three instances at 80% cannot absorb one failing. Redundancy without headroom is a slower outage.
- Does it share a dependency? Two regions behind one DNS provider, one identity provider, one certificate authority. Find the shared root.
- Does redundancy add a failure mode? Split brain, replication lag, and the cost of coordination are all new risks.
worked numbers
correlated failure is the one that gets you: independent: 3 instances, P(all fail) = p³ correlated : 3 instances, same bad deploy → P = p the maths of redundancy assumes independence, and most real outages are correlated failures. which is why canary deploys and staged rollouts matter more than instance count: they break the correlation.
12
Timeouts, retries, and the retry storm
the rules
- Every network call has a timeout. No exceptions. A call without one inherits the OS default, which is minutes.
- Timeouts derive from the budget, not from a guess. Part 8 of the CBA module propagated a deadline so each step got what was actually left.
- Retry only what is safe. Idempotent operations, and only on retryable errors. Never on
INVALID_ARGUMENT. - Exponential backoff with full jitter. The jitter matters more than the backoff, because without it every client retries at the same instant.
- A retry budget. Cap retries at a percentage of total requests, typically 10%. Above that, stop retrying: the dependency is down and you are the load.
- Never retry across layers. Three layers each retrying three times is 27 requests for one call.
code
// full jitter. the point is that clients DESYNCHRONISE.
const delay = Math.random() * Math.min(cap, base * 2 ** attempt);
// and a budget, so retries cannot become the load
if (retryBudget.consumed() > 0.1 * totalRequests) {
return fail('retry budget exhausted'); // the dependency is down.
}retry storm
how a brief blip becomes a sustained outage
swipe the figure sideways, or tap expand for full screen
1/8
steady
A dependency running comfortably below capacity. Load is steady.
13
Circuit breakers, bulkheads, and load shedding
| Pattern | What it protects | Mechanism |
|---|---|---|
| Circuit breaker | You, from a failing dependency | Stop calling it. Fail fast. Probe occasionally |
| Bulkhead | Other paths, from one saturating | A bounded pool per dependency |
| Load shedding | You, from your own callers | Reject early when overloaded |
| Rate limiting | You, from one noisy caller | Per-caller quotas |
| Backpressure | The whole chain | Signal upstream to slow down |
the distinction people blur
A circuit breaker protects you from downstream; load shedding protects you from upstream. They are different directions and you need both. Part 6 of the CBA module put breakers on payment providers, and Part 13 shed analytics before transfers. A system with breakers and no shedding still falls over under its own success.
shedding well
- Shed early, at the edge, before work has been done. Rejecting after the database call wasted the expensive part.
- Shed by priority, using the Part 13 order: analytics first, then notifications, then history, then outbound, and never card authorisation.
- Return a clear, retryable error with
Retry-After. A fast 503 beats a 30-second hang. - Shed before you saturate, not after. Once queues are full the system is already in the non-linear region and recovery is slow.
14
Backpressure as a first-class design element
worked numbers
without backpressure, a fast producer and slow consumer give you: unbounded queue → memory growth → OOM → total loss with backpressure: bounded queue → producer blocks or sheds → degraded, alive an unbounded queue is not a design, it is the absence of one.
| Where | Signal | Response |
|---|---|---|
| HTTP | 503 plus Retry-After | Client backs off |
| Kafka | Consumer lag | Alert, scale consumers, or shed producers |
| TCP | Window size | Automatic, and invisible until you look |
| Database | Pool exhaustion | Fail fast, never queue unboundedly |
| WebSocket | Send buffer full | Drop, collapse or disconnect, per Part 14 |
| Internal queue | Depth threshold | Reject at the edge |
the question to ask of every queue
"What happens when this fills?" If the answer is "it grows", there is no backpressure and the failure mode is memory exhaustion, which takes the whole process rather than one path. Part 14 of the CBA module made this concrete for the balance feed: the semantics of the data decide the policy, since a balance can be collapsed to the newest value and a transaction list cannot.
15
Idempotency as a reliability primitive
Usually filed under correctness. It is equally a reliability mechanism, because it is what makes retry safe, and retry is what makes a distributed system survive a network.
worked numbers
without idempotency: timeout → cannot retry safely → must resolve manually → every transient failure becomes an incident with idempotency: timeout → retry freely → resolves itself → transient failures become invisible idempotency converts a class of incidents into a non-event.
where it appeared in the CBA design, and it was always the same idea
- Part 2: a unique idempotency key on the journal row, so a client retry posts once.
- Part 4: idempotent consumers, so at-least-once delivery is safe.
- Part 5: accrual keyed by loan and date, so a re-run is a no-op.
- Part 6: our own reference on every outbound payment, so a provider deduplicates.
- Part 8: the clearing-file key, so a redelivered file posts nothing twice.
- Part 18: claim allocation keyed on the triggering credit entry.
the general statement
Design every mutating operation so that performing it twice is indistinguishable from performing it once. That single property is what lets you retry, replay, reprocess and re-run a batch without fear, and each of those is a reliability mechanism. A system without it has no safe recovery action, which is why its incidents all require a human.
16
Chaos engineering, proportionate to the system
a ladder, and most teams should stop partway up
- Rung 1: drills. Force a failover, flush the cache, kill an instance, in business hours, announced. Most of the value is here, and most teams have not done it.
- Rung 2: game days. Break something deliberately and have the on-call team respond as if real. Tests the humans and the runbooks, not just the system.
- Rung 3: continuous, scoped. Automated instance termination in non-critical services during working hours.
- Rung 4: production chaos on critical paths. Only with mature observability, automated rollback, and a genuine organisational appetite. For a ledger, rarely justified.
the honest position for a bank
Rungs 1 and 2 are mandatory and rung 4 usually is not. Part 13 listed six drills and the cloud module added restore verification; that is already more than most teams do. Randomly killing processes on the posting path is a poor trade when a scheduled, observed failover drill gives you most of the information without the risk. Saying that is better than performing sophistication.
17
The reliability review, as a checklist
before a service goes on the critical path
- Every external call has a timeout derived from a budget, and the budget is documented.
- Retries are jittered, bounded and budgeted, and only on idempotent operations.
- A circuit breaker per dependency, with separate breakers for read and write paths where their failure costs differ.
- A bounded pool per dependency, so one cannot starve the others.
- Every queue is bounded, with a stated policy for what happens when it fills.
- Load shedding at the edge, by priority, with a clear retryable error.
- Every mutating operation is idempotent, keyed on the caller's intent.
- Failure modes are documented: what happens when each dependency crashes, slows, and returns wrong data.
- Degradation is designed, with a stated order of what is shed.
- A failover drill has been run, and the measured RTO is published.
- Alerts are on symptoms, and each has a runbook with a do-not list.
using a checklist without it becoming a ritual
A checklist is a memory aid for people who already understand the items, not a substitute for understanding them. The useful form is a conversation where each line is answered with a specific mechanism rather than a yes. "Do you have timeouts?" answered with "yes" is worthless; answered with "250ms on the ledger call, derived from the 1000ms authorisation budget" is a design review.