Part 1 · 9 chapters · ~55 min

Reliability: designing for the failure you will get

Reliability is not uptime and it is not redundancy. It is the set of decisions that determine what happens when a dependency is slow rather than dead, when every client retries at the same instant, and when the queue that was supposed to absorb a spike grows without bound instead. This part is those decisions, with the mechanisms named.

9

Availability arithmetic, and why nines mislead

worked numbers
  99%       = 3.65 days a year    = 7.2 hours a month
  99.9%    = 8.77 hours a year   = 43 minutes a month
  99.99%   = 52.6 minutes a year = 4.3 minutes a month
  99.999% = 5.3 minutes a year  = 26 seconds a month

but availability composes badly:
  6 dependencies at 99.9% in series = 0.999⁶ = 99.4%
  = 52 hours a year, which is Part 4 chapter 43 again
why the number misleads on its own
  1. It averages over time. Four minutes down at 03:00 on a Sunday and four minutes down on salary day are the same number and completely different events.
  2. It averages over users. 99.9% overall can be 100% for most customers and 0% for one shard, which is 20,000 people with a totally broken service.
  3. It hides partial failure. A system where transfers work and balances do not is "up" by most measurements.
  4. It says nothing about correctness. Part 13: the four golden signals can be green while the money is wrong.
what to use instead
Availability per critical user journey, measured as the fraction of attempts that succeeded, sliced by shard and region. "99.94% of transfer attempts succeeded, worst shard 99.81%" is actionable. A single organisation-wide uptime figure is a number for a slide.
10

Failure modes: crash, slow, wrong, and partitioned

ModeWhat happensDifficultyDefence
CrashProcess dies, connection refusedEasiestHealth checks, restart, redundancy
SlowResponds, eventuallyMuch harderTimeouts, breakers, bulkheads
WrongResponds fast, with bad dataHardestValidation, invariants, reconciliation
PartitionedSome can reach it, some cannotHardQuorum, fencing, designed CAP choice
ByzantineInconsistent answers to different callersHardestRarely defended against outside consensus systems
worked numbers
why slow is worse than crash:

  a crashed dependency → fails fast, breaker opens, you degrade
  a slow dependency   → holds your threads, fills your pool,
                               and takes down paths that never touched it

this is why every call has a timeout, and why the cloud module
capped autoscaling: slow dependencies amplify.
the mode nobody designs for
Wrong. A service returning plausible bad data passes every health check and every latency alert. In a bank the defence is structural rather than operational: the zero-sum invariant, the conditional write, and Part 10 reconciliation exist to catch "wrong" specifically, because no amount of redundancy detects it.
11

Redundancy that actually helps

the questions that separate real redundancy from the appearance of it
  1. Is the failure independent? Three instances in one AZ share a power domain. Three replicas of the same buggy build share the bug.
  2. Does failover work, measured? Cloud module Part 1: the documented RTO is a fraction of the real one. Untested redundancy is a hypothesis.
  3. Is there capacity to absorb the loss? Three instances at 80% cannot absorb one failing. Redundancy without headroom is a slower outage.
  4. Does it share a dependency? Two regions behind one DNS provider, one identity provider, one certificate authority. Find the shared root.
  5. Does redundancy add a failure mode? Split brain, replication lag, and the cost of coordination are all new risks.
worked numbers
correlated failure is the one that gets you:

  independent: 3 instances, P(all fail) = p³
  correlated : 3 instances, same bad deploy → P = p

  the maths of redundancy assumes independence,
  and most real outages are correlated failures.

which is why canary deploys and staged rollouts matter more
than instance count: they break the correlation.
12

Timeouts, retries, and the retry storm

the rules
  1. Every network call has a timeout. No exceptions. A call without one inherits the OS default, which is minutes.
  2. Timeouts derive from the budget, not from a guess. Part 8 of the CBA module propagated a deadline so each step got what was actually left.
  3. Retry only what is safe. Idempotent operations, and only on retryable errors. Never on INVALID_ARGUMENT.
  4. Exponential backoff with full jitter. The jitter matters more than the backoff, because without it every client retries at the same instant.
  5. A retry budget. Cap retries at a percentage of total requests, typically 10%. Above that, stop retrying: the dependency is down and you are the load.
  6. Never retry across layers. Three layers each retrying three times is 27 requests for one call.
code
// full jitter. the point is that clients DESYNCHRONISE.
const delay = Math.random() * Math.min(cap, base * 2 ** attempt);

// and a budget, so retries cannot become the load
if (retryBudget.consumed() > 0.1 * totalRequests) {
  return fail('retry budget exhausted');   // the dependency is down.
}
retry storm
how a brief blip becomes a sustained outage
swipe the figure sideways, or tap expand for full screen
1/8
steady
A dependency running comfortably below capacity. Load is steady.
13

Circuit breakers, bulkheads, and load shedding

PatternWhat it protectsMechanism
Circuit breakerYou, from a failing dependencyStop calling it. Fail fast. Probe occasionally
BulkheadOther paths, from one saturatingA bounded pool per dependency
Load sheddingYou, from your own callersReject early when overloaded
Rate limitingYou, from one noisy callerPer-caller quotas
BackpressureThe whole chainSignal upstream to slow down
the distinction people blur
A circuit breaker protects you from downstream; load shedding protects you from upstream. They are different directions and you need both. Part 6 of the CBA module put breakers on payment providers, and Part 13 shed analytics before transfers. A system with breakers and no shedding still falls over under its own success.
shedding well
  1. Shed early, at the edge, before work has been done. Rejecting after the database call wasted the expensive part.
  2. Shed by priority, using the Part 13 order: analytics first, then notifications, then history, then outbound, and never card authorisation.
  3. Return a clear, retryable error with Retry-After. A fast 503 beats a 30-second hang.
  4. Shed before you saturate, not after. Once queues are full the system is already in the non-linear region and recovery is slow.
14

Backpressure as a first-class design element

worked numbers
without backpressure, a fast producer and slow consumer give you:

  unbounded queue → memory growth → OOM → total loss

with backpressure:
  bounded queue → producer blocks or sheds → degraded, alive

an unbounded queue is not a design, it is the absence of one.
WhereSignalResponse
HTTP503 plus Retry-AfterClient backs off
KafkaConsumer lagAlert, scale consumers, or shed producers
TCPWindow sizeAutomatic, and invisible until you look
DatabasePool exhaustionFail fast, never queue unboundedly
WebSocketSend buffer fullDrop, collapse or disconnect, per Part 14
Internal queueDepth thresholdReject at the edge
the question to ask of every queue
"What happens when this fills?" If the answer is "it grows", there is no backpressure and the failure mode is memory exhaustion, which takes the whole process rather than one path. Part 14 of the CBA module made this concrete for the balance feed: the semantics of the data decide the policy, since a balance can be collapsed to the newest value and a transaction list cannot.
15

Idempotency as a reliability primitive

Usually filed under correctness. It is equally a reliability mechanism, because it is what makes retry safe, and retry is what makes a distributed system survive a network.

worked numbers
without idempotency:
  timeout → cannot retry safely → must resolve manually
  → every transient failure becomes an incident

with idempotency:
  timeout → retry freely → resolves itself
  → transient failures become invisible

idempotency converts a class of incidents into a non-event.
where it appeared in the CBA design, and it was always the same idea
  1. Part 2: a unique idempotency key on the journal row, so a client retry posts once.
  2. Part 4: idempotent consumers, so at-least-once delivery is safe.
  3. Part 5: accrual keyed by loan and date, so a re-run is a no-op.
  4. Part 6: our own reference on every outbound payment, so a provider deduplicates.
  5. Part 8: the clearing-file key, so a redelivered file posts nothing twice.
  6. Part 18: claim allocation keyed on the triggering credit entry.
the general statement
Design every mutating operation so that performing it twice is indistinguishable from performing it once. That single property is what lets you retry, replay, reprocess and re-run a batch without fear, and each of those is a reliability mechanism. A system without it has no safe recovery action, which is why its incidents all require a human.
16

Chaos engineering, proportionate to the system

a ladder, and most teams should stop partway up
  1. Rung 1: drills. Force a failover, flush the cache, kill an instance, in business hours, announced. Most of the value is here, and most teams have not done it.
  2. Rung 2: game days. Break something deliberately and have the on-call team respond as if real. Tests the humans and the runbooks, not just the system.
  3. Rung 3: continuous, scoped. Automated instance termination in non-critical services during working hours.
  4. Rung 4: production chaos on critical paths. Only with mature observability, automated rollback, and a genuine organisational appetite. For a ledger, rarely justified.
the honest position for a bank
Rungs 1 and 2 are mandatory and rung 4 usually is not. Part 13 listed six drills and the cloud module added restore verification; that is already more than most teams do. Randomly killing processes on the posting path is a poor trade when a scheduled, observed failover drill gives you most of the information without the risk. Saying that is better than performing sophistication.
17

The reliability review, as a checklist

before a service goes on the critical path
  1. Every external call has a timeout derived from a budget, and the budget is documented.
  2. Retries are jittered, bounded and budgeted, and only on idempotent operations.
  3. A circuit breaker per dependency, with separate breakers for read and write paths where their failure costs differ.
  4. A bounded pool per dependency, so one cannot starve the others.
  5. Every queue is bounded, with a stated policy for what happens when it fills.
  6. Load shedding at the edge, by priority, with a clear retryable error.
  7. Every mutating operation is idempotent, keyed on the caller's intent.
  8. Failure modes are documented: what happens when each dependency crashes, slows, and returns wrong data.
  9. Degradation is designed, with a stated order of what is shed.
  10. A failover drill has been run, and the measured RTO is published.
  11. Alerts are on symptoms, and each has a runbook with a do-not list.
using a checklist without it becoming a ritual
A checklist is a memory aid for people who already understand the items, not a substitute for understanding them. The useful form is a conversation where each line is answered with a specific mechanism rather than a yes. "Do you have timeouts?" answered with "yes" is worthless; answered with "250ms on the ledger call, derived from the 1000ms authorisation budget" is a design review.