Part 11 · 4 chapters · ~35 min

The operating picture

Eleven parts assembled into one table, one realistic year, and one ordering. The through-line is that the three properties which fail silently are the three that get neglected, two of them are the ones that make everything else fixable, and the difference between a good year and a bad one is whether the mechanisms existed before they were needed.

109

The seven properties, and their mechanisms

PropertyFails howPrimary mechanismsMeasured by
ReliabilityLoudlyTimeouts, jittered retry, breakers, bulkheads, shedding, backpressure, idempotencySuccess rate, error budget burn
PerformanceVisiblyBudgets, headroom at 65%, caching, tail-aware measurementp99 per journey
SecurityEventually, badlyThreat models, least privilege, four-eyes, audit, rotationOpen criticals, patch MTTR
ObservabilitySilentlyWide events, tail sampling, one correlation id, correctness signalIncidents diagnosed without a deploy
AdaptabilitySilentlyExpand-contract, small changes, tested rollback, deletionLead time, change failure rate
SupportabilityThrough people leavingPage budget, runbooks, agent tooling, contact-driven fixesPages per night, escalation rate
SustainabilitySilently, then all at onceMaintenance budget, interest-rate prioritisation, EOL trackingMaintenance share
the pattern across the table
The three that fail silently are the three that get neglected, and two of them, adaptability and sustainability, are the ones that make everything else fixable. That is the central argument of this module: the properties with no alarm are the ones that need a deliberate mechanism, because nothing else will make them visible.
110

A year in the life of this system

worked numbers
what a realistic year actually contains:

  ~3 Sev 1 or Sev 2 incidents, each with a postmortem
  ~40 Sev 3, mostly provider degradation
  ~600 pages, of which perhaps 480 actionable
  2 database major upgrades
  1 framework major, taking longer than planned
  ~15 dependency CVEs needing out-of-cycle patching
  1 provider outage lasting hours
  1 regulatory change requiring a limit or reporting change
  2 to 4 people joining, 1 to 2 leaving
  ~250 deploys

none of that is failure. that is a healthy system, running.
what the year looks like if the mechanisms are absent
  1. The same three incidents, but each takes four times as long to diagnose because observability was never built.
  2. The same pages, but 30% actionable instead of 80%, so the real one is missed and becomes a Sev 1.
  3. The upgrades slip a year, then become an emergency when support ends.
  4. The two leavers become five, and the CVE patching stops because nobody has capacity.
  5. 250 deploys becomes 40, because each one is risky, which makes each one larger, which makes it riskier.
the honest summary
The difference between the two years is not talent or effort. It is whether the mechanisms in this module existed before they were needed. Every one of them is cheap to build in advance and expensive to build during the incident that demonstrates why it was necessary.
111

What to fix first, in any system you inherit

the order, and the reasoning
  1. 1. Can you tell what is happening? Observability first, because everything else is guesswork without it. If incidents require a deploy to diagnose, fix that before anything.
  2. 2. Can you deploy and roll back safely? Because every subsequent fix requires shipping, and if shipping is dangerous nothing improves.
  3. 3. Is anyone being woken unnecessarily? Page volume, because a tired team fixes nothing and leaves. The alert review is usually a week of work and the highest immediate return.
  4. 4. What is the correctness story? For a money system, before availability: does anything check that the data is right?
  5. 5. What is the single biggest recurring incident cause? One fix, largest reduction.
  6. 6. Only now: architecture. The interesting work, and it is sixth because doing it before the first five means changing a system you cannot observe, deploy safely, or staff sustainably.
the mistake almost everyone makes
Starting at six. A new senior or staff engineer arrives, sees the architectural problems immediately, and proposes a redesign. The architecture is usually not the constraint: the constraint is that the team cannot observe, deploy, or sustain what they already have. Fixing one to five makes the architecture work possible; doing six first usually produces an abandoned migration and a reputation for not understanding the real problems.
112

Defending it: the fifteen hardest questions

01How do you justify time on things no customer asked for?
By naming what breaks without them, in the language of the business. "Six weeks to ship what used to take three days" is a business problem; "the code is messy" is an aesthetic complaint. A property nobody can describe the absence of will never be funded, and that is usually an articulation failure rather than a prioritisation one.
02What is your availability target and why?
Per journey, not per organisation. Card authorisation 99.99% because the network answers for us if we go quiet, transfers 99.9% because customers retry, notifications 99% because one channel always works. And correctness has no target because it has no acceptable failure rate, which is the distinction most SRE material misses.
03Your dependency slows down. What happens?
Slow is worse than dead. A crashed dependency fails fast and the breaker opens; a slow one holds threads, fills the pool and takes down paths that never touched it. So: per-dependency timeouts derived from a budget, bounded pools so it cannot starve others, a breaker, and a capped autoscaling maximum, because scaling into a slow dependency is a positive feedback loop.
04Why did a two-second blip become a twenty-minute outage?
Synchronised retries. Every client had the same timeout, so they all gave up at the same instant and retried immediately, tripling load on a service already struggling. Full jitter and a retry budget: jitter desynchronises, and the budget stops retries at around 10% of traffic because above that the dependency is down and you are the load.
05Mean latency is 45 ms. Is that good?
It tells me almost nothing. A bimodal distribution of 10 ms cache hits and 2 s misses has a comfortable mean and a terrible experience. I want p99 and p99.9, and I want to know the fan-out: twenty backend calls each at p99 of one second means 18% of pages hit a slow call. The tail is amplified by architecture.
06Is the system secure?
Unanswerable as asked. The answerable version is "what can go wrong and what stops it", which is a forty-minute threat model per trust boundary. And the boundary I would examine first is staff to customer data, because it is crossed thousands of times a day by people who are supposed to, which is why insider loss consistently outweighs external intrusion in financial services.
07Your observability bill doubled. Why?
Almost certainly cardinality. Someone added a dimension to a metric, and metrics cost per unique label combination. Adding account_id to a metric with four other dimensions turns 768 series into 15 billion. The fix is structural: low-cardinality dimensions on metrics, high-cardinality fields in wide structured events, and a correlation id bridging them.
08You have 1% trace sampling. Find me the failing request.
I probably cannot, which is why head-based uniform sampling is the wrong default. It discards 99% of errors along with 99% of successes. Tail-based sampling decides after the trace completes: keep every error, everything slower than the SLO, everything above a value threshold, and 0.1% of ordinary successes for a baseline.
09Why does a change take three weeks now when it took three days?
Adaptability decayed and nothing alerted, because it is the property that fails silently. I would look at coupling first: how many services must deploy together to ship one change. If that is regularly more than one, the boundaries are wrong regardless of the architecture diagram, and you have a distributed monolith with all the operational cost and none of the independence.
10How do you change the ledger schema without downtime?
Expand-contract, always, in four deploys. Add the nullable column, deploy code writing both, backfill in throttled resumable batches, deploy code reading the new one, then contract. No blocking DDL and no single UPDATE over ten billion rows, and the migration is tested against production-sized data because 200 ms on 10,000 rows can be six hours on 10 billion.
11Three pages a night. Is that a problem?
Yes, and not primarily as a morale issue. Above roughly two, responders stop reading pages carefully and the real one gets missed, so alert fatigue is a detection failure. I would categorise last month's pages, delete or demote everything not actionable, which is usually a third, and fix the top repeat cause. And treat sustained breach like an error budget breach: feature work stops.
12Your postmortem says "be more careful". Good enough?
No, that is not a mechanism and it will not survive a busy week. Blameless is a mechanism rather than a tone: ask what made the wrong action reasonable. What did the dashboard show, what did the runbook say, what did the tooling permit without warning. "The deploy tool now blocks when the canary correctness metric moves" is a finding; "be careful" is a wish.
13The team wants a rewrite. Should they get one?
One question decides it: can you run both in parallel and prove equivalence before cutting over? If yes, it is a strangler migration with manageable risk, which is exactly what the cloud module's ledger migration did with reconciliation as the gate. If no, it is a big-bang rewrite, the old system encodes years of undocumented edge cases, and the historical success rate is poor.
14How much capacity goes to maintenance?
A stated share, defended, around 20%. The failure mode is smuggling it into feature estimates, which makes estimates look padded, invites pressure, and squeezes maintenance out without anyone ever deciding to stop doing it. Made explicit, reducing it becomes a decision someone takes rather than an accident of estimation.
15What would you fix first in a system you just inherited?
Not the architecture, which is what everyone proposes and it is sixth. First: can you tell what is happening. Second: can you deploy and roll back safely. Third: is anyone being woken unnecessarily. Fourth, for a money system: does anything check the data is right. The architecture is usually not the constraint; the constraint is that the team cannot observe, ship or sustain what they already have.
the through-line of the whole module
Every part of this ends in the same place: a property that nobody measures is a property nobody has, and a mechanism that exists before the incident is worth ten times one built during it. The engineering is not difficult. Noticing, naming and funding the work whose success is an absence is the difficult part, and that is the job this module describes.