Part 0 · 8 chapters · ~50 min

Non-functionals are the job

Every property that decides whether a system survives has no ticket, no demo and no natural sponsor. That asymmetry, rather than any technical difficulty, is why reliability and adaptability decay while features ship. This part is about making them measurable, owned and budgeted, and about the under-investment that is engineering rather than negligence because it was written down.

1

Why the interesting requirements have no feature

the question

“The system works. Why is half the engineering effort going into things no customer asked for?”

Because the properties that decide whether a system survives are the ones with no ticket, no demo and no obvious owner.

worked numbers
a feature is discrete, visible, and done
  "customers can send money to a saved beneficiary"  → shipped

a property is continuous, invisible, and never done
  "the system is reliable"                            → when?

so features have sponsors and properties have volunteers,
and that asymmetry is why properties decay by default.
the four reasons they get neglected
  1. No natural sponsor. A product manager owns a feature. Nobody's objectives contain "the system remains changeable in 2028".
  2. The cost of neglect is deferred. Skipping reliability work this quarter costs nothing this quarter, which is precisely the problem.
  3. Success is invisible. Nobody notices the outage that did not happen, so the work that prevented it has no evidence.
  4. They are expensive to retrofit. Part 15 of the CBA module made the point about multi-tenancy: some properties are cheap at design time and a rewrite later.
the reframing that works
Stop calling them non-functional. Say what breaks without them, in the language of the business: "unavailable on salary day", "a breach we cannot detect", "six weeks to ship what used to take three days". A property nobody can describe the absence of will never be funded, and that is usually an articulation failure rather than a prioritisation one.
2

The seven properties, and how they conflict

PropertyThe question it answersPrimary mechanism
ReliabilityDoes it work when things fail?Redundancy, isolation, and designed degradation
PerformanceIs it fast enough, at the tail?Budgets, measurement, and capacity headroom
SecurityCan it be made to do the wrong thing?Threat models, least privilege, and defence in depth
ObservabilityCan you tell what it is doing?Structured events, traces, and correctness signals
AdaptabilityCan it be changed safely?Contracts, tests, and loose coupling
SupportabilityCan humans operate it?Runbooks, tooling, and page budgets
SustainabilityDoes it stay viable?Debt management and budgeted maintenance
worked numbers
the conflicts, stated honestly:

  reliability ↔ adaptability    every change risks the thing that works
  performance ↔ observability  instrumentation costs latency and money
  security ↔ supportability    least privilege slows an engineer at 3am
  performance ↔ reliability    Part 1: fsync costs 5ms and buys durability
  everything ↔ cost                   all of it is paid for

a design that claims to maximise all seven has not been costed.
the useful move
Rank them for this system, explicitly, and write it down. For our bank: correctness first and unconditionally, then reliability of the posting path, then security, then performance, then observability, then adaptability, then cost. A stated ranking turns a hundred future arguments into a lookup, and the arguments happen anyway if you do not write it down.
3

Making a property measurable, or abandoning it

The single most useful discipline in this module. An unmeasured property is an opinion, and opinions lose to deadlines.

WishMeasurementNow it is real
"It should be reliable"99.9% of transfers succeed, 30-day windowAn error budget with a policy
"It should be fast"p99 under 300 ms for a transferA regression test and an alert
"It should be secure"Zero criticals over 30 days; MTTR for patching under 7 daysA queue with a deadline
"It should be observable"Any transfer traceable from one id; 90% of incidents diagnosed without a new deployA capability you can test
"It should be maintainable"Lead time under 2 days; change failure rate under 15%DORA metrics with a trend
"It should be supportable"Under 2 pages per on-call nightA budget that can be breached
the honest corollary
If a property cannot be measured, either find a proxy or stop claiming it. "Maintainable" resisted measurement for decades until lead time and change failure rate turned it into something observable. An unmeasurable property is indistinguishable from a property you do not have, and pretending otherwise is how teams believe they are doing well right up until an incident.
4

Who owns a property nobody asked for

the four models, and what each actually produces
  1. Nobody. The default. Produces steady decay punctuated by incident-driven panic.
  2. A specialist team. An SRE or platform group owns reliability. Produces expertise and a moral hazard: product teams stop feeling responsible.
  3. Every team owns it for their service. Produces genuine ownership and inconsistency, because twelve teams solve it twelve ways.
  4. Platform provides, teams consume, a named person stewards. The one that works: paved roads make the right thing easy, teams own outcomes, and one person tracks the property across the organisation.
the stewardship role, described
Not a team and not a title. One named engineer per property whose job is to know the current state, maintain the scorecard, and raise it when it slips. Perhaps two hours a week. The value is that someone notices, which is the entire difference between a property that decays and one that does not. This is a classic staff-engineer responsibility, and Part 9 returns to it.
5

Budgets: latency, error, cost, and change

worked numbers
a budget turns a property into an allowance you can spend:

  latency budget  300 ms for a transfer, itemised per step
  error budget    0.1% over 30 days = 300,000 failures
  cost budget     $45k per region per month
  change budget  how much risk you take per deploy

the point of a budget is that spending it is ALLOWED.
an unspent error budget means the target is too loose or
you are over-investing in reliability. both are worth knowing.
what a budget makes possible that a target does not
  1. A pre-agreed policy. Budget exhausted means feature work stops until it recovers, agreed in advance rather than negotiated during an incident.
  2. Permission to take risk. Budget remaining means ship, experiment, take the deploy. Without that, every team defaults to caution and velocity dies quietly.
  3. A shared language with product. "This launch will cost about a third of the quarterly error budget" is a conversation. "It might be risky" is not.
  4. Burn-rate alerting. Burning a month of budget in an hour pages immediately; burning it slowly over a week raises a ticket. Part 13 of the CBA module made this point.
6

The conversation that turns a wish into a target

the five questions, in order
  1. "What breaks if we do not?" In customer and money terms. If there is no answer, the property may not matter here.
  2. "How would we know?" The measurement. If nobody can propose one, that is the first work item.
  3. "What is good enough?" Not perfect. The number where the next increment costs more than it returns.
  4. "What does that cost?" In engineering time and in running cost, because every nine is roughly an order of magnitude more expensive.
  5. "Who decides when we breach it?" The escalation path, agreed while nothing is on fire.
worked numbers
the shape of the cost curve, which people underestimate:

  99%      → a competent single-region deployment
  99.9%    → redundancy, failover, and drills
  99.99%   → multi-AZ, stand-in paths, deep investment
  99.999% → a different company, and probably not honest

each nine is roughly 10x the cost of the one before.
Part 13 gave card auth four nines and notifications two, deliberately.
the sentence that ends the argument
"What would we give up to get there?" Reliability, performance and security are not free and are not infinite. Asking what the organisation will trade converts an aspiration into a decision, and frequently reveals that the stated target was never meant seriously.
7

When to deliberately under-invest

The chapter that makes the rest credible. An engineer who wants maximum everything has not understood that the resources are finite.

legitimate reasons to under-invest, and how to do it honestly
  1. The component is genuinely not critical. The marketing site does not need the ledger's availability. Say so, write it down, and give it a lower target rather than no target.
  2. You are still finding product-market fit. Hardening something you may delete is waste. Timebox it explicitly: "we accept this until X, and revisit".
  3. The failure is cheap and visible. A batch job that fails loudly and reruns needs far less than a path where failure is silent.
  4. You have a compensating control. Part 10 reconciliation catches a class of errors, which justifies less defensiveness elsewhere.
what makes it under-investment rather than negligence
  1. It is a decision, recorded, with a rationale and a date.
  2. The consequence is understood by whoever will be woken by it.
  3. It is revisited on a schedule, not when it fails.
  4. It is not applied to correctness. Part 13 was explicit: availability gets a budget, correctness does not. There is no acceptable rate of creating money.
the distinction to hold
Deliberate under-investment is engineering. Undiscussed under-investment is decay. The difference is entirely whether it was written down, and that is also the difference between a postmortem that improves the system and one that finds somebody to blame.
8

The scorecard, and using it without theatre

PropertyMetricTargetNowTrend
ReliabilityTransfer success, 30d99.9%99.94%flat
Card auth availability99.99%99.992%flat
PerformanceTransfer p99< 300 ms268 msrising
Balance read p99< 100 ms41 msflat
CorrectnessInvariant violations00flat
Unresolved break value< ₦500k₦180kfalling
SecurityOpen criticals01new
Patch MTTR< 7 days4.2 daysflat
ObservabilityIncidents needing a new deploy< 10%18%rising
AdaptabilityLead time for change< 2 days3.1 daysrising
Change failure rate< 15%11%flat
SupportabilityPages per on-call night< 23.4rising
SustainabilityMaintenance share> 20%12%falling
how to read that table
Four rising trends and one falling maintenance share, which is the causal story. Maintenance dropped to 12%, so lead time grew, observability gaps widened, and pages per night rose. The scorecard is useful precisely because it makes that chain visible, and none of the individual numbers would have shown it alone.
using it without it becoming theatre
  1. Review it monthly, with the team, for twenty minutes. Not a quarterly presentation to leadership.
  2. Trend matters more than value. A number inside target and moving the wrong way is the actionable signal.
  3. One owner per row, per chapter 4.
  4. A breach produces a decision, which may legitimately be "accept and lower the target", recorded per chapter 7.
  5. Never tie it to individual performance review. The moment it does, the numbers become managed rather than measured, and you have lost the instrument.