Non-functionals are the job
Every property that decides whether a system survives has no ticket, no demo and no natural sponsor. That asymmetry, rather than any technical difficulty, is why reliability and adaptability decay while features ship. This part is about making them measurable, owned and budgeted, and about the under-investment that is engineering rather than negligence because it was written down.
Why the interesting requirements have no feature
“The system works. Why is half the engineering effort going into things no customer asked for?”
Because the properties that decide whether a system survives are the ones with no ticket, no demo and no obvious owner.
a feature is discrete, visible, and done "customers can send money to a saved beneficiary" → shipped a property is continuous, invisible, and never done "the system is reliable" → when? so features have sponsors and properties have volunteers, and that asymmetry is why properties decay by default.
- No natural sponsor. A product manager owns a feature. Nobody's objectives contain "the system remains changeable in 2028".
- The cost of neglect is deferred. Skipping reliability work this quarter costs nothing this quarter, which is precisely the problem.
- Success is invisible. Nobody notices the outage that did not happen, so the work that prevented it has no evidence.
- They are expensive to retrofit. Part 15 of the CBA module made the point about multi-tenancy: some properties are cheap at design time and a rewrite later.
The seven properties, and how they conflict
| Property | The question it answers | Primary mechanism |
|---|---|---|
| Reliability | Does it work when things fail? | Redundancy, isolation, and designed degradation |
| Performance | Is it fast enough, at the tail? | Budgets, measurement, and capacity headroom |
| Security | Can it be made to do the wrong thing? | Threat models, least privilege, and defence in depth |
| Observability | Can you tell what it is doing? | Structured events, traces, and correctness signals |
| Adaptability | Can it be changed safely? | Contracts, tests, and loose coupling |
| Supportability | Can humans operate it? | Runbooks, tooling, and page budgets |
| Sustainability | Does it stay viable? | Debt management and budgeted maintenance |
the conflicts, stated honestly: reliability ↔ adaptability every change risks the thing that works performance ↔ observability instrumentation costs latency and money security ↔ supportability least privilege slows an engineer at 3am performance ↔ reliability Part 1: fsync costs 5ms and buys durability everything ↔ cost all of it is paid for a design that claims to maximise all seven has not been costed.
Making a property measurable, or abandoning it
The single most useful discipline in this module. An unmeasured property is an opinion, and opinions lose to deadlines.
| Wish | Measurement | Now it is real |
|---|---|---|
| "It should be reliable" | 99.9% of transfers succeed, 30-day window | An error budget with a policy |
| "It should be fast" | p99 under 300 ms for a transfer | A regression test and an alert |
| "It should be secure" | Zero criticals over 30 days; MTTR for patching under 7 days | A queue with a deadline |
| "It should be observable" | Any transfer traceable from one id; 90% of incidents diagnosed without a new deploy | A capability you can test |
| "It should be maintainable" | Lead time under 2 days; change failure rate under 15% | DORA metrics with a trend |
| "It should be supportable" | Under 2 pages per on-call night | A budget that can be breached |
Who owns a property nobody asked for
- Nobody. The default. Produces steady decay punctuated by incident-driven panic.
- A specialist team. An SRE or platform group owns reliability. Produces expertise and a moral hazard: product teams stop feeling responsible.
- Every team owns it for their service. Produces genuine ownership and inconsistency, because twelve teams solve it twelve ways.
- Platform provides, teams consume, a named person stewards. The one that works: paved roads make the right thing easy, teams own outcomes, and one person tracks the property across the organisation.
Budgets: latency, error, cost, and change
a budget turns a property into an allowance you can spend: latency budget 300 ms for a transfer, itemised per step error budget 0.1% over 30 days = 300,000 failures cost budget $45k per region per month change budget how much risk you take per deploy the point of a budget is that spending it is ALLOWED. an unspent error budget means the target is too loose or you are over-investing in reliability. both are worth knowing.
- A pre-agreed policy. Budget exhausted means feature work stops until it recovers, agreed in advance rather than negotiated during an incident.
- Permission to take risk. Budget remaining means ship, experiment, take the deploy. Without that, every team defaults to caution and velocity dies quietly.
- A shared language with product. "This launch will cost about a third of the quarterly error budget" is a conversation. "It might be risky" is not.
- Burn-rate alerting. Burning a month of budget in an hour pages immediately; burning it slowly over a week raises a ticket. Part 13 of the CBA module made this point.
The conversation that turns a wish into a target
- "What breaks if we do not?" In customer and money terms. If there is no answer, the property may not matter here.
- "How would we know?" The measurement. If nobody can propose one, that is the first work item.
- "What is good enough?" Not perfect. The number where the next increment costs more than it returns.
- "What does that cost?" In engineering time and in running cost, because every nine is roughly an order of magnitude more expensive.
- "Who decides when we breach it?" The escalation path, agreed while nothing is on fire.
the shape of the cost curve, which people underestimate: 99% → a competent single-region deployment 99.9% → redundancy, failover, and drills 99.99% → multi-AZ, stand-in paths, deep investment 99.999% → a different company, and probably not honest each nine is roughly 10x the cost of the one before. Part 13 gave card auth four nines and notifications two, deliberately.
When to deliberately under-invest
The chapter that makes the rest credible. An engineer who wants maximum everything has not understood that the resources are finite.
- The component is genuinely not critical. The marketing site does not need the ledger's availability. Say so, write it down, and give it a lower target rather than no target.
- You are still finding product-market fit. Hardening something you may delete is waste. Timebox it explicitly: "we accept this until X, and revisit".
- The failure is cheap and visible. A batch job that fails loudly and reruns needs far less than a path where failure is silent.
- You have a compensating control. Part 10 reconciliation catches a class of errors, which justifies less defensiveness elsewhere.
- It is a decision, recorded, with a rationale and a date.
- The consequence is understood by whoever will be woken by it.
- It is revisited on a schedule, not when it fails.
- It is not applied to correctness. Part 13 was explicit: availability gets a budget, correctness does not. There is no acceptable rate of creating money.
The scorecard, and using it without theatre
| Property | Metric | Target | Now | Trend |
|---|---|---|---|---|
| Reliability | Transfer success, 30d | 99.9% | 99.94% | flat |
| Card auth availability | 99.99% | 99.992% | flat | |
| Performance | Transfer p99 | < 300 ms | 268 ms | rising |
| Balance read p99 | < 100 ms | 41 ms | flat | |
| Correctness | Invariant violations | 0 | 0 | flat |
| Unresolved break value | < ₦500k | ₦180k | falling | |
| Security | Open criticals | 0 | 1 | new |
| Patch MTTR | < 7 days | 4.2 days | flat | |
| Observability | Incidents needing a new deploy | < 10% | 18% | rising |
| Adaptability | Lead time for change | < 2 days | 3.1 days | rising |
| Change failure rate | < 15% | 11% | flat | |
| Supportability | Pages per on-call night | < 2 | 3.4 | rising |
| Sustainability | Maintenance share | > 20% | 12% | falling |
- Review it monthly, with the team, for twenty minutes. Not a quarterly presentation to leadership.
- Trend matters more than value. A number inside target and moving the wrong way is the actionable signal.
- One owner per row, per chapter 4.
- A breach produces a decision, which may legitimately be "accept and lower the target", recorded per chapter 7.
- Never tie it to individual performance review. The moment it does, the numbers become managed rather than measured, and you have lost the instrument.