Part 5 · 10 chapters · ~50 min
Adaptability: changing it without breaking it
Reliability fails loudly, performance fails visibly, security fails eventually. Adaptability fails silently: no alert, no incident, just estimates that keep growing and nobody able to say why. It is also the property that makes the other six fixable, because every one of them is maintained by changing the system.
47
The property that decays silently
the question
“Why does a change that took three days last year take three weeks now?”
worked numbers
reliability fails loudly: an outage, a page, a postmortem performance fails visibly: a graph moves security fails eventually, and then catastrophically adaptability fails silently. no alert. no incident. no graph. just estimates that keep growing and nobody can say why. and it is the property that makes the other six fixable.
why it is the load-bearing one
- Every other property is maintained by change. You fix reliability by changing the system, improve performance by changing it, patch security by changing it.
- A system that cannot be changed safely cannot be improved, so every other property freezes at whatever level it happened to reach.
- It compounds. Each unsafe change adds coupling, which makes the next change less safe.
- It is the one properly measurable by DORA metrics: lead time, deploy frequency, change failure rate, time to restore. Part 0 made that the proxy.
the diagnostic question
"How long from a one-line change to it running in production, including everything?" If the honest answer is days, adaptability has already degraded, and every conversation about reliability or security is constrained by it. The number is usually known and rarely said out loud, because it sounds like an accusation rather than a measurement.
48
Coupling, cohesion, and where change propagates
| Coupling type | Example | Cost of change |
|---|---|---|
| Content | Reaching into another service’s database | Catastrophic. Their schema is now your API |
| Common | Shared mutable state, a shared table | Very high. Change affects unknown consumers |
| Control | Passing a flag that changes their behaviour | High. Their logic is now yours |
| Stamp | Passing a whole object when one field is needed | Moderate. The whole shape becomes a contract |
| Data | Passing exactly what is needed | Low. The target. |
| Message | Publishing a fact, no knowledge of consumers | Lowest. Part 4 of the CBA module |
worked numbers
the test for a boundary, from Part 14: can this service enforce its invariant using only data it owns? loans: outstanding = disbursed − repaid both are facts loans recorded from events it consumed → the boundary is real if loans needed to read the ledger's tables to know this, → the boundary is fictional and the coupling is content-level
the practical signal
Count how many services must be deployed together to ship a change. If the answer is regularly more than one, the boundaries are in the wrong place regardless of how the architecture diagram looks. A distributed monolith has all the operational cost of microservices and none of the independence, and it is diagnosed by deploy coupling rather than by inspection.
49
Contracts, versioning, and expand-contract
Part 14 of the CBA module covered protobuf versioning. The general pattern matters more than the format.
worked numbers
expand-contract, the only safe way to change a contract: 1. expand add the new thing, keep the old 2. migrate move consumers, one at a time, at their pace 3. verify confirm zero usage of the old thing 4. contract remove it four deploys instead of one, and no coordinated release. the alternative is a flag day, which does not scale past two teams.
what makes step 3 possible
- Instrument usage of the deprecated thing, by caller. Part 14: because every call carries a service identity, the metric is
deprecated_field_use{field, caller}. - That converts a broadcast announcement into three specific conversations.
- Never remove on a date alone. Remove on evidence of zero usage, sustained.
- Reserve the identifier permanently, so it cannot be reused with a different meaning later.
50
Database migrations on a live money path
the rules, which are stricter than for ordinary systems
- Every migration is expand-contract. No exceptions on the ledger.
- No blocking DDL. A table lock on
entriesstops the bank. Use concurrent index creation, and check what each statement actually locks in your engine and version. - Backfills are batched, throttled and resumable. A single
UPDATEover 10 billion rows is an outage and a replication lag incident. - Never add a NOT NULL column with a default to a huge table without checking whether your version rewrites the whole table. This has caused real outages.
- Test the migration against production-sized data, per Part 2. A migration that takes 200 ms on 10,000 rows can take six hours on 10 billion.
- Have a rollback for each of the four phases, and know which phase is the point of no return.
code
-- expand: add nullable, no rewrite, no lock ALTER TABLE entries ADD COLUMN value_date DATE; -- backfill: batched, throttled, resumable, off-peak UPDATE entries SET value_date = created_at::date WHERE value_date IS NULL AND id BETWEEN $lo AND $hi; -- loop, 10k rows at a time, with a pause and a lag check -- verify, THEN constrain, and even then: NOT VALID first, -- then VALIDATE separately, so the lock is brief.
51
Feature flags, and their half-life
| Kind | Lifespan | Risk |
|---|---|---|
| Release flag | Days to weeks | Becomes permanent if not removed |
| Experiment flag | Weeks | Combinatorial explosion with other flags |
| Ops flag / kill switch | Permanent, by design | Must be tested or it does not work when needed |
| Permission flag | Permanent | Not a flag. It is authorisation, and belongs in the model |
worked numbers
a codebase with 30 release flags has, in principle, 2³⁰ = over a billion configurations of which you have tested perhaps three flags are debt with a very short half-life. every release flag needs an expiry date and an owner at creation, and a build that warns when one is past it.
the rule that keeps them useful
Kill switches are permanent and must be exercised. Part 13 of the CBA module built per-rail and per-provider switches, and a switch that has never been flipped in production is a hypothesis. Flip one in a drill, quarterly. Release flags are the opposite: they should be deleted aggressively, and a flag older than its expiry should fail the build rather than generate a ticket nobody reads.
52
Deployment strategies, and rollback as a first thought
what makes deployment safe rather than fast
- Small changes. The single largest factor in change failure rate. A one-line change that breaks is diagnosed in minutes; a two-week merge is not.
- Automated rollback triggers, on symptoms. Error rate, latency, and for us the correctness signal: if invariant violations move during a canary, roll back immediately regardless of the other metrics.
- Rollback tested, not assumed. An untested rollback is a hypothesis, exactly as Part 13 said about backups.
- Decoupled deploy and release. Ship the code dark behind a flag, then enable separately. Turning a flag off is faster and safer than a rollback.
- Database and code deployed separately, per chapter 50, so a code rollback never requires a schema rollback.
the question that reveals the real state
"What is the fastest we can get a change to production, and what is the fastest we can undo one?" If undo is slower than do, the system will accumulate bad changes because reverting is more painful than patching forward. Rollback speed is a feature of the deployment system, and it is the one people build last.
53
Testing: the pyramid, and where it lies to you
worked numbers
the classic pyramid: many unit, fewer integration, few end-to-end the reasoning is sound: fast and cheap at the bottom where it misleads for a system like ours: most of our risk is at boundaries and in concurrency unit tests with mocks assume away the interesting failures a mock never times out, never returns stale, never deadlocks the shape should follow where the risk is, not a diagram.
| Test type | Catches | Misses |
|---|---|---|
| Unit | Logic errors, edge cases in pure functions | Everything about integration |
| Integration, real dependencies | Contract mismatches, transaction behaviour | Production scale and concurrency |
| Contract | Provider and consumer drifting apart | Runtime behaviour |
| Property-based | Cases you did not imagine | Anything outside the generated space |
| Load | Contention, queueing, scale-dependent bugs | Correctness |
| Chaos / drill | Failure handling, human process | Anything not exercised |
where to actually spend, for a ledger
Integration tests against a real database, because the transaction semantics, the conditional write and the constraint behaviour are the thing being tested and a mock tests none of it. Plus property-based tests on the money logic, which is the next chapter. A mocked unit test of a posting function verifies that you wrote the function you wrote.
54
Testing money: property tests and invariant checks
the properties worth asserting, which hold for every input
- Entries sum to zero, per journal per currency, for any generated posting.
- Balance equals the sum of entries, for any sequence of operations in any order.
- Idempotency: applying the same operation twice with the same key equals applying it once.
- Commutativity where it should hold: two independent transfers in either order produce the same final balances.
- No operation can push a balance below its floor, for any interleaving.
- Round-trip: format then parse an amount returns the original, for every currency exponent.
code
// property-based: the framework generates thousands of cases,
// including the ones you would never write by hand.
test.prop([arb.postings()])('entries always sum to zero', (p) => {
const r = buildJournal(p);
for (const ccy of currencies(r)) {
expect(sumBy(r.entries, ccy)).toBe(0n);
}
});
// and the one that finds real bugs: any interleaving, any order
test.prop([arb.operations(), arb.permutation()])(
'balance is order-independent for independent ops', (ops, perm) => {
expect(runAll(ops)).toEqual(runAll(perm(ops)));
});why property tests are worth more here than anywhere else
The invariants are already written down. Part 1 of the CBA module defined them precisely, which means the tests almost write themselves, and they check the thing that actually matters rather than the implementation. Property-based testing finds the JPY zero-exponent bug, the rounding residual and the off-by-one in the tier cap, which are exactly the cases a human writes no test for.
55
Making the codebase legible to the next person
what actually helps, in order of value
- Names that match the domain. If the business says "post-no-debit", the code should say
postNoDebit, notflagB. A shared vocabulary between code and conversation removes an entire translation layer. - Explicit over clever. The reader is a tired person at 3am who did not write it, and that person is frequently you in eight months.
- Comments that explain why, never what. The code says what. "We fsync here despite the 5ms cost because an acknowledged transfer must survive node loss" is the comment worth writing.
- Architecture decision records. Short, dated, with the alternatives considered. The value is in preventing the same debate every eighteen months, and in explaining a decision that looks wrong without its context.
- A README that says how to run it and where the risky parts are, kept current because onboarding uses it, per Part 10.
the test for legibility
Onboarding time. How long before a new engineer ships a change to the posting path unsupervised? That number is a direct measurement of legibility and it is one of the few things that improves reliably when someone pays attention to it. It also degrades silently, exactly like adaptability, which is not a coincidence: they are the same property viewed from different sides.
56
Deletion as an engineering discipline
worked numbers
code that exists has ongoing cost: · it is read during every investigation · it is a dependency during every refactor · it must be patched, tested and understood · it is a surface for bugs and for attackers unused code is not free. it is a permanent tax with no return.
what to delete, and how to know it is safe
- Feature flags past their expiry, per chapter 51.
- Endpoints with zero traffic. Instrument, wait, confirm, remove. The instrumentation from chapter 49 makes this a fact rather than a guess.
- Dead experiments that were never cleaned up, which accumulate silently.
- Deprecated fields after confirmed zero usage, with the identifier reserved.
- Whole services that were replaced but never turned off, which is more common than it sounds and costs real money.
the cultural obstacle
Nobody gets credit for deleting things, and deletion feels risky in a way that addition does not. Both are wrong: deletion reduces surface, and instrumented deletion is far safer than most additions. Making "lines removed" as visible as "lines added" in how work is discussed is a small change that shifts behaviour, and a periodic explicit deletion pass is one of the cheapest adaptability investments available.