Part 5 · 10 chapters · ~50 min

Adaptability: changing it without breaking it

Reliability fails loudly, performance fails visibly, security fails eventually. Adaptability fails silently: no alert, no incident, just estimates that keep growing and nobody able to say why. It is also the property that makes the other six fixable, because every one of them is maintained by changing the system.

47

The property that decays silently

the question

“Why does a change that took three days last year take three weeks now?”

worked numbers
reliability fails loudly: an outage, a page, a postmortem
performance fails visibly: a graph moves
security fails eventually, and then catastrophically

adaptability fails silently.
  no alert. no incident. no graph.
  just estimates that keep growing and nobody can say why.

and it is the property that makes the other six fixable.
why it is the load-bearing one
  1. Every other property is maintained by change. You fix reliability by changing the system, improve performance by changing it, patch security by changing it.
  2. A system that cannot be changed safely cannot be improved, so every other property freezes at whatever level it happened to reach.
  3. It compounds. Each unsafe change adds coupling, which makes the next change less safe.
  4. It is the one properly measurable by DORA metrics: lead time, deploy frequency, change failure rate, time to restore. Part 0 made that the proxy.
the diagnostic question
"How long from a one-line change to it running in production, including everything?" If the honest answer is days, adaptability has already degraded, and every conversation about reliability or security is constrained by it. The number is usually known and rarely said out loud, because it sounds like an accusation rather than a measurement.
48

Coupling, cohesion, and where change propagates

Coupling typeExampleCost of change
ContentReaching into another service’s databaseCatastrophic. Their schema is now your API
CommonShared mutable state, a shared tableVery high. Change affects unknown consumers
ControlPassing a flag that changes their behaviourHigh. Their logic is now yours
StampPassing a whole object when one field is neededModerate. The whole shape becomes a contract
DataPassing exactly what is neededLow. The target.
MessagePublishing a fact, no knowledge of consumersLowest. Part 4 of the CBA module
worked numbers
the test for a boundary, from Part 14:

  can this service enforce its invariant using only data it owns?

loans: outstanding = disbursed − repaid
  both are facts loans recorded from events it consumed
  → the boundary is real

if loans needed to read the ledger's tables to know this,
  → the boundary is fictional and the coupling is content-level
the practical signal
Count how many services must be deployed together to ship a change. If the answer is regularly more than one, the boundaries are in the wrong place regardless of how the architecture diagram looks. A distributed monolith has all the operational cost of microservices and none of the independence, and it is diagnosed by deploy coupling rather than by inspection.
49

Contracts, versioning, and expand-contract

Part 14 of the CBA module covered protobuf versioning. The general pattern matters more than the format.

worked numbers
expand-contract, the only safe way to change a contract:

  1. expand   add the new thing, keep the old
  2. migrate  move consumers, one at a time, at their pace
  3. verify   confirm zero usage of the old thing
  4. contract remove it

four deploys instead of one, and no coordinated release.
the alternative is a flag day, which does not scale past two teams.
what makes step 3 possible
  1. Instrument usage of the deprecated thing, by caller. Part 14: because every call carries a service identity, the metric is deprecated_field_use{field, caller}.
  2. That converts a broadcast announcement into three specific conversations.
  3. Never remove on a date alone. Remove on evidence of zero usage, sustained.
  4. Reserve the identifier permanently, so it cannot be reused with a different meaning later.
50

Database migrations on a live money path

the rules, which are stricter than for ordinary systems
  1. Every migration is expand-contract. No exceptions on the ledger.
  2. No blocking DDL. A table lock on entries stops the bank. Use concurrent index creation, and check what each statement actually locks in your engine and version.
  3. Backfills are batched, throttled and resumable. A single UPDATE over 10 billion rows is an outage and a replication lag incident.
  4. Never add a NOT NULL column with a default to a huge table without checking whether your version rewrites the whole table. This has caused real outages.
  5. Test the migration against production-sized data, per Part 2. A migration that takes 200 ms on 10,000 rows can take six hours on 10 billion.
  6. Have a rollback for each of the four phases, and know which phase is the point of no return.
code
-- expand: add nullable, no rewrite, no lock
ALTER TABLE entries ADD COLUMN value_date DATE;

-- backfill: batched, throttled, resumable, off-peak
UPDATE entries SET value_date = created_at::date
 WHERE value_date IS NULL AND id BETWEEN $lo AND $hi;
-- loop, 10k rows at a time, with a pause and a lag check

-- verify, THEN constrain, and even then: NOT VALID first,
-- then VALIDATE separately, so the lock is brief.
51

Feature flags, and their half-life

KindLifespanRisk
Release flagDays to weeksBecomes permanent if not removed
Experiment flagWeeksCombinatorial explosion with other flags
Ops flag / kill switchPermanent, by designMust be tested or it does not work when needed
Permission flagPermanentNot a flag. It is authorisation, and belongs in the model
worked numbers
a codebase with 30 release flags has, in principle,
  2³⁰ = over a billion configurations
  of which you have tested perhaps three

flags are debt with a very short half-life.
every release flag needs an expiry date and an owner at creation,
and a build that warns when one is past it.
the rule that keeps them useful
Kill switches are permanent and must be exercised. Part 13 of the CBA module built per-rail and per-provider switches, and a switch that has never been flipped in production is a hypothesis. Flip one in a drill, quarterly. Release flags are the opposite: they should be deleted aggressively, and a flag older than its expiry should fail the build rather than generate a ticket nobody reads.
52

Deployment strategies, and rollback as a first thought

what makes deployment safe rather than fast
  1. Small changes. The single largest factor in change failure rate. A one-line change that breaks is diagnosed in minutes; a two-week merge is not.
  2. Automated rollback triggers, on symptoms. Error rate, latency, and for us the correctness signal: if invariant violations move during a canary, roll back immediately regardless of the other metrics.
  3. Rollback tested, not assumed. An untested rollback is a hypothesis, exactly as Part 13 said about backups.
  4. Decoupled deploy and release. Ship the code dark behind a flag, then enable separately. Turning a flag off is faster and safer than a rollback.
  5. Database and code deployed separately, per chapter 50, so a code rollback never requires a schema rollback.
the question that reveals the real state
"What is the fastest we can get a change to production, and what is the fastest we can undo one?" If undo is slower than do, the system will accumulate bad changes because reverting is more painful than patching forward. Rollback speed is a feature of the deployment system, and it is the one people build last.
53

Testing: the pyramid, and where it lies to you

worked numbers
the classic pyramid: many unit, fewer integration, few end-to-end

the reasoning is sound: fast and cheap at the bottom

where it misleads for a system like ours:
  most of our risk is at boundaries and in concurrency
  unit tests with mocks assume away the interesting failures
  a mock never times out, never returns stale, never deadlocks

the shape should follow where the risk is, not a diagram.
Test typeCatchesMisses
UnitLogic errors, edge cases in pure functionsEverything about integration
Integration, real dependenciesContract mismatches, transaction behaviourProduction scale and concurrency
ContractProvider and consumer drifting apartRuntime behaviour
Property-basedCases you did not imagineAnything outside the generated space
LoadContention, queueing, scale-dependent bugsCorrectness
Chaos / drillFailure handling, human processAnything not exercised
where to actually spend, for a ledger
Integration tests against a real database, because the transaction semantics, the conditional write and the constraint behaviour are the thing being tested and a mock tests none of it. Plus property-based tests on the money logic, which is the next chapter. A mocked unit test of a posting function verifies that you wrote the function you wrote.
54

Testing money: property tests and invariant checks

the properties worth asserting, which hold for every input
  1. Entries sum to zero, per journal per currency, for any generated posting.
  2. Balance equals the sum of entries, for any sequence of operations in any order.
  3. Idempotency: applying the same operation twice with the same key equals applying it once.
  4. Commutativity where it should hold: two independent transfers in either order produce the same final balances.
  5. No operation can push a balance below its floor, for any interleaving.
  6. Round-trip: format then parse an amount returns the original, for every currency exponent.
code
// property-based: the framework generates thousands of cases,
// including the ones you would never write by hand.
test.prop([arb.postings()])('entries always sum to zero', (p) => {
  const r = buildJournal(p);
  for (const ccy of currencies(r)) {
    expect(sumBy(r.entries, ccy)).toBe(0n);
  }
});

// and the one that finds real bugs: any interleaving, any order
test.prop([arb.operations(), arb.permutation()])(
  'balance is order-independent for independent ops', (ops, perm) => {
    expect(runAll(ops)).toEqual(runAll(perm(ops)));
  });
why property tests are worth more here than anywhere else
The invariants are already written down. Part 1 of the CBA module defined them precisely, which means the tests almost write themselves, and they check the thing that actually matters rather than the implementation. Property-based testing finds the JPY zero-exponent bug, the rounding residual and the off-by-one in the tier cap, which are exactly the cases a human writes no test for.
55

Making the codebase legible to the next person

what actually helps, in order of value
  1. Names that match the domain. If the business says "post-no-debit", the code should say postNoDebit, not flagB. A shared vocabulary between code and conversation removes an entire translation layer.
  2. Explicit over clever. The reader is a tired person at 3am who did not write it, and that person is frequently you in eight months.
  3. Comments that explain why, never what. The code says what. "We fsync here despite the 5ms cost because an acknowledged transfer must survive node loss" is the comment worth writing.
  4. Architecture decision records. Short, dated, with the alternatives considered. The value is in preventing the same debate every eighteen months, and in explaining a decision that looks wrong without its context.
  5. A README that says how to run it and where the risky parts are, kept current because onboarding uses it, per Part 10.
the test for legibility
Onboarding time. How long before a new engineer ships a change to the posting path unsupervised? That number is a direct measurement of legibility and it is one of the few things that improves reliably when someone pays attention to it. It also degrades silently, exactly like adaptability, which is not a coincidence: they are the same property viewed from different sides.
56

Deletion as an engineering discipline

worked numbers
code that exists has ongoing cost:
  · it is read during every investigation
  · it is a dependency during every refactor
  · it must be patched, tested and understood
  · it is a surface for bugs and for attackers

unused code is not free. it is a permanent tax with no return.
what to delete, and how to know it is safe
  1. Feature flags past their expiry, per chapter 51.
  2. Endpoints with zero traffic. Instrument, wait, confirm, remove. The instrumentation from chapter 49 makes this a fact rather than a guess.
  3. Dead experiments that were never cleaned up, which accumulate silently.
  4. Deprecated fields after confirmed zero usage, with the identifier reserved.
  5. Whole services that were replaced but never turned off, which is more common than it sounds and costs real money.
the cultural obstacle
Nobody gets credit for deleting things, and deletion feels risky in a way that addition does not. Both are wrong: deletion reduces surface, and instrumented deletion is far safer than most additions. Making "lines removed" as visible as "lines added" in how work is discussed is a small change that shifts behaviour, and a periodic explicit deletion pass is one of the cheapest adaptability investments available.