12 parts · 112 chapters
Running It, And The Team
The system exists and it is deployed. Now it has to stay reliable, secure, fast, observable and changeable for years, while people who did not build it operate it at 3am and a different set of people extend it. None of those properties has a feature, a ticket or an obvious owner, which is exactly why they decay.
The first half treats each property as an engineering discipline with mechanisms rather than as an aspiration. The second half is the human system that keeps them alive: on-call, support, technical debt, leading the work, and owning the outcome.
reliabilityDesigning for the failure you will actually get, rather than the one in the diagram.
performanceLatency and throughput as separate problems, and why the tail is what customers feel.
securityThreat models instead of checklists, and the insider threat that is the real one.
observabilityBeing able to ask a question nobody anticipated, bounded by cardinality and cost.
adaptabilityThe property that decays silently, and the only one that makes the others fixable.
supportabilityOn-call, support volume as a product signal, and the loop back into priorities.
sustainabilityTechnical debt named, priced and paid, and maintenance budgeted rather than smuggled.
00
Non-functionals are the job
Why the interesting requirements have no feature · The seven properties, and how they conflict · Making a property measurable, or abandoning it · Who owns a property nobody asked for · Budgets: latency, error, cost, and change · The conversation that turns a wish into a target · When to deliberately under-invest · The scorecard, and using it without theatre
8 ch · ~50 min01Reliability: designing for the failure you will get
Availability arithmetic, and why nines mislead · Failure modes: crash, slow, wrong, and partitioned · Redundancy that actually helps · Timeouts, retries, and the retry storm · Circuit breakers, bulkheads, and load shedding · Backpressure as a first-class design element · Idempotency as a reliability primitive · Chaos engineering, proportionate to the system · The reliability review, as a checklist
9 ch · ~55 min02Performance: latency, throughput, and the tail
Latency and throughput are different problems · Little’s Law, and using it in a room · Why the tail matters more than the mean · Where latency actually goes: a budget, itemised · Queueing theory, the useful ten percent · The four things that make software slow · Measuring before optimising, properly · Load testing that resembles production · Capacity planning, and headroom as policy · Performance as a regression test
10 ch · ~55 min03Security: threat models, not checklists
Threat modelling in forty minutes · The trust boundaries in our bank · Authentication, authorisation, and the gap between · Secrets: the lifecycle, not the store · Cryptography choices a non-specialist must get right · The insider threat, which is the real one · Supply chain: dependencies and build integrity · Security in the pipeline, not after it · Incident response when it is an attack · The security review, as a checklist
10 ch · ~55 min04Observability: being able to ask new questions
Monitoring answers known questions; observability answers new ones · The three signals, and the fourth for money · Cardinality: the constraint that shapes everything · Structured events over metrics, where it matters · Sampling strategies that keep the interesting traces · Dashboards people actually use · Alerting: the rules that stop alert fatigue · Debugging a system you cannot reproduce · The cost of observability, and containing it
9 ch · ~45 min05Adaptability: changing it without breaking it
The property that decays silently · Coupling, cohesion, and where change propagates · Contracts, versioning, and expand-contract · Database migrations on a live money path · Feature flags, and their half-life · Deployment strategies, and rollback as a first thought · Testing: the pyramid, and where it lies to you · Testing money: property tests and invariant checks · Making the codebase legible to the next person · Deletion as an engineering discipline
10 ch · ~50 min06On-call: the human side of reliability
What on-call is actually for · Rotation design, and the humane constraints · Severity levels that mean something · The page budget, and treating it as a limit · Handover, and the shift report · Incident command: roles during an incident · Communicating during an incident, internally and out · The postmortem, and making blamelessness real · Follow-up actions that actually get done · Measuring on-call health
10 ch · ~50 min07Support: the feedback loop you are ignoring
Support volume as a product signal · Tiering, escalation, and the engineering interface · Tooling: what agents need and never get · The read model for support, and its access controls · Contact-driven development · Deflection, and when it is hostile · Complaints, regulators, and the paper trail · Closing the loop into engineering priorities
8 ch · ~40 min08Technical debt: naming it, pricing it, paying it
The metaphor, and where it breaks down · A taxonomy: deliberate, accidental, and rot · Pricing debt in terms the business understands · The interest rate: how to tell what is urgent · Refactoring strategies that survive review · The rewrite, and when it is genuinely right · Migration as a permanent condition · Budgeting maintenance, explicitly · Making decay visible
9 ch · ~45 min09Leading the work: staff engineering in practice
What changes between senior and staff · Technical strategy, written down · The design document that gets read · Reviewing designs and code as leverage · Influence without authority, concretely · Disagreeing well, and committing after · Growing other engineers deliberately · Working with product and with finance · Estimation, and being honest about uncertainty · Choosing what not to work on · Visibility without self-promotion
11 ch · ~50 min10Engineering management: when you own the outcome
The transition, and what you give up · Team topology, and Conway’s Law as a tool · Hiring for a money system · Onboarding into a system this size · One-to-ones that are not status meetings · Feedback, performance, and the hard conversation · Planning that survives contact with reality · Metrics for engineering teams, used carefully · Protecting focus, and the cost of interruption · Incidents, blame, and psychological safety · Budget, headcount, and arguing for both · Retention, and why people actually leave · Managing in a regulated environment · The manager’s version of technical judgement
14 ch · ~55 min11The operating picture
The seven properties, and their mechanisms · A year in the life of this system · What to fix first, in any system you inherit · Defending it: the fifteen hardest questions
4 ch · ~35 minWritten against the bank from the other modulesEvery example comes from the system built in
Core Banking Architecture and deployed in
Deploying The Core Bank. The principles are
general; the concreteness is deliberate, because reliability advice without a
system attached is indistinguishable from a poster.