12 parts · 112 chapters

Running It, And The Team

The system exists and it is deployed. Now it has to stay reliable, secure, fast, observable and changeable for years, while people who did not build it operate it at 3am and a different set of people extend it. None of those properties has a feature, a ticket or an obvious owner, which is exactly why they decay.

The first half treats each property as an engineering discipline with mechanisms rather than as an aspiration. The second half is the human system that keeps them alive: on-call, support, technical debt, leading the work, and owning the outcome.

non-functionals · people
reliabilityDesigning for the failure you will actually get, rather than the one in the diagram.
performanceLatency and throughput as separate problems, and why the tail is what customers feel.
securityThreat models instead of checklists, and the insider threat that is the real one.
observabilityBeing able to ask a question nobody anticipated, bounded by cardinality and cost.
adaptabilityThe property that decays silently, and the only one that makes the others fixable.
supportabilityOn-call, support volume as a product signal, and the loop back into priorities.
sustainabilityTechnical debt named, priced and paid, and maintenance budgeted rather than smuggled.
00

Non-functionals are the job

Why the interesting requirements have no feature · The seven properties, and how they conflict · Making a property measurable, or abandoning it · Who owns a property nobody asked for · Budgets: latency, error, cost, and change · The conversation that turns a wish into a target · When to deliberately under-invest · The scorecard, and using it without theatre
8 ch · ~50 min
01

Reliability: designing for the failure you will get

Availability arithmetic, and why nines mislead · Failure modes: crash, slow, wrong, and partitioned · Redundancy that actually helps · Timeouts, retries, and the retry storm · Circuit breakers, bulkheads, and load shedding · Backpressure as a first-class design element · Idempotency as a reliability primitive · Chaos engineering, proportionate to the system · The reliability review, as a checklist
9 ch · ~55 min
02

Performance: latency, throughput, and the tail

Latency and throughput are different problems · Little’s Law, and using it in a room · Why the tail matters more than the mean · Where latency actually goes: a budget, itemised · Queueing theory, the useful ten percent · The four things that make software slow · Measuring before optimising, properly · Load testing that resembles production · Capacity planning, and headroom as policy · Performance as a regression test
10 ch · ~55 min
03

Security: threat models, not checklists

Threat modelling in forty minutes · The trust boundaries in our bank · Authentication, authorisation, and the gap between · Secrets: the lifecycle, not the store · Cryptography choices a non-specialist must get right · The insider threat, which is the real one · Supply chain: dependencies and build integrity · Security in the pipeline, not after it · Incident response when it is an attack · The security review, as a checklist
10 ch · ~55 min
04

Observability: being able to ask new questions

Monitoring answers known questions; observability answers new ones · The three signals, and the fourth for money · Cardinality: the constraint that shapes everything · Structured events over metrics, where it matters · Sampling strategies that keep the interesting traces · Dashboards people actually use · Alerting: the rules that stop alert fatigue · Debugging a system you cannot reproduce · The cost of observability, and containing it
9 ch · ~45 min
05

Adaptability: changing it without breaking it

The property that decays silently · Coupling, cohesion, and where change propagates · Contracts, versioning, and expand-contract · Database migrations on a live money path · Feature flags, and their half-life · Deployment strategies, and rollback as a first thought · Testing: the pyramid, and where it lies to you · Testing money: property tests and invariant checks · Making the codebase legible to the next person · Deletion as an engineering discipline
10 ch · ~50 min
06

On-call: the human side of reliability

What on-call is actually for · Rotation design, and the humane constraints · Severity levels that mean something · The page budget, and treating it as a limit · Handover, and the shift report · Incident command: roles during an incident · Communicating during an incident, internally and out · The postmortem, and making blamelessness real · Follow-up actions that actually get done · Measuring on-call health
10 ch · ~50 min
07

Support: the feedback loop you are ignoring

Support volume as a product signal · Tiering, escalation, and the engineering interface · Tooling: what agents need and never get · The read model for support, and its access controls · Contact-driven development · Deflection, and when it is hostile · Complaints, regulators, and the paper trail · Closing the loop into engineering priorities
8 ch · ~40 min
08

Technical debt: naming it, pricing it, paying it

The metaphor, and where it breaks down · A taxonomy: deliberate, accidental, and rot · Pricing debt in terms the business understands · The interest rate: how to tell what is urgent · Refactoring strategies that survive review · The rewrite, and when it is genuinely right · Migration as a permanent condition · Budgeting maintenance, explicitly · Making decay visible
9 ch · ~45 min
09

Leading the work: staff engineering in practice

What changes between senior and staff · Technical strategy, written down · The design document that gets read · Reviewing designs and code as leverage · Influence without authority, concretely · Disagreeing well, and committing after · Growing other engineers deliberately · Working with product and with finance · Estimation, and being honest about uncertainty · Choosing what not to work on · Visibility without self-promotion
11 ch · ~50 min
10

Engineering management: when you own the outcome

The transition, and what you give up · Team topology, and Conway’s Law as a tool · Hiring for a money system · Onboarding into a system this size · One-to-ones that are not status meetings · Feedback, performance, and the hard conversation · Planning that survives contact with reality · Metrics for engineering teams, used carefully · Protecting focus, and the cost of interruption · Incidents, blame, and psychological safety · Budget, headcount, and arguing for both · Retention, and why people actually leave · Managing in a regulated environment · The manager’s version of technical judgement
14 ch · ~55 min
11

The operating picture

The seven properties, and their mechanisms · A year in the life of this system · What to fix first, in any system you inherit · Defending it: the fifteen hardest questions
4 ch · ~35 min
Written against the bank from the other modulesEvery example comes from the system built in Core Banking Architecture and deployed in Deploying The Core Bank. The principles are general; the concreteness is deliberate, because reliability advice without a system attached is indistinguishable from a poster.