Part 8 · 9 chapters · ~45 min
Technical debt: naming it, pricing it, paying it
Calling everything technical debt flattens three different things with three different remedies, which is why a debt sprint rarely works. This part separates deliberate shortcuts from accidental complexity from rot, prices each in terms a business will fund, and argues that the skill is not avoiding debt but ranking it by interest rate and leaving the cheap parts alone permanently.
75
The metaphor, and where it breaks down
worked numbers
the metaphor: shipping imperfect code borrows time,
and you repay with interest later.
where it works:
· it makes the cost ongoing rather than one-off
· it makes deliberate borrowing legitimate
· it gives a financial frame the business understands
where it breaks:
· real debt has a known principal and rate. this does not
· real debt can be ignored without growing worse. this compounds
· "debt" implies someone chose to borrow, and most of it
was never a decision at allwhy the distinction matters practically
Calling everything "tech debt" flattens three different things with three different remedies. A deliberate shortcut needs scheduling. Accidental complexity needs learning. Rot needs continuous maintenance. Treating all three as one backlog item is why "we will do a debt sprint" almost never works.
76
A taxonomy: deliberate, accidental, and rot
| Type | Origin | Remedy | Urgency signal |
|---|---|---|---|
| Deliberate, prudent | "Ship now, fix after launch" | Schedule it, and actually do it | The date you set |
| Deliberate, reckless | "No time for tests" | Stop doing it. Cultural | Change failure rate |
| Accidental | "Now we understand the domain better" | Refactor as you touch it | Lead time growth |
| Rot | The world moved, the code did not | Continuous maintenance | CVEs, EOL dates |
| Bit rot in knowledge | The people who understood it left | Documentation, pairing | Bus factor |
the one that is not optional
Rot. A dependency reaching end of life, a TLS version being deprecated, a cloud service being retired: these have external deadlines and no amount of prioritisation changes them. Rot work is not debt repayment, it is maintenance, and budgeting it separately from improvement work is what stops it being perpetually deferred until it becomes an emergency.
77
Pricing debt in terms the business understands
worked numbers
"the payments module is messy" → loses to a feature, always "a change to payment routing takes 9 days instead of 2, and we make roughly 20 such changes a year. that is 140 engineer-days annually, and it is growing." the second is a business case. the first is an aesthetic complaint.
the pricing method
- Identify the recurring cost. Slower changes, more incidents, more support load, more onboarding time.
- Measure it, even roughly. Lead time on that component versus elsewhere. Incident count attributable to it. An estimate you can defend beats a feeling.
- Multiply by frequency. A tax paid twenty times a year is a different argument from one paid twice.
- Estimate the fix and the payback period. "Six weeks of work, saving 140 days a year" is a conversation that ends quickly.
- Name what happens if it is deferred another year, because compounding is the part the metaphor gets right.
78
The interest rate: how to tell what is urgent
| Signal | Interest rate | Example |
|---|---|---|
| Touched constantly | High | The posting path everyone changes |
| Touched rarely | Near zero | An ugly script run twice a year |
| Blocks other work | Very high | A schema that prevents a whole feature area |
| Causes incidents | Very high | It shows up in postmortems repeatedly |
| Security or compliance | External deadline | EOL dependency, deprecated TLS |
| Blocks hiring or onboarding | High | Nobody can work on it without a long ramp |
worked numbers
the prioritisation rule that follows: fix what you touch, in proportion to how often you touch it. ugly code nobody opens has an interest rate near zero, and refactoring it is a cost with no return. "it offends me" is not an interest rate.
the corollary engineers resist
Some bad code should be left alone permanently. A component that works, is rarely changed and is not a security concern has a low interest rate regardless of how it looks. Engineering judgement includes deciding what not to fix, and a team that refactors on aesthetics rather than on interest rate spends its credibility on work the business correctly perceives as low value.
79
Refactoring strategies that survive review
what works, in order of safety
- The boy scout rule. Improve what you touch, slightly, in the same change. No separate permission needed and it compounds.
- Refactor then change, as two commits. First make the change easy, then make the easy change. The review is comprehensible because each commit does one thing.
- Strangler, for larger surfaces. New implementation alongside, route incrementally, delete the old. The cloud module used exactly this for the migration.
- Branch by abstraction. Introduce a seam, put both implementations behind it, switch, remove. Safer than a long-lived branch by a wide margin.
- Never: the long-lived refactor branch. It diverges, the merge is enormous, and it is usually abandoned. This is the most common way refactoring effort is wasted.
the review point
A pull request that mixes refactoring and behaviour change cannot be reviewed properly, because the reviewer cannot tell which diff lines are supposed to change behaviour. Separating them is not bureaucracy; it is what makes the review capable of catching the bug. Two commits, or two pull requests, always.
80
The rewrite, and when it is genuinely right
why rewrites usually fail
- The old system encodes years of edge cases that nobody remembers and nobody documented. The rewrite rediscovers them one production incident at a time.
- The business does not stop. You are rewriting a moving target while the original team keeps shipping.
- Estimates are wrong by more than usual, because the unknown work is precisely the part nobody understands.
- No value until the end. A two-year rewrite delivers nothing for two years, and organisations lose patience at about month nine.
when it is genuinely right
- A fundamental constraint cannot be changed incrementally. Part 15 named multi-tenancy: retrofitting a tenant dimension into a sharded ledger is a rewrite by another name.
- The platform is ending. A runtime or database with no support path and a date.
- The component is small and well-bounded, so the rewrite is weeks rather than years.
- You can run both and compare. This is the decisive one: if the old system can validate the new one, the risk collapses. The cloud module migration did exactly this with reconciliation as the gate.
the test
"Can I run both in parallel and prove equivalence before cutting over?" If yes, it is a strangler migration with manageable risk. If no, it is a big-bang rewrite and the historical success rate is poor. That single question is worth more than any amount of debate about whether the old code is bad.
81
Migration as a permanent condition
worked numbers
the fantasy: migrate, finish, be modern the reality for any system that lives: a runtime version upgrade every year a framework major every two a database version every three a cloud service deprecation whenever they decide an internal platform change, constantly there is no steady state. there is only the current migration.
what follows from accepting that
- Budget it permanently, not per project. Part 0 measured maintenance share and Part 8 argues for a floor.
- Invest in migration capability. Good contracts, feature flags, expand-contract habits and parallel-run tooling make every future migration cheaper.
- Never fall more than one major version behind. The cost of catching up grows superlinearly, and three versions behind is a project rather than a chore.
- Track EOL dates as a first-class backlog, with the date visible. These are the items with an external deadline and no negotiation available.
82
Budgeting maintenance, explicitly
worked numbers
the common failure: maintenance is smuggled into feature estimates → estimates look padded → pressure to reduce them → maintenance gets squeezed out silently → and nobody ever decided to stop doing it the alternative: a stated share of capacity, defended 20% is a common and defensible number the point is that reducing it becomes a DECISION, taken by someone, rather than an accident of estimation.
what the budget covers, and what it does not
- Covers: dependency upgrades, EOL migrations, refactoring on the high-interest paths, tooling improvements, and postmortem follow-ups.
- Does not cover: incident response, which is unplanned and separate, or feature work dressed up as refactoring.
- Is protected: it does not get borrowed for a launch, because borrowing it once establishes that it is borrowable.
- Is reported: what was done with it, at the monthly review, so it is visibly producing value rather than disappearing.
the number to watch
Part 0's scorecard had maintenance share at 12% against a 20% target, alongside four rising trend lines. That is the causal story of a system degrading, and the maintenance number is the leading indicator. Lead time and page count are lagging: by the time they move, the under-investment is already a year old.
83
Making decay visible
the instruments that make an invisible property visible
- Lead time for change, per component. Divergence between components localises the problem.
- Change failure rate. Rising means the safety net is degrading.
- Incidents attributable to a component, from postmortems, which turns anecdote into a count.
- Dependency age and EOL distance, which is the easiest to automate and the most ignored.
- Onboarding time to first change, which measures legibility per Part 5.
- Maintenance share of capacity, which is the leading indicator for all of the above.
the closing argument for this part
Technical debt is not a moral failing and it is not free. It is a recurring cost that compounds, and the engineering skill is not avoiding it but pricing it, ranking it by interest rate, and paying down the expensive parts deliberately while leaving the cheap parts alone forever. A team that refactors everything and a team that refactors nothing are both failing, and the difference between them and a good team is measurement rather than discipline.