Part 10 · 14 chapters · ~55 min
Engineering management: when you own the outcome
A different job rather than a promotion, and the hardest adjustment is the feedback loop: as an engineer you know by Friday whether the week worked, and as a manager you may not know for two quarters. This part is about the mechanisms that actually move a team, and about the fact that every practice in the preceding ten parts depends on people telling the truth about the system.
95
The transition, and what you give up
the question
“Should I move into management?”
| Staff engineer | Engineering manager | |
|---|---|---|
| Primary lever | Technical judgement | People and structure |
| Output | Decisions and designs | A team that produces |
| Feedback loop | Days to weeks | Months to quarters |
| Hardest part | Ambiguity | Situations with no good option |
| What you lose | Deep focus, and being the one who builds it | |
| What you gain | Shaping who does what, and how it feels to work here | |
| Reversible? | Yes, and it gets harder each year |
the honest reasons to do it, and not to
- Good reason: you find the people and structure problems genuinely interesting, and you want to change how a team works rather than what it builds.
- Good reason: you are already doing most of it informally and want the authority to do it properly.
- Bad reason: it is the only way to progress. If your organisation has no staff track, that is information about the organisation.
- Bad reason: you are frustrated with your manager and think you would do better. Possibly true, and not a foundation.
- Bad reason: status. The job is largely invisible and its wins are other people's wins.
what nobody says out loud
The feedback loop is the hardest adjustment. As an engineer you know by Friday whether the week worked. As a manager you may not know for two quarters whether a hiring decision, a team split or a coaching conversation was right. Engineers who need frequent evidence that they are effective find this genuinely difficult, and it is worth knowing about yourself before making the move.
96
Team topology, and Conway’s Law as a tool
worked numbers
Conway's Law: systems mirror the communication structure of the organisation that builds them. most people quote it as a warning. the useful version is that it works in reverse: design the team structure you want the architecture to have. one team per bounded context → services with clean seams one shared team for everything → a distributed monolith
| Team type | Owns | Example here |
|---|---|---|
| Stream-aligned | A product flow, end to end | Payments, lending, cards |
| Platform | Paved roads other teams consume | Deployment, observability, the ledger SDK |
| Enabling | Capability uplift, temporarily | A security team embedding for a quarter |
| Complicated subsystem | Something genuinely specialist | The ISO 8583 gateway, the risk models |
applying it to our bank
The ledger is a complicated subsystem and a platform simultaneously, which is why Part 14 of the CBA module gave it one narrow posting API and refused to let product knowledge in. That API is an organisational boundary as much as a technical one: it is what allows five product teams to ship without coordinating with the ledger team. If the boundary is wrong, the teams will feel it long before the architecture diagram shows it.
97
Hiring for a money system
what actually predicts success here
- Care about correctness. Does the candidate think about the failure case unprompted? In a ledger this matters more than raw speed.
- Comfort with ambiguity. Most of the hard problems in this module had no single right answer.
- Ability to explain. They will write design documents and review others. An engineer who cannot explain a tradeoff cannot lead one.
- Curiosity about the domain. The people who thrive in fintech find double-entry interesting rather than tedious.
- Evidence of learning from being wrong, which is the single best predictor of growth and is easy to probe.
what the interview should and should not be
- Should: a realistic problem resembling the actual work, discussed collaboratively, where the candidate can ask questions.
- Should: a real code sample or a pairing session on something like production code.
- Should not: algorithmic puzzles with no relationship to the job, which mostly measure recent interview preparation.
- Should not: a take-home larger than four hours, which selects for people with free time rather than for skill.
- Should: be consistent across candidates, with a written rubric, because unstructured interviews mostly measure similarity to the interviewer.
the thing managers underweight
The interview is also the candidate evaluating you, and the strongest candidates have options. A process that is slow, disorganised or rude loses people you wanted, and you never find out why. Speed and clarity in a hiring process are a competitive advantage that costs nothing.
98
Onboarding into a system this size
worked numbers
a bank with 20 services, a ledger, six rails and four regions is not learnable by reading the code. time to first meaningful change, typical: unstructured: 6 to 10 weeks, and much of it lost structured: 1 to 2 weeks that difference, across every hire, is the return on two days of someone writing an onboarding path.
what a good onboarding contains
- A working environment on day one. Not day five. Every hour lost here is pure waste and it sets a tone.
- A guided tour of the money path, following a single transfer through every component. The Part 20 trace is exactly this document.
- A small, real, shippable change in the first week. Scoped by their manager, reviewed properly, deployed by them.
- A named buddy, distinct from the manager, for the questions people do not want to ask their manager.
- The written-down things: architecture decision records, the strategy, the scorecard, the runbooks. Onboarding is the forcing function that keeps them current.
- On-call shadowing before on-call. Two rotations observing, then primary with a strong secondary.
the diagnostic value
Onboarding time is the most honest measurement of a system's legibility, per Part 5. If it is getting worse, the codebase is getting harder to understand and every other estimate in the organisation is quietly inflating. It is one of the few metrics that improves reliably when someone simply pays attention to it.
99
One-to-ones that are not status meetings
what the time is for
- Their agenda first. If you fill it, you have a status meeting and they will stop bringing anything real.
- The things that do not fit elsewhere: frustration, ambiguity about direction, interpersonal friction, career questions.
- Feedback in both directions, and asking for it explicitly, because most people will not volunteer criticism upward.
- Context transfer. What you know about where the company is going, that they cannot see from where they sit.
- Not status. Status belongs in writing where everyone can read it.
the questions that produce real conversation
- "What is frustrating you that I could remove?" Concrete and actionable.
- "What are you learning?" A flat answer over several months is an early signal of disengagement.
- "What do you want to be doing in a year that you are not doing now?" The retention question, asked early enough to act on.
- "What am I getting wrong?" Uncomfortable, and the single highest-value question you can ask.
- Silence. Ask, then wait. The important thing is usually said after a pause.
the failure mode
Cancelling them when busy. It is the most common thing managers do and it communicates precisely what it looks like. Thirty minutes every two weeks, protected, is the floor, and the weeks you are too busy are the weeks the conversation matters most.
100
Feedback, performance, and the hard conversation
the rules for feedback that lands
- Specific and behavioural. "In the design review you interrupted twice before the proposal was finished" rather than "you can be dismissive".
- Close to the event. Feedback in a performance review about something from four months ago is useless and feels like an ambush.
- Separate the observation from the interpretation. "Here is what I saw, here is how it landed for me, is that what you intended?" leaves room for a different explanation.
- Positive feedback is equally specific. "Good job" teaches nothing. "The way you scoped that migration into four reversible steps is why it went well" teaches something repeatable.
the hard conversation, done properly
- Nothing in a formal review should be a surprise. If it is, the failure is the manager's.
- Be direct about the gap and the consequence. Softening it to the point of ambiguity is unkind, because the person does not know they are in trouble until it is too late to fix.
- Be specific about what changed behaviour looks like, with a timeframe and support.
- Write it down afterwards, and share it, so both people have the same record.
- Follow through. A stated consequence that does not happen destroys your credibility with everyone watching, which is the whole team.
the principle underneath
Clarity is kindness. Managers avoid hard conversations because they feel cruel, and the actual cruelty is letting someone continue failing without knowing, until the outcome is decided and irreversible. The uncomfortable conversation in month two is a favour; the one in month ten is a formality.
101
Planning that survives contact with reality
what makes a plan useful rather than decorative
- Plan outcomes, not tasks. "Reduce transfer p99 below 250 ms" survives learning; "implement connection pooling" does not.
- Leave slack, explicitly. A plan at 100% of capacity fails on contact with the first incident. 70 to 80% committed is realistic, and the rest is absorbed by reality.
- Include the maintenance budget, per Part 8, as a line item rather than as an assumption.
- Fewer things. Three outcomes delivered beats eight started.
- Re-plan on a rhythm. A quarterly plan reviewed monthly is a tool; one reviewed at the end is a report card.
worked numbers
where the capacity actually goes, measured: planned feature work ~50% maintenance and upgrades ~20% incidents and support ~15% reviews, interviews, meetings ~15% planning as if the first row is 100% is why every plan slips, and the other three rows do not stop happening because you left them out of the spreadsheet.
102
Metrics for engineering teams, used carefully
| Metric | Useful for | Dangerous when |
|---|---|---|
| DORA four | Trend of delivery health | Compared between teams with different contexts |
| Incident count and MTTR | Reliability trend | Used to discourage declaring incidents |
| Page volume | On-call health | Gamed by silencing alerts |
| Cycle time | Finding bottlenecks | Applied to individuals |
| Lines of code, commits, PRs | Nothing | Always |
| Onboarding time | System legibility | Rarely dangerous, rarely measured |
Goodhart’s law, which is not optional
When a measure becomes a target, it ceases to be a good measure. Measure deploy frequency and you get more, smaller, emptier deploys. Measure incidents and you get fewer declared incidents rather than fewer incidents. The defence is to use metrics for direction rather than for evaluation, and never to attach them to individual performance, which Part 0 said about the scorecard and which applies doubly to people.
how to use them honestly
- Team level, never individual.
- Trend, not absolute, and compared against the team's own history rather than another team's.
- As a prompt for a conversation, not as a verdict. "Lead time doubled, what changed?" is the use.
- Paired with qualitative input. The team usually knows why before the metric does.
103
Protecting focus, and the cost of interruption
worked numbers
a deep-focus task needs roughly 20 to 30 minutes to re-enter
after an interruption.
4 interruptions across a day = ~2 hours of recovery
meetings at 10:00 and 14:00 = three fragments,
none long enough for hard work
calendar shape matters more than calendar volume.what a manager can actually do
- Cluster meetings. Four meetings in one block costs far less than four spread across a day.
- Protect at least two days a week with no scheduled meetings for the team, as policy rather than aspiration.
- Take the interrupt yourself. Being the buffer between the team and the rest of the organisation is one of the most valuable things a manager does and one of the least visible.
- Rotate the support and interrupt duty, per Part 7, so it lands on one person rather than fragmenting everyone.
- Default to asynchronous. A written question answered in two hours usually beats a meeting tomorrow, and costs nobody their afternoon.
104
Incidents, blame, and psychological safety
what psychological safety actually is, and is not
- Is: the belief that you can raise a problem, admit a mistake or disagree without it being held against you.
- Is not: absence of accountability, low standards, or everyone being nice. High-safety teams frequently have very high standards, which is what makes them work.
- The measurable consequence: how quickly bad news travels upward. In a low-safety team, problems surface late, which in a bank means they surface as incidents.
what a manager does that builds or destroys it
- How you respond to the first mistake sets the pattern for everything after. Everyone is watching, whether or not you notice.
- Admit your own errors, specifically and unprompted. It is the single most effective signal available.
- Ask for criticism and then act on some of it visibly, because asking and ignoring is worse than not asking.
- Defend the team externally, and address problems internally. Reversing that is how trust is destroyed in one meeting.
- Never let a blameless postmortem become blame with better manners, per Part 6. The mechanism is asking what made the wrong action reasonable.
the connection to the rest of this module
Every mechanism in this module depends on people telling the truth about the system. Scorecards, postmortems, error budgets and threat models are all worthless if the reported numbers are managed rather than measured. Psychological safety is not a soft topic sitting beside the engineering; it is the precondition for the engineering practices working at all.
105
Budget, headcount, and arguing for both
how to make the argument
- In outcomes, not headcount. "Two engineers on support tooling removes 30% of escalations, which is worth ₦x and frees a third of a person permanently" beats "we need two more people".
- With the counterfactual. What happens if we do not: which roadmap items slip, which risks stay open, what the on-call load becomes.
- Costed against alternatives. Managed service versus engineers, contractor versus hire, buy versus build. Showing you considered them is what makes the ask credible.
- With a ramp. A new hire is net negative for two to three months and everyone knows it; saying so makes the rest of your estimate believable.
- Early. Headcount is usually decided in a planning cycle you are not in the room for, and arguing after it is decided is too late.
the infrastructure version
The cloud module priced the platform at roughly $180k a month. That is a budget line somebody defends, and if it is not you, it will be cut by someone who does not know which 20% is load-bearing. Part 19's unit economics is the argument: cost per transaction, trending, with the drivers named. An engineering manager who cannot explain their infrastructure bill will eventually be told to halve it.
106
Retention, and why people actually leave
the real reasons, roughly in order
- Their manager. Consistently the largest single factor, and the one most managers discount.
- No growth. Doing the same thing for two years with no expanding scope. This is the most preventable one.
- Toil. On-call load, support interruption, and maintenance that never ends. Part 6 and Part 8 are retention topics as much as engineering ones.
- Not being heard. Raising the same problem repeatedly and watching nothing change.
- Compensation, which matters and is usually the stated reason rather than the actual one.
- Life. Sometimes people leave for reasons that have nothing to do with you, and treating every departure as a failure is its own error.
the intervention that works
Ask the retention question before they are leaving. "What would make you want to still be here in two years?" asked in a one-to-one, acted on where possible. By the time someone resigns, the decision is usually three months old and counter-offers mostly delay it. The signals are visible earlier: flat answers about learning, disengagement in design discussions, and the quiet withdrawal from things they used to care about.
107
Managing in a regulated environment
what is different
- Some decisions are not yours. A regulator's requirement is a constraint, not an input to a tradeoff, and the team needs to understand that distinction or it feels arbitrary.
- Evidence matters as much as outcome. Doing the right thing and being unable to demonstrate it is a finding. This is genuinely frustrating for engineers and needs explaining rather than enforcing.
- Change control is real. Part 5 of this module argued for fast deployment; a regulated environment adds approval and audit trails. The goal is making the controls automatic rather than manual, so speed and evidence are not in conflict.
- Four-eyes and separation of duties constrain how you allocate work, per Part 3.
- Your team will be audited, and preparing them for that, including that it is normal and not an accusation, is part of the job.
the framing that helps
Compliance is a design constraint like latency, not an enemy. The CBA module treated residency, retention and audit as inputs that shaped the architecture, and the result was better rather than worse: append-only ledgers, immutable audit logs and reproducible reports are good engineering that compliance happened to require. Teams that frame regulation as obstruction fight it continuously; teams that frame it as a constraint design around it once.
108
The manager’s version of technical judgement
what you still need to be able to do
- Tell a good design review from a bad one, without doing the design yourself.
- Ask the question that reveals the gap. "What happens when the provider times out?" is a manager's question and it is worth more than a week of oversight.
- Judge an estimate's honesty, per Part 9. Not its accuracy, its honesty.
- Know which risks are existential and which are inconvenient. In this system: correctness is existential, availability is serious, latency is a quality problem.
- Recognise when the team is wrong and let them proceed anyway, because the cost of being overruled exceeds the cost of the mistake more often than managers believe.
what you must stop doing
- Making the technical decision yourself because you would be faster. It is true and it is corrosive.
- Reviewing everything. You become a bottleneck and you signal distrust.
- Keeping the interesting work. Same trap as Part 9 chapter 90, with more authority to indulge it.
- Pretending to know. Your credibility rests on being accurate about what you do and do not understand.
the closing thought
Technical judgement does not disappear in management; its application changes. You stop using it to choose the design and start using it to judge whether the process that produced the design was sound, whether the risk is correctly classified, and whether the person presenting it has thought about the failure case. That is a real skill, it decays if you stop engaging with the system entirely, and keeping enough contact to retain it without taking the work back is the balance nobody teaches.