Part 13 · 11 chapters · ~20 min

Round thirteen: “run it at 3am”

Twelve rounds have produced a system with dozens of components, six external dependencies and four regions. Now it is 03:12 and something is wrong. This round is about the difference between a system that can be operated and one that merely works, and the distinguishing question is whether a tired person with no context can tell what is broken and what to do about it.

141

The pressure: it is broken and you are asleep

interviewer

“It is 3am. Transfers are failing for some customers but not others. You are on call and you know nothing else. Walk me through the next ten minutes.”

This question is not about monitoring tools. It tests whether the design produces answers, and the honest measure of an architecture is how fast it can be diagnosed by someone who did not build it.

what the first ten minutes must establish
  1. Is it real? A customer-visible failure, or a noisy check?
  2. How bad? What fraction of transfers, which regions, which rails, trending which way?
  3. Is money at risk? The only question that changes the urgency of everything else.
  4. What changed? A deploy, a config flag, a provider, a traffic shift.
  5. Can I stop the bleeding without a full diagnosis?
the ordering that matters
Mitigate before you diagnose. At 3am the goal is to stop customer harm, not to understand the bug. A kill switch that disables one rail in ten seconds is worth more than a perfect root cause at 5am. Understanding is a daytime activity, and designing for that separation is what chapters 148 and 149 are about.
non-functional, new
  1. Every customer-facing flow has a named SLO with an error budget.
  2. Alerts fire on symptoms customers feel, never on causes.
  3. Any transfer can be traced end to end from a single identifier.
  4. Every paging alert has a runbook reachable from the alert itself.
  5. Degradation is designed, with a stated order of what is shed.
  6. Recovery from region loss is drilled, not documented.
142

SLI, SLO, error budget for a money path

worked numbers
SLI  a measurement of one user-visible property

               "fraction of transfers completing in under 300 ms"


SLO  a target for that measurement over a window

               "99.9% of transfers, measured over 30 rolling days"


error budget  what the SLO permits you to spend

               100% − 99.9% = 0.1% × 300M transfers/month

               = 300,000 failed transfers a month, allowed
FlowSLISLOWhy this number
Balance readp99 latency99.95% under 100 msOn every screen. Cheap to serve from cache, so the target is high
Internal transferSuccess rate and p9999.9% under 300 msFully within our control, so we own every failure
Card authorisationAvailability99.99% under 500 msThe network answers for us if we time out. Highest target in the bank
Outbound to a railReaching a terminal state99.5% within 1 hourDepends on providers we do not control, so the target is honest about that
NotificationDelivery within 10 s99% p95Deferrable, and one channel is always available
Ledger correctnessInvariant violationsZero. No budget.Not an SLO. A correctness property with no acceptable failure rate
the distinction that matters most in a bank
Availability gets an error budget. Correctness does not. A transfer that fails is inside the budget and the customer retries. A transfer that posts the wrong amount is not a budget item at any rate, because there is no tolerable frequency for creating money. Separating budgeted SLOs from unbudgeted invariants is the single most important idea in this chapter, and most generic SRE material does not make it.

What the budget is actually for

the error budget as a decision-making tool
  1. Budget remaining: ship features, take risks, run experiments.
  2. Budget exhausted: the team stops feature work and spends on reliability until it recovers. Agreed in advance, so it is not a negotiation during an incident.
  3. Budget consistently unused: the SLO is too loose or you are over-investing in reliability. Both are worth knowing.
  4. Burn rate alerts beat threshold alerts: burning a month's budget in an hour pages immediately, and burning it slowly over a week raises a ticket.
143

The four golden signals, and the fifth for banks

SignalQuestionWhere it bites here
LatencyHow long do requests take?Split successful from failed. A fast failure and a slow success are different problems and averaging them hides both
TrafficHow much demand?Transfers/s, authorisations/s. A sudden drop is as alarming as a spike
ErrorsWhat fraction fails?By type: insufficient funds is not an error, a provider timeout is
SaturationHow full is the system?Connection pools, queue depth, outbox lag, consumer lag, disk
CorrectnessIs the money right?The fifth signal, and the one only a bank has. Invariant violations, suspense ageing, reconciliation breaks
the fifth signal is the answer to a common interview question
Asked "what would you monitor", nearly everyone recites the four golden signals. Adding correctness as a first-class signal, and pointing at the Part 10 checks as its implementation, demonstrates that you understand what makes this domain different. A bank can be fast, available and completely wrong, and the first four signals would all be green.

The correctness dashboard

code
# the five metrics that say "the money is right", from Part 10
ledger_invariant_violations_total          # must be 0. any value pages.
ledger_suspense_oldest_age_seconds{kind}   # stuck operations
ledger_suspense_balance_minor{kind}        # value at risk right now
recon_unresolved_break_value_minor{cp}     # by counterparty
recon_oldest_break_age_hours{cp}           # age, not count

# and the pipeline health that correctness depends on
outbox_oldest_unpublished_age_seconds{shard}
kafka_consumer_lag_seconds{group}          # seconds, not records
posting_state_accepted_not_posted_count
alert on symptoms, not causes
  1. Bad: "CPU above 80% on ledger-7". Might be fine, might be a nightly job, and it does not tell you whether anybody is affected.
  2. Good: "Transfer success rate below 99.5% for 5 minutes in region ng". A customer is feeling this right now.
  3. Cause metrics belong on dashboards for diagnosis, and symptom metrics belong on pages.
  4. The test: if this fires and no customer is affected, it should not have paged. Alerts that fail this test train people to ignore alerts.
144

Business metrics that page

The most valuable alerts in a payments system are business metrics, because they detect whole categories of failure that no technical metric sees.

MetricDetectsComparison basis
Transactions per minuteA silent path failure, an app release bug, a rail outageSame time last week, never an absolute threshold
Total value per minuteHigh-value flows failing while small ones succeedSame time last week
Success rate by railA single provider degradingPer-rail baseline
Authorisation approval rateA fraud rule that is too aggressive. This is the alert that catches your own mistakeRolling baseline per merchant category
New account creationsOnboarding broken, or an automated attackSame time last week, both directions
Support contact rateEverything your monitoring missedRolling baseline
why "same time last week" and not a threshold
Transaction volume has strong daily and weekly seasonality, plus salary-day spikes. A fixed threshold either pages every night at 3am when volume legitimately drops, or is set so low it never catches a real 40% drop at midday. Compare against the same hour of the same weekday, and alert on the deviation. It is a small implementation detail that makes the difference between an alert people trust and one they mute.

The approval-rate alert, which watches us rather than the customers

code
# Part 9 built automated blocking with a rate limit on itself.
# this is the business-level counterpart: if approval rate falls,
# the most likely explanation is OUR rule, not a fraud wave.
- alert: AuthApprovalRateDrop
  expr: |
    (sum(rate(auth_decisions_total{outcome="approved"}[10m]))
     / sum(rate(auth_decisions_total[10m])))
    < 0.92
  for: 5m
  labels:   { severity: page }
  annotations:
    summary: "Authorisation approval rate {{ $value | humanizePercentage }}"
    runbook: "https://runbooks/auth-approval-drop"
    # the first line of that runbook: check for a recently activated
    # fraud rule, and be ready to use the kill switch.
the silent failure
every technical signal green, transaction volume collapsed
swipe the figure sideways, or tap expand for full screen
1/7
normal
Normal operation. Four golden signal panels: latency, error rate, CPU and saturation.
145

Distributed tracing across the posting path

A transfer now touches a gateway, the ledger service, a shard, an outbox relay, Kafka, four consumers, a provider and a notification channel. Without tracing, diagnosing a slow transfer means correlating logs across nine systems by hand.

worked numbers
trace  one end-to-end operation. one trace id.

span   one unit of work within it, with a parent.

context propagation  passing the ids across every boundary


          the boundaries that are usually missed:

            · into Kafka      → trace id in message headers

            · out of the outbox → store it in the outbox row

            · into a batch job   → a trace per item, linked to the run

            · across a webhook  → correlate on our own reference


a trace that stops at the queue boundary is the common failure,

and the async half is exactly where the latency usually hides.
code
-- the outbox row carries the trace context, so the asynchronous
-- half of the operation joins the same trace as the synchronous half.
ALTER TABLE outbox
  ADD COLUMN trace_id TEXT,
  ADD COLUMN span_id  TEXT;

-- and the relay restores it when publishing:
--   headers: { traceparent: `00-${trace_id}-${span_id}-01` }
-- so a notification sent 4 seconds later appears under the same trace
-- as the transfer that caused it.
sampling, which must be biased rather than uniform
  1. Always sample anything that errored. Errors are rare and maximally informative.
  2. Always sample anything slower than the SLO threshold, since those are the traces you will want.
  3. Always sample high-value transactions above a threshold.
  4. Sample 0.1% of ordinary successful transfers, for baselines.
  5. Prefer tail-based sampling: decide after the trace completes, when you know whether it was interesting. Costs buffering, and is worth it.
the identifier that makes everything else work
One identifier must appear in every log line, every span, every event and every external call for a given operation. We already have a natural candidate: the journal_id. Using the business identifier as the correlation key means support, engineering and reconciliation are all searching the same string, and a customer quoting a transaction reference lands you directly on the trace.
146

Structured logs, correlation, and PII

code
// unsearchable, unaggregatable, and it leaks a name and an amount
logger.info(`Transfer failed for Adeniji: 50000 to 0123456789`);

// structured, correlatable, and PII-free
logger.info('transfer.rejected', {
  journal_id: 'j-8f1a',        // the correlation key
  trace_id: '4bf92f35...',
  account_id: 'acct-771',      // an internal id, not a NUBAN
  reason_code: 'INSUFFICIENT_FUNDS',
  amount_bucket: '10k_50k',     // bucketed, not exact
  currency: 'NGN',
  shard: 7, region: 'ng', rail: 'internal'
});
Never in logsLog instead
Full PAN, CVV, PINNothing. Not even masked, per Part 8
Customer name, phone, email, addressInternal customer id
Full account numbers, NUBAN, IBANInternal account id, or last 4
Exact balancesBucketed ranges, where needed at all
Auth tokens, session ids, API keysA hash prefix, if correlation is needed
Full request or response bodiesA field allowlist
why an allowlist rather than a denylist
A denylist redacts the fields somebody remembered. When a new field is added upstream, it is logged in full by default and nobody notices until an audit. An allowlist fails closed: a new field is omitted until someone deliberately permits it. Same instinct as the cross-region schema in Part 7, which had no field for data that may not travel: make the violation structurally impossible rather than forbidden by policy.
retention, split by purpose
  1. Debug logs: 7 days. High volume, low long-term value.
  2. Application logs: 30 to 90 days. Incident investigation.
  3. Audit logs of who did what to money: 7 years, append-only, separate store, different access controls. These are evidence rather than telemetry.
  4. Traces: 7 to 30 days for sampled, longer for errored.
147

The runbook, and what makes one useful

Most runbooks are useless because they describe the system rather than the decisions. A useful one is written for someone with no context, at 3am, under stress.

OutboxRelayLagHigh oldest unpublished outbox row older than 120 s · severity: page
why Events are not reaching consumers. Money is posting correctly, but notifications, fraud scoring and loan updates are all blind. Not a correctness problem yet; it becomes one if the lag exceeds Kafka retention.
1 Check the blast radius: outbox_oldest_unpublished_age_seconds by shard. One shard or all of them? One shard points at that database or relay instance; all of them points at Kafka or a bad deploy.
2 Check whether the relay is running at all, and whether it is erroring or idle. Idle with a backlog means it cannot claim rows, which usually means lock contention or an exhausted connection pool.
3 Check Kafka: is the broker accepting writes? An ISR below min.insync.replicas makes every publish fail, which is the Part 4 failure mode and it is deliberate.
4 Mitigate: if a recent deploy is implicated, roll back. If Kafka is degraded, page the platform team. Do not delete outbox rows to clear the backlog, ever: that loses events permanently and there is no recovery.
5 Escalate if lag exceeds 50% of Kafka retention. Beyond retention, consumers cannot be rebuilt from the topic and a replay from the ledger is needed, which is a much larger operation.
no Do not increase the polling rate to "catch up" without checking database load first: the usual cause is contention, and polling harder makes it worse.
the anatomy of a runbook that works
  1. What this means for customers, in one sentence, first. Everything else follows from severity.
  2. The blast-radius question before the diagnosis. One shard or all, one region or all.
  3. Mitigation before root cause. Explicitly permit rollback without understanding.
  4. The escalation threshold, as a number rather than "if it gets bad".
  5. An explicit do-not list. The destructive action somebody will reach for under pressure, named and forbidden.
  6. Linked from the alert, because a runbook nobody can find during an incident does not exist.
the item most runbooks omit
The do-not list. At 3am, deleting the backlog looks like a fix, and it destroys events irrecoverably. Naming the tempting destructive action and forbidding it explicitly is often the single highest-value line in the document, and it is the one written only after someone has already done it once.
148

Graceful degradation: what to shed first

Under load or partial failure, the system must give something up. Deciding the order in advance is the difference between degradation and an outage.

PriorityCapabilityShed when
P0, never shedCard authorisation. The ledger invariant. Posting correctness.Never. We fail requests before we compromise these
P1Balance reads, internal transfers, inbound fundingOnly under severe saturation, and by rate limiting rather than by error
P2Outbound payments to external railsQueue them rather than rejecting. Tell the customer it is pending, honestly
P3Transaction history, statements, searchServe stale from cache, with an explicit "as of" timestamp
P4Notifications beyond in-appDefer. In-app remains, per Part 12
P5, shed firstAnalytics refresh, marketing, recommendations, non-urgent batchImmediately. Nobody notices within an hour
the three ways to degrade, in order of preference
  1. Serve stale. A balance from 30 seconds ago with a visible timestamp beats an error, for reads.
  2. Queue and promise. Accept the instruction, return 202 with an honest pending state, and process when capacity returns. The journal-first design from Part 3 makes this possible.
  3. Reject cleanly. A fast, clear, retryable error beats a timeout. The worst outcome is a request that hangs for 30 seconds and then fails.
the sentence that captures the priority
"Under saturation I would shed analytics and marketing immediately, defer notifications, serve history stale, queue outbound payments, and rate-limit transfers. The two things I would never degrade are card authorisation, because the network answers for us if we go quiet, and posting correctness, because I would rather reject a transfer than record it wrongly. Every degradation step is a decision made before the incident, not during it."
149

Kill switches and feature flags on money

The 3am tool. A kill switch is how you stop customer harm in seconds without a deployment, and its design deserves the same rigour as the money path.

the switches worth having before you need them
  1. Per rail. Disable NIBSS outbound while leaving everything else running.
  2. Per provider. Route around one PSP without touching the rail.
  3. Per product. Suspend new loan disbursement while repayments continue.
  4. Per fraud rule. Deactivate one rule, which is the Part 9 counterpart.
  5. Automated blocking, globally. One flag stops all automated enforcement.
  6. Per region. Fail a region out entirely.
  7. New account creation. Stop an onboarding attack without affecting existing customers.
what makes a kill switch safe to use at 3am
  1. Effective in under 10 seconds, with a short cache TTL and no deployment.
  2. Fail-safe default. If the flag service is unreachable, the cached last-known value is used, never a default that silently enables something.
  3. Narrow scope. A switch that disables more than intended is not usable under pressure.
  4. Audited. Who flipped it, when, and why, in the append-only audit log.
  5. Reversible and tested. If turning it back on has never been exercised, you have a one-way door.
  6. Never affects correctness. A switch may stop accepting work; it may never change how money is recorded.
code
// the flag check on the money path: cached, fail-safe, and it can only
// ever REJECT work, never alter how a posting is recorded.
async function railEnabled(rail: string): Promise<boolean> {
  try {
    return await flags.get(`rail.${rail}.outbound`, { maxStaleMs: 10_000 });
  } catch {
    // the flag service is down. use the last known value from local
    // cache. NEVER default to enabled: that turns a flag outage into
    // re-enabling something an operator deliberately switched off.
    return localCache.lastKnown(`rail.${rail}.outbound`) ?? false;
  }
}
the constraint that keeps flags from becoming a hazard
A flag may change what we accept; it may never change how we record. A flag that switched between two posting schemes would mean the ledger's meaning depends on configuration at the time of the write, and reconstructing history would require knowing every flag's value at every instant. Flags gate the front door, never the accounting.
150

Disaster recovery: RTO, RPO, and the drill

worked numbers
RTO recovery time objective   how long until service resumes

RPO recovery point objective  how much data may be lost


          for our ledger:

            RPO = zero     ← non-negotiable. no acknowledged transfer is lost.

            RTO = 15 minutes ← for a regional failure


RPO zero is what forced synchronous replication in round one,

and the ~5 ms it costs on every commit. this is where that

decision is finally cashed in.
FailureDetectionResponseRTO
One application instanceHealth checkLoad balancer removes itSeconds
One shard's primaryReplication monitorPromote the synchronous replicaUnder 60 s
One availability zoneCloud provider plus our own checksTraffic shifts to remaining zonesMinutes
An entire regionMultiple signals, with a human decisionDeclared failover, with data residency constraints respected15 minutes
Data corruptionThe Part 10 invariant checksPoint-in-time recovery, then replayHours
A provider outageCircuit breakerRoute to the alternate, or queueSeconds
the regional failover constraint unique to this design
Part 7 established that a customer's region is a legal boundary. So failing the EU region over to Nigeria is not available to us, however convenient it would be during an outage. Regional resilience must therefore be built within each region, across availability zones, which is more expensive than a global active-active design. A legal constraint became an architectural cost, and naming that tradeoff explicitly is exactly the kind of connection interviewers are listening for.
why the drill is the only part that matters
  1. An untested recovery procedure has an unknown RTO, which is the same as having no RTO.
  2. Game days: deliberately fail a component in production, during business hours, with the team watching.
  3. Restore drills: actually restore a backup to a fresh cluster and verify the invariant holds. A backup that has never been restored is a hypothesis.
  4. Failover drills: promote a replica on a schedule, so it is routine rather than terrifying.
  5. Measure the real RTO during the drill and publish it. If the drill takes 40 minutes, the RTO is 40 minutes, regardless of the documented target.
151

Sketch v13: the observable core

The 3am question, answered

the ten minutes, concretely
  1. 0:00 Page: "transfer success rate below 99.5% in region ng for 5 minutes". A symptom, so a customer is affected.
  2. 0:30 Open the SLO dashboard. Failures are 100% on the nibss rail and 0% elsewhere. Blast radius established.
  3. 1:00 Correctness panel is green: invariant clean, suspense ageing normal. Money is not at risk, so this is availability rather than an emergency.
  4. 2:00 Open a sampled failing trace. The span for the NIBSS dispatch is timing out at 30 s. The circuit breaker has opened as designed.
  5. 3:00 Deploy timeline shows no change from us in 6 hours. It is them, not us.
  6. 4:00 Mitigate: flip routing to the alternate provider for that rail. Success rate recovers.
  7. 6:00 Confirm queued payments are draining and nothing is stuck in settle_suspense past its threshold.
  8. 8:00 Notify the provider, update the incident channel, and go back to bed. Root cause is a daytime activity.

What changed, and the cost accepted

ChangeDriven byCost accepted
SLOs with error budgets per flowReliability must be a negotiated targetMeasurement infrastructure, and an agreed budget policy
Correctness as a fifth signalA bank can be fast, available and wrongThe Part 10 checks become paging alerts
Business metrics that pageSilent failures are invisible to technical signalsSeasonality-aware baselines rather than thresholds
Tracing across async boundariesNine components per transferTrace context in outbox rows and Kafka headers
Allowlist-based structured logsPII must fail closedLog schemas to maintain per event type
Runbooks with do-not listsPeople reach for destructive fixes at 3amWriting and maintaining them, and they decay if not used
Tiered degradation, decided in advanceSomething must be given up under loadEvery capability needs a priority assigned and honoured
Kill switches, front door onlyMitigation must not require a deployA flag service on the money path, fail-safe by design
how to close round thirteen
"v13 is what makes the previous twelve rounds operable. Three things I would defend: correctness is a fifth golden signal, because the first four can all be green while the money is wrong; business metrics compared against the same hour last week, because a silent volume collapse is invisible to latency and error rate; and mitigate before diagnose, which is why kill switches are per-rail and per-provider and effective in ten seconds. And the one constraint I would flag is that Part 7's legal regions mean I cannot fail the EU over to Nigeria, so resilience has to be built inside each region, which costs more."
architecture v13
five signals, one correlation id, tiered degradation
swipe the figure sideways, or tap expand for full screen
1/8
the system
Thirteen rounds have produced a system with dozens of components. Operability is now its own design problem.