Round thirteen: “run it at 3am”
Twelve rounds have produced a system with dozens of components, six external dependencies and four regions. Now it is 03:12 and something is wrong. This round is about the difference between a system that can be operated and one that merely works, and the distinguishing question is whether a tired person with no context can tell what is broken and what to do about it.
The pressure: it is broken and you are asleep
“It is 3am. Transfers are failing for some customers but not others. You are on call and you know nothing else. Walk me through the next ten minutes.”
This question is not about monitoring tools. It tests whether the design produces answers, and the honest measure of an architecture is how fast it can be diagnosed by someone who did not build it.
- Is it real? A customer-visible failure, or a noisy check?
- How bad? What fraction of transfers, which regions, which rails, trending which way?
- Is money at risk? The only question that changes the urgency of everything else.
- What changed? A deploy, a config flag, a provider, a traffic shift.
- Can I stop the bleeding without a full diagnosis?
- Every customer-facing flow has a named SLO with an error budget.
- Alerts fire on symptoms customers feel, never on causes.
- Any transfer can be traced end to end from a single identifier.
- Every paging alert has a runbook reachable from the alert itself.
- Degradation is designed, with a stated order of what is shed.
- Recovery from region loss is drilled, not documented.
SLI, SLO, error budget for a money path
SLI a measurement of one user-visible property
"fraction of transfers completing in under 300 ms"
SLO a target for that measurement over a window
"99.9% of transfers, measured over 30 rolling days"
error budget what the SLO permits you to spend
100% − 99.9% = 0.1% × 300M transfers/month
= 300,000 failed transfers a month, allowed| Flow | SLI | SLO | Why this number |
|---|---|---|---|
| Balance read | p99 latency | 99.95% under 100 ms | On every screen. Cheap to serve from cache, so the target is high |
| Internal transfer | Success rate and p99 | 99.9% under 300 ms | Fully within our control, so we own every failure |
| Card authorisation | Availability | 99.99% under 500 ms | The network answers for us if we time out. Highest target in the bank |
| Outbound to a rail | Reaching a terminal state | 99.5% within 1 hour | Depends on providers we do not control, so the target is honest about that |
| Notification | Delivery within 10 s | 99% p95 | Deferrable, and one channel is always available |
| Ledger correctness | Invariant violations | Zero. No budget. | Not an SLO. A correctness property with no acceptable failure rate |
What the budget is actually for
- Budget remaining: ship features, take risks, run experiments.
- Budget exhausted: the team stops feature work and spends on reliability until it recovers. Agreed in advance, so it is not a negotiation during an incident.
- Budget consistently unused: the SLO is too loose or you are over-investing in reliability. Both are worth knowing.
- Burn rate alerts beat threshold alerts: burning a month's budget in an hour pages immediately, and burning it slowly over a week raises a ticket.
The four golden signals, and the fifth for banks
| Signal | Question | Where it bites here |
|---|---|---|
| Latency | How long do requests take? | Split successful from failed. A fast failure and a slow success are different problems and averaging them hides both |
| Traffic | How much demand? | Transfers/s, authorisations/s. A sudden drop is as alarming as a spike |
| Errors | What fraction fails? | By type: insufficient funds is not an error, a provider timeout is |
| Saturation | How full is the system? | Connection pools, queue depth, outbox lag, consumer lag, disk |
| Correctness | Is the money right? | The fifth signal, and the one only a bank has. Invariant violations, suspense ageing, reconciliation breaks |
The correctness dashboard
# the five metrics that say "the money is right", from Part 10
ledger_invariant_violations_total # must be 0. any value pages.
ledger_suspense_oldest_age_seconds{kind} # stuck operations
ledger_suspense_balance_minor{kind} # value at risk right now
recon_unresolved_break_value_minor{cp} # by counterparty
recon_oldest_break_age_hours{cp} # age, not count
# and the pipeline health that correctness depends on
outbox_oldest_unpublished_age_seconds{shard}
kafka_consumer_lag_seconds{group} # seconds, not records
posting_state_accepted_not_posted_count- Bad: "CPU above 80% on ledger-7". Might be fine, might be a nightly job, and it does not tell you whether anybody is affected.
- Good: "Transfer success rate below 99.5% for 5 minutes in region ng". A customer is feeling this right now.
- Cause metrics belong on dashboards for diagnosis, and symptom metrics belong on pages.
- The test: if this fires and no customer is affected, it should not have paged. Alerts that fail this test train people to ignore alerts.
Business metrics that page
The most valuable alerts in a payments system are business metrics, because they detect whole categories of failure that no technical metric sees.
| Metric | Detects | Comparison basis |
|---|---|---|
| Transactions per minute | A silent path failure, an app release bug, a rail outage | Same time last week, never an absolute threshold |
| Total value per minute | High-value flows failing while small ones succeed | Same time last week |
| Success rate by rail | A single provider degrading | Per-rail baseline |
| Authorisation approval rate | A fraud rule that is too aggressive. This is the alert that catches your own mistake | Rolling baseline per merchant category |
| New account creations | Onboarding broken, or an automated attack | Same time last week, both directions |
| Support contact rate | Everything your monitoring missed | Rolling baseline |
The approval-rate alert, which watches us rather than the customers
# Part 9 built automated blocking with a rate limit on itself.
# this is the business-level counterpart: if approval rate falls,
# the most likely explanation is OUR rule, not a fraud wave.
- alert: AuthApprovalRateDrop
expr: |
(sum(rate(auth_decisions_total{outcome="approved"}[10m]))
/ sum(rate(auth_decisions_total[10m])))
< 0.92
for: 5m
labels: { severity: page }
annotations:
summary: "Authorisation approval rate {{ $value | humanizePercentage }}"
runbook: "https://runbooks/auth-approval-drop"
# the first line of that runbook: check for a recently activated
# fraud rule, and be ready to use the kill switch.Distributed tracing across the posting path
A transfer now touches a gateway, the ledger service, a shard, an outbox relay, Kafka, four consumers, a provider and a notification channel. Without tracing, diagnosing a slow transfer means correlating logs across nine systems by hand.
trace one end-to-end operation. one trace id.
span one unit of work within it, with a parent.
context propagation passing the ids across every boundary
the boundaries that are usually missed:
· into Kafka → trace id in message headers
· out of the outbox → store it in the outbox row
· into a batch job → a trace per item, linked to the run
· across a webhook → correlate on our own reference
a trace that stops at the queue boundary is the common failure,
and the async half is exactly where the latency usually hides.-- the outbox row carries the trace context, so the asynchronous
-- half of the operation joins the same trace as the synchronous half.
ALTER TABLE outbox
ADD COLUMN trace_id TEXT,
ADD COLUMN span_id TEXT;
-- and the relay restores it when publishing:
-- headers: { traceparent: `00-${trace_id}-${span_id}-01` }
-- so a notification sent 4 seconds later appears under the same trace
-- as the transfer that caused it.- Always sample anything that errored. Errors are rare and maximally informative.
- Always sample anything slower than the SLO threshold, since those are the traces you will want.
- Always sample high-value transactions above a threshold.
- Sample 0.1% of ordinary successful transfers, for baselines.
- Prefer tail-based sampling: decide after the trace completes, when you know whether it was interesting. Costs buffering, and is worth it.
journal_id. Using the business identifier as the correlation
key means support, engineering and reconciliation are all searching the same
string, and a customer quoting a transaction reference lands you directly on the
trace.Structured logs, correlation, and PII
// unsearchable, unaggregatable, and it leaks a name and an amount
logger.info(`Transfer failed for Adeniji: 50000 to 0123456789`);
// structured, correlatable, and PII-free
logger.info('transfer.rejected', {
journal_id: 'j-8f1a', // the correlation key
trace_id: '4bf92f35...',
account_id: 'acct-771', // an internal id, not a NUBAN
reason_code: 'INSUFFICIENT_FUNDS',
amount_bucket: '10k_50k', // bucketed, not exact
currency: 'NGN',
shard: 7, region: 'ng', rail: 'internal'
});| Never in logs | Log instead |
|---|---|
| Full PAN, CVV, PIN | Nothing. Not even masked, per Part 8 |
| Customer name, phone, email, address | Internal customer id |
| Full account numbers, NUBAN, IBAN | Internal account id, or last 4 |
| Exact balances | Bucketed ranges, where needed at all |
| Auth tokens, session ids, API keys | A hash prefix, if correlation is needed |
| Full request or response bodies | A field allowlist |
- Debug logs: 7 days. High volume, low long-term value.
- Application logs: 30 to 90 days. Incident investigation.
- Audit logs of who did what to money: 7 years, append-only, separate store, different access controls. These are evidence rather than telemetry.
- Traces: 7 to 30 days for sampled, longer for errored.
The runbook, and what makes one useful
Most runbooks are useless because they describe the system rather than the decisions. A useful one is written for someone with no context, at 3am, under stress.
outbox_oldest_unpublished_age_seconds by
shard. One shard or all of them? One shard points at that database or
relay instance; all of them points at Kafka or a bad deploy.
min.insync.replicas makes every publish fail, which is the Part 4
failure mode and it is deliberate.
- What this means for customers, in one sentence, first. Everything else follows from severity.
- The blast-radius question before the diagnosis. One shard or all, one region or all.
- Mitigation before root cause. Explicitly permit rollback without understanding.
- The escalation threshold, as a number rather than "if it gets bad".
- An explicit do-not list. The destructive action somebody will reach for under pressure, named and forbidden.
- Linked from the alert, because a runbook nobody can find during an incident does not exist.
Graceful degradation: what to shed first
Under load or partial failure, the system must give something up. Deciding the order in advance is the difference between degradation and an outage.
| Priority | Capability | Shed when |
|---|---|---|
| P0, never shed | Card authorisation. The ledger invariant. Posting correctness. | Never. We fail requests before we compromise these |
| P1 | Balance reads, internal transfers, inbound funding | Only under severe saturation, and by rate limiting rather than by error |
| P2 | Outbound payments to external rails | Queue them rather than rejecting. Tell the customer it is pending, honestly |
| P3 | Transaction history, statements, search | Serve stale from cache, with an explicit "as of" timestamp |
| P4 | Notifications beyond in-app | Defer. In-app remains, per Part 12 |
| P5, shed first | Analytics refresh, marketing, recommendations, non-urgent batch | Immediately. Nobody notices within an hour |
- Serve stale. A balance from 30 seconds ago with a visible timestamp beats an error, for reads.
- Queue and promise. Accept the instruction, return
202with an honest pending state, and process when capacity returns. The journal-first design from Part 3 makes this possible. - Reject cleanly. A fast, clear, retryable error beats a timeout. The worst outcome is a request that hangs for 30 seconds and then fails.
Kill switches and feature flags on money
The 3am tool. A kill switch is how you stop customer harm in seconds without a deployment, and its design deserves the same rigour as the money path.
- Per rail. Disable NIBSS outbound while leaving everything else running.
- Per provider. Route around one PSP without touching the rail.
- Per product. Suspend new loan disbursement while repayments continue.
- Per fraud rule. Deactivate one rule, which is the Part 9 counterpart.
- Automated blocking, globally. One flag stops all automated enforcement.
- Per region. Fail a region out entirely.
- New account creation. Stop an onboarding attack without affecting existing customers.
- Effective in under 10 seconds, with a short cache TTL and no deployment.
- Fail-safe default. If the flag service is unreachable, the cached last-known value is used, never a default that silently enables something.
- Narrow scope. A switch that disables more than intended is not usable under pressure.
- Audited. Who flipped it, when, and why, in the append-only audit log.
- Reversible and tested. If turning it back on has never been exercised, you have a one-way door.
- Never affects correctness. A switch may stop accepting work; it may never change how money is recorded.
// the flag check on the money path: cached, fail-safe, and it can only
// ever REJECT work, never alter how a posting is recorded.
async function railEnabled(rail: string): Promise<boolean> {
try {
return await flags.get(`rail.${rail}.outbound`, { maxStaleMs: 10_000 });
} catch {
// the flag service is down. use the last known value from local
// cache. NEVER default to enabled: that turns a flag outage into
// re-enabling something an operator deliberately switched off.
return localCache.lastKnown(`rail.${rail}.outbound`) ?? false;
}
}Disaster recovery: RTO, RPO, and the drill
RTO recovery time objective how long until service resumes
RPO recovery point objective how much data may be lost
for our ledger:
RPO = zero ← non-negotiable. no acknowledged transfer is lost.
RTO = 15 minutes ← for a regional failure
RPO zero is what forced synchronous replication in round one,
and the ~5 ms it costs on every commit. this is where that
decision is finally cashed in.| Failure | Detection | Response | RTO |
|---|---|---|---|
| One application instance | Health check | Load balancer removes it | Seconds |
| One shard's primary | Replication monitor | Promote the synchronous replica | Under 60 s |
| One availability zone | Cloud provider plus our own checks | Traffic shifts to remaining zones | Minutes |
| An entire region | Multiple signals, with a human decision | Declared failover, with data residency constraints respected | 15 minutes |
| Data corruption | The Part 10 invariant checks | Point-in-time recovery, then replay | Hours |
| A provider outage | Circuit breaker | Route to the alternate, or queue | Seconds |
- An untested recovery procedure has an unknown RTO, which is the same as having no RTO.
- Game days: deliberately fail a component in production, during business hours, with the team watching.
- Restore drills: actually restore a backup to a fresh cluster and verify the invariant holds. A backup that has never been restored is a hypothesis.
- Failover drills: promote a replica on a schedule, so it is routine rather than terrifying.
- Measure the real RTO during the drill and publish it. If the drill takes 40 minutes, the RTO is 40 minutes, regardless of the documented target.
Sketch v13: the observable core
The 3am question, answered
- 0:00 Page: "transfer success rate below 99.5% in region ng for 5 minutes". A symptom, so a customer is affected.
- 0:30 Open the SLO dashboard. Failures are 100% on the
nibssrail and 0% elsewhere. Blast radius established. - 1:00 Correctness panel is green: invariant clean, suspense ageing normal. Money is not at risk, so this is availability rather than an emergency.
- 2:00 Open a sampled failing trace. The span for the NIBSS dispatch is timing out at 30 s. The circuit breaker has opened as designed.
- 3:00 Deploy timeline shows no change from us in 6 hours. It is them, not us.
- 4:00 Mitigate: flip routing to the alternate provider for that rail. Success rate recovers.
- 6:00 Confirm queued payments are draining and nothing is stuck in
settle_suspensepast its threshold. - 8:00 Notify the provider, update the incident channel, and go back to bed. Root cause is a daytime activity.
What changed, and the cost accepted
| Change | Driven by | Cost accepted |
|---|---|---|
| SLOs with error budgets per flow | Reliability must be a negotiated target | Measurement infrastructure, and an agreed budget policy |
| Correctness as a fifth signal | A bank can be fast, available and wrong | The Part 10 checks become paging alerts |
| Business metrics that page | Silent failures are invisible to technical signals | Seasonality-aware baselines rather than thresholds |
| Tracing across async boundaries | Nine components per transfer | Trace context in outbox rows and Kafka headers |
| Allowlist-based structured logs | PII must fail closed | Log schemas to maintain per event type |
| Runbooks with do-not lists | People reach for destructive fixes at 3am | Writing and maintaining them, and they decay if not used |
| Tiered degradation, decided in advance | Something must be given up under load | Every capability needs a priority assigned and honoured |
| Kill switches, front door only | Mitigation must not require a deploy | A flag service on the money path, fail-safe by design |