| psp_receivable:paystack | −1000000 | Paystack owes us this cash |
| wallet A | +985000 | customer credited now |
| psp_fee_expense | +15000 | our cost of collection |
| Σ | 0 | we now hold a receivable against Paystack, visible on the balance sheet |
Round six: “connect to the outside world”
Until now every transfer stayed inside a database we control. This round money starts leaving, into six systems with different protocols, different settlement windows, different failure modes, and no obligation to answer us promptly. The architectural problem is not integration. It is that a payment can be in a state where we genuinely do not know whether it happened, and the ledger still has to sum to zero.
The pressure: rails you do not control
“Customers need to send money to other Nigerian banks, receive from Paystack, pay a supplier in Germany, and settle into a US account. Four rails, and next year six. Design it.”
The first thing to establish is what changes fundamentally, rather than what changes mechanically. Adding an HTTP client is mechanical. What changes fundamentally:
- Atomicity. No transaction spans our database and NIBSS. The dual-write problem returns, in a harsher form.
- Immediate certainty. A request can time out having succeeded. Unknown is a real state, not an error.
- Reversibility. An outbound transfer that has settled cannot be rolled back. Only compensated, if the counterparty agrees.
- Our own availability. A provider that is down or slow becomes our problem, and their latency lands inside our budget.
- Send money out to any supported rail, and receive money in from any of them.
- Report a payment's state accurately, including pending and unknown.
- Never double-send, even across retries, restarts and provider duplicates.
- Adding a rail requires no change to the ledger or the product services.
- Route a payment across multiple providers for the same rail, by cost and health.
- A provider outage degrades that rail only, never the whole bank.
- No provider call ever happens inside a database transaction, per Part 2's lock discipline.
- Every outbound payment is reconciled against the provider's own record daily. Part 10.
- An unknown-state payment is resolved within one settlement cycle, automatically where possible.
The connector interface, and why it must be narrow
Six rails, one interface. The design goal is that the rest of the bank cannot tell which rail a payment used.
interface PaymentConnector {
readonly id: string; // 'nibss' | 'paystack' | 'sepa' ...
readonly capabilities: Capability[]; // push | pull | mandate | instant
// validate BEFORE any money moves. cheap, no side effects.
validate(d: Destination): Promise<ValidationResult>;
// resolve a destination to a human name, for confirmation screens.
resolve(d: Destination): Promise<{ accountName: string } | null>;
// the only method that moves money. MUST be idempotent on our reference.
dispatch(p: OutboundPayment): Promise<DispatchResult>;
// the answer to "did that actually happen?". the unknown-state resolver.
query(ourReference: string): Promise<PaymentStatus>;
// parse and VERIFY an inbound callback. returns null if unauthentic.
parseWebhook(raw: Buffer, headers: Headers): InboundEvent | null;
}
// the result type that makes the hard case explicit in the type system
type DispatchResult =
| { state: 'accepted'; providerRef: string }
| { state: 'rejected'; reason: string; retryable: boolean }
| { state: 'unknown'; detail: string }; // ← the one that matterscatch and treat it as a failure, which is the single most
expensive bug in payments: you mark it failed, refund the customer, and the
money leaves anyway. Making unknown a value forces every caller to handle
it. The compiler becomes the thing that stops you losing money.Why the interface must stay narrow
The temptation is to expose each provider's richness: Paystack's split payments, NIBSS's name enquiry semantics, SEPA's mandate types. Resist it. Every capability added to the interface must be implemented or stubbed by all six connectors, and the abstraction stops paying for itself.
- The interface carries only what is common to all rails.
- Rail-specific behaviour goes behind a capability flag the caller can query, never a method everyone must implement.
- Rail-specific data travels in an opaque metadata field the connector alone interprets.
- If a feature cannot be expressed this way, it gets its own service rather than distorting the interface.
NIBSS and the Nigerian rails
- settlement
- Near-instant to the beneficiary, with net settlement between banks at the central bank later the same day.
- identifier
- 10-digit NUBAN plus a 3-digit bank code. The NUBAN has a check digit, so a typo is usually catchable locally before any network call.
- name enquiry
- A separate call that returns the account holder's name. Mandatory in practice: the customer confirms the name before sending, which prevents a large class of misdirected payments.
- limits
- Per-transaction and daily caps set by KYC tier, enforced by us and by the CBN.
- failure mode
- Timeouts are common at peak. A timeout does not mean failure, which is what makes the query method essential.
- reversal
- No protocol-level reversal. A wrong payment is recovered by a manual interbank process that can take weeks.
The name-enquiry-then-transfer sequence
Two calls, and the gap between them is a design decision:
// 1. name enquiry. cheap, read-only, no money at risk.
const ne = await nibss.resolve({ nuban: '0123456789', bankCode: '058' });
// → { accountName: 'ADENIJI OLUWAFERANMI' }
// 2. the customer SEES the name and confirms. this is a product
// requirement, not a nicety: it is the last chance to catch an error
// on a rail with no reversal.
// 3. transfer, carrying the name-enquiry session so NIBSS can correlate
const r = await nibss.dispatch({
ourReference: journalId, // our idempotency key, end to end
sessionId: ne.sessionId,
amount: 50000n, currency: 'NGN',
beneficiary: { nuban: '0123456789', bankCode: '058', name: ne.accountName }
});The other Nigerian rails, briefly
| Rail | Use | Note |
|---|---|---|
| NIP | Instant transfers | The default. Everything below is a special case. |
| NEFT | Batch transfers | Cheaper, settles in windows. Correct for bulk payroll where instant is not required. |
| RTGS | High-value | Gross settlement at the central bank, for large amounts. Different limits and controls. |
| Direct debit | Mandated pulls | Needs a stored mandate, and a dispute process that can reverse a collection. |
| USSD / POS | Channels, not rails | They originate transactions that then travel over the rails above. Part 8 handles POS. |
Paystack and Flutterwave as PSPs
A payment service provider is a different kind of counterparty from NIBSS: an aggregator sitting between us and many underlying rails, with its own float, its own settlement schedule, and its own balance we must track.
- direction
- Strong at collections (money in): cards, bank transfer to a virtual account, USSD. Also do payouts.
- virtual accounts
- A dedicated NUBAN per customer, so an inbound transfer is automatically attributable. This is the mechanism most Nigerian fintechs use for funding, and it removes the reference-matching problem entirely.
- settlement
- T+1 typically. The customer's wallet is credited immediately on webhook, but the cash reaches our bank account the next day. That gap is a real balance-sheet position.
- fees
- Percentage with a cap, often plus a flat component. Varies by channel, which is what makes routing worth doing.
- failure mode
- Webhook delivery is at-least-once and occasionally never. A poller is mandatory, not optional.
- chargebacks
- Card collections can be reversed weeks later. The funds must be recoverable or provisioned for.
The settlement gap, and why it needs its own account
A customer funds their wallet with ₦10,000 via Paystack. We credit them immediately because the product requires it. But Paystack holds the cash until tomorrow. The ledger must represent that honestly:
| bank_account:gtb_main | −985000 | cash actually arrives (asset increases) |
| psp_receivable:paystack | +985000 | receivable clears to zero |
| Σ | 0 | the receivable account is now the reconciliation surface |
SEPA, IBAN, and European settlement
- variants
- SCT credit transfer, settles next business day. SCT Inst instant, under 10 seconds, up to €100,000. SDD direct debit with a mandate and an 8-week no-questions refund right.
- identifier
- IBAN, up to 34 characters, with a mod-97 checksum. Validate locally before any network call: it catches most typos for free.
- message format
- ISO 20022 XML,
pain.001for initiation andpacs.008for interbank. Verbose, strongly typed, and far richer than the Nigerian rails. - the refund right
- SDD collections can be unconditionally refunded for 8 weeks. Any product built on SDD must provision for this; it is a business risk expressed in the ledger.
- business days
- Settlement follows TARGET2 calendar, which is neither weekends nor national holidays. A payment submitted Friday evening settles Monday or later.
// IBAN check: move the first 4 chars to the end, map letters to numbers,
// then the whole thing mod 97 must equal 1. no network call needed.
function validIban(iban: string): boolean {
const s = iban.replace(/\s/g, '').toUpperCase();
if (!/^[A-Z]{2}[0-9]{2}[A-Z0-9]{,30}$/.test(s)) return false;
const re = s.slice(4) + s.slice(0, 4);
const num = [...re].map(c =>
/[A-Z]/.test(c) ? (c.charCodeAt(0) - 55).toString() : c).join('');
// the number is far bigger than a float, so reduce in chunks
let r = 0;
for (const d of num) r = (r * 10 + +d) % 97;
return r === 1;
}ACH, wire, RTP in the United States
| ACH | Fedwire | RTP / FedNow | |
|---|---|---|---|
| Speed | 1 to 3 business days; same-day windows exist | Same day, within operating hours | Seconds, 24/7/365 |
| Cost | Cents. Batch economics | Tens of dollars | Low, a few cents |
| Reversibility | Reversible for specific return codes, days later | Final. Effectively irrevocable | Final. Credit push only |
| Identifier | Routing number plus account number | Routing number plus account number | Routing plus account, or a token |
| Direction | Push and pull | Push only | Push only, with request-for-pay messaging |
| Use for | Payroll, bill pay, funding | Large, urgent, irreversible | Instant consumer payouts |
ACH returns, the failure mode that defines the rail
An ACH debit can be accepted, appear to settle, and then be returned days later with a reason code. This is not an edge case; it is routine, and it determines how the product must behave.
day 0: ACH debit submitted, customer's wallet credited
day 1: settles. everything looks complete.
day 4: R01 insufficient funds → the credit is clawed back
if the customer already spent it → we are short
this is exactly what the CLEARED balance from Part 5 is for- Hold the funds until the return window closes. Safest, and the worst customer experience.
- Release a portion immediately based on risk score, and hold the rest. What most good products do.
- Release fully and absorb returns as a modelled loss, priced into fees. Viable only with strong underwriting.
Whichever is chosen, the mechanism is the same: the credit posts to the ledger immediately and is excluded from the cleared balance until the window closes. The three-balance model from Part 5 was built for precisely this.
SWIFT MT103 and correspondent banking
Cross-border, outside a common scheme. The important thing to understand is that SWIFT is messaging, not settlement: it tells banks what to do, and the money moves through accounts those banks hold with each other.
our bank (Lagos) beneficiary bank (Frankfurt)
↓ MT103 ↑
correspondent A → correspondent B
(our nostro) (their vostro)
nostro: "our account with them", in their currency
vostro: "their account with us", in our currency
each hop deducts a fee and adds a day.
the beneficiary often receives less than was sent, which the
product must disclose up front rather than explain afterwards.| Property | Reality |
|---|---|
| Speed | 1 to 5 business days, depending on hops and compliance review |
| Cost | $15 to $50, plus deductions at each correspondent |
| Traceability | Improved by gpi with a UETR tracking reference, though not universal |
| Identifier | BIC for the bank, plus IBAN or a local account format |
| Compliance | Heavy. Sanctions screening at every hop; a payment can be frozen mid-route for days |
| Failure mode | A payment can sit in limbo at a correspondent with no automated status. Investigation is a human process |
The saga: hold, dispatch, confirm, or compensate
The core flow of the round. Money must leave our ledger before we know the external payment succeeded, and the ledger must stay balanced throughout.
Step 1: reserve, locally and atomically
BEGIN;
INSERT INTO journal (id, kind, idempotency_key)
VALUES ($jid, 'payout.initiated', $clientKey);
-- debit the customer, checked against their available balance
INSERT INTO entries ... ($wallet, -50000);
-- credit a per-rail settlement suspense account
INSERT INTO entries ... ('settle_suspense:nibss', +50000);
-- the dispatch intent, committed with the money. no dual write.
INSERT INTO outbox (event_type, payload)
VALUES ('payout.dispatch.requested', $payload);
COMMIT;
-- ledger sums to zero. money is in a named account. nothing external yet.Step 2: dispatch, outside any transaction
async function dispatch(p: OutboundPayment) {
// our journal id IS the idempotency key, end to end. a retry of this
// call, or a duplicate outbox delivery, cannot double-send.
const r = await connector.dispatch({ ...p, ourReference: p.journalId });
switch (r.state) {
case 'accepted':
// NOT settled. just accepted. record the provider ref and wait.
return markAwaitingConfirmation(p, r.providerRef);
case 'rejected':
// authoritative no. safe to compensate immediately.
return compensate(p, r.reason);
case 'unknown':
// DO NOT compensate. DO NOT retry the dispatch.
// the money stays in suspense and the resolver takes over.
return scheduleStatusQuery(p);
}
}Step 3a: confirmed
| settle_suspense:nibss | −50000 | suspense clears |
| nostro:nibss_settlement | +50000 | against our settlement position |
| Σ | 0 | the payment is now complete and reconcilable |
Step 3b: rejected, so compensate
| settle_suspense:nibss | −50000 | suspense clears |
| wallet A | +50000 | customer made whole |
| Σ | 0 | a new journal linked to the original. nothing was edited or deleted |
settle_suspense:nibss is precisely "money we have taken from customers
and not yet settled", and anything sitting there too long is an alert. One
pattern, reused: park value in a named place while the outcome is uncertain.Webhooks: signing, replay, and the poller you still need
Providers tell us about state changes by calling us. Webhooks are a public, unauthenticated-by-default HTTP endpoint that moves money, which makes them the most security-sensitive surface in the system.
- Verify the signature over the raw bytes, using a constant-time comparison. Not the parsed body; a re-serialised body has different bytes.
- Check the timestamp is within a few minutes, to bound replay.
- Deduplicate on the provider's event id. Webhooks are at-least-once.
- Verify it refers to a payment we initiated, in a state where this transition is legal.
- Check the amount matches what we dispatched. Never trust the amount in the callback.
function verify(raw: Buffer, sig: string, ts: string, secret: string): boolean {
// bound replay: reject anything older than 5 minutes
if (Math.abs(Date.now() / 1000 - +ts) > 300) return false;
// sign the timestamp WITH the body, so a captured signature cannot be
// replayed later against a different timestamp header.
const expected = hmacSha512(`${ts}.` + raw, secret);
// constant time: a fast-failing comparison leaks the correct prefix
// and lets an attacker forge a signature byte by byte.
return crypto.timingSafeEqual(Buffer.from(sig), Buffer.from(expected));
}Why a poller is mandatory
Webhooks fail in ways that are entirely outside our control, and every one of these has happened to real payment systems:
| Failure | Consequence without a poller |
|---|---|
| Our endpoint was down during a deploy | The provider retries for a while, then gives up. The payment is stuck forever. |
| The provider's webhook queue backs up | Hours of delay, with customers seeing pending. |
| The provider never sends one for this case | Certain terminal states are simply not notified by some providers. |
| A network partition drops the callback | Silent, permanent loss of the state change. |
-- the poller's work queue: anything non-terminal, oldest first,
-- with backoff so a stuck payment is not queried every second forever.
SELECT * FROM outbound_payments
WHERE state IN ('awaiting_confirmation', 'unknown')
AND next_query_at <= now()
ORDER BY next_query_at
LIMIT 200
FOR UPDATE SKIP LOCKED; -- many pollers, disjoint workTimeouts, retries, and the unknown-state problem
The hardest problem in payments, and the one most designs get wrong. Worth slowing down for.
we send a dispatch request. the socket times out after 30 s.
what actually happened? one of:
a) the request never arrived → no money moved
b) it arrived and was rejected → no money moved
c) it arrived, succeeded, response lost → MONEY MOVED
d) it is still being processed → money may yet move
we cannot distinguish these from our side. ever.- Treat it as failure and refund. If the truth was (c), we refunded a customer whose money also left. We are short, and it is our loss.
- Treat it as success. If the truth was (a), the customer is debited and nobody was paid. They will call, and we will have to find it.
- Blindly retry the dispatch. If the truth was (c) and the provider is not idempotent, we have sent the money twice.
- Record the state as
unknown. It is a legitimate state, and the money stays in suspense. - Query the provider by our own reference, with backoff. Most rails resolve within minutes.
- If the provider supports idempotent dispatch, retrying with the same reference is safe and doubles as a query. This is why
ourReferenceis in the interface. - If unresolved by the next settlement cycle, the daily reconciliation in Part 10 resolves it against the provider's own file, which is authoritative.
- Only then, if still unresolved, does a human decide, with full context from the suspense account.
Timeout budgets per rail
| Rail | Connect | Read | Retry dispatch on timeout? |
|---|---|---|---|
| NIP | 2 s | 30 s | Only with our reference, which NIBSS deduplicates on |
| PSP | 2 s | 20 s | Yes. Both support idempotency keys properly |
| SEPA | 3 s | 60 s | No. Query instead; batch semantics make duplicates costly |
| SWIFT | 5 s | 120 s | Never. No reliable dedup and no reversal |
Circuit breakers and provider health
A provider that is down and a provider that is slow need different handling, and slow is the more dangerous of the two because it consumes our resources while failing.
provider p99 goes from 300 ms → 25 s
with 200 concurrent request slots:
at 300 ms → 660 req/s of capacity
at 25 s → 8 req/s. every slot is occupied waiting.
a slow dependency exhausts our threads and takes down
rails that were perfectly healthy. this is the cascade.
The three states
| State | Behaviour | Transition |
|---|---|---|
| Closed | Requests flow. Failures are counted in a rolling window. | To open when the failure rate exceeds the threshold over a minimum volume. |
| Open | Fail immediately without calling. Preserves our capacity and stops hammering a struggling provider. | To half-open after a cool-down. |
| Half-open | Allow a small number of probes through. | To closed if they succeed, back to open if they do not. |
// bulkheads: a bounded, SEPARATE pool per provider, so one provider's
// slowness cannot consume the capacity the others need.
const pools = {
nibss: semaphore(120),
paystack: semaphore(60),
flutterwave:semaphore(60),
sepa: semaphore(40),
swift: semaphore(20)
};
// this is the difference between a degraded rail and a degraded bank.Routing: cost, reliability, and failover
Several providers can often reach the same destination. Choosing between them is a real optimisation with real money attached, and it is the reason the connector interface was worth building.
- Capability. Can this provider reach this destination at all? A hard filter.
- Health. Is its circuit closed, and what is its current success rate and latency?
- Cost. The fee for this amount and channel. Tiered and non-linear, so it must be computed rather than looked up.
- Limits. Our remaining balance or credit with that provider, and their per-transaction caps.
- Speed. Does this payment need to be instant, or is it a batch payout where cheaper and slower is better?
function route(p: OutboundPayment, providers: Provider[]): Provider | null {
const viable = providers
.filter(x => x.canReach(p.destination)) // capability
.filter(x => x.breaker.state !== 'open') // health
.filter(x => x.remainingFloat >= p.amount) // funding
.filter(x => p.speed !== 'instant' || x.isInstant);
if (!viable.length) return null; // queue it, alert, do NOT guess
// score: cost matters, but a cheap provider that fails 5% of the time
// is more expensive than it looks once you price the support load.
return viable.sort((a, b) =>
(a.feeFor(p) / a.successRate) - (b.feeFor(p) / b.successRate)
)[0];
}Failover, and the rule that keeps it safe
If a dispatch fails, when may we try another provider? The answer follows directly from chapter 75:
- Rejected by provider A: safe to try provider B. The rejection is authoritative, so no money moved.
- Circuit open before the call: safe to try B. We never called A.
- Unknown from A: NEVER try B. We might send twice. Resolve A's state first, always.
unknown a distinct
state that the routing code cannot silently treat as failure. This is the payoff
for putting unknown in the type system in chapter 67, and drawing
that line back is the kind of coherence interviewers remember.Sketch v6: pluggable rails
The state machine every outbound payment follows
initiated → dispatching → awaiting_confirmation → settled
↓ ↓
unknown returned
↓ ↓
(resolver) (compensate)
↓
settled or reversed every path ends in a terminal state
manual_review ← the escape hatch when automation cannot decide
What changed, and the cost accepted
| Change | Driven by | Cost accepted |
|---|---|---|
| Narrow connector interface | Six rails, more coming | Rail-specific features live behind capability flags rather than in the interface |
| Per-rail suspense accounts | Money must be somewhere real during uncertainty | More accounts to reconcile, which Part 10 turns into an advantage |
unknown as a first-class state | Timeouts are ambiguous by nature | A resolver to build, and a state customers can see as pending |
| Poller plus webhooks | Webhook delivery is unreliable | Extra provider API calls and cost, which is cheaper than stuck payments |
| Breakers and bulkheads per provider | A slow provider exhausts shared capacity | Tuning per provider, and separate breakers for dispatch and query |
| Cost and health-aware routing | Several providers reach the same destination | Failover is forbidden on unknown, which must be enforced in code |