Round fourteen: “how does anything else talk to it?”
Thirteen rounds have treated service boundaries as arrows on a whiteboard. This round makes them real: which protocol, and why, for each kind of conversation. Then it designs the single interface every other service in the bank uses to move money, and walks through exactly how loans, overdraft and savings call it, including the question of who is allowed to exceed a limit and how that permission travels.
The pressure: nine services, one ledger
“Loans, savings, cards, overdraft, working capital, fraud, notifications, support and the mobile backend all need to talk to the ledger. What protocol, what interface, and how do you stop nine teams breaking it?”
Three separate questions wearing one coat, and separating them is the first move again.
- Transport. Which protocol for which conversation shape, and why.
- Contract. What operations exist, and what does the ledger refuse to know about.
- Governance. Authentication, authorisation, versioning, and blast radius when a caller is wrong.
the conversations, by shape: request/response, low latency post entries, check a balance request/response, slow statement generation, bulk export fire and forget an event happened, react if you care server → client push live balance on an open app screen persistent bidirectional card network authorisation traffic five shapes. picking one protocol for all of them is the mistake.
- Internal posting call adds under 5 ms of transport overhead.
- A misbehaving caller cannot degrade the ledger for other callers.
- No caller can violate the invariant, even with a bug.
- The API can evolve without coordinated deployment of nine services.
- Every call is authenticated, authorised and attributed to a service identity.
Choosing a protocol: the honest decision table
The answer that earns credit maps conversation shape to protocol, with a reason for each. Here is the whole table before the chapters that justify it.
TCP, and what it guarantees you already rely on
Everything above except UDP runs on TCP. Being able to say what TCP does and does not promise is what makes the later choices defensible.
| TCP guarantees | TCP does not guarantee |
|---|---|
| Bytes arrive in order | That the application processed them |
| Bytes are not corrupted, per its checksum | That the peer is still alive right now |
| Lost segments are retransmitted | Message boundaries. It is a byte stream |
| Flow control, so a fast sender cannot drown a slow receiver | Bounded latency. A retransmit can take seconds |
write() means the bytes are in a kernel buffer, not
that the peer received or processed them. This is precisely why Part 6's
unknown state exists: a TCP write can succeed, the connection can
drop, and you genuinely cannot tell whether the payment was processed. The
transport's guarantees stop well short of the application's needs, and every
idempotency key in this design exists to cover that gap.The settings that matter on a money path
// a dead peer with no traffic is indistinguishable from an idle one. // without keepalive, a half-open connection can sit for HOURS. TCP_KEEPIDLE = 30 // start probing after 30 s idle TCP_KEEPINTVL = 10 // probe every 10 s TCP_KEEPCNT = 3 // 3 failures ⇒ dead. detected in ~60 s. // Nagle batches small writes to save packets, adding up to 40 ms. // on a request/response path that is pure added latency. TCP_NODELAY = true // and a connect timeout, because the OS default can be over a minute, // which is far longer than any budget we set in Part 8. connectTimeoutMs = 2000
HTTP/1.1, HTTP/2, and head-of-line blocking
| HTTP/1.1 | HTTP/2 | |
|---|---|---|
| Concurrency per connection | One request at a time | Many interleaved streams |
| Head-of-line blocking | At the application layer | Removed at the application layer, remains at TCP |
| Framing | Text, newline-delimited | Binary frames with stream ids |
| Headers | Repeated in full on every request | HPACK compressed, incrementally encoded |
| Server push | No | Yes, though largely deprecated in practice |
| Connection count | 6 to 8 per origin, by convention | One, which is why connection pooling changes shape |
The practical consequence for us
HTTP/1.1, 8 connections, one slow 900 ms request:
→ 1 of 8 connections blocked. 12% capacity loss.
HTTP/2, 1 connection, one slow 900 ms stream:
→ other streams continue. no application-level loss.
→ but one lost packet stalls all of them briefly.
so: HTTP/2 for internal service calls, and keep per-stream
deadlines, because multiplexing removes queueing but not slowness.gRPC internals: HTTP/2 streams and protobuf framing
gRPC = HTTP/2 + protobuf + a calling convention
one RPC = one HTTP/2 stream
request = HEADERS frame + DATA frames
response = HEADERS + DATA + trailers carrying the status
each message in a DATA frame is length-prefixed:
[1 byte compressed flag][4 byte length][payload]
the status in TRAILERS is why gRPC can stream a response
and still report a failure after the body has started.Why protobuf rather than JSON, concretely
| Property | JSON | Protobuf |
|---|---|---|
| A typical posting message | ~420 bytes | ~90 bytes |
| Parse cost | String scanning, allocation per field | Binary, length-prefixed, minimal allocation |
| Schema | Convention, or a separate validator | Compiled. The wire format is the contract |
| Field evolution | By name, forgiving and untracked | By field number, with explicit rules |
| Large integers | Dangerous. Parsed as doubles, so money loses precision | int64 is exact |
| Debuggability | Readable in any tool | Needs the schema to decode |
9007199254740993 silently becomes
...992. Protobuf has a real int64, so an amount is an
integer on the wire with no encoding workaround. For a system whose entire
correctness argument rests on exact integer arithmetic, that is not a minor
convenience.// the four call types, and where each is used here
service Ledger {
// 1. unary: the overwhelming majority. post, query, resolve.
rpc Post(PostRequest) returns (PostResponse);
// 2. server streaming: a large export without buffering it all.
rpc StreamEntries(StreamRequest) returns (stream Entry);
// 3. client streaming: bulk upload, e.g. a payroll batch.
rpc PostBatch(stream PostRequest) returns (BatchSummary);
// 4. bidirectional: rare internally. used by the card authoriser
// to keep a warm channel with continuous health signalling.
rpc AuthChannel(stream AuthMsg) returns (stream AuthReply);
}- Deadlines propagate automatically. A deadline set by the caller travels with the request and every downstream hop sees the remaining time. This is the Part 8 budget mechanism, built into the protocol.
- Cancellation propagates. If the caller gives up, downstream work stops rather than continuing to hold locks.
- Status codes are structured, so
RESOURCE_EXHAUSTEDandFAILED_PRECONDITIONare distinguishable without parsing a message string. - Interceptors give one place for auth, tracing, metrics and idempotency handling across every method.
When gRPC is wrong, and REST is right
Being able to argue against your own default is what makes the default credible.
| Situation | Use | Why |
|---|---|---|
| Internal service to service | gRPC | We control both ends, deploy both toolchains, and want typed contracts and deadlines |
| A public or partner API | REST | A partner should not need protoc, a code generator or an HTTP/2-capable client to send us a payment |
| Webhooks we receive | REST | The sender decides, and every provider sends JSON over HTTP/1.1 |
| Browser clients | REST or SSE | Browsers cannot speak raw gRPC without a proxy such as gRPC-Web |
| Anything an operator debugs by hand | REST | curl at 3am beats a schema-dependent binary decoder, and Part 13 cares about that |
| Through unknown middleboxes | REST | Old proxies mangle HTTP/2 or trailers, and diagnosing that at a partner is painful |
WebSockets: the balance feed, and backpressure
A customer has the app open. Their balance should update when money moves. That is a server-to-client push problem, and WebSockets are the tool most people reach for first.
upgrade handshake: HTTP GET + Upgrade: websocket
→ 101 Switching Protocols
→ the TCP connection is now a bidirectional frame stream
costs, at 20M customers:
· ~2% concurrently connected = 400,000 open sockets
· each holds file descriptors, buffers, and server memory
· connections are stateful, so a deploy disconnects everyone
· load balancers must support long-lived upgrades
~50k sockets per node ⇒ 8 nodes purely to hold connectionsBackpressure, which is where naive implementations fail
- A client is on a slow mobile connection and stops reading.
- The server keeps writing balance updates into the socket.
- The kernel send buffer fills, then the application's own queue grows.
- Server memory climbs, one buffer per slow client.
- With thousands of slow clients, the node runs out of memory and drops every connection it holds, including the healthy ones.
// the fix: a bounded per-connection queue with an explicit policy for
// what to do when it fills. never an unbounded buffer.
class BalanceFeed {
private queue: Update[] = [];
private static MAX = 32;
push(u: Update) {
if (this.queue.length >= BalanceFeed.MAX) {
// for a BALANCE, only the latest value matters: collapse.
// this is the key insight: the semantics of the data decide
// the shedding policy.
this.queue = [this.queue[0], u];
this.markResyncNeeded(); // tell the client to refetch
return;
}
this.queue.push(u);
}
}Server-sent events, and why they often win
For our actual requirement, which is server to client only, SSE is simpler and better in almost every dimension.
| WebSocket | SSE | |
|---|---|---|
| Direction | Bidirectional | Server to client only |
| Protocol | A separate protocol after upgrade | Plain HTTP. Works through every proxy |
| Reconnection | You implement it | Automatic, built into the browser |
| Missed messages | You implement resumption | Last-Event-ID header, built in |
| Compression | Per-message extension | Standard HTTP gzip |
| Auth | Awkward: no custom headers on upgrade in browsers | Normal headers and cookies |
| Binary | Yes | Text only, which is fine for JSON |
// SSE with resumption. the id: field is what makes reconnect lossless,
// and we already have a perfect monotonic id: the entry id from Part 1.
res.writeHead(200, {
'Content-Type': 'text/event-stream',
'Cache-Control': 'no-cache',
'X-Accel-Buffering': 'no' // stop nginx buffering the stream
});
// the browser sends Last-Event-ID on reconnect, so we resume exactly.
const resumeFrom = req.headers['last-event-id'] ?? '0';
for await (const e of entriesSince(accountId, BigInt(resumeFrom))) {
res.write(`id: ${e.id}\n`);
res.write(`event: balance\n`);
res.write(`data: ${JSON.stringify({ balance: e.balanceAfter.toString() })}\n\n`);
}BIGSERIAL entry id. The same property that made
balance snapshots safe in Part 3 and made ordering work in Part 4 now makes a
reconnecting client lossless. Monotonic ids keep paying off, which is why
"prefer monotonic over random" appeared in the Part 0 vocabulary.- SSE for the balance and transaction feed. Simpler, resumable, works everywhere.
- WebSocket only where the client genuinely streams upward too, such as support chat.
- Polling is still correct for low-frequency updates. A 30-second poll costs nothing and holds no state, and 400,000 idle sockets to deliver an update every few minutes is a bad trade.
UDP, and the two places a bank actually uses it
UDP: no handshake, no ordering, no retransmission, no flow control
the single property that makes it fast
fire and forget: no round trip, no connection state
the same property, restated
the sender cannot know whether it arrived
for a payment, that is the entire problem we spent Part 6 solving.- Metrics emission. StatsD over UDP. Losing 0.1% of metric samples changes no decision, and the application must never block on its own telemetry. Metrics over TCP can make an observability outage into a service outage.
- DNS. Every service call resolves a name first, over UDP, with a TCP fallback for large responses.
- QUIC, and therefore HTTP/3. Reliability is reimplemented in userspace above UDP, per stream, which is how it escapes TCP's head-of-line blocking.
- NTP. Clock sync, which matters more than it seems: Part 10's booking dates and Part 9's velocity windows both depend on clocks being close.
ISO 8583 and the persistent socket reality of cards
Part 8 treated the card network as an arrow. Here is what is actually on the wire, and it is unlike everything else in this design.
[2 or 4 byte length][MTI][bitmap][data elements…]
MTI message type: 0100 auth request, 0110 auth response,
0200 financial, 0400 reversal, 0800 network echo
bitmap 64 or 128 bits. bit N set ⇒ data element N present.
DE2 = PAN, DE4 = amount, DE39 = response code,
DE7 = transmission time, DE11 = trace number
a compact binary format designed in the 1980s for expensive links,
and still carrying most card traffic on earth.- Persistent sockets. A small number of long-lived TCP connections to the scheme, not a connection per request. Establishing one involves certificates and, historically, paperwork.
- Correlation by field, not by connection. Responses can arrive out of order on the same socket, matched by the STAN (DE11) plus the terminal and date. It is multiplexing, hand-rolled, decades before HTTP/2.
- Echo tests. A 0800 network management message every 30 seconds proves the link is alive. Missing echoes mean the link is down even with no traffic.
- Reversals are a message type. A 0400 reversal is first-class, because a timed-out authorisation must be explicitly unwound rather than left ambiguous.
- Strict timeouts, enforced by the scheme. Exceed them and the network stands in on your behalf, using rules you configured weeks earlier. This is the Part 8 constraint, in protocol form.
The posting API: one interface, every caller
The contract at the centre of the bank. Part 5 sketched it; here it is properly specified, and every decision in it is defensive.
service Ledger {
rpc Post(PostRequest) returns (PostResponse);
rpc GetBalance(BalanceRequest) returns (BalanceResponse);
rpc PlaceHold(HoldRequest) returns (HoldResponse);
rpc ResolveHold(ResolveHoldRequest) returns (PostResponse);
rpc GetJournal(JournalRequest) returns (Journal);
}
message PostRequest {
// REQUIRED. the caller's own natural key for this intent.
// "loan-disb-{loanId}", "accrual-{loanId}-{date}". see Part 5.
string idempotency_key = 1;
// what KIND of business event. drives GL mapping and reporting,
// and it is an enum so a typo is a compile error.
JournalKind kind = 2;
// the entries. the ledger validates they sum to zero PER CURRENCY
// and REFUSES otherwise. a caller bug cannot break the invariant.
repeated EntrySpec entries = 3;
// opaque to the ledger. stored, indexed, never interpreted.
map<string, string> metadata = 4;
// links a correction to what it corrects. Part 10.
optional string reverses_journal_id = 5;
}
message EntrySpec {
string account_id = 1;
int64 amount_minor = 2; // signed. negative = debit. int64, exact.
string currency = 3;
// the floor for THIS entry: 0 normally, negative for overdraft.
// the ledger enforces it inside the insert, per Part 1.
optional int64 min_balance_minor = 4;
}- Entries in, not a product operation. The ledger has no
disburseLoan(), so it never accumulates product knowledge. Prevents the ledger becoming the place every product rule lives. - Invariant validated server-side. Unbalanced entries are refused. Prevents a caller bug from creating money.
- Idempotency key required, not optional. Prevents the most common integration error, which is a retry that double-posts.
- The floor is per entry. Prevents an overdraft-aware caller from accidentally granting overdraft on an unrelated account.
- Metadata is opaque. Prevents callers from pressuring the ledger schema to grow a column per product.
Loan disbursement, as a ledger call
The full path, end to end, showing where each responsibility sits.
// LOANS SERVICE owns: eligibility, pricing, schedule, product rules.
// LEDGER owns: that the entries balance and are durably recorded.
const res = await ledger.post({
idempotencyKey: `loan-disb-${loan.id}`, // natural, one per loan
kind: JournalKind.LOAN_DISBURSEMENT,
entries: [
{ accountId: `loan_receivable:${loan.id}`, amountMinor: -50_000_000n,
currency: 'NGN' },
{ accountId: customer.ngnWalletId, amountMinor: 49_500_000n,
currency: 'NGN' },
{ accountId: 'fee_income:origination', amountMinor: 500_000n,
currency: 'NGN' }
],
metadata: { loan_id: loan.id, product: 'salary_advance', tenor_days: '30' }
});
// if the loans service has a bug and these do not sum to zero,
// the ledger returns INVALID_ARGUMENT and writes nothing.- Loans service: is this customer eligible, what rate, what schedule, what fee. All product policy.
- Ledger: do these entries balance, does this key already exist, is the account real, is it restricted. All correctness.
- Neither: whether the customer wanted a loan. That is the app, and it is upstream of both.
- The event published afterwards is what activates the repayment schedule, so the loans service reacts to the confirmed fact rather than assuming its own call succeeded.
Loan recovery, sweeps, and partial repayment
Getting the money back is harder than lending it, and the mechanisms are worth knowing because they are where lending products actually differ.
| Mechanism | How it works | Risk |
|---|---|---|
| Scheduled debit | On the due date, debit the wallet for the instalment | Fails if the balance is short. Needs a retry policy |
| Sweep | Whenever a credit lands, take a portion toward the loan | Very effective, and the most customer-hostile. Needs clear consent and a floor |
| Salary assignment | The employer remits directly before the customer sees it | Strongest recovery. Requires an employer relationship |
| Direct debit on an external account | Pull from another bank via the Part 6 rails | Reversible for a period, so the money is not certain |
| Overdraft absorption | Let the repayment push the wallet into overdraft | Converts one debt into another, and is sometimes correct |
The sweep, and the floor that makes it acceptable
// a sweep consumes the Part 4 event stream and reacts to credits.
// the floor is the whole ethics of the feature: never take a customer
// to zero, or they cannot eat and they default anyway.
async function onCredit(e: LedgerEvent) {
const loan = await loans.activeFor(e.accountId);
if (!loan || !loan.sweepConsented) return;
const available = await ledger.getBalance(e.accountId);
const sweepable = available - loan.protectedFloorMinor; // e.g. ₦5,000
if (sweepable <= 0n) return;
const amount = min(sweepable, loan.outstandingMinor, loan.maxSweepPerEvent);
await ledger.post({
// the key includes the triggering entry, so a redelivered event
// cannot sweep twice.
idempotencyKey: `sweep-${loan.id}-${e.entryId}`,
kind: JournalKind.LOAN_REPAYMENT,
entries: allocateRepayment(loan, amount) // fees → interest → principal
});
}Overdraft: how a limit is communicated and enforced
Your specific question, answered precisely. The overdraft service grants a facility; the ledger enforces a floor. The interesting part is how permission travels between them without either side trusting the other too much.
the wrong design
caller sends: min_balance = −50,000
ledger trusts it.
any service with posting access can grant itself overdraft.
the right design
caller sends: min_balance = −50,000
ledger looks up the facility itself and takes the
more conservative of the two.
the caller can only ever request LESS headroom than granted.-- the ledger's own check, inside the posting transaction.
-- effective_floor = MAX(requested, actually_granted)
-- both are negative, so MAX is the more conservative.
WITH facility AS (
SELECT COALESCE(-limit_amount, 0) AS granted_floor
FROM overdraft_facilities
WHERE account_id = $acct
AND state = 'active'
AND (expires_at IS NULL OR expires_at > now())
-- the kind gate from Part 5: this transaction type may draw it
AND $kind = ANY(allowed_kinds)
)
INSERT INTO entries (journal_id, account_id, amount, currency)
SELECT $jid, $acct, $amt, $ccy
FROM facility f
WHERE (SELECT COALESCE(SUM(amount),0) FROM entries
WHERE account_id = $acct AND currency = $ccy) + $amt
>= GREATEST($requested_floor, f.granted_floor);- Grant. The overdraft service writes a facility row: limit, currency, allowed transaction kinds, expiry. This is the only write that creates permission.
- Publish. A
facility.grantedevent lets the app show available headroom and lets the card authoriser cache it for the Part 8 budget. - Request. A posting caller may pass a floor, which is a request rather than an instruction.
- Enforce. The ledger reads the facility inside the same transaction that checks the balance, and applies the stricter of the two floors.
Savings, interest accrual, and scheduled posting
The simplest integration in the round, and useful precisely because it shows how little a well-designed product service needs to do.
- Hold product configuration: rate, compounding frequency, day-count convention, minimum balance, notice period.
- Run the daily accrual batch from Part 5, sharded and idempotent per day.
- Run the monthly capitalisation, moving accrued interest into the wallet.
- Enforce withdrawal rules, such as notice periods or penalty interest.
- Everything else is a ledger call.
// accrual: the key carries the date, so a re-run posts nothing twice.
await ledger.post({
idempotencyKey: `sav-accrual-${account.id}-${date}`,
kind: JournalKind.INTEREST_ACCRUAL,
entries: [
{ accountId: 'interest_expense', amountMinor: -21_917n, currency: 'NGN' },
{ accountId: `interest_payable:${account.id}`, amountMinor: 21_917n, currency: 'NGN' }
]
});
// capitalisation: monthly, moving accrued interest into spendable money.
await ledger.post({
idempotencyKey: `sav-capitalise-${account.id}-${yearMonth}`,
kind: JournalKind.INTEREST_CAPITALISATION,
entries: [
{ accountId: `interest_payable:${account.id}`, amountMinor: -657_510n, currency: 'NGN' },
{ accountId: account.walletId, amountMinor: 657_510n, currency: 'NGN' }
]
});Which transactions may exceed a limit, and who decides
Your other specific question, and it deserves its own chapter because "who may override a limit" is a governance problem expressed in code.
| Limit | May it be exceeded? | Who authorises | Mechanism |
|---|---|---|---|
| Account balance | Yes, up to the overdraft floor | The facility, granted in advance | Floor in the posting check |
| Per-transaction limit | Yes, with step-up authentication | The customer, by authenticating | A token proving the step-up, single use |
| Daily cumulative | Rarely. Usually KYC-tier bound | Compliance, by tier upgrade | Tier change, not an override |
| Regulatory limit | Never | Nobody. Not the CEO | Hard-coded refusal, no override path exists |
| Velocity / fraud limit | Yes, by review | A fraud analyst, with four-eyes above a threshold | Time-boxed exception on the account |
| Lien / court freeze | Never by us | The issuing authority only | Lift requires a legal instruction |
// an override is a SIGNED, SCOPED, SINGLE-USE, EXPIRING grant.
// it is never a boolean flag on a request, because a boolean can be
// set by anyone who can call the API.
interface LimitOverride {
overrideId: string;
accountId: string;
limitType: 'per_transaction' | 'velocity' | 'daily_cumulative';
maxAmountMinor: bigint; // bounded. not unlimited.
currency: string;
grantedBy: string; // 'customer:step_up' | 'analyst:u-441'
approvedBy?: string; // four-eyes where required
reason: string; // recorded for audit, always
expiresAt: Date; // minutes, not days
singleUse: boolean;
// signed by the granting service so the ledger can verify it
// without trusting the caller that presented it.
signature: string;
}- Bounded. It raises a limit to a specific number, never removes it.
- Scoped. One account, one limit type, one transaction.
- Expiring. Minutes, so a leaked override is nearly worthless.
- Attributed. Who granted it and why, in the append-only audit log from Part 13.
- Verified, not trusted. Signed by the granting service, and consumed atomically by the ledger exactly like the FX quote in Part 7.
Service decomposition: where the seams go
Part 0 asserted that a service owns the entities whose invariants it is responsible for. Here is that rule applied to the whole bank.
| Service | Owns the invariant | Owns the data |
|---|---|---|
| Ledger | Entries sum to zero per journal per currency; balances never breach their floor | journal, entries, accounts, holds |
| Loans | Outstanding equals disbursed minus repaid; schedules are consistent | loans, schedules, facilities |
| Cards | A capture never exceeds its authorisation beyond tolerance | cards, authorisations, disputes |
| Payments | Every outbound payment reaches a terminal state exactly once | outbound_payments, connector state |
| Risk | Every decision is explainable and attributable | rules, decisions, cases, restrictions |
| Customer | Identity and KYC state are consistent and current | customers, kyc, documents |
- No shared database. A service that reads another's tables is coupled to its schema forever, and the boundary is fictional.
- Own your invariant, or do not own the data. If you cannot enforce a rule, the data belongs to whoever can.
- Ask for facts, react to events. Synchronous when you need an answer now, asynchronous when you need to know something happened.
- The ledger is called, never a caller on the money path. It publishes events and answers questions; it does not orchestrate products.
- A cycle in the call graph is a design smell. If A calls B and B calls A, the boundary is in the wrong place or an event should replace one direction.
disbursed − repaid, both of which are
facts loans recorded itself from events it consumed. A boundary that requires
reaching into another service's data is not a boundary.Service-to-service auth: mTLS, SPIFFE, and scoped tokens
Nine services can move money. Knowing which one made each call, and limiting what each may do, is the difference between a compromised service and a compromised bank.
two separate questions, two separate mechanisms:
authentication which service is this? → mTLS
authorisation what may it do? → scoped token
a shared API key answers neither well: it is copyable,
rarely rotated, identical across instances, and grants everything.
- Every service gets a short-lived certificate with a SPIFFE identity such as
spiffe://bank/ns/prod/sa/loans. - Certificates are automatically rotated, typically hourly, so a stolen one expires before it is useful.
- Both sides verify. The ledger verifies the caller, and the caller verifies the ledger, which prevents a rogue service impersonating the ledger.
- The identity is cryptographic, not a header, so it cannot be forged by anything that can reach the network.
- Usually terminated by a service mesh sidecar, so application code does not implement it.
// authorisation: what this identity may do, as explicit policy.
// note it is scoped to ACCOUNT KINDS, not just to methods.
{
"spiffe://bank/ns/prod/sa/loans": {
"ledger.Post": {
"kinds": ["LOAN_DISBURSEMENT", "LOAN_REPAYMENT", "INTEREST_ACCRUAL"],
// may touch loan accounts and customer wallets, and nothing else.
"account_kinds": ["loan_receivable", "interest_receivable",
"customer_wallet", "fee_income"],
"max_amount_minor": 500000000, // ₦5m per posting
"rate_limit_per_sec": 2000
},
"ledger.GetBalance": { "account_kinds": ["customer_wallet"] }
}
}Post. Restricting which account kinds a service may touch is
what contains a compromise. A compromised loans service can move money between loan
accounts and wallets, which is bad; it cannot touch the FX position, the nostro
accounts or the card settlement accounts, which is the difference between an
incident and an existential one. Blast radius is set by the authorisation
model, and that is worth saying explicitly.The client SDK, and what belongs inside it
Nine teams integrating independently will make the same nine mistakes. An SDK is how you make the correct usage the default one.
- Retry with jittered backoff, and critically only on retryable status codes. Never on
INVALID_ARGUMENTorALREADY_EXISTS. - Idempotency key enforcement. The client refuses to send a posting without one, so the most common integration bug becomes impossible.
- Deadline propagation, defaulted sanely and always set.
- Client-side balance validation, so an unbalanced posting fails locally with a clear message instead of as a server error.
- Tracing and metrics, so every caller is instrumented identically and comparably.
- Connection pooling and mTLS configured correctly, once.
- Business logic. If the SDK computes a fee, the fee logic now ships in nine different versions across the bank.
- Client-side caching of balances. Part 3 was explicit: no authorisation decision from a cached balance, and an SDK cache would make that invisible.
- Silent fallbacks. An SDK that quietly returns a stale value on error hides failures from the caller's own monitoring.
- Anything that cannot be upgraded independently. An SDK that must be updated in lockstep with the server has recreated the coupling it was meant to remove.
// the SDK makes the correct thing the easy thing
const ledger = new LedgerClient({
identity: 'spiffe://bank/ns/prod/sa/loans',
defaultDeadlineMs: 250,
retry: { maxAttempts: 3, jitter: true,
retryOn: ['UNAVAILABLE', 'DEADLINE_EXCEEDED', 'RESOURCE_EXHAUSTED'] }
});
// compile error without an idempotency key. not a runtime warning.
await ledger.post({ idempotencyKey, kind, entries });
// and it validates the sum locally first, so a caller bug surfaces
// as "entries do not balance: NGN sums to -1000" rather than as a 400.Versioning a money API without breaking anyone
Nine teams on nine release schedules, and the ledger cannot have a coordinated deployment. The rules are the same as Part 4's schema evolution, applied to RPC.
| Change | Safe? | Rule |
|---|---|---|
| Add an optional field | Yes | New field number, never reuse an old one |
| Add an RPC method | Yes | Old clients simply do not call it |
| Add an enum value | Careful | Old clients see it as unknown. They must handle a default rather than crash |
| Rename a field | Yes on the wire | Protobuf uses field numbers, so the name is cosmetic. It does break generated code |
| Remove a field | No | Deprecate, reserve the number, remove after confirming no caller reads it |
| Change a field's type | Never | Add a new field instead |
| Change a field's meaning | Never, and it is the worst one | Silent and invisible in review. Changing an amount from major to minor units would be catastrophic and compile cleanly |
message EntrySpec {
string account_id = 1;
int64 amount_minor = 2;
string currency = 3;
optional int64 min_balance_minor = 4;
// removed in v2.3. the number is reserved so it can NEVER be reused
// for a different meaning, which would silently corrupt old callers.
reserved 5;
reserved "legacy_floor";
// added in v2.4, optional with a safe default.
optional string value_date = 6; // Part 10's second date
}- Announce, with a migration guide and a date.
- Instrument. Count calls to the deprecated field or method by service identity, which mTLS makes possible.
- Chase the specific teams still using it. The metric names them, so this is a short list rather than a broadcast.
- Wait for zero usage for a sustained period. Never remove on a date alone; remove on evidence.
- Reserve the field number permanently and remove the code.
deprecated_field_use{field, caller}, and the migration becomes
three specific conversations instead of an announcement nobody reads. The auth
design from chapter 169 turns out to be what makes API evolution tractable,
which is a connection worth pointing out.Sketch v14: the integration surface
What changed, and the cost accepted
| Change | Driven by | Cost accepted |
|---|---|---|
| gRPC internally, REST at the edge | Typed contracts where we control both ends; accessibility where we do not | Two API surfaces to maintain and document |
| One posting API taking balanced entries | The ledger must never learn product rules | Callers must construct entries, so the SDK matters |
| Server-side invariant validation | A caller bug must not create money | A validation pass on every posting |
| Floors verified against the facility | A caller must not grant itself overdraft | A facility lookup inside the posting transaction |
| Signed, expiring, single-use overrides | Limit exceptions need governance, not a boolean | A grant-and-consume flow, like the Part 7 FX quote |
| mTLS plus scope by account kind | Blast radius of a compromised service | A service mesh, and per-service policy to maintain |
| SSE for live updates | Simpler and resumable versus WebSocket | One-directional only, which is all we needed |
| ISO 8583 gateway isolated | Stateful sockets inside a stateless architecture | A component that cannot autoscale or deploy casually |