Part 14 · 21 chapters · ~20 min

Round fourteen: “how does anything else talk to it?”

Thirteen rounds have treated service boundaries as arrows on a whiteboard. This round makes them real: which protocol, and why, for each kind of conversation. Then it designs the single interface every other service in the bank uses to move money, and walks through exactly how loans, overdraft and savings call it, including the question of who is allowed to exceed a limit and how that permission travels.

152

The pressure: nine services, one ledger

interviewer

“Loans, savings, cards, overdraft, working capital, fraud, notifications, support and the mobile backend all need to talk to the ledger. What protocol, what interface, and how do you stop nine teams breaking it?”

Three separate questions wearing one coat, and separating them is the first move again.

what is actually being asked
  1. Transport. Which protocol for which conversation shape, and why.
  2. Contract. What operations exist, and what does the ledger refuse to know about.
  3. Governance. Authentication, authorisation, versioning, and blast radius when a caller is wrong.
worked numbers
          the conversations, by shape:


request/response, low latency   post entries, check a balance

request/response, slow        statement generation, bulk export

fire and forget                an event happened, react if you care

server → client push           live balance on an open app screen

persistent bidirectional       card network authorisation traffic


five shapes. picking one protocol for all of them is the mistake.
non-functional, new
  1. Internal posting call adds under 5 ms of transport overhead.
  2. A misbehaving caller cannot degrade the ledger for other callers.
  3. No caller can violate the invariant, even with a bug.
  4. The API can evolve without coordinated deployment of nine services.
  5. Every call is authenticated, authorised and attributed to a service identity.
153

Choosing a protocol: the honest decision table

The answer that earns credit maps conversation shape to protocol, with a reason for each. Here is the whole table before the chapters that justify it.

gRPCservice to service The posting API, balance queries, every internal call. Binary framing, HTTP/2 multiplexing, generated clients, and a schema that is enforced rather than documented. The latency and the type safety both matter on a path nine teams call.
REST / JSONexternal and public Partner APIs, webhooks we receive, anything a third party integrates with. Debuggable with curl, cacheable, universally understood, and no toolchain requirement on the caller.
Kafkafan-out of facts Everything from Part 4. One publish, many independent consumers, replayable. Used when the producer must not know or care who is listening.
SSEserver to browser Live balance and transaction feed on an open screen. One-directional, plain HTTP, automatic reconnection with event ids. Simpler than WebSockets and sufficient, which is why it usually wins.
WebSocketbidirectional, stateful Support chat, trading screens, anything where the client sends a continuous stream too. Only when the bidirectionality is genuinely needed, because it costs connection state and a backpressure design.
ISO 8583 over TCPcard networks Not a choice. The scheme dictates it: persistent sockets, binary bitmap-framed messages, and echo tests. Part 8's authorisation path lives here.
UDPtwo narrow uses Metrics emission and DNS. Never for money. Chapter 160 explains why the one property that makes UDP fast is exactly the one a payment cannot tolerate.
the meta-point about protocol questions
An interviewer asking "REST or gRPC?" is rarely asking which is better. They are checking whether you have a reason. "gRPC internally for typed, low-latency calls between services we control, REST externally because partners should not need our toolchain" is a complete answer in one sentence, and it is complete because it names the different constraints on each side.
154

TCP, and what it guarantees you already rely on

Everything above except UDP runs on TCP. Being able to say what TCP does and does not promise is what makes the later choices defensible.

TCP guaranteesTCP does not guarantee
Bytes arrive in orderThat the application processed them
Bytes are not corrupted, per its checksumThat the peer is still alive right now
Lost segments are retransmittedMessage boundaries. It is a byte stream
Flow control, so a fast sender cannot drown a slow receiverBounded latency. A retransmit can take seconds
the second column is where payment bugs come from
A successful write() means the bytes are in a kernel buffer, not that the peer received or processed them. This is precisely why Part 6's unknown state exists: a TCP write can succeed, the connection can drop, and you genuinely cannot tell whether the payment was processed. The transport's guarantees stop well short of the application's needs, and every idempotency key in this design exists to cover that gap.

The settings that matter on a money path

code
// a dead peer with no traffic is indistinguishable from an idle one.
// without keepalive, a half-open connection can sit for HOURS.
TCP_KEEPIDLE  = 30   // start probing after 30 s idle
TCP_KEEPINTVL = 10   // probe every 10 s
TCP_KEEPCNT   = 3    // 3 failures ⇒ dead. detected in ~60 s.

// Nagle batches small writes to save packets, adding up to 40 ms.
// on a request/response path that is pure added latency.
TCP_NODELAY = true

// and a connect timeout, because the OS default can be over a minute,
// which is far longer than any budget we set in Part 8.
connectTimeoutMs = 2000
155

HTTP/1.1, HTTP/2, and head-of-line blocking

HTTP/1.1HTTP/2
Concurrency per connectionOne request at a timeMany interleaved streams
Head-of-line blockingAt the application layerRemoved at the application layer, remains at TCP
FramingText, newline-delimitedBinary frames with stream ids
HeadersRepeated in full on every requestHPACK compressed, incrementally encoded
Server pushNoYes, though largely deprecated in practice
Connection count6 to 8 per origin, by conventionOne, which is why connection pooling changes shape
the honesty that impresses: HTTP/2 does not remove all of it
HTTP/2 removes head-of-line blocking at the application layer. TCP still delivers bytes in order, so a single lost packet stalls every stream on that connection until it is retransmitted. That is why HTTP/3 moved to QUIC over UDP, implementing per-stream reliability above an unordered transport. Saying "HTTP/2 fixes head-of-line blocking" is the common half-answer; naming which layer it fixes and which it does not is the full one.

The practical consequence for us

worked numbers
          HTTP/1.1, 8 connections, one slow 900 ms request:

            → 1 of 8 connections blocked. 12% capacity loss.


          HTTP/2, 1 connection, one slow 900 ms stream:

            → other streams continue. no application-level loss.

            → but one lost packet stalls all of them briefly.


so: HTTP/2 for internal service calls, and keep per-stream

deadlines, because multiplexing removes queueing but not slowness.
head-of-line blocking
HTTP/1.1 pipelining vs HTTP/2 multiplexing
swipe the figure sideways, or tap expand for full screen
1/7
four requests
Four requests share one HTTP/1.1 connection. The first is slow, because it builds a statement.
156

gRPC internals: HTTP/2 streams and protobuf framing

worked numbers
          gRPC = HTTP/2 + protobuf + a calling convention


          one RPC = one HTTP/2 stream

            request  = HEADERS frame + DATA frames

            response = HEADERS + DATA + trailers carrying the status


          each message in a DATA frame is length-prefixed:

            [1 byte compressed flag][4 byte length][payload]


the status in TRAILERS is why gRPC can stream a response

and still report a failure after the body has started.

Why protobuf rather than JSON, concretely

PropertyJSONProtobuf
A typical posting message~420 bytes~90 bytes
Parse costString scanning, allocation per fieldBinary, length-prefixed, minimal allocation
SchemaConvention, or a separate validatorCompiled. The wire format is the contract
Field evolutionBy name, forgiving and untrackedBy field number, with explicit rules
Large integersDangerous. Parsed as doubles, so money loses precisionint64 is exact
DebuggabilityReadable in any toolNeeds the schema to decode
the money-specific reason, which settles it internally
Part 4 had to transmit amounts as strings in JSON, because JSON numbers are IEEE 754 doubles in most parsers and 9007199254740993 silently becomes ...992. Protobuf has a real int64, so an amount is an integer on the wire with no encoding workaround. For a system whose entire correctness argument rests on exact integer arithmetic, that is not a minor convenience.
code
// the four call types, and where each is used here
service Ledger {
  // 1. unary: the overwhelming majority. post, query, resolve.
  rpc Post(PostRequest) returns (PostResponse);

  // 2. server streaming: a large export without buffering it all.
  rpc StreamEntries(StreamRequest) returns (stream Entry);

  // 3. client streaming: bulk upload, e.g. a payroll batch.
  rpc PostBatch(stream PostRequest) returns (BatchSummary);

  // 4. bidirectional: rare internally. used by the card authoriser
  //    to keep a warm channel with continuous health signalling.
  rpc AuthChannel(stream AuthMsg) returns (stream AuthReply);
}
gRPC features that matter on a money path
  1. Deadlines propagate automatically. A deadline set by the caller travels with the request and every downstream hop sees the remaining time. This is the Part 8 budget mechanism, built into the protocol.
  2. Cancellation propagates. If the caller gives up, downstream work stops rather than continuing to hold locks.
  3. Status codes are structured, so RESOURCE_EXHAUSTED and FAILED_PRECONDITION are distinguishable without parsing a message string.
  4. Interceptors give one place for auth, tracing, metrics and idempotency handling across every method.
157

When gRPC is wrong, and REST is right

Being able to argue against your own default is what makes the default credible.

SituationUseWhy
Internal service to servicegRPCWe control both ends, deploy both toolchains, and want typed contracts and deadlines
A public or partner APIRESTA partner should not need protoc, a code generator or an HTTP/2-capable client to send us a payment
Webhooks we receiveRESTThe sender decides, and every provider sends JSON over HTTP/1.1
Browser clientsREST or SSEBrowsers cannot speak raw gRPC without a proxy such as gRPC-Web
Anything an operator debugs by handRESTcurl at 3am beats a schema-dependent binary decoder, and Part 13 cares about that
Through unknown middleboxesRESTOld proxies mangle HTTP/2 or trailers, and diagnosing that at a partner is painful
the summary to give
"gRPC inside the trust boundary, REST at the edges. The deciding question is whether I control the client. If I do, a typed contract with automatic deadline propagation is worth the toolchain. If I do not, forcing my toolchain on a partner is a cost I am imposing to solve a problem I do not have."
158

WebSockets: the balance feed, and backpressure

A customer has the app open. Their balance should update when money moves. That is a server-to-client push problem, and WebSockets are the tool most people reach for first.

worked numbers
          upgrade handshake: HTTP GET + Upgrade: websocket

          → 101 Switching Protocols

          → the TCP connection is now a bidirectional frame stream


          costs, at 20M customers:

            · ~2% concurrently connected = 400,000 open sockets

            · each holds file descriptors, buffers, and server memory

            · connections are stateful, so a deploy disconnects everyone

            · load balancers must support long-lived upgrades


~50k sockets per node ⇒ 8 nodes purely to hold connections

Backpressure, which is where naive implementations fail

the failure, step by step
  1. A client is on a slow mobile connection and stops reading.
  2. The server keeps writing balance updates into the socket.
  3. The kernel send buffer fills, then the application's own queue grows.
  4. Server memory climbs, one buffer per slow client.
  5. With thousands of slow clients, the node runs out of memory and drops every connection it holds, including the healthy ones.
code
// the fix: a bounded per-connection queue with an explicit policy for
// what to do when it fills. never an unbounded buffer.
class BalanceFeed {
  private queue: Update[] = [];
  private static MAX = 32;

  push(u: Update) {
    if (this.queue.length >= BalanceFeed.MAX) {
      // for a BALANCE, only the latest value matters: collapse.
      // this is the key insight: the semantics of the data decide
      // the shedding policy.
      this.queue = [this.queue[0], u];
      this.markResyncNeeded();   // tell the client to refetch
      return;
    }
    this.queue.push(u);
  }
}
the general principle
Backpressure policy is decided by what the data means. For a balance, only the newest value matters, so old updates can be dropped and the client told to resync. For a transaction list, every item matters, so you must slow down or disconnect rather than drop. Deciding "drop, buffer, or disconnect" per data type is the design, and an unbounded queue is the absence of a decision.
159

Server-sent events, and why they often win

For our actual requirement, which is server to client only, SSE is simpler and better in almost every dimension.

WebSocketSSE
DirectionBidirectionalServer to client only
ProtocolA separate protocol after upgradePlain HTTP. Works through every proxy
ReconnectionYou implement itAutomatic, built into the browser
Missed messagesYou implement resumptionLast-Event-ID header, built in
CompressionPer-message extensionStandard HTTP gzip
AuthAwkward: no custom headers on upgrade in browsersNormal headers and cookies
BinaryYesText only, which is fine for JSON
code
// SSE with resumption. the id: field is what makes reconnect lossless,
// and we already have a perfect monotonic id: the entry id from Part 1.
res.writeHead(200, {
  'Content-Type': 'text/event-stream',
  'Cache-Control': 'no-cache',
  'X-Accel-Buffering': 'no'    // stop nginx buffering the stream
});

// the browser sends Last-Event-ID on reconnect, so we resume exactly.
const resumeFrom = req.headers['last-event-id'] ?? '0';

for await (const e of entriesSince(accountId, BigInt(resumeFrom))) {
  res.write(`id: ${e.id}\n`);
  res.write(`event: balance\n`);
  res.write(`data: ${JSON.stringify({ balance: e.balanceAfter.toString() })}\n\n`);
}
the connection back to round one
SSE resumption needs a monotonic, unique event id, and we have had one since chapter 17: the BIGSERIAL entry id. The same property that made balance snapshots safe in Part 3 and made ordering work in Part 4 now makes a reconnecting client lossless. Monotonic ids keep paying off, which is why "prefer monotonic over random" appeared in the Part 0 vocabulary.
the recommendation, stated as a tradeoff
  1. SSE for the balance and transaction feed. Simpler, resumable, works everywhere.
  2. WebSocket only where the client genuinely streams upward too, such as support chat.
  3. Polling is still correct for low-frequency updates. A 30-second poll costs nothing and holds no state, and 400,000 idle sockets to deliver an update every few minutes is a bad trade.
160

UDP, and the two places a bank actually uses it

worked numbers
          UDP: no handshake, no ordering, no retransmission, no flow control


the single property that makes it fast

            fire and forget: no round trip, no connection state


the same property, restated

            the sender cannot know whether it arrived


for a payment, that is the entire problem we spent Part 6 solving.
where UDP is correct here
  1. Metrics emission. StatsD over UDP. Losing 0.1% of metric samples changes no decision, and the application must never block on its own telemetry. Metrics over TCP can make an observability outage into a service outage.
  2. DNS. Every service call resolves a name first, over UDP, with a TCP fallback for large responses.
where UDP appears without being chosen
  1. QUIC, and therefore HTTP/3. Reliability is reimplemented in userspace above UDP, per stream, which is how it escapes TCP's head-of-line blocking.
  2. NTP. Clock sync, which matters more than it seems: Part 10's booking dates and Part 9's velocity windows both depend on clocks being close.
the answer to "would you use UDP for payments"
"No, and the reason is not performance. UDP's defining property is that the sender learns nothing about delivery, which is exactly the ambiguity Part 6 spent a whole round making safe. I would be reimplementing acknowledgement, ordering and retransmission in userspace to reach TCP's guarantees, which is what QUIC does, and QUIC exists to solve head-of-line blocking rather than to avoid reliability. Where I do use UDP is metrics, because losing a sample is free and blocking on telemetry is not."
161

ISO 8583 and the persistent socket reality of cards

Part 8 treated the card network as an arrow. Here is what is actually on the wire, and it is unlike everything else in this design.

worked numbers
[2 or 4 byte length][MTI][bitmap][data elements…]


MTI  message type: 0100 auth request, 0110 auth response,

                 0200 financial, 0400 reversal, 0800 network echo


bitmap  64 or 128 bits. bit N set ⇒ data element N present.

                    DE2 = PAN, DE4 = amount, DE39 = response code,

                    DE7 = transmission time, DE11 = trace number


a compact binary format designed in the 1980s for expensive links,

and still carrying most card traffic on earth.
what makes it operationally different from everything else here
  1. Persistent sockets. A small number of long-lived TCP connections to the scheme, not a connection per request. Establishing one involves certificates and, historically, paperwork.
  2. Correlation by field, not by connection. Responses can arrive out of order on the same socket, matched by the STAN (DE11) plus the terminal and date. It is multiplexing, hand-rolled, decades before HTTP/2.
  3. Echo tests. A 0800 network management message every 30 seconds proves the link is alive. Missing echoes mean the link is down even with no traffic.
  4. Reversals are a message type. A 0400 reversal is first-class, because a timed-out authorisation must be explicitly unwound rather than left ambiguous.
  5. Strict timeouts, enforced by the scheme. Exceed them and the network stands in on your behalf, using rules you configured weeks earlier. This is the Part 8 constraint, in protocol form.
the design consequence worth stating
The ISO 8583 gateway is a stateful, connection-oriented service inside an otherwise stateless architecture. It cannot autoscale on CPU, it cannot be redeployed casually, and it holds scheme credentials. So it is isolated, given its own deployment lifecycle and its own PCI boundary, and it exposes ordinary gRPC to the rest of the bank. Wrap the protocol you cannot change in an interface you can, which is the same move as the Part 6 connector.
162

The posting API: one interface, every caller

The contract at the centre of the bank. Part 5 sketched it; here it is properly specified, and every decision in it is defensive.

code
service Ledger {
  rpc Post(PostRequest) returns (PostResponse);
  rpc GetBalance(BalanceRequest) returns (BalanceResponse);
  rpc PlaceHold(HoldRequest) returns (HoldResponse);
  rpc ResolveHold(ResolveHoldRequest) returns (PostResponse);
  rpc GetJournal(JournalRequest) returns (Journal);
}

message PostRequest {
  // REQUIRED. the caller's own natural key for this intent.
  // "loan-disb-{loanId}", "accrual-{loanId}-{date}". see Part 5.
  string idempotency_key = 1;

  // what KIND of business event. drives GL mapping and reporting,
  // and it is an enum so a typo is a compile error.
  JournalKind kind = 2;

  // the entries. the ledger validates they sum to zero PER CURRENCY
  // and REFUSES otherwise. a caller bug cannot break the invariant.
  repeated EntrySpec entries = 3;

  // opaque to the ledger. stored, indexed, never interpreted.
  map<string, string> metadata = 4;

  // links a correction to what it corrects. Part 10.
  optional string reverses_journal_id = 5;
}

message EntrySpec {
  string account_id = 1;
  int64  amount_minor = 2;    // signed. negative = debit. int64, exact.
  string currency = 3;
  // the floor for THIS entry: 0 normally, negative for overdraft.
  // the ledger enforces it inside the insert, per Part 1.
  optional int64 min_balance_minor = 4;
}
the five defensive properties, and what each prevents
  1. Entries in, not a product operation. The ledger has no disburseLoan(), so it never accumulates product knowledge. Prevents the ledger becoming the place every product rule lives.
  2. Invariant validated server-side. Unbalanced entries are refused. Prevents a caller bug from creating money.
  3. Idempotency key required, not optional. Prevents the most common integration error, which is a retry that double-posts.
  4. The floor is per entry. Prevents an overdraft-aware caller from accidentally granting overdraft on an unrelated account.
  5. Metadata is opaque. Prevents callers from pressuring the ledger schema to grow a column per product.
the sentence that captures the whole design
"The ledger exposes one write operation that accepts balanced entries and validates the invariant before writing. It knows debits, credits, currencies and zero, and nothing about loans, cards or savings. That is why five products in Part 5 and cards in Part 8 required no change to the entries table: product knowledge lives in product services, and the ledger's contract is arithmetic."
163

Loan disbursement, as a ledger call

The full path, end to end, showing where each responsibility sits.

code
// LOANS SERVICE owns: eligibility, pricing, schedule, product rules.
// LEDGER owns: that the entries balance and are durably recorded.
const res = await ledger.post({
  idempotencyKey: `loan-disb-${loan.id}`,   // natural, one per loan
  kind: JournalKind.LOAN_DISBURSEMENT,
  entries: [
    { accountId: `loan_receivable:${loan.id}`, amountMinor: -50_000_000n,
      currency: 'NGN' },
    { accountId: customer.ngnWalletId,          amountMinor:  49_500_000n,
      currency: 'NGN' },
    { accountId: 'fee_income:origination',      amountMinor:     500_000n,
      currency: 'NGN' }
  ],
  metadata: { loan_id: loan.id, product: 'salary_advance', tenor_days: '30' }
});
// if the loans service has a bug and these do not sum to zero,
// the ledger returns INVALID_ARGUMENT and writes nothing.
where each responsibility lives, and why
  1. Loans service: is this customer eligible, what rate, what schedule, what fee. All product policy.
  2. Ledger: do these entries balance, does this key already exist, is the account real, is it restricted. All correctness.
  3. Neither: whether the customer wanted a loan. That is the app, and it is upstream of both.
  4. The event published afterwards is what activates the repayment schedule, so the loans service reacts to the confirmed fact rather than assuming its own call succeeded.
why the schedule activates on the event, not on the response
If the loans service activated the schedule on the RPC response, a response lost in transit would leave a disbursed loan with no repayment schedule. Reacting to the published event instead means the schedule is created from a fact the ledger confirmed, and a redelivered event is handled by the consumer idempotency from Part 4. React to events, not to your own optimism.
loan disbursement
who decides what, and what crosses the wire
swipe the figure sideways, or tap expand for full screen
1/8
request
A customer requests a loan in the app, which reaches the loans service.
164

Loan recovery, sweeps, and partial repayment

Getting the money back is harder than lending it, and the mechanisms are worth knowing because they are where lending products actually differ.

MechanismHow it worksRisk
Scheduled debitOn the due date, debit the wallet for the instalmentFails if the balance is short. Needs a retry policy
SweepWhenever a credit lands, take a portion toward the loanVery effective, and the most customer-hostile. Needs clear consent and a floor
Salary assignmentThe employer remits directly before the customer sees itStrongest recovery. Requires an employer relationship
Direct debit on an external accountPull from another bank via the Part 6 railsReversible for a period, so the money is not certain
Overdraft absorptionLet the repayment push the wallet into overdraftConverts one debt into another, and is sometimes correct

The sweep, and the floor that makes it acceptable

code
// a sweep consumes the Part 4 event stream and reacts to credits.
// the floor is the whole ethics of the feature: never take a customer
// to zero, or they cannot eat and they default anyway.
async function onCredit(e: LedgerEvent) {
  const loan = await loans.activeFor(e.accountId);
  if (!loan || !loan.sweepConsented) return;

  const available = await ledger.getBalance(e.accountId);
  const sweepable = available - loan.protectedFloorMinor;   // e.g. ₦5,000
  if (sweepable <= 0n) return;

  const amount = min(sweepable, loan.outstandingMinor, loan.maxSweepPerEvent);

  await ledger.post({
    // the key includes the triggering entry, so a redelivered event
    // cannot sweep twice.
    idempotencyKey: `sweep-${loan.id}-${e.entryId}`,
    kind: JournalKind.LOAN_REPAYMENT,
    entries: allocateRepayment(loan, amount)   // fees → interest → principal
  });
}
the two details that matter most
The idempotency key includes the triggering entry id, so a redelivered event sweeps once rather than repeatedly. And the protected floor is the difference between a recovery mechanism and a harmful one: sweeping a customer to zero maximises today's recovery and produces tomorrow's default. Both are one-line decisions that define whether the product is defensible.
165

Overdraft: how a limit is communicated and enforced

Your specific question, answered precisely. The overdraft service grants a facility; the ledger enforces a floor. The interesting part is how permission travels between them without either side trusting the other too much.

worked numbers
the wrong design

            caller sends: min_balance = −50,000

            ledger trusts it.

            any service with posting access can grant itself overdraft.


the right design

            caller sends: min_balance = −50,000

            ledger looks up the facility itself and takes the

            more conservative of the two.


the caller can only ever request LESS headroom than granted.
code
-- the ledger's own check, inside the posting transaction.
-- effective_floor = MAX(requested, actually_granted)
-- both are negative, so MAX is the more conservative.
WITH facility AS (
  SELECT COALESCE(-limit_amount, 0) AS granted_floor
    FROM overdraft_facilities
   WHERE account_id = $acct
     AND state = 'active'
     AND (expires_at IS NULL OR expires_at > now())
     -- the kind gate from Part 5: this transaction type may draw it
     AND $kind = ANY(allowed_kinds)
)
INSERT INTO entries (journal_id, account_id, amount, currency)
SELECT $jid, $acct, $amt, $ccy
  FROM facility f
 WHERE (SELECT COALESCE(SUM(amount),0) FROM entries
          WHERE account_id = $acct AND currency = $ccy) + $amt
       >= GREATEST($requested_floor, f.granted_floor);
how the limit travels, in four steps
  1. Grant. The overdraft service writes a facility row: limit, currency, allowed transaction kinds, expiry. This is the only write that creates permission.
  2. Publish. A facility.granted event lets the app show available headroom and lets the card authoriser cache it for the Part 8 budget.
  3. Request. A posting caller may pass a floor, which is a request rather than an instruction.
  4. Enforce. The ledger reads the facility inside the same transaction that checks the balance, and applies the stricter of the two floors.
the security principle, stated generally
A caller may request a restriction, never an expansion. Any parameter that could loosen a rule must be verified against authoritative state on the server side, inside the same transaction. This is the same instinct as the Part 8 webhook rule about never trusting the amount in a callback, and it generalises to every API that enforces a policy.
166

Savings, interest accrual, and scheduled posting

The simplest integration in the round, and useful precisely because it shows how little a well-designed product service needs to do.

the savings service's entire job
  1. Hold product configuration: rate, compounding frequency, day-count convention, minimum balance, notice period.
  2. Run the daily accrual batch from Part 5, sharded and idempotent per day.
  3. Run the monthly capitalisation, moving accrued interest into the wallet.
  4. Enforce withdrawal rules, such as notice periods or penalty interest.
  5. Everything else is a ledger call.
code
// accrual: the key carries the date, so a re-run posts nothing twice.
await ledger.post({
  idempotencyKey: `sav-accrual-${account.id}-${date}`,
  kind: JournalKind.INTEREST_ACCRUAL,
  entries: [
    { accountId: 'interest_expense',               amountMinor: -21_917n, currency: 'NGN' },
    { accountId: `interest_payable:${account.id}`, amountMinor:  21_917n, currency: 'NGN' }
  ]
});

// capitalisation: monthly, moving accrued interest into spendable money.
await ledger.post({
  idempotencyKey: `sav-capitalise-${account.id}-${yearMonth}`,
  kind: JournalKind.INTEREST_CAPITALISATION,
  entries: [
    { accountId: `interest_payable:${account.id}`, amountMinor: -657_510n, currency: 'NGN' },
    { accountId: account.walletId,                  amountMinor:  657_510n, currency: 'NGN' }
  ]
});
the pattern across all three products
Loans, overdraft and savings each amount to configuration plus a schedule plus balanced entries. None of them writes to the ledger database, none of them needs a schema change, and all three use natural idempotency keys that make re-runs free. When a fourth product arrives, the work is a service and a set of posting recipes, which is the return on the round-one decision, restated one final time.
167

Which transactions may exceed a limit, and who decides

Your other specific question, and it deserves its own chapter because "who may override a limit" is a governance problem expressed in code.

LimitMay it be exceeded?Who authorisesMechanism
Account balanceYes, up to the overdraft floorThe facility, granted in advanceFloor in the posting check
Per-transaction limitYes, with step-up authenticationThe customer, by authenticatingA token proving the step-up, single use
Daily cumulativeRarely. Usually KYC-tier boundCompliance, by tier upgradeTier change, not an override
Regulatory limitNeverNobody. Not the CEOHard-coded refusal, no override path exists
Velocity / fraud limitYes, by reviewA fraud analyst, with four-eyes above a thresholdTime-boxed exception on the account
Lien / court freezeNever by usThe issuing authority onlyLift requires a legal instruction
code
// an override is a SIGNED, SCOPED, SINGLE-USE, EXPIRING grant.
// it is never a boolean flag on a request, because a boolean can be
// set by anyone who can call the API.
interface LimitOverride {
  overrideId: string;
  accountId: string;
  limitType: 'per_transaction' | 'velocity' | 'daily_cumulative';
  maxAmountMinor: bigint;      // bounded. not unlimited.
  currency: string;
  grantedBy: string;           // 'customer:step_up' | 'analyst:u-441'
  approvedBy?: string;         // four-eyes where required
  reason: string;              // recorded for audit, always
  expiresAt: Date;             // minutes, not days
  singleUse: boolean;
  // signed by the granting service so the ledger can verify it
  // without trusting the caller that presented it.
  signature: string;
}
the five properties every override must have
  1. Bounded. It raises a limit to a specific number, never removes it.
  2. Scoped. One account, one limit type, one transaction.
  3. Expiring. Minutes, so a leaked override is nearly worthless.
  4. Attributed. Who granted it and why, in the append-only audit log from Part 13.
  5. Verified, not trusted. Signed by the granting service, and consumed atomically by the ledger exactly like the FX quote in Part 7.
the line that should never have an override path
A regulatory limit has no override mechanism at all, and that is a design decision rather than an oversight. If an override path exists, it will eventually be used under commercial pressure, and the existence of the capability is itself the audit finding. Some rules are enforced by not building the ability to break them, which is the same instinct as the cross-region schema in Part 7 and the revoked privileges in Part 2.
168

Service decomposition: where the seams go

Part 0 asserted that a service owns the entities whose invariants it is responsible for. Here is that rule applied to the whole bank.

ServiceOwns the invariantOwns the data
LedgerEntries sum to zero per journal per currency; balances never breach their floorjournal, entries, accounts, holds
LoansOutstanding equals disbursed minus repaid; schedules are consistentloans, schedules, facilities
CardsA capture never exceeds its authorisation beyond tolerancecards, authorisations, disputes
PaymentsEvery outbound payment reaches a terminal state exactly onceoutbound_payments, connector state
RiskEvery decision is explainable and attributablerules, decisions, cases, restrictions
CustomerIdentity and KYC state are consistent and currentcustomers, kyc, documents
the seam rules
  1. No shared database. A service that reads another's tables is coupled to its schema forever, and the boundary is fictional.
  2. Own your invariant, or do not own the data. If you cannot enforce a rule, the data belongs to whoever can.
  3. Ask for facts, react to events. Synchronous when you need an answer now, asynchronous when you need to know something happened.
  4. The ledger is called, never a caller on the money path. It publishes events and answers questions; it does not orchestrate products.
  5. A cycle in the call graph is a design smell. If A calls B and B calls A, the boundary is in the wrong place or an event should replace one direction.
the test to apply to any proposed boundary
"Can this service enforce its invariant using only data it owns?" If loans needed the ledger's tables to know an outstanding balance, the boundary is wrong. It does not: outstanding is disbursed − repaid, both of which are facts loans recorded itself from events it consumed. A boundary that requires reaching into another service's data is not a boundary.
169

Service-to-service auth: mTLS, SPIFFE, and scoped tokens

Nine services can move money. Knowing which one made each call, and limiting what each may do, is the difference between a compromised service and a compromised bank.

worked numbers
          two separate questions, two separate mechanisms:


authentication  which service is this?   → mTLS

authorisation   what may it do?      → scoped token


a shared API key answers neither well: it is copyable,

          rarely rotated, identical across instances, and grants everything.
        
mTLS with workload identity
  1. Every service gets a short-lived certificate with a SPIFFE identity such as spiffe://bank/ns/prod/sa/loans.
  2. Certificates are automatically rotated, typically hourly, so a stolen one expires before it is useful.
  3. Both sides verify. The ledger verifies the caller, and the caller verifies the ledger, which prevents a rogue service impersonating the ledger.
  4. The identity is cryptographic, not a header, so it cannot be forged by anything that can reach the network.
  5. Usually terminated by a service mesh sidecar, so application code does not implement it.
code
// authorisation: what this identity may do, as explicit policy.
// note it is scoped to ACCOUNT KINDS, not just to methods.
{
  "spiffe://bank/ns/prod/sa/loans": {
    "ledger.Post": {
      "kinds": ["LOAN_DISBURSEMENT", "LOAN_REPAYMENT", "INTEREST_ACCRUAL"],
      // may touch loan accounts and customer wallets, and nothing else.
      "account_kinds": ["loan_receivable", "interest_receivable",
                        "customer_wallet", "fee_income"],
      "max_amount_minor": 500000000,     // ₦5m per posting
      "rate_limit_per_sec": 2000
    },
    "ledger.GetBalance": { "account_kinds": ["customer_wallet"] }
  }
}
why scoping by account kind is the important part
Restricting methods is obvious and weak: every product service needs Post. Restricting which account kinds a service may touch is what contains a compromise. A compromised loans service can move money between loan accounts and wallets, which is bad; it cannot touch the FX position, the nostro accounts or the card settlement accounts, which is the difference between an incident and an existential one. Blast radius is set by the authorisation model, and that is worth saying explicitly.
170

The client SDK, and what belongs inside it

Nine teams integrating independently will make the same nine mistakes. An SDK is how you make the correct usage the default one.

what belongs in the SDK
  1. Retry with jittered backoff, and critically only on retryable status codes. Never on INVALID_ARGUMENT or ALREADY_EXISTS.
  2. Idempotency key enforcement. The client refuses to send a posting without one, so the most common integration bug becomes impossible.
  3. Deadline propagation, defaulted sanely and always set.
  4. Client-side balance validation, so an unbalanced posting fails locally with a clear message instead of as a server error.
  5. Tracing and metrics, so every caller is instrumented identically and comparably.
  6. Connection pooling and mTLS configured correctly, once.
what must never be in the SDK
  1. Business logic. If the SDK computes a fee, the fee logic now ships in nine different versions across the bank.
  2. Client-side caching of balances. Part 3 was explicit: no authorisation decision from a cached balance, and an SDK cache would make that invisible.
  3. Silent fallbacks. An SDK that quietly returns a stale value on error hides failures from the caller's own monitoring.
  4. Anything that cannot be upgraded independently. An SDK that must be updated in lockstep with the server has recreated the coupling it was meant to remove.
code
// the SDK makes the correct thing the easy thing
const ledger = new LedgerClient({
  identity: 'spiffe://bank/ns/prod/sa/loans',
  defaultDeadlineMs: 250,
  retry: { maxAttempts: 3, jitter: true,
           retryOn: ['UNAVAILABLE', 'DEADLINE_EXCEEDED', 'RESOURCE_EXHAUSTED'] }
});

// compile error without an idempotency key. not a runtime warning.
await ledger.post({ idempotencyKey, kind, entries });

// and it validates the sum locally first, so a caller bug surfaces
// as "entries do not balance: NGN sums to -1000" rather than as a 400.
the principle
An SDK encodes the operational contract, never the business contract. Retry policy, deadlines, idempotency and instrumentation are operational and belong in one shared place. Fees, eligibility and pricing are business logic and belong in the service that owns them. Mixing the two is how an SDK becomes a distributed monolith.
171

Versioning a money API without breaking anyone

Nine teams on nine release schedules, and the ledger cannot have a coordinated deployment. The rules are the same as Part 4's schema evolution, applied to RPC.

ChangeSafe?Rule
Add an optional fieldYesNew field number, never reuse an old one
Add an RPC methodYesOld clients simply do not call it
Add an enum valueCarefulOld clients see it as unknown. They must handle a default rather than crash
Rename a fieldYes on the wireProtobuf uses field numbers, so the name is cosmetic. It does break generated code
Remove a fieldNoDeprecate, reserve the number, remove after confirming no caller reads it
Change a field's typeNeverAdd a new field instead
Change a field's meaningNever, and it is the worst oneSilent and invisible in review. Changing an amount from major to minor units would be catastrophic and compile cleanly
code
message EntrySpec {
  string account_id = 1;
  int64  amount_minor = 2;
  string currency = 3;
  optional int64 min_balance_minor = 4;

  // removed in v2.3. the number is reserved so it can NEVER be reused
  // for a different meaning, which would silently corrupt old callers.
  reserved 5;
  reserved "legacy_floor";

  // added in v2.4, optional with a safe default.
  optional string value_date = 6;   // Part 10's second date
}
the deprecation process, as a sequence with gates
  1. Announce, with a migration guide and a date.
  2. Instrument. Count calls to the deprecated field or method by service identity, which mTLS makes possible.
  3. Chase the specific teams still using it. The metric names them, so this is a short list rather than a broadcast.
  4. Wait for zero usage for a sustained period. Never remove on a date alone; remove on evidence.
  5. Reserve the field number permanently and remove the code.
why step 2 is the one that makes this work
Most deprecations fail because nobody knows who is still calling. Because every call carries a cryptographic service identity, the metric is deprecated_field_use{field, caller}, and the migration becomes three specific conversations instead of an announcement nobody reads. The auth design from chapter 169 turns out to be what makes API evolution tractable, which is a connection worth pointing out.
172

Sketch v14: the integration surface

What changed, and the cost accepted

ChangeDriven byCost accepted
gRPC internally, REST at the edgeTyped contracts where we control both ends; accessibility where we do notTwo API surfaces to maintain and document
One posting API taking balanced entriesThe ledger must never learn product rulesCallers must construct entries, so the SDK matters
Server-side invariant validationA caller bug must not create moneyA validation pass on every posting
Floors verified against the facilityA caller must not grant itself overdraftA facility lookup inside the posting transaction
Signed, expiring, single-use overridesLimit exceptions need governance, not a booleanA grant-and-consume flow, like the Part 7 FX quote
mTLS plus scope by account kindBlast radius of a compromised serviceA service mesh, and per-service policy to maintain
SSE for live updatesSimpler and resumable versus WebSocketOne-directional only, which is all we needed
ISO 8583 gateway isolatedStateful sockets inside a stateless architectureA component that cannot autoscale or deploy casually
how to close round fourteen
"v14 makes the arrows real. The through-line is that protocol follows conversation shape: gRPC for typed internal calls with deadline propagation, REST where a partner should not need my toolchain, Kafka where the producer must not know its consumers, SSE for one-directional push, and ISO 8583 because the scheme decided. The two things I would defend hardest are that the ledger validates the invariant server-side, so nine teams cannot create money between them, and that authorisation is scoped by account kind rather than by method, because that is what bounds the blast radius of a compromise. What is left is ten years of data."
architecture v14
protocol per conversation shape
swipe the figure sideways, or tap expand for full screen
1/8
ledger at centre
The ledger sits at the centre, and its contract is arithmetic: it validates that entries sum to zero and knows nothing about products.