Round twelve: “tell the customer”
Ten million notifications a day, each one about somebody's money, each one arriving on a phone that may be off, in a country with a different quiet-hours rule, through a provider that will occasionally accept a message and never deliver it. The interesting constraint is not the volume. It is that a duplicate debit alert makes a customer believe they have been robbed twice.
The pressure: 10M notifications, ordered, deduped
“Every transaction generates an alert. Customers have preferences, some are asleep, SMS costs money, and your Kafka consumer will occasionally see the same event twice. Design it.”
Every element of that sentence is a real constraint, and together they make this less trivial than it first appears.
10M transactions/day × ~1.6 notifiable parties ≈ 16M notifications/day
÷ 86,400 = 185/s average
× 4 peak = 740/s at peak
and SMS costs roughly ₦4 each:
16M × ₦4 = ₦64m/day if everything went by SMS
≈ ₦23bn/year
so channel selection is a cost decision worth billions,
not a user-preference nicety.- Notify on every notifiable event, across push, SMS, email and in-app.
- Respect per-customer, per-event-type channel preferences.
- Respect quiet hours, in the customer's own timezone.
- Never send the same notification twice.
- Preserve per-customer ordering: a debit alert before its balance update.
- Track delivery outcome, and fail over between providers.
- p95 under 10 seconds from posting to device for transaction alerts.
- A provider outage degrades that channel only.
- A marketing blast cannot delay a fraud alert. Priority is structural.
- Cost per notification is measured and attributed per event type.
Fan-out from the event log
Notifications are a consumer of the Part 4 log, and the pipeline is deliberately a sequence of narrowing stages rather than one handler.
1. consume ledger.entry.posted, keyed by account 2. is it notifiable? most events are not. filter early. 3. who cares? resolve the parties: holder, joint, guardian 4. preferences which channels, and are they suppressed? 5. dedup have we already sent this exact thing? 6. render template + locale + currency formatting 7. enqueue per channel, per priority 8. deliver provider call, with failover 9. record receipt, cost, and outcome each stage discards work. by stage 7 we are sending far fewer messages than stage 1 produced, which is the point.
- An internal fee posting to a revenue account notifies nobody, and roughly half of all entries are internal legs.
- A customer with push enabled and SMS disabled costs ₦0 instead of ₦4.
- A suppressed duplicate saves both the cost and the customer's trust.
- Filtering at stage 2 is a cheap predicate; filtering at stage 8 is a paid API call. Cheapest checks first, exactly as in Part 8's authorisation ordering.
Ordering, which comes free from the partition key
Part 4 partitioned ledger.entry.posted by account_id,
which means one account's events arrive at one consumer in order. That is precisely
the ordering notifications need, and it was not a coincidence:
Preference resolution and quiet hours
Preferences are a resolution problem with a precedence order, and the order matters because some rules may never be overridden.
// resolved in strict precedence order. the first four are not
// customer preferences at all: they are obligations.
function resolveChannels(ev: Event, cust: Customer): Channel[] {
// 1. MANDATORY: regulatory or security. no opt-out exists.
if (ev.type === 'security.credential_changed') return ['sms', 'email'];
if (ev.type === 'account.restricted') return ['sms', 'email'];
// 2. LEGAL SUPPRESSION: deceased, closed, or a DND registry entry.
if (cust.legalSuppression) return [];
// 3. GLOBAL OPT-OUT, for non-mandatory only.
if (cust.optedOutAll) return ['in_app']; // still visible if they look
// 4. PER-EVENT-TYPE preference, the actual customer choice.
let chans = cust.prefs[ev.type] ?? defaultsFor(ev.type);
// 5. QUIET HOURS, in THEIR timezone, and only for deferrable events.
if (inQuietHours(cust) && !ev.urgent) {
chans = chans.filter(c => c === 'in_app'); // silent channels only
scheduleDigest(cust, ev); // batch it for the morning
}
// 6. CAPABILITY: no point queueing push with no registered device.
return chans.filter(c => cust.canReceive(c));
}- The customer's timezone, not the server's. A Lagos customer travelling in California still has Lagos quiet hours unless they changed them, because it is a preference rather than a location.
- Urgent events ignore quiet hours. A fraud alert at 3am is the entire point of a fraud alert.
- Suppressed does not mean discarded. It becomes a morning digest, so the customer is not simply uninformed.
- Store the timezone, not an offset. Offsets change with daylight saving;
Africa/Lagosdoes not.
Per-channel workers and blast radius
Four channels with wildly different latency, cost and failure characteristics. One shared worker pool would let the slowest and cheapest channel starve the most important one.
- latency
- ~1 s
- cost
- ~free
- reliability
- Silent failure if the token is stale
- use for
- The default for everything
- latency
- 2 to 30 s, sometimes minutes
- cost
- ₦4 each
- reliability
- High reach, opaque delivery
- use for
- Mandatory and urgent only
- latency
- Seconds to minutes
- cost
- Fractions of a naira
- reliability
- Spam filtering, reputation-dependent
- use for
- Statements, receipts, digests
- latency
- Instant when they look
- cost
- Free
- reliability
- Always succeeds. It is our own store
- use for
- Everything, always, as the record
- A queue per channel per priority. Eight queues, so a marketing email backlog cannot delay a fraud SMS.
- A bounded worker pool per queue, sized to the provider's rate limit rather than to our own volume.
- A circuit breaker per provider, as in Part 6, with the same reasoning.
- Shed the lowest priority first under saturation. Marketing is dropped long before transaction alerts degrade.
- In-app is written first and synchronously, so the notification exists in our own system regardless of what every external provider does.
Deduplication windows
Two distinct problems both called deduplication, and conflating them produces either duplicate alerts or missing ones.
| Technical duplicate | Semantic repetition | |
|---|---|---|
| Cause | Kafka redelivery, relay retry, consumer restart | The customer genuinely did the same thing twice |
| Correct behaviour | Suppress absolutely. Send once | Send both. They are different events |
| Key | event_id, which is unique per event | Not applicable: the events are distinct |
| Window | Longer than the topic's retention | None |
| Getting it wrong | "I was debited twice" panic | A real transaction silently unreported |
-- technical dedup: keyed on the event id, which is unique per event. -- NEVER key on (customer, amount, beneficiary): two genuine transfers -- of the same amount to the same person would collapse into one, and -- the customer would never learn about the second. CREATE TABLE notification_sent ( event_id TEXT NOT NULL, channel TEXT NOT NULL, recipient TEXT NOT NULL, sent_at TIMESTAMPTZ NOT NULL DEFAULT now(), -- the uniqueness that does the work. one send per event per -- channel per recipient, enforced by the database. PRIMARY KEY (event_id, channel, recipient) ) PARTITION BY RANGE (sent_at); -- DROP old partitions, no mass DELETE
The third thing, which is neither
Rate limiting. A merchant receiving 400 payments an hour does not want 400 push notifications, and every one of them is a genuine, distinct event. That is not deduplication; it is aggregation:
// each is real and must be reported. the FORM changes, not the fact.
if (await rate.exceeded(customer, ev.type, { window: '1h', max: 20 })) {
// switch to a rolling digest rather than dropping anything
await digest.add(customer, ev);
return { deferred: true, reason: 'rate_limited_to_digest' };
}Provider abstraction and failover
The same shape as the Part 6 connector interface, with one important difference in the failure posture.
interface NotificationProvider {
readonly id: string;
readonly channel: Channel;
readonly costPerMessage: bigint; // minor units, for routing
readonly rateLimit: number;
send(m: RenderedMessage): Promise<SendResult>;
parseReceipt(raw: Buffer, h: Headers): DeliveryReceipt | null;
}
type SendResult =
| { state: 'accepted'; providerRef: string } // accepted ≠ delivered
| { state: 'rejected'; retryable: boolean }
| { state: 'unknown' };- Two providers minimum per channel. A single SMS aggregator is a single point of failure for mandatory notifications.
- Route by cost × delivery rate, as in Part 6. The cheapest provider with a 90% delivery rate is not the cheapest provider.
- Route by destination. Local aggregators are usually cheaper and more reliable for their own country.
- Failover on rejection, timeout, or an open breaker, with the priority queue preserved.
- Cross-channel escalation for mandatory notifications: if push fails, fall back to SMS. This is the only case where we willingly spend ₦4 to replace a free message.
Delivery receipts and the feedback loop
"Accepted by the provider" and "read by a human" are separated by four distinct states, and confusing them is how a bank believes it informed someone it did not.
queued we have it
sent the provider accepted it ← most systems stop here
delivered the device or inbox received it
read a human actually saw it
"sent" is the weakest possible claim and the easiest to
mistake for success. a stale push token accepts and
silently never delivers, indefinitely.
- Cross-channel escalation. No delivery receipt for a mandatory push within 60 seconds, so send the SMS.
- Provider scoring. Real delivery rates per provider per destination feed the routing decision from the previous chapter.
- Token hygiene. A push token that fails repeatedly is retired, which stops us paying to send into a void.
- Dispute evidence. "We notified you, here is the delivery receipt with its timestamp" is the answer to a category of complaint.
- Cost attribution. Spend per event type, so a product team proposing a new notification sees its own bill.
// receipts arrive as webhooks, so they get the full Part 6 treatment: // verify the signature over raw bytes, bound replay by timestamp, // deduplicate on the provider's event id, and confirm the message is // one we actually sent before trusting anything in it. CREATE TABLE delivery_receipts ( notification_id UUID NOT NULL, provider_ref TEXT NOT NULL, state TEXT NOT NULL, -- sent|delivered|failed|read failure_code TEXT, -- drives token retirement cost_minor BIGINT, -- what it actually cost received_at TIMESTAMPTZ NOT NULL DEFAULT now(), PRIMARY KEY (notification_id, provider_ref, state) );
Sketch v12: notification subsystem
What changed, and the cost accepted
| Change | Driven by | Cost accepted |
|---|---|---|
| Narrowing pipeline, filter first | Half of all entries notify nobody, and SMS costs ₦4 | More stages to observe and instrument |
Dedup on event_id | Our own relay produces duplicates by design | A partitioned dedup store per channel and recipient |
| Rate limiting into digests | Merchants receive hundreds of genuine events | Digest rendering and scheduling, per timezone |
| Queue per channel per priority | Marketing must not delay fraud alerts | Eight queues to size and monitor |
| Two providers per channel | A single aggregator is a single point of failure | Two integrations, two contracts, routing logic |
| Retry on unknown, unlike Part 6 | A duplicate SMS costs ₦4; a missing alert costs trust | Occasional duplicate messages, accepted deliberately |
| In-app written first, synchronously | One channel we own must always work | A write on the path before any external call |
| Receipts, scoring, cost attribution | "Sent" is the weakest claim available | Webhook ingestion with full verification |