Part 12 · 8 chapters · ~50 min

Round twelve: “tell the customer”

Ten million notifications a day, each one about somebody's money, each one arriving on a phone that may be off, in a country with a different quiet-hours rule, through a provider that will occasionally accept a message and never deliver it. The interesting constraint is not the volume. It is that a duplicate debit alert makes a customer believe they have been robbed twice.

133

The pressure: 10M notifications, ordered, deduped

interviewer

“Every transaction generates an alert. Customers have preferences, some are asleep, SMS costs money, and your Kafka consumer will occasionally see the same event twice. Design it.”

Every element of that sentence is a real constraint, and together they make this less trivial than it first appears.

worked numbers
          10M transactions/day × ~1.6 notifiable parties ≈ 16M notifications/day

          ÷ 86,400                                      = 185/s average

          × 4 peak                                    = 740/s at peak


          and SMS costs roughly ₦4 each:

            16M × ₦4 = ₦64m/day if everything went by SMS

            ≈ ₦23bn/year


so channel selection is a cost decision worth billions,

not a user-preference nicety.
functional, new
  1. Notify on every notifiable event, across push, SMS, email and in-app.
  2. Respect per-customer, per-event-type channel preferences.
  3. Respect quiet hours, in the customer's own timezone.
  4. Never send the same notification twice.
  5. Preserve per-customer ordering: a debit alert before its balance update.
  6. Track delivery outcome, and fail over between providers.
non-functional, new
  1. p95 under 10 seconds from posting to device for transaction alerts.
  2. A provider outage degrades that channel only.
  3. A marketing blast cannot delay a fraud alert. Priority is structural.
  4. Cost per notification is measured and attributed per event type.
the requirement people forget
Never send twice is harder than it sounds, because our own design guarantees duplicates. The Part 4 relay publishes before marking published, so it produces duplicates deliberately, and Kafka delivers at least once. The notification service is the place where at-least-once delivery has to become exactly-once-looking behaviour for a human being, and a human noticing a duplicate debit alert concludes they have been robbed twice.
134

Fan-out from the event log

Notifications are a consumer of the Part 4 log, and the pipeline is deliberately a sequence of narrowing stages rather than one handler.

worked numbers
1. consume         ledger.entry.posted, keyed by account

2. is it notifiable?  most events are not. filter early.

3. who cares?      resolve the parties: holder, joint, guardian

4. preferences       which channels, and are they suppressed?

5. dedup            have we already sent this exact thing?

6. render           template + locale + currency formatting

7. enqueue          per channel, per priority

8. deliver          provider call, with failover

9. record           receipt, cost, and outcome


each stage discards work. by stage 7 we are sending far

fewer messages than stage 1 produced, which is the point.
why filtering early matters so much
  1. An internal fee posting to a revenue account notifies nobody, and roughly half of all entries are internal legs.
  2. A customer with push enabled and SMS disabled costs ₦0 instead of ₦4.
  3. A suppressed duplicate saves both the cost and the customer's trust.
  4. Filtering at stage 2 is a cheap predicate; filtering at stage 8 is a paid API call. Cheapest checks first, exactly as in Part 8's authorisation ordering.

Ordering, which comes free from the partition key

Part 4 partitioned ledger.entry.posted by account_id, which means one account's events arrive at one consumer in order. That is precisely the ordering notifications need, and it was not a coincidence:

the payoff for choosing the partition key carefully
Because events for an account are totally ordered within their partition, a customer cannot receive "your balance is now X" before "you were debited Y". The ordering requirement was satisfied four rounds before it was stated, by picking the smallest ordering unit that anything needed. Choosing the partition key well pays off in rounds you have not reached yet.
135

Preference resolution and quiet hours

Preferences are a resolution problem with a precedence order, and the order matters because some rules may never be overridden.

code
// resolved in strict precedence order. the first four are not
// customer preferences at all: they are obligations.
function resolveChannels(ev: Event, cust: Customer): Channel[] {
  // 1. MANDATORY: regulatory or security. no opt-out exists.
  if (ev.type === 'security.credential_changed') return ['sms', 'email'];
  if (ev.type === 'account.restricted')         return ['sms', 'email'];

  // 2. LEGAL SUPPRESSION: deceased, closed, or a DND registry entry.
  if (cust.legalSuppression) return [];

  // 3. GLOBAL OPT-OUT, for non-mandatory only.
  if (cust.optedOutAll) return ['in_app'];   // still visible if they look

  // 4. PER-EVENT-TYPE preference, the actual customer choice.
  let chans = cust.prefs[ev.type] ?? defaultsFor(ev.type);

  // 5. QUIET HOURS, in THEIR timezone, and only for deferrable events.
  if (inQuietHours(cust) && !ev.urgent) {
    chans = chans.filter(c => c === 'in_app');   // silent channels only
    scheduleDigest(cust, ev);                    // batch it for the morning
  }

  // 6. CAPABILITY: no point queueing push with no registered device.
  return chans.filter(c => cust.canReceive(c));
}
the four rules about quiet hours that are easy to get wrong
  1. The customer's timezone, not the server's. A Lagos customer travelling in California still has Lagos quiet hours unless they changed them, because it is a preference rather than a location.
  2. Urgent events ignore quiet hours. A fraud alert at 3am is the entire point of a fraud alert.
  3. Suppressed does not mean discarded. It becomes a morning digest, so the customer is not simply uninformed.
  4. Store the timezone, not an offset. Offsets change with daylight saving; Africa/Lagos does not.
the classification that resolves most arguments
Every notification type is tagged mandatory, urgent or deferrable, and that tag decides preference handling, quiet hours, priority queue and provider choice all at once. A debit alert is urgent, a monthly statement is deferrable, and a credential change is mandatory. Without that taxonomy every new notification type restarts the same debate.
136

Per-channel workers and blast radius

Four channels with wildly different latency, cost and failure characteristics. One shared worker pool would let the slowest and cheapest channel starve the most important one.

push
latency
~1 s
cost
~free
reliability
Silent failure if the token is stale
use for
The default for everything
SMS
latency
2 to 30 s, sometimes minutes
cost
₦4 each
reliability
High reach, opaque delivery
use for
Mandatory and urgent only
email
latency
Seconds to minutes
cost
Fractions of a naira
reliability
Spam filtering, reputation-dependent
use for
Statements, receipts, digests
in-app
latency
Instant when they look
cost
Free
reliability
Always succeeds. It is our own store
use for
Everything, always, as the record
the isolation rules
  1. A queue per channel per priority. Eight queues, so a marketing email backlog cannot delay a fraud SMS.
  2. A bounded worker pool per queue, sized to the provider's rate limit rather than to our own volume.
  3. A circuit breaker per provider, as in Part 6, with the same reasoning.
  4. Shed the lowest priority first under saturation. Marketing is dropped long before transaction alerts degrade.
  5. In-app is written first and synchronously, so the notification exists in our own system regardless of what every external provider does.
why in-app first is the quietly important decision
Writing the in-app notification before attempting any external channel means the customer can always open the app and see what happened, even during a total provider outage. It also makes the in-app record the source of truth for what we told them, which matters in a dispute. One cheap, always-available channel that we own makes every other channel optional, and that is a resilience pattern worth stating explicitly.
137

Deduplication windows

Two distinct problems both called deduplication, and conflating them produces either duplicate alerts or missing ones.

Technical duplicateSemantic repetition
CauseKafka redelivery, relay retry, consumer restartThe customer genuinely did the same thing twice
Correct behaviourSuppress absolutely. Send onceSend both. They are different events
Keyevent_id, which is unique per eventNot applicable: the events are distinct
WindowLonger than the topic's retentionNone
Getting it wrong"I was debited twice" panicA real transaction silently unreported
code
-- technical dedup: keyed on the event id, which is unique per event.
-- NEVER key on (customer, amount, beneficiary): two genuine transfers
-- of the same amount to the same person would collapse into one, and
-- the customer would never learn about the second.
CREATE TABLE notification_sent (
  event_id   TEXT NOT NULL,
  channel    TEXT NOT NULL,
  recipient  TEXT NOT NULL,
  sent_at    TIMESTAMPTZ NOT NULL DEFAULT now(),
  -- the uniqueness that does the work. one send per event per
  -- channel per recipient, enforced by the database.
  PRIMARY KEY (event_id, channel, recipient)
) PARTITION BY RANGE (sent_at);   -- DROP old partitions, no mass DELETE

The third thing, which is neither

Rate limiting. A merchant receiving 400 payments an hour does not want 400 push notifications, and every one of them is a genuine, distinct event. That is not deduplication; it is aggregation:

code
// each is real and must be reported. the FORM changes, not the fact.
if (await rate.exceeded(customer, ev.type, { window: '1h', max: 20 })) {
  // switch to a rolling digest rather than dropping anything
  await digest.add(customer, ev);
  return { deferred: true, reason: 'rate_limited_to_digest' };
}
the distinction to state clearly
Deduplicate on identity, aggregate on volume, and never confuse the two. Suppressing by content is the bug that makes a bank miss a real transaction, and it looks sensible in code review. Suppressing by event id is always safe. Aggregation changes the form of the report; deduplication removes a repetition that never happened.
two kinds of duplicate
redelivery versus genuine repetition
swipe the figure sideways, or tap expand for full screen
1/7
event sent
Case A: an event arrives and a notification is sent. Straightforward.
138

Provider abstraction and failover

The same shape as the Part 6 connector interface, with one important difference in the failure posture.

code
interface NotificationProvider {
  readonly id: string;
  readonly channel: Channel;
  readonly costPerMessage: bigint;   // minor units, for routing
  readonly rateLimit: number;

  send(m: RenderedMessage): Promise<SendResult>;
  parseReceipt(raw: Buffer, h: Headers): DeliveryReceipt | null;
}

type SendResult =
  | { state: 'accepted'; providerRef: string }   // accepted ≠ delivered
  | { state: 'rejected'; retryable: boolean }
  | { state: 'unknown' };
the crucial difference from Part 6
In Part 6, unknown forbade retrying on another provider, because a duplicate payment moves real money. Here, a duplicate SMS costs ₦4 and mild confusion. So the posture inverts: on unknown we do retry, possibly on a second provider, because a missing debit alert is worse than a duplicated one. Same interface shape, opposite retry policy, and the reason is the cost of being wrong in each direction. Being able to explain why the same pattern gets different policies is exactly the judgement an interviewer is probing for.
routing across providers
  1. Two providers minimum per channel. A single SMS aggregator is a single point of failure for mandatory notifications.
  2. Route by cost × delivery rate, as in Part 6. The cheapest provider with a 90% delivery rate is not the cheapest provider.
  3. Route by destination. Local aggregators are usually cheaper and more reliable for their own country.
  4. Failover on rejection, timeout, or an open breaker, with the priority queue preserved.
  5. Cross-channel escalation for mandatory notifications: if push fails, fall back to SMS. This is the only case where we willingly spend ₦4 to replace a free message.
139

Delivery receipts and the feedback loop

"Accepted by the provider" and "read by a human" are separated by four distinct states, and confusing them is how a bank believes it informed someone it did not.

worked numbers
queued      we have it

sent        the provider accepted it  ← most systems stop here

delivered  the device or inbox received it

read        a human actually saw it


"sent" is the weakest possible claim and the easiest to

mistake for success. a stale push token accepts and

          silently never delivers, indefinitely.
        
what receipts are actually used for
  1. Cross-channel escalation. No delivery receipt for a mandatory push within 60 seconds, so send the SMS.
  2. Provider scoring. Real delivery rates per provider per destination feed the routing decision from the previous chapter.
  3. Token hygiene. A push token that fails repeatedly is retired, which stops us paying to send into a void.
  4. Dispute evidence. "We notified you, here is the delivery receipt with its timestamp" is the answer to a category of complaint.
  5. Cost attribution. Spend per event type, so a product team proposing a new notification sees its own bill.
code
// receipts arrive as webhooks, so they get the full Part 6 treatment:
// verify the signature over raw bytes, bound replay by timestamp,
// deduplicate on the provider's event id, and confirm the message is
// one we actually sent before trusting anything in it.
CREATE TABLE delivery_receipts (
  notification_id UUID NOT NULL,
  provider_ref    TEXT NOT NULL,
  state           TEXT NOT NULL,   -- sent|delivered|failed|read
  failure_code    TEXT,               -- drives token retirement
  cost_minor      BIGINT,             -- what it actually cost
  received_at     TIMESTAMPTZ NOT NULL DEFAULT now(),
  PRIMARY KEY (notification_id, provider_ref, state)
);
the cost-attribution point worth raising unprompted
Reporting notification spend per event type back to the team that owns it changes behaviour more than any policy document. When a product manager sees that their new "we miss you" SMS costs ₦18m a year, the conversation about whether push would suffice resolves itself. Making cost visible at the point of the decision is a design choice, not an accounting one.
140

Sketch v12: notification subsystem

What changed, and the cost accepted

ChangeDriven byCost accepted
Narrowing pipeline, filter firstHalf of all entries notify nobody, and SMS costs ₦4More stages to observe and instrument
Dedup on event_idOur own relay produces duplicates by designA partitioned dedup store per channel and recipient
Rate limiting into digestsMerchants receive hundreds of genuine eventsDigest rendering and scheduling, per timezone
Queue per channel per priorityMarketing must not delay fraud alertsEight queues to size and monitor
Two providers per channelA single aggregator is a single point of failureTwo integrations, two contracts, routing logic
Retry on unknown, unlike Part 6A duplicate SMS costs ₦4; a missing alert costs trustOccasional duplicate messages, accepted deliberately
In-app written first, synchronouslyOne channel we own must always workA write on the path before any external call
Receipts, scoring, cost attribution"Sent" is the weakest claim availableWebhook ingestion with full verification
how to close round twelve
"v12 is mostly about not sending things: filtering early, deduplicating on event id rather than on content, and aggregating genuine volume into digests. Three things I would highlight: ordering came free from partitioning by account in round four; in-app is written first and synchronously so one channel always works; and the retry posture is the opposite of round six's, because a duplicate message costs ₦4 while a duplicate payment costs real money. What is left is that all twelve rounds of this have to be operated by someone at 3am."
architecture v12
narrowing stages, isolated channels
swipe the figure sideways, or tap expand for full screen
1/8
consume
Notifications are just another consumer of the Part 4 event log, keyed by account, which means one account’s events already arrive in order.