Part 9 · 13 chapters · ~20 min

Round nine: “stop the fraud”

Every previous round dealt with systems that fail. This one deals with people who are trying to make it fail, and who adapt when you stop them. That changes the engineering: a static rule decays the moment it becomes known, a model trained on last month's fraud misses this month's, and the cost of a false positive is a real customer locked out of their own money.

99

The pressure: adversaries, not just load

interviewer

“Someone is draining accounts through a mule network, and someone else is testing stolen cards against your authorisation endpoint. Stop them, without freezing legitimate customers on salary day.”

The structuring insight: fraud is not one problem. Each pattern has a different detection signal, a different acceptable latency, and a different cost of being wrong.

Account takeover Credentials stolen; the real account used by someone else. Typically preceded by a device change, then a beneficiary addition, then a rapid drain. signal: device and behavioural change · window: seconds · false positive cost: high
Card testing Thousands of small authorisations against stolen card numbers to find which are live. Low value individually, and the precursor to real loss. signal: velocity per merchant and per BIN · window: milliseconds · false positive cost: low
Money mule networks Funds fragmented through many newly-created accounts to obscure origin before cashing out. Each individual transfer looks completely normal. signal: graph structure, not transaction attributes · window: minutes to hours · false positive cost: medium
Authorised push payment fraud The customer is deceived into sending the money themselves. Every authentication check passes, because it genuinely is them. signal: beneficiary reputation and behavioural anomaly · window: seconds · false positive cost: high
First-party fraud The customer takes a loan or overdraft with no intention of repaying, often after building a thin but clean history. signal: application and behavioural patterns · window: days · false positive cost: medium
Insider abuse A staff member using support tooling to move money or lift a lien. Every action is technically authorised. signal: privileged-action audit anomalies · window: hours · false positive cost: low
non-functional, new, and in tension
  1. Synchronous scoring adds under 150 ms at p99, fitting inside the card budget from Part 8.
  2. False positive rate under 0.1% on the block action. At 10M transactions a day, 0.1% is 10,000 wrongly blocked customers daily.
  3. Detect and contain a draining attack within 60 seconds.
  4. The scoring path fails open, because a fraud outage must not become a payments outage.
  5. Every decision is explainable, for the customer and for the regulator.
the number that reframes the whole round
10,000 wrongly blocked customers a day at a 0.1% false positive rate. That is a call centre overwhelmed, a reputation problem, and a regulatory complaint trail. Fraud systems fail far more often by being too aggressive than by being too permissive, and saying that early shows you understand the real operating constraint rather than just the detection problem.
100

Rules engine on the synchronous path

Rules are unfashionable and they are where every real system starts, for three reasons worth being able to state.

why rules, before any model
  1. Explainable. "Blocked because a new beneficiary received more than ₦500,000 within an hour of a device change" is a sentence a customer, an agent and a regulator can all act on.
  2. Instantly deployable. A new attack pattern can be blocked in minutes, where retraining a model takes days.
  3. Cheap and predictable. Microseconds, with no inference infrastructure and no tail latency.
code
// rules are DATA, versioned and effective-dated, exactly like the
// compliance policies in Part 7. same reasoning: they change on an
// attacker's timetable, not a release schedule.
interface Rule {
  id: string;
  version: number;
  when: Condition;          // a tree over features, no arbitrary code
  then: 'allow' | 'review' | 'block' | 'step_up';
  weight: number;           // contribution when combined with a model score
  // shadow mode: evaluate and record, but do not act. every new rule
  // runs here first so its true false-positive rate is measured on
  // real traffic BEFORE it can decline anybody.
  mode: 'shadow' | 'active';
  effectiveFrom: Date;
  effectiveTo?: Date;
}
shadow mode is the most valuable idea in this chapter
Never activate a rule straight into enforcement. Run it in shadow for days, recording what it would have done, then measure precision against confirmed fraud. A rule that looks obviously correct routinely turns out to catch 40 genuine customers for every fraudster. Shadow mode is how you find that out without the incident, and proposing it unprompted is a strong signal.

The rules that earn their place on the hot path

RuleCatchesAction
New beneficiary plus amount above a thresholdAccount takeover, APP fraudStep up authentication, rather than block
Device change within 24 h of a large outbound transferAccount takeoverStep up
More than N declined authorisations per card in 10 minutesCard testingBlock the card, narrow and safe
Amount just below a reporting threshold, repeatedlyStructuringReview. Never auto-block, since it is often legitimate
Beneficiary on an internal mule listMule networksBlock
Transfer to a newly created account from many sourcesMule collectionReview

Note how few of these block. Step-up authentication is usually the right action: it stops the fraudster, who cannot pass it, and mildly inconveniences the genuine customer, who can. Reaching for block first is the most common design error in this space.

101

Feature stores and the freshness problem

A rule or a model needs features, and features have wildly different freshness requirements and computation costs. That mismatch is the core engineering problem of a fraud platform.

FeatureFreshness neededWhere it lives
Transactions in the last 60 secondsSub-secondRedis counters, incremented on the write path
Is this beneficiary new to this customer?Sub-secondRedis set, checked and updated inline
Average transaction size over 90 daysDaily is fineBatch-computed, loaded into the online store
Device reputation scoreMinutesStreaming aggregation from the event log
Graph distance to a known muleHoursBatch graph job, materialised per account
Account age, KYC tierStaticRead from the account record

The training and serving skew problem

worked numbers
          a model trained on "average transaction size over 90 days"

          computed by a batch job over the warehouse…


          …is served in production by a different query, written by a

          different engineer, against a different store, at a different time.


if the two definitions differ even slightly, the model sees

features it was never trained on and its accuracy silently collapses.


this is the single most common cause of an ML system that

performs well offline and badly in production.
what a feature store is actually for
  1. One definition of each feature, used for both training and serving. This is the whole point; the rest is plumbing.
  2. Point-in-time correctness for training: the feature values as they were when the decision was made, never as they are now.
  3. An online store for low-latency serving and an offline store for training, populated from the same computation.
  4. Versioning, so a model is pinned to the feature versions it was trained against.
point-in-time correctness, and why it is easy to get wrong
If you train on "account's total lifetime fraud count" computed today, that feature already includes the fraud you are trying to predict. The model scores beautifully in testing and is useless in production, because at decision time that count was zero. Label leakage through non-point-in-time features is the classic fraud-modelling failure, and naming it demonstrates real ML engineering experience rather than familiarity with the vocabulary.
102

Behavioural baselines per account

Population-level thresholds cannot work in a bank. A ₦2,000,000 transfer is alarming for a student and routine for a business, so "normal" has to be defined per account.

code
// a compact per-account baseline, cheap to store for 20M accounts
interface Baseline {
  accountId: string;
  // distribution, not just a mean: the percentiles are what matter
  amountP50: bigint; amountP95: bigint; amountP99: bigint;
  txPerDayP95: number;
  // the shape of their week, learned rather than assumed
  activeHours: number[];          // 24 buckets of relative frequency
  activeDays: number[];           // 7 buckets
  knownBeneficiaries: number;
  knownDevices: number;
  typicalChannels: string[];      // app | ussd | pos | card | api
  // a baseline built on 3 transactions is not a baseline. say so.
  sampleSize: number;
  confidence: 'low' | 'medium' | 'high';
  computedAt: Date;
}
worked numbers
          deviation score, per dimension:


            amount_z = (amount − p50) ÷ (p95 − p50)

            hour_unusual = 1 − activeHours[hour_of_day]

            new_beneficiary = beneficiary ∉ known ? 1 : 0

            new_device = device ∉ known ? 1 : 0


          combined, weighted, and scaled by confidence


a low-confidence baseline must not drive a block
the three traps in behavioural baselining
  1. Cold start. A new account has no baseline, so it cannot be scored against one. New accounts get tier-based limits instead, until enough history exists.
  2. Drift. People legitimately change: a new job, a new city, a business that grows. The baseline must decay old data and re-fit continuously, or every genuine life change becomes a fraud alert.
  3. Poisoning. A patient fraudster establishes a "normal" of increasing transfers over weeks, so the eventual drain looks in-distribution. Mitigated by absolute caps that no baseline can override, and by weighting recent behaviour less on young accounts.
salary day, which the interviewer explicitly mentioned
On the 25th of the month, millions of accounts simultaneously deviate from their baseline: a large credit, then unusual spending. A naive per-account anomaly detector fires on the entire customer base at once. The fix is population-level context as a feature: is this deviation happening to everyone right now? If so, it is not anomalous. Anomaly detection without population context is how fraud systems take themselves down on payday.
103

Velocity checks in a sliding window

The cheapest high-value signal in fraud, and the implementation detail matters because the obvious approach has an exploitable hole.

Fixed window, and the boundary attack

worked numbers
          fixed hourly window, limit 5 transactions:


            09:59 → 5 transactions  allowed

            10:00 → counter resets

            10:01 → 5 more         allowed


10 transactions in 2 minutes, and the limit was never breached.

every attacker knows where your window boundary is.

Sliding window with a sorted set

code
-- one atomic Lua script: trim the window, count, then add.
-- doing this as separate calls is the Part 2 lost update, again.
local key, now, window, limit = KEYS[1], ARGV[1], ARGV[2], ARGV[3]

-- drop everything older than the window
redis.call('ZREMRANGEBYSCORE', key, 0, now - window)

local count = redis.call('ZCARD', key)
if count >= tonumber(limit) then
  return {0, count}                      -- denied
end

-- member must be unique per event, or two events at the same
-- millisecond collapse into one and the count under-reports.
redis.call('ZADD', key, now, ARGV[4])
redis.call('EXPIRE', key, math.ceil(window / 1000) + 60)
return {1, count + 1}
ApproachMemory per keyAccuracyUse for
Fixed window counterOne integerExploitable at the boundaryNothing security-relevant
Sliding window logOne entry per eventExactLow-volume, high-value checks. Our default
Sliding window counterA few integersApproximate, weighted across bucketsHigh-volume checks where exactness is not required
Token bucketTwo valuesExact for rate, permits burstsAPI rate limiting rather than fraud velocity

The dimensions worth counting

velocity keys, and what each catches
  1. vel:acct:{id} transactions per account. Draining.
  2. vel:card:{token} authorisations per card. Card testing.
  3. vel:dev:{deviceId} across accounts. One device operating many accounts is the strongest mule signal available.
  4. vel:ben:{beneficiary} inbound sources. Mule collection points.
  5. vel:ip:{subnet} for onboarding and login. Automated account creation.
  6. vel:bin:{bin} per card BIN. A breach at another issuer being tested against us.

The third one deserves emphasis. Transaction-level signals see each mule transfer as normal; a single device fingerprint appearing across forty accounts is unambiguous, and it is a counter rather than a model.

104

Graph signals: money mules and rings

Some fraud is invisible at the transaction level by construction. Mule networks fragment money precisely so that no individual movement looks unusual, which means the signal lives in the structure rather than in any row.

Graph signalWhat it indicatesCost to compute
Fan-out then fan-inClassic layering: one source splits, then reconvergesModerate. A bounded traversal
Shared device across accountsStrongest single signal. One operator, many identitiesCheap. A Redis set
Shared beneficiary detailsSame phone, address or next-of-kin across unrelated accountsCheap. An index lookup
Short path to a confirmed muleGuilt by proximity, useful within 2 hopsModerate, if materialised per account
Velocity of new edgesAn account suddenly transacting with many new counterpartiesCheap. A counter
Community detectionDense clusters that transact mostly internallyExpensive. Batch only

The architecture: cheap signals online, expensive ones batch

code
// ONLINE, on the synchronous path, single-digit milliseconds.
// these are precomputed or counter-based, never traversals.
interface OnlineGraphFeatures {
  beneficiaryIsFlaggedMule: boolean;      // set membership
  deviceAccountCount: number;             // counter
  hopsToNearestConfirmedMule: number;     // materialised nightly
  beneficiaryInboundSources24h: number;   // counter
}

// BATCH, hours, over the whole graph. results are materialised back
// into the online store so the fast path never traverses anything.
//   - connected components over the transfer graph
//   - community detection, e.g. Louvain
//   - PageRank-style centrality to find collection points
//   - path distances to every confirmed mule
the decision to justify
Do not put a graph database on the authorisation path. A traversal has unpredictable latency and the budget is 150 ms. Instead, run graph analysis as a batch job over the event log from Part 4 and materialise the results as scalar features. Precompute the expensive structure; serve a number. That principle recurs throughout the design, from balance snapshots onward.
mule network
invisible per transaction, obvious as a graph
swipe the figure sideways, or tap expand for full screen
1/8
victim
A victim account, freshly compromised through account takeover. The money needs to leave in a way that cannot be traced or recalled.
105

Scoring: rules, gradient boosting, and LLM triage

Three mechanisms with genuinely different strengths. The design is to use each where it wins rather than to pick one.

RulesGradient boosted treesLLM
LatencyMicroseconds1 to 10 msSeconds
On the sync path?YesYesNo
ExplainabilityPerfectGood, via SHAP valuesPlausible narrative, not a true explanation
Handles novel patternsNoSomewhatYes
Handles unstructured inputNoNoYes. Chat logs, documents, narratives
DeterminismTotalTotalVariable, needs pinning and low temperature
Best roleHard blocks and step-upsThe primary scoreAnalyst assistance in case review

Why gradient boosted trees rather than a neural network

the honest reasons
  1. They win on tabular data. For structured features with no spatial or sequential structure, boosted trees consistently match or beat deep models.
  2. Fast inference on CPU, which matters at 5,600 decisions per second.
  3. Feature importance and SHAP values come nearly free, and explainability is a hard requirement here.
  4. Robust to unscaled, mixed-type, partially missing features, which describes real fraud features exactly.
  5. They handle extreme class imbalance well, and fraud is often under 0.1% of transactions.

Where the LLM genuinely helps

LLM in the review queue, not on the money path
  1. Case summarisation. Turn 200 transactions, 3 device changes and a support thread into a paragraph an analyst reads in 20 seconds instead of 10 minutes.
  2. Unstructured evidence. Read the chat transcript where a customer was socially engineered, and identify the APP fraud pattern no tabular feature encodes.
  3. Narrative generation for suspicious activity reports, drafted for human review and sign-off.
  4. Pattern discovery across confirmed cases, proposing candidate rules for a human to evaluate in shadow mode.
the boundary to state clearly
An LLM never makes an automated decision to block money. It is non-deterministic, its latency does not fit the budget, it cannot be validated to a regulator's satisfaction, and its stated reasoning is a plausible narrative rather than the actual cause of its output. It is an analyst's assistant that makes humans faster, and drawing that line unprompted is the difference between sounding current and sounding careless.
106

Three-tier decisions: allow, review, block

A binary decision forces a bad tradeoff. Adding intermediate actions is what makes the false positive budget achievable.

worked numbers
          score 0.00 → 0.70   ALLOW            ~99.2% of traffic

          score 0.70 → 0.90   STEP UP          ~0.6%

          score 0.90 → 0.97   REVIEW           ~0.15%

          score 0.97 → 1.00   BLOCK            ~0.05%


only 0.05% is auto-blocked, which is how a 0.1% false

positive budget on blocks becomes achievable at all.
ActionCustomer experienceFraudster experience
AllowNothingSucceeds
Step upOne extra factor: OTP, biometric, in-app confirm. Mild frictionUsually stopped. They lack the second factor
Delay plus review"Processing, up to 30 minutes." Tolerable for a large transferStopped, and the delay itself is the defence: it gives humans time
BlockCannot use their own money. A call, a complaint, possibly a lost customerStopped
why delay is underrated
For fraud, time is the most valuable resource. A 30-minute hold on a suspicious large transfer costs a genuine customer very little and destroys most fraud economics, because it gives a human, or the customer's own second thought, a chance to intervene. Many systems reach for block when delay would achieve the same protection at a fraction of the customer cost.

Choosing the thresholds honestly

worked numbers
          expected cost of a threshold t:


            cost(t) = FP(t) × cost_of_false_positive

                    + FN(t) × average_fraud_loss


          cost_of_false_positive is NOT just a support call. it includes

          churn probability, reputational effect and regulatory complaints.


the threshold is a business decision with an engineering input,

and presenting it that way is the senior move.
107

Automated blocking and the blast radius

Automated blocking is necessary and dangerous. A bug or a poisoned signal can lock out a large fraction of the customer base in minutes, and that has happened to real banks.

the controls that make automation safe
  1. A global rate limit on the block action itself. If the system tries to block more than N accounts per minute, it stops blocking and pages a human. This single control prevents the worst outcome available.
  2. Proportionality. Block the narrowest thing that works: the transaction, then the card, then the channel, then the account. Full account freeze is the last resort.
  3. Automatic expiry. An automated block expires after a defined period unless a human confirms it. Failure mode becomes unblocking rather than permanent lockout.
  4. A kill switch. One flag disables automated blocking entirely, leaving rules in shadow. Part 13 covers the operational side.
  5. Never block the whole bank. No rule may match on a property shared by all customers. This must be validated when the rule is authored, not discovered in production.
code
// the circuit breaker on our own enforcement. it is the difference
// between a bad rule and a bank-wide incident.
async function enforceBlock(target: Target, reason: Reason) {
  const rate = await counters.increment('blocks:global', { window: 60 });

  if (rate > BLOCK_RATE_CEILING) {
    // we are probably the problem, not the customers.
    await alerts.page('fraud.block_rate_exceeded', { rate });
    await flags.disable('fraud.auto_block');
    return { blocked: false, reason: 'enforcement_suspended' };
  }

  // narrowest effective scope, and an expiry by default
  return blocks.create({
    scope: narrowestScopeFor(reason),
    target,
    expiresAt: addHours(now(), 24),   // requires human confirmation to persist
    createdBy: 'automated',
    reason
  });
}
the sentence to say out loud
"I would put a rate limit on my own blocking action. If the system starts blocking more than a few hundred accounts a minute, the most likely explanation is my rule is wrong, not that fraud increased a thousandfold. So it stops enforcing and pages a human. The system's default assumption under anomaly should be that the system is broken."
108

Account and wallet flagging state machine

"Blocked" is not one state. Different restrictions serve different purposes, and conflating them either over-restricts customers or fails to contain attacks.

StateCredits inDebits outUsed for
activeYesYesNormal
watchYesYesElevated monitoring, no customer impact. The most useful state, and the most underused
debit_restrictedYesNoSuspected compromise. Contains the attack while salary still arrives
credit_restrictedNoYesSuspected mule collection point. Stops it receiving more
frozenNoNoConfirmed fraud, or a legal instruction
closedNoNoTerminal. Residual balance handled by a defined process
why debit_restricted is the state that matters
In an account takeover, the goal is to stop money leaving. Freezing the account entirely also stops the customer's salary arriving, their standing orders landing, and refunds reaching them, which turns a security action into a serious customer harm. Restricting debits contains the attack completely while causing almost no collateral damage, and reaching for it instead of a freeze is exactly the kind of judgement this round is testing.
code
CREATE TABLE account_restrictions (
  id           UUID PRIMARY KEY,
  account_id   UUID NOT NULL,
  state        TEXT NOT NULL,
  -- WHY, in a form that can be shown to a customer and an auditor
  reason_code  TEXT NOT NULL,
  reason_detail TEXT,
  -- WHO: automation, an analyst, or a court. drives who may lift it.
  imposed_by   TEXT NOT NULL,   -- automated | analyst:id | legal
  -- automated restrictions EXPIRE. legal ones do not.
  expires_at   TIMESTAMPTZ,
  lifted_at    TIMESTAMPTZ,
  lifted_by    TEXT,
  created_at   TIMESTAMPTZ NOT NULL DEFAULT now()
);

-- the restriction is checked on the posting path, and it is checked
-- against the DIRECTION of the entry, not merely its existence.
CREATE INDEX restr_active ON account_restrictions (account_id)
  WHERE lifted_at IS NULL;
rules about lifting
  1. Automation may impose a restriction; only a human may lift a confirmed one.
  2. A legal restriction cannot be lifted by an analyst at all, and the system must enforce that rather than rely on process.
  3. Four-eyes on lifting anything above a value threshold, because lifting restrictions is a prime insider-abuse vector.
  4. Every imposition and lifting is immutably logged, with the actor and the reason. This log is itself monitored for insider anomalies.
109

Sanctions, PEP, AML transaction monitoring

Adjacent to fraud and legally distinct. Fraud protects the bank and its customers from loss. Financial crime compliance is a statutory obligation whose failure brings fines and criminal liability, not just losses.

ControlWhat it checksWhen
Sanctions screeningNames against OFAC, UN, EU, UK and local listsOnboarding and every cross-border payment. Blocking, not advisory
PEP screeningPolitically exposed persons and their associatesOnboarding and periodically. Triggers enhanced due diligence, not a block
Adverse mediaPublic reporting of financial crimeOnboarding and periodically
Transaction monitoringStructuring, rapid movement, high-risk corridors, unexplained activityContinuous
Threshold reportingTransactions above a statutory amountPer jurisdiction, per Part 7's policy engine

Fuzzy name matching, and why it is genuinely hard

worked numbers
          sanctions list:  "Mohammed Al-Hassan"


          must also match:  Muhammad Al Hassan  · Mohamad AlHassan

                                Hassan, Mohammed  · محمد الحسن


          must not match every one of the millions of people

          legitimately named Mohammed Hassan.


false negative = sanctions violation, fines, criminal exposure

false positive = a real customer blocked from their money


so the tuning is deliberately conservative and the review queue

is staffed accordingly. this is a people problem as much as a

technical one, and the design must acknowledge that.
the architectural point
Sanctions screening is blocking and synchronous on cross-border payments, because releasing a payment to a sanctioned party is a violation the moment it settles. It therefore needs its own latency budget and availability target, and unlike fraud scoring it must fail closed: if the screening service is unavailable, cross-border payments queue rather than proceed. That is the opposite posture from fraud, for a clear legal reason, and being able to articulate why the two differ is the point of this chapter.
110

Case management and the feedback loop

The part that is always underbuilt, and the part that determines whether the system improves. Without labels there is no learning, and labels come from human decisions.

worked numbers
          detection → case → analyst decision → label → training data

                                                     ↓

                                              retrain, re-tune


without the loop, the model degrades from the day it ships

          because attackers adapt and the model does not.
        
what a case must contain to be decidable in minutes
  1. The triggering decision, its score, and the top contributing features by SHAP value.
  2. Which rules fired, at which versions, including those in shadow mode.
  3. The account's baseline and precisely how this deviated from it.
  4. A transaction timeline with the graph neighbourhood rendered visually.
  5. Linked cases on the same device, beneficiary or network cluster.
  6. An LLM-drafted summary, clearly marked as assistance rather than a finding.

The label problem, which is subtler than it looks

three biases that corrupt the training set
  1. Blocked transactions have no outcome. We prevented it, so we never learn whether it was truly fraud. Mitigated by letting a small random sample of high-score transactions through and observing, which is uncomfortable and necessary.
  2. Confirmation lag. Fraud is often confirmed weeks later via a dispute, so recent data is systematically under-labelled and a model trained on it underestimates fraud.
  3. Analyst inconsistency. Different analysts label the same case differently. Needs calibration sets and inter-rater measurement, exactly like any other annotation pipeline.
the uncomfortable idea worth raising
Deliberately allowing a small random sample of high-scoring transactions through is how you measure your own precision. Without it you have no idea whether your block threshold catches fraud or catches customers, because blocked transactions generate no ground truth. The sample is small, capped by value, and treated as a measurement cost. Proposing it shows you have thought about how the system knows whether it is working.
111

Sketch v9: risk in the path, not beside it

The two paths, and why both are necessary

Fast pathDeep path
Latency budget150 msSeconds to hours
InputsOnline features, counters, precomputed scalarsThe whole event log, the graph, unstructured evidence
TechniquesRules plus boosted treesGraph algorithms, baselining, LLM triage
Can block a live transaction?YesNo, but it can restrict the account for the next one
Failure postureFails openDegrades: detection is delayed, not lost
CatchesCard testing, known patterns, velocityMule networks, slow-burn fraud, novel patterns

What changed, and the cost accepted

ChangeDriven byCost accepted
Rules engine, rules as versioned dataAttackers adapt faster than releasesA rule authoring and governance process
Mandatory shadow modeRules that look obvious have terrible precisionDays of delay before a rule can enforce
Feature store, online and offlineTraining and serving skew silently destroys accuracyReal infrastructure, and feature governance
Graph analysis in batch, materialisedMule structure is invisible per transactionHours of detection latency on that signal
Four-tier actions including step-up and delayA 0.1% false positive budget on blocksStep-up and delay flows to build in the product
Rate limit on our own blockingA bad rule can lock out the bankUnder a real attack, enforcement may suspend and page
Six-state restriction modelFreezing an account harms customers unnecessarilyEvery money path must check direction-aware restrictions
Sanctions fails closedA statutory obligation, unlike fraudCross-border payments queue when screening is down
how to close round nine
"v9 splits risk into a fast path that can decline and a deep path that can restrict, because the signals that catch card testing and the signals that catch a mule network have latency requirements three orders of magnitude apart. The three things I would defend hardest: fraud fails open and sanctions fails closed, for different and clear reasons; a rate limit on my own blocking action, because a bad rule is more likely than a thousandfold fraud spike; and step-up before block, because at this volume a 0.1% false positive rate is 10,000 wrongly blocked customers every day. What I still cannot do is prove any of the money is right."
architecture v9
fast path synchronous, deep path asynchronous
swipe the figure sideways, or tap expand for full screen
1/9
two paths
Risk splits into two paths, because the signals that catch card testing and the signals that catch a mule network have latency requirements three orders of magnitude apart.