Part 0 · 8 chapters · ~45 min

How to design a system in a room

Before the bank, the method. A system design interview is not a test of whether you know what Kafka is. It is a test of whether you can take an ambiguous sentence, turn it into a specification, choose between imperfect options, and defend the choice when the ground shifts. This part is the frame every later round runs inside, so when the interviewer says "now ten million a day", you already know which four things to update and in what order.

1

What the interviewer is actually measuring

The prompt is always underspecified. "Design a wallet." Four words against forty-five minutes. This is deliberate, and it is the first measurement: a junior engineer starts drawing boxes, a senior engineer starts asking what the boxes are for.

There are five things being assessed, and they are weighted very unevenly.

What is measuredWeightHow it shows
ScopingHighDo you convert ambiguity into a written spec, or start solving the wrong problem confidently?
Tradeoff reasoningHighestCan you name two viable options, state what each costs, and pick one for a reason tied to the requirements?
Depth on demandHighWhen pushed on one box, can you go three levels down, and stop before you waffle?
Failure thinkingHighDo you volunteer what breaks, or only discuss it when asked?
Component knowledgeLowKnowing Kafka has partitions is table stakes. It earns almost nothing on its own.

That last row is the one candidates over-invest in. Reciting that DynamoDB has partition keys is not a signal. Saying "I'll key by account_id rather than transaction_id, because transaction_id spreads a single account's history across every shard and every balance read becomes a scatter-gather": that is the signal. Same knowledge, applied under constraint.

The asymmetry worth internalising

A candidate who designs a merely adequate system but narrates every decision cleanly will usually out-score a candidate who draws a more sophisticated system in silence. The interviewer cannot grade what they cannot hear. Your reasoning is the artefact; the diagram is a by-product.

2

The seven-step frame, and why order matters

Run every design through the same seven steps, in this order, out loud. The order is not decoration. Each step constrains the next, and skipping forward means you make choices you cannot justify.

  1. Clarify

    Convert the four-word prompt into a bounded problem. Three to five questions, chosen for the ones whose answers would change your design.

    3–5 min
  2. Scope

    State what you will design and what you will explicitly leave out. This is not hedging; it is how you buy permission to go deep instead of wide.

    1–2 min
  3. Requirements

    Functional as verbs. Non-functional as numbers. If a non-functional requirement has no number attached, it is a wish, not a requirement.

    4–5 min
  4. Estimate

    Four numbers: write rate, read rate, storage growth, and the tightest latency budget. These decide whether you need one database or twelve.

    3–4 min
  5. Data model

    Entities and their relationships, before any service boxes. Data outlives every service that reads it. Get this wrong and nothing downstream is fixable.

    5–7 min
  6. Design

    Components, then the flow of one concrete request end to end. A design you cannot trace a single request through is not a design yet.

    10–15 min
  7. Stress

    Bottlenecks, failure modes, and the tradeoffs you accepted. Volunteer these. The candidate who raises the weakness first controls how it is discussed.

    8–10 min

Why data model before components

Because the shape of the data decides the shape of the system. If you draw services first, you invent a service topology and then bend the data to fit it. If you model the data first, the service boundaries fall out of it almost mechanically: a service owns the entities whose invariants it is responsible for. In the ledger we are about to build, the entire architecture is downstream of one modelling decision: whether a balance is a stored number or a derived one.

✓ "Let me model the data first, because the balance-storage decision constrains everything about the write path."
✗ "So we'll have an API gateway, then a load balancer, then some microservices…", topology with no justification behind it.
3

Functional vs non-functional, stated properly

Most candidates state requirements badly, in one of two ways: functional requirements that are really features, or non-functional requirements with no numbers.

Functional: what the system does

Write them as verbs with an actor and an object. Not "wallets" but "a customer can hold a balance in a given currency". The verb form forces you to notice missing pieces. Hold implies create, and create implies who is allowed to.

Non-functional: the qualities it must have while doing it

These are the ones that actually pick your technology. And they are useless without numbers. Compare:

✗ "It should be highly available and fast."
Unfalsifiable. Every design satisfies it. Decides nothing.
✓ "99.99% availability on the posting path, 52 minutes of budget a year. p99 under 400ms for a transfer. Zero tolerance for a lost or duplicated posting; we will trade latency for correctness, never the reverse."
Now the design is constrained: 52 minutes rules out single-region-single-AZ, and the zero-tolerance clause rules out fire-and-forget async posting.

The six non-functional dimensions to price

DimensionState it asBanking example
AvailabilityNines, per path99.99% posting, 99.9% reporting
LatencyPercentile, not averagep99 < 400ms transfer; p99 < 2s card auth
ThroughputPeak, not mean10M/day, 8× peak concentration
ConsistencyPer operationStrong on balance; eventual on statement
DurabilityLoss toleranceZero committed postings lost, ever
ComplianceNamed regimeCBN retention 7y; PCI DSS scope; GDPR residency

Notice that consistency is stated per operation. This is a senior move and it comes up repeatedly in this module: "strong where money moves, eventual where money is described." A system that is strongly consistent everywhere is slow for no benefit; a system that is eventually consistent everywhere loses money.

4

Back-of-envelope: the four numbers to derive

You need four numbers, and you should derive them out loud in under four minutes. They are the difference between "we'll need sharding" as an assertion and as a conclusion.

Our scale for this module: 20 million customers, 10 million transactions per day. Derive from there.

1 · Write rate

worked numbers
          10,000,000 txn/day ÷ 86,400 s = ~116 txn/s mean


          Traffic is never flat. Salary day, 6pm, month-end:

          peak multiplier 8× → ~930 txn/s peak


          But one transaction is not one write. In double-entry:

          2 ledger entries + 1 journal row + 2 balance updates + 1 outbox event

          = 6 writes per transaction


930 × 6 ≈ 5,600 writes/s at peak

That last number is the one that matters. A well-tuned single Postgres primary on good NVMe will do somewhere in the region of 10–20k simple writes/s, so 5,600/s is survivable on one node, but with very little headroom and no room for the next product launch. That is precisely the conversation Part 3 has.

2 · Read rate

worked numbers
          Balance checks dominate. Assume 20 per customer per day

          (app opens, card auths, pre-transfer checks):


          20,000,000 × 20 = 400,000,000 reads/day

          ÷ 86,400 = ~4,600 reads/s mean → ~37,000 reads/s peak


Read:write ratio ≈ 40:1

A 40:1 ratio is the single most design-shaping number here. It tells you immediately that the read path deserves its own treatment: caching, replicas, a materialised balance. And that optimising the write path for read convenience would be backwards.

3 · Storage growth

worked numbers
          Per ledger entry: ids, amount, currency, timestamps, refs ≈ 300 bytes

          2 entries per txn → 600 B, plus journal + indexes ≈ 1.5 KB/txn


          10M × 1.5 KB = 15 GB/day

          × 365 = 5.5 TB/year

          × 7 (CBN retention) = ~38 TB of immutable history
        

38 TB is not large for a warehouse and is very large for a hot OLTP primary. This one number is why Part 11 exists: the money path keeps months, the warehouse keeps years.

4 · Latency budget

worked numbers
          Card authorisation, network-mandated: 2,000 ms end to end


          scheme + acquirer network round trip  −600 ms

          our edge, auth, TLS  −100 ms

          fraud scoring  −150 ms

          safety margin  −250 ms

          ────────────────────────────

          left for the ledger decision: ~900 ms


          Within that, p99 target for the posting itself: < 300 ms

Budgets are subtractive, and you state them that way. "Two seconds" is the ceiling someone else set; 300ms is what is actually left for you. Part 8 spends a whole round inside this budget.

The numbers worth memorising

QuantityOrder of magnitude
Seconds in a day86,400 ≈ 10⁵
1M/day≈ 12/s
1M/s sustained≈ 86 billion/day, you are Visa
Postgres simple writes, one primary10k–20k/s
Redis GET, single node~100k/s
Kafka, one partition~10 MB/s, tens of thousands msg/s
Row in a hot cachetens of bytes to a few KB
5

Latency numbers every engineer should know

These are the physical constants of the job. You do not need them precisely; you need the ratios, because the ratios tell you what is worth avoiding.

OperationTimeRelative
L1 cache reference~1 ns1×
Main memory reference~100 ns100×
Read 1 MB sequentially from memory~10 µs10,000×
NVMe SSD random read~100 µs100,000×
Round trip within one datacentre~500 µs500,000×
Read 1 MB sequentially from SSD~1 ms10⁶×
Disk seek, spinning rust~10 ms10⁷×
Lagos → Frankfurt round trip~120 ms10⁸×
Lagos → Virginia round trip~180 ms~2×10⁸×

The three consequences that matter here

Cross-region is 100,000× a memory read. This is why Part 7 puts jurisdictional data in-region rather than making a European posting wait for a Nigerian primary. One cross-Atlantic hop eats 180ms of a 300ms budget.

A cache hit and a database hit differ by ~10×, not 1000×. Redis is not magic; it is one network hop to memory instead of one network hop to memory-plus-maybe-disk. Caching buys you throughput and tail-latency stability far more than it buys raw speed. Say it that way and you sound like you have measured something.

Sequential beats random by ~10×, at every level. Which is why the journal is an append and why monotonic ids beat random ones, a theme that runs from Part 1 to Part 11.

6

The vocabulary of tradeoffs

Every design decision in this module is an instance of a small number of recurring tensions. Learning them by name means you recognise the shape of a problem before you have solved it.

TensionWhat you gainWhat you payWhere it appears
Consistency ↔ availabilityCorrect reads alwaysUnavailability during partitionP2 isolation, P7 regions
Latency ↔ durabilityFast acknowledgementWindow where a crash loses dataP1 fsync, P4 acks
Normalised ↔ denormalisedOne source of truthReads need joins or fan-outP1 balances, P11 warehouse
Sync ↔ asyncImmediate certaintyCoupling and cascading failureP4 events, P6 rails
Push ↔ pullFreshnessBackpressure and thundering herdsP6 webhooks, P12 fan-out
Precision ↔ recallFewer false alarmsMore missed fraudP9 risk thresholds
Simplicity ↔ scaleComprehensible systemA ceiling you will hitEvery single round

The three-part sentence

There is a sentence shape that makes a tradeoff legible. Use it every time:

the tradeoff sentence
I'm choosing X over Y, because the requirement says Z. The cost is W, and I'd accept it because V. If Z changed, I'd revisit.

Filled in, from Part 3: "I'm choosing to shard by account_id over transaction_id, because the requirement says balance reads are 40× writes and must be p99 sub-100ms. The cost is that cross-account transfers now span two shards and need a saga instead of one transaction. I'd accept that because transfers are 6% of volume and reads are the hot path. If this were a write-dominated ledger with rare reads, I'd revisit."

The final clause, if Z changed, I'd revisit, is what separates a decision from a dogma, and interviewers hear the difference.

7

How to draw so the room follows you

The diagram is a shared workspace, not an artwork. Four rules make it legible.

Left to right is the flow of a request

Client on the left, storage on the right. Every arrow points the direction data travels. When you later add a component, its position on the page already tells the room where in the request it sits.

Label every arrow with protocol and sync-ness

HTTP sync, Kafka async, gRPC sync, CDC stream. An unlabelled arrow hides the single most important property of an integration: whether the caller waits. Most cascading-failure discussions start from an arrow someone assumed was async.

Draw the datastore with its access pattern, not just its name

Not Postgres but Postgres: ledger, append-only, sharded by account. The annotation is where the design lives; the box is just a rectangle.

Redraw, do not patch

When a round changes the architecture materially, draw the new version beside the old one rather than scribbling over it. The interviewer gets to see the delta, which is the thing they are grading. Every part in this module ends with a numbered sketch for exactly this reason: v1 … v13 is a record of pressure applied and absorbed.

8

Signals of seniority, and of its absence

Interviewers are pattern-matching against a small set of behaviours. These are the ones that move the needle, in both directions.

Positive signals

✓ Volunteering the weakness. "The obvious problem with v1 is the hot row on the company account, since every transaction touches it. Want me to fix that now or keep going?" You have now framed the next twenty minutes.
✓ Naming the number that decided it. "40:1 read:write is why the balance is cached rather than computed."
✓ Scoping out loud. "I'll leave KYC onboarding out unless you want it, it doesn't touch the money path."
✓ Saying what you would measure. "I'd want to see the p99 on the posting path per shard before deciding whether this needs partitioning."
✓ Admitting the edge of your knowledge, then reasoning anyway. "I haven't run Bigtable in production. From the model I'd expect the row key to be the whole game, so I'd design it as account-reversed-timestamp and validate that assumption early."

Negative signals

✗ Technology first. "I'd use Kafka and Cassandra" before any requirement exists to justify either.
✗ Uniform consistency. Strong everywhere, or eventual everywhere. Both mean the consistency question was never actually considered.
✗ Microservices by reflex. Eleven services for a problem that has one invariant. Boundaries with no reasoning behind them.
✗ Silence while thinking. Thirty seconds of nothing reads as stuck. "Let me think about the write path for a moment" costs you nothing and buys you the same time.
✗ Defending instead of updating. When the interviewer applies pressure, they are giving you the next requirement. Absorb it. Arguing that v1 was fine is the wrong instinct.

The reframe that helps most

The interviewer is not an examiner. They are the product manager who keeps remembering things they forgot to mention, and you are the engineer who has to keep the system coherent as they do. Every round in this module is written in that voice, because that is the job, in the room and afterwards.

With the method in place, we can start. Part 1 opens with four words and a blank page.