How to design a system in a room
Before the bank, the method. A system design interview is not a test of whether you know what Kafka is. It is a test of whether you can take an ambiguous sentence, turn it into a specification, choose between imperfect options, and defend the choice when the ground shifts. This part is the frame every later round runs inside, so when the interviewer says "now ten million a day", you already know which four things to update and in what order.
What the interviewer is actually measuring
The prompt is always underspecified. "Design a wallet." Four words against forty-five minutes. This is deliberate, and it is the first measurement: a junior engineer starts drawing boxes, a senior engineer starts asking what the boxes are for.
There are five things being assessed, and they are weighted very unevenly.
| What is measured | Weight | How it shows |
|---|---|---|
| Scoping | High | Do you convert ambiguity into a written spec, or start solving the wrong problem confidently? |
| Tradeoff reasoning | Highest | Can you name two viable options, state what each costs, and pick one for a reason tied to the requirements? |
| Depth on demand | High | When pushed on one box, can you go three levels down, and stop before you waffle? |
| Failure thinking | High | Do you volunteer what breaks, or only discuss it when asked? |
| Component knowledge | Low | Knowing Kafka has partitions is table stakes. It earns almost nothing on its own. |
That last row is the one candidates over-invest in. Reciting that DynamoDB has partition keys is not a signal. Saying "I'll key by account_id rather than transaction_id, because transaction_id spreads a single account's history across every shard and every balance read becomes a scatter-gather": that is the signal. Same knowledge, applied under constraint.
The asymmetry worth internalising
A candidate who designs a merely adequate system but narrates every decision cleanly will usually out-score a candidate who draws a more sophisticated system in silence. The interviewer cannot grade what they cannot hear. Your reasoning is the artefact; the diagram is a by-product.
The seven-step frame, and why order matters
Run every design through the same seven steps, in this order, out loud. The order is not decoration. Each step constrains the next, and skipping forward means you make choices you cannot justify.
Clarify
Convert the four-word prompt into a bounded problem. Three to five questions, chosen for the ones whose answers would change your design.
3–5 minScope
State what you will design and what you will explicitly leave out. This is not hedging; it is how you buy permission to go deep instead of wide.
1–2 minRequirements
Functional as verbs. Non-functional as numbers. If a non-functional requirement has no number attached, it is a wish, not a requirement.
4–5 minEstimate
Four numbers: write rate, read rate, storage growth, and the tightest latency budget. These decide whether you need one database or twelve.
3–4 minData model
Entities and their relationships, before any service boxes. Data outlives every service that reads it. Get this wrong and nothing downstream is fixable.
5–7 minDesign
Components, then the flow of one concrete request end to end. A design you cannot trace a single request through is not a design yet.
10–15 minStress
Bottlenecks, failure modes, and the tradeoffs you accepted. Volunteer these. The candidate who raises the weakness first controls how it is discussed.
8–10 min
Why data model before components
Because the shape of the data decides the shape of the system. If you draw services first, you invent a service topology and then bend the data to fit it. If you model the data first, the service boundaries fall out of it almost mechanically: a service owns the entities whose invariants it is responsible for. In the ledger we are about to build, the entire architecture is downstream of one modelling decision: whether a balance is a stored number or a derived one.
Functional vs non-functional, stated properly
Most candidates state requirements badly, in one of two ways: functional requirements that are really features, or non-functional requirements with no numbers.
Functional: what the system does
Write them as verbs with an actor and an object. Not "wallets" but "a customer can hold a balance in a given currency". The verb form forces you to notice missing pieces. Hold implies create, and create implies who is allowed to.
Non-functional: the qualities it must have while doing it
These are the ones that actually pick your technology. And they are useless without numbers. Compare:
Unfalsifiable. Every design satisfies it. Decides nothing.
Now the design is constrained: 52 minutes rules out single-region-single-AZ, and the zero-tolerance clause rules out fire-and-forget async posting.
The six non-functional dimensions to price
| Dimension | State it as | Banking example |
|---|---|---|
| Availability | Nines, per path | 99.99% posting, 99.9% reporting |
| Latency | Percentile, not average | p99 < 400ms transfer; p99 < 2s card auth |
| Throughput | Peak, not mean | 10M/day, 8× peak concentration |
| Consistency | Per operation | Strong on balance; eventual on statement |
| Durability | Loss tolerance | Zero committed postings lost, ever |
| Compliance | Named regime | CBN retention 7y; PCI DSS scope; GDPR residency |
Notice that consistency is stated per operation. This is a senior move and it comes up repeatedly in this module: "strong where money moves, eventual where money is described." A system that is strongly consistent everywhere is slow for no benefit; a system that is eventually consistent everywhere loses money.
Back-of-envelope: the four numbers to derive
You need four numbers, and you should derive them out loud in under four minutes. They are the difference between "we'll need sharding" as an assertion and as a conclusion.
Our scale for this module: 20 million customers, 10 million transactions per day. Derive from there.
1 · Write rate
10,000,000 txn/day ÷ 86,400 s = ~116 txn/s mean
Traffic is never flat. Salary day, 6pm, month-end:
peak multiplier 8× → ~930 txn/s peak
But one transaction is not one write. In double-entry:
2 ledger entries + 1 journal row + 2 balance updates + 1 outbox event
= 6 writes per transaction
930 × 6 ≈ 5,600 writes/s at peakThat last number is the one that matters. A well-tuned single Postgres primary on good NVMe will do somewhere in the region of 10–20k simple writes/s, so 5,600/s is survivable on one node, but with very little headroom and no room for the next product launch. That is precisely the conversation Part 3 has.
2 · Read rate
Balance checks dominate. Assume 20 per customer per day
(app opens, card auths, pre-transfer checks):
20,000,000 × 20 = 400,000,000 reads/day
÷ 86,400 = ~4,600 reads/s mean → ~37,000 reads/s peak
Read:write ratio ≈ 40:1A 40:1 ratio is the single most design-shaping number here. It tells you immediately that the read path deserves its own treatment: caching, replicas, a materialised balance. And that optimising the write path for read convenience would be backwards.
3 · Storage growth
Per ledger entry: ids, amount, currency, timestamps, refs ≈ 300 bytes
2 entries per txn → 600 B, plus journal + indexes ≈ 1.5 KB/txn
10M × 1.5 KB = 15 GB/day
× 365 = 5.5 TB/year
× 7 (CBN retention) = ~38 TB of immutable history
38 TB is not large for a warehouse and is very large for a hot OLTP primary. This one number is why Part 11 exists: the money path keeps months, the warehouse keeps years.
4 · Latency budget
Card authorisation, network-mandated: 2,000 ms end to end
scheme + acquirer network round trip −600 ms
our edge, auth, TLS −100 ms
fraud scoring −150 ms
safety margin −250 ms
────────────────────────────
left for the ledger decision: ~900 ms
Within that, p99 target for the posting itself: < 300 msBudgets are subtractive, and you state them that way. "Two seconds" is the ceiling someone else set; 300ms is what is actually left for you. Part 8 spends a whole round inside this budget.
The numbers worth memorising
| Quantity | Order of magnitude |
|---|---|
| Seconds in a day | 86,400 ≈ 10⁵ |
| 1M/day | ≈ 12/s |
| 1M/s sustained | ≈ 86 billion/day, you are Visa |
| Postgres simple writes, one primary | 10k–20k/s |
| Redis GET, single node | ~100k/s |
| Kafka, one partition | ~10 MB/s, tens of thousands msg/s |
| Row in a hot cache | tens of bytes to a few KB |
Latency numbers every engineer should know
These are the physical constants of the job. You do not need them precisely; you need the ratios, because the ratios tell you what is worth avoiding.
| Operation | Time | Relative |
|---|---|---|
| L1 cache reference | ~1 ns | 1× |
| Main memory reference | ~100 ns | 100× |
| Read 1 MB sequentially from memory | ~10 µs | 10,000× |
| NVMe SSD random read | ~100 µs | 100,000× |
| Round trip within one datacentre | ~500 µs | 500,000× |
| Read 1 MB sequentially from SSD | ~1 ms | 10⁶× |
| Disk seek, spinning rust | ~10 ms | 10⁷× |
| Lagos → Frankfurt round trip | ~120 ms | 10⁸× |
| Lagos → Virginia round trip | ~180 ms | ~2×10⁸× |
The three consequences that matter here
Cross-region is 100,000× a memory read. This is why Part 7 puts jurisdictional data in-region rather than making a European posting wait for a Nigerian primary. One cross-Atlantic hop eats 180ms of a 300ms budget.
A cache hit and a database hit differ by ~10×, not 1000×. Redis is not magic; it is one network hop to memory instead of one network hop to memory-plus-maybe-disk. Caching buys you throughput and tail-latency stability far more than it buys raw speed. Say it that way and you sound like you have measured something.
Sequential beats random by ~10×, at every level. Which is why the journal is an append and why monotonic ids beat random ones, a theme that runs from Part 1 to Part 11.
The vocabulary of tradeoffs
Every design decision in this module is an instance of a small number of recurring tensions. Learning them by name means you recognise the shape of a problem before you have solved it.
| Tension | What you gain | What you pay | Where it appears |
|---|---|---|---|
| Consistency ↔ availability | Correct reads always | Unavailability during partition | P2 isolation, P7 regions |
| Latency ↔ durability | Fast acknowledgement | Window where a crash loses data | P1 fsync, P4 acks |
| Normalised ↔ denormalised | One source of truth | Reads need joins or fan-out | P1 balances, P11 warehouse |
| Sync ↔ async | Immediate certainty | Coupling and cascading failure | P4 events, P6 rails |
| Push ↔ pull | Freshness | Backpressure and thundering herds | P6 webhooks, P12 fan-out |
| Precision ↔ recall | Fewer false alarms | More missed fraud | P9 risk thresholds |
| Simplicity ↔ scale | Comprehensible system | A ceiling you will hit | Every single round |
The three-part sentence
There is a sentence shape that makes a tradeoff legible. Use it every time:
Filled in, from Part 3: "I'm choosing to shard by account_id over transaction_id, because the requirement says balance reads are 40× writes and must be p99 sub-100ms. The cost is that cross-account transfers now span two shards and need a saga instead of one transaction. I'd accept that because transfers are 6% of volume and reads are the hot path. If this were a write-dominated ledger with rare reads, I'd revisit."
The final clause, if Z changed, I'd revisit, is what separates a decision from a dogma, and interviewers hear the difference.
How to draw so the room follows you
The diagram is a shared workspace, not an artwork. Four rules make it legible.
Left to right is the flow of a request
Client on the left, storage on the right. Every arrow points the direction data travels. When you later add a component, its position on the page already tells the room where in the request it sits.
Label every arrow with protocol and sync-ness
HTTP sync, Kafka async, gRPC sync,
CDC stream. An unlabelled arrow hides the single most important
property of an integration: whether the caller waits. Most cascading-failure
discussions start from an arrow someone assumed was async.
Draw the datastore with its access pattern, not just its name
Not Postgres but Postgres: ledger, append-only, sharded by
account. The annotation is where the design lives; the box is just a
rectangle.
Redraw, do not patch
When a round changes the architecture materially, draw the new version beside the old one rather than scribbling over it. The interviewer gets to see the delta, which is the thing they are grading. Every part in this module ends with a numbered sketch for exactly this reason: v1 … v13 is a record of pressure applied and absorbed.
Signals of seniority, and of its absence
Interviewers are pattern-matching against a small set of behaviours. These are the ones that move the needle, in both directions.
Positive signals
Negative signals
The reframe that helps most
The interviewer is not an examiner. They are the product manager who keeps remembering things they forgot to mention, and you are the engineer who has to keep the system coherent as they do. Every round in this module is written in that voice, because that is the job, in the room and afterwards.
With the method in place, we can start. Part 1 opens with four words and a blank page.