Part 7 · 10 chapters · ~55 min
The migration: moving a running bank
The most realistic problem in this module and the one most likely to be asked, because greenfield cloud architecture is a design exercise while migrating a running ledger is an engineering one. The good news is that the decisions made in rounds one and three of the CBA module, append-only entries with monotonic ids and application-level sharding, turn out to be exactly what makes a migration provable and incremental.
67
The premise: you already have a system
interviewer
“You have a working bank on self-managed Postgres and Kafka in a datacentre. Move it to the cloud without losing a transaction, and without stopping.”
The hardest problem in this module, and the most realistic. Greenfield cloud architecture is a design exercise; migrating a running ledger is an engineering one.
what makes this different from a normal migration
- You cannot stop. A bank has no maintenance window long enough, and a card network does not accept one.
- You cannot lose a transaction. Not one, not ever, and the Part 10 invariant will tell you if you did.
- You cannot corrupt. A migration that duplicates a posting is worse than one that fails.
- You must be able to go back. Until a point of no return you choose deliberately and announce.
worked numbers
the property that makes it tractable: the ledger is append-only and every entry has a monotonic id. so "has everything up to N been copied" is answerable exactly, and "do both systems agree" is a query, not an inference. the round-one design decisions are what make the migration provable.
68
What you move first, and what you move last
| Order | Component | Why here |
|---|---|---|
| 1 | Analytics and the lake | No customer impact. Wrong answers are visible and fixable. Builds confidence and cloud muscle |
| 2 | Notifications | Failure is annoying, not financial. Exercises the event consumption path |
| 3 | Serving layer | Read-only derived copy. Rebuild it in the cloud and compare against the old one |
| 4 | Event backbone | Now the hard part begins. Must not lose an event |
| 5 | Stateless services | Posting service, connectors, risk. Move compute before state |
| 6 | The ledger | Last, and alone. Everything else is already proven |
| never | The card gateway, possibly | Scheme connectivity may be contractually tied to a location |
the ordering principle
Move in increasing order of consequence-of-failure. By the time the ledger moves, the team has migrated six things, the tooling is proven, the runbooks exist and the monitoring is in place. Teams that start with the database because it is the hard part do the riskiest step with the least experience, which is exactly backwards.
69
The strangler pattern applied to a ledger
how it works for a sharded ledger specifically
- The shard map is already the routing layer. Part 3 built it, and it can point some logical shards at the old system and some at the new.
- Migrate one logical shard range at a time. 512 shards of 4,096 is 12.5% of accounts, which is a real test and a bounded blast radius.
- Each shard is independently correct. The invariant holds per shard, so a migrated shard proves itself without reference to the others.
- Cross-shard transfers span both systems during the transition, which is the Part 3 saga with a longer, slower leg. It already handles that case.
- Rollback is a routing change, until the point where the new system has accepted writes that the old one has not seen.
why this is not luck
The migration is tractable because Part 3 sharded the system. An unsharded monolithic ledger has to move all at once, which is a big-bang cutover with no incremental proof. The sharding was done for write throughput and it pays off here as migration granularity, which is a good example of a structural decision paying dividends in a dimension it was not chosen for.
strangler
routing shifts while both systems stay correct
swipe the figure sideways, or tap expand for full screen
1/8
legacy only
A working bank on self-managed infrastructure. All traffic routes through the shard map to the legacy ledger.
70
Dual writing, and why it is the dual-write problem again
worked numbers
the tempting approach: write to the old ledger and the new one, compare, cut over this is Part 4 chapter 48, exactly. two systems, two commits, no shared transaction a crash between them leaves them permanently divergent and now the divergence is in the ledger itself the same fix applies: write to ONE system and derive the other.
the correct approach
- One system is authoritative for a given shard at a given time. Never both.
- The other is fed by replication, either CDC from the old system or the event log, so it is derived rather than independently written.
- Compare continuously while the new one is passive, using the Part 10 machinery pointed at the two systems.
- Cut over by changing the routing, at which point the new system becomes authoritative and the old one becomes the derived copy for a defined bake period.
- Then stop replicating, which is the point of no return for that shard.
the sentence to say
"I would not dual-write. Dual writing a ledger is the dual-write problem from Part 4 applied to the most critical table in the company, and its failure mode is two divergent versions of the truth. One writer, replication to the other, continuous comparison, and cut over by routing. The comparison is exactly the reconciliation machinery from Part 10, with the old system as the counterparty."
71
Migrating the event backbone without losing an event
the sequence
- Stand up the new cluster with the same topics and partition counts.
- Mirror from old to new with MirrorMaker 2 or an equivalent, so the new cluster has full history and stays current.
- Move consumers first, one at a time. A consumer reading from the new cluster with translated offsets, verified by comparing its output against the old one.
- Move producers last, because once a producer writes to the new cluster, the old one is incomplete.
- Bake with both clusters running and mirroring stopped, watching consumer lag and output equivalence.
- Decommission only after the retention window has passed, so a replay is still possible from the old cluster if needed.
the offset translation problem
Offsets are not portable between clusters. MirrorMaker maintains a translation but it is approximate, so a consumer restarted at a translated offset may reprocess or, worse, skip. This is why consumer idempotency from Part 4 chapter 52 is what makes the migration safe: reprocessing is harmless. If your consumers are not idempotent, fix that before migrating, not during.
72
Migrating data: DMS, Datastream, and the cutover
AWS
GCP
DMS
with full load plus CDC: bulk copy the existing rows, then stream ongoing changes until the gap is near zero.
Validation task compares source and target row counts and checksums.
Validation task compares source and target row counts and checksums.
Datastream for Postgres to Postgres or into BigQuery, same model.
Or native logical replication, which for Postgres to Postgres is often the cleanest path of all.
Or native logical replication, which for Postgres to Postgres is often the cleanest path of all.
worked numbers
the cutover, per shard, with the gap closed: 1. replication lag < 1 second, sustained 2. stop accepting writes for this shard ~10-30 s 3. wait for lag to reach zero 4. verify: max(entry_id) matches on both 5. verify: the invariant holds on the new system 6. flip the shard map 7. resume writes, now against the new system a 10 to 30 second write pause per shard, not a maintenance window. and queued requests drain afterwards, per Part 13 degradation.
why the pause is acceptable
Part 13 designed graceful degradation and Part 3 designed journal-first acceptance. A 20-second pause on 12.5% of accounts, during off-peak, with requests queued rather than rejected, is a degradation the system was built to absorb. Stating it that way, with the mechanism named, is much stronger than promising zero downtime and discovering the truth during the cutover.
73
Proving the migration: reconciliation as the gate
the gates, each of which must pass before proceeding
- Row-level equivalence. Every entry id in the old system exists in the new one with identical amount, account and currency.
- The invariant on the new system. Sums to zero per journal per currency, per Part 10 chapter 113.
- Balance equivalence. Computed balance per account matches between systems, sampled across every account type and the Part 18 tier and restriction states.
- Aggregate equivalence. The trial balance from Part 19 produces identical figures on both.
- External reconciliation still passes against every counterparty from the new system.
- Serving equivalence, per Part 19 chapter 225: what a customer would see matches on both.
the framing that makes this a strength
The migration is provable because the system was built to prove itself. Part 10 built reconciliation to catch drift against external counterparties; pointing the same machinery at the legacy system turns a migration from an act of faith into a series of passed checks. "How do you know the migration was correct?" has the same answer as "how do you know the ledger is correct?", which is a satisfying thing to be able to say.
74
Rollback, and the point of no return
worked numbers
reversible, while the new system is derived: flip the shard map back. seconds. no data loss. reversible with effort, during the bake period: reverse replication is running, so flip back and let it catch up this is why reverse replication is configured BEFORE cutover the point of no return: when reverse replication stops, or when the old system's retention has passed, whichever is first after this, going back means restoring from backup and losing everything since.
the discipline around it
- Configure reverse replication before cutting over, not after. Setting it up during an incident is not a plan.
- Announce the point of no return explicitly, with a date, so nobody assumes a rollback is still available.
- Define the rollback decision criteria in advance, in numbers, so the decision during an incident is a lookup rather than a debate.
- Rehearse the rollback on a non-production shard. An untested rollback is a hypothesis, exactly as Part 13 said about backups.
75
What you would refuse to move
the honest list
- The HSM, possibly. If the regulator requires keys in a specific physical location under your control, a cloud HSM may not satisfy it. GCP's External Key Manager is the best answer; sometimes the answer is that it stays.
- Scheme connectivity, sometimes. Card scheme links are contractual and location-bound. Moving them can require the scheme's agreement and a new certification.
- Anything a regulator has not approved. Part 0 chapter 4: material outsourcing may need notification or approval, and moving before that is a compliance breach regardless of how well it works.
- Data the residency rules forbid. Part 7 and Part 0 chapter 7: for CBN-regulated data with no in-country region, the honest answer is that it does not move.
- Nothing else. Everything else is an engineering problem with a known shape.
the value of having a refusal list
An architect who says everything can move has not read the contracts. An architect who says nothing can move has not read the technology. A specific, short, reasoned list of exceptions is what a credible migration plan looks like, and it is also what a regulator wants to see in the outsourcing notification.
76
The migration plan, as one page
phases, with gates
- Phase 0, foundations. Accounts, networks, IAM, IaC, observability, and one non-critical service deployed end to end. Gate: a drill passes.
- Phase 1, analytics. Lake, warehouse, CDC. Gate: reports match the legacy ones for a full month.
- Phase 2, derived systems. Serving layer and notifications. Gate: serving equivalence passes on a sample.
- Phase 3, event backbone. Mirror, move consumers, move producers, bake. Gate: no lost events across a full retention window.
- Phase 4, stateless services. Posting service, connectors, risk. Gate: latency and error SLOs met for two weeks.
- Phase 5, the ledger, shard by shard. 512 logical shards at a time. Gate per batch: all six checks from chapter 73.
- Phase 6, decommission. After the retention window and the announced point of no return.
worked numbers
realistic duration for a bank at this scale:
phase 0 2 to 3 months
phases 1-2 3 to 4 months
phase 3 2 months
phase 4 2 to 3 months
phase 5 3 to 6 months, deliberately slow
phase 6 1 to 2 months
──────────────
13 to 20 months
anyone promising six months has not migrated a ledger.