Part 8 · 8 chapters · ~45 min
Both boards, side by side
Eight parts of comparison, assembled. Eleven of fourteen layers are equivalent and three are not, which is a more useful conclusion than either platform winning. The closing argument is that almost nothing in the CBA module changed: the cloud decided who patches the database and what the bill looks like, and it decided nothing at all about what makes the money correct.
77
The full AWS architecture
| Layer | Service | Why |
|---|---|---|
| Edge | ALB, API Gateway, CloudFront | Standard. WAF in front of public endpoints |
| Posting, auth, relay | ECS Fargate | Warm pools, warm caches, steady load |
| Spiky work | Lambda | Webhooks, stream verifier, scheduled sweeps |
| Ledger | Aurora PostgreSQL, 8 shards | Fastest failover, which is the RTO |
| Connections | RDS Proxy | Pooling and failover-aware routing |
| Events | MSK, 48 partitions | Real partitions, 30-day retention, the Kafka ecosystem |
| Serving | DynamoDB | Per-account history at any scale |
| Cache | ElastiCache for Redis | Derived data only, per Part 3 |
| Lake and archive | S3 + Iceberg + Glue | One copy for compliance and analytics |
| Warehouse | Athena, Redshift if justified | Per-byte, no cluster to size |
| CDC | DMS | Into S3 for the analytics path |
| Keys | KMS, CloudHSM in the PCI account | Regional keys enforce residency |
| Observability | CloudWatch + EMF, X-Ray | Watch cardinality |
| Guardrails | Organizations + SCPs | Region pinning that an admin cannot escape |
both boards
the same design, two platforms
swipe the figure sideways, or tap expand for full screen
1/9
edge
The edge: load balancing and a web application firewall. Equivalent on both, and nobody chooses a cloud for this.
78
The full GCP architecture
| Layer | Service | Why |
|---|---|---|
| Edge | Cloud Load Balancing, Cloud Armor | Global anycast, integrated WAF |
| Posting, auth, relay | Cloud Run, min instances | Per-instance concurrency solves the connection problem |
| Spiky work | Cloud Run or Functions | Same runtime, so one mental model |
| Ledger | AlloyDB, 8 shards | Aurora-equivalent failover characteristics |
| Connections | Auth Proxy plus app pooling | Cloud Run concurrency does much of the work |
| Events | Managed Kafka, or Pub/Sub | Kafka for the ecosystem; Pub/Sub if retention is designed around |
| Serving | Bigtable | Globally sorted keys, per-account prefix scans |
| Cache | Memorystore | Equivalent |
| Lake and archive | GCS + BigLake | Same one-copy property |
| Warehouse | BigQuery | The strongest single product in this table |
| CDC | Datastream | Straight into BigQuery, very clean |
| Keys | Cloud KMS, Cloud HSM or EKM | EKM is the best answer to a regulator |
| Observability | Cloud Monitoring, Logging, Trace | Better log query language |
| Guardrails | Organization Policy | Cleaner residency control of the two |
the shape of the difference
Eleven of fourteen layers are equivalent. The three that are not: Cloud Run's concurrency model genuinely simplifies the connection story, BigQuery is a better warehouse than anything on the AWS side, and Organization Policy plus EKM are cleaner compliance controls. AWS wins on ecosystem breadth and on how many auditors and engineers have seen it before, which is not a technical argument and is a real one.
79
The service mapping table, complete
| Component | AWS | GCP | Equivalent? |
|---|---|---|---|
| Relational ledger | Aurora PostgreSQL | AlloyDB / Cloud SQL | Yes |
| Horizontally-scaled SQL | (none) | Spanner | No GCP-only |
| Event log | MSK | Managed Kafka | Yes |
| Serverless messaging | SNS + SQS, Kinesis | Pub/Sub | Different model |
| Containers, simple | ECS Fargate | Cloud Run | Cloud Run is simpler |
| Kubernetes | EKS | GKE | GKE is better run |
| Functions | Lambda | Cloud Functions | Yes, Lambda deeper |
| Wide column | DynamoDB | Bigtable | Similar, different model |
| Warehouse | Redshift / Athena | BigQuery | BigQuery ahead |
| Object storage | S3 | GCS | Yes |
| Cache | ElastiCache | Memorystore | Yes |
| CDC | DMS | Datastream | Yes |
| Orchestration | Step Functions | Workflows | Step Functions richer |
| Secrets | Secrets Manager | Secret Manager | Yes |
| Keys | KMS + CloudHSM | Cloud KMS + HSM + EKM | EKM is GCP-only |
| Identity | IAM roles, IRSA | Service accounts, Workload Identity | Yes |
| Org guardrails | SCPs | Organization Policy | GCP cleaner |
| Observability | CloudWatch, X-Ray | Monitoring, Trace | GCP logs better |
| Data pipelines | Glue, EMR | Dataflow | Dataflow stronger |
| ML platform | SageMaker | Vertex AI | Comparable |
80
The monthly bill, itemised and compared
worked numbers
at 20M customers and 10M transactions a day, one region:
AWS GCP
ledger (8 shards + replicas) $16k $15k
event backbone $2.8k $2.4k
compute $5.2k $4.6k
serving layer $7k $8k
warehouse + lake $3.4k $3.8k
cache $1.4k $1.3k
observability $4k $3.6k
security, keys, HSM $2.4k $2.0k
network + egress $2.6k $2.2k
──────────────────
per region, per month ~$45k ~$43k
× 4 regions ~$180k ~$172kwhat the numbers are and are not
These are order-of-magnitude figures for shaping a conversation, not a quote. Real pricing depends on commitments, region, negotiated discounts and instance families, and at this spend both providers negotiate. The useful observations are structural: the two platforms are within about 5% of each other, the ledger is roughly 35% of the bill, and observability and egress together are about 15%, which is more than most teams expect and is almost entirely under engineering control.
81
Where each platform is genuinely better
| AWS is better | GCP is better | |
|---|---|---|
| Breadth | More services, more integrations, more third parties | Narrower, more coherent |
| Data warehouse | BigQuery, clearly | |
| Simple containers | Cloud Run | |
| Kubernetes | GKE is better operated | |
| Functions | Lambda is deeper and more integrated | |
| Networking | More control, more primitives | Simpler, global by default |
| Compliance controls | Organization Policy, EKM | |
| Data pipelines | Glue is adequate | Dataflow is stronger |
| Hiring pool | Substantially larger | |
| Auditor familiarity | More auditors have seen it | |
| Payments ecosystem | More fintech vendors integrate first |
the pattern in that table
GCP wins on individual product quality in the places where it competes; AWS wins on breadth, ecosystem and familiarity. For a bank, breadth and familiarity are worth more than they sound: six payment integrations, a dozen vendors, auditors, and the ability to hire. That is why the recommendation lands on AWS despite GCP having the better warehouse and the cleaner compliance controls, and being able to hold both of those thoughts at once is the point.
82
The hybrid question, and when it is not madness
when a split is defensible
- Core on one, analytics on the other. Ledger and payments on AWS, warehouse on BigQuery. The data crosses once, in one direction, in batch, and the egress is a known cost. This is the most common defensible split.
- Regulatory concentration risk. A supervisor concerned that a whole sector runs on one provider may push for critical capability on a second, which is a compliance-driven rather than engineering-driven decision.
- Acquisition. You bought a company on the other platform and integrating is cheaper than migrating, at least for now.
- A genuinely unique service. Spanner or EKM, where nothing equivalent exists and the requirement is real.
when it is a mistake
- "For resilience." Running the money path across two providers roughly doubles the operational surface, and a team stretched across two platforms is worse at both. It usually reduces reliability.
- "To avoid lock-in." Building to the intersection of two platforms means using the weaker feature set of both, permanently, to buy an option you will probably never exercise.
- Without a platform team. Two clouds is two sets of IAM, networking, observability and runbooks. Part 0 chapter 3 identified operational muscle memory as the deepest lock-in, and a split divides it.
the answer
"Hybrid for a reason, never hybrid on principle. The split I would actually consider is core on AWS and warehouse on BigQuery, because the data crosses once in batch, the egress cost is knowable, and BigQuery is genuinely better. What I would not do is run the posting path across two providers for resilience, because that doubles the operational surface to protect against a failure mode that is rarer than the ones the complexity would introduce."
83
What stays the same on both
every design decision from the CBA module, unchanged
- Double-entry with a zero-sum invariant per currency. No cloud has an opinion about this.
- Integer minor units with the exponent from a currency table.
- Append-only entries, corrections as new linked journals,
UPDATEandDELETErevoked. - Idempotency on the intent, enforced by a unique constraint.
- Suspense accounts holding value whenever an outcome is uncertain.
- The transactional outbox, because the dual-write problem is not a cloud problem.
- Available versus ledger balance, holds, liens and the Part 18 entitlement layer.
- Continuous invariant checks and external reconciliation.
- Designed degradation, with kill switches on the front door only.
worked numbers
what the cloud actually changed: who patches the database how fast you can add a shard whether you run a Kafka cluster yourself what the bill looks like what it did not change: every single thing that makes the money correct. that is the whole lesson of this module.
the framing for an interview
"The architecture is the architecture. Moving it to a cloud changed who operates the components and what the failure modes look like, and it changed nothing about double-entry, idempotency, suspense accounts or reconciliation. An engineer who can only design on one platform learned the catalogue; an engineer who can design the system and then map it learned the system."
84
Defending it: the fifteen hardest questions
01Why not serverless everything? It scales to zero.
Because nothing here scales to zero. Part 3 derived a steady 464 transfers per second at peak, so there is no idle to save. And the posting path is connection-heavy and latency-sensitive, which is exactly the profile functions handle worst. I use Lambda for webhooks and stream verifiers, where the work genuinely is spiky and stateless.
02Why not Spanner? It removes your sharding entirely.
It removes sharding I already have working, and it charges for global external consistency inside a single legal region. Part 7 made regions a legal boundary, so cross-region consistency is not available to use. Paying for a property you cannot exercise is the most common cloud mistake.
03Aurora is proprietary. Is that not lock-in?
It speaks the Postgres wire protocol, so the application is portable. The deeper lock-in is operational, per Part 0: runbooks, IAM, dashboards and what the team knows at 3am. I accept Aurora because the failover duration is the RTO and it is consistently the fastest, and because the escape path to plain Postgres exists.
04Pub/Sub has no partitions. Is that disqualifying?
No. Ordering keys give per-account ordering, which is all Part 4 required. The real tradeoff is retention and replay precision: 31 days maximum and seek-by-timestamp rather than by offset. That is acceptable if the rebuild-from-warehouse path is designed and drilled, and dangerous if discovered during an incident.
05Your RTO is 60 seconds. Prove it.
I cannot from documentation, only from a drill. Part 1 decomposed a real failover into detection, promotion, endpoint propagation, pool recovery and backlog drain, which totals 30 seconds to 4 minutes against a documented 30 seconds. The RTO is what the monthly drill measured, and that number is published.
06What is the single biggest cost surprise?
Observability, at around 9% of the bill, above the warehouse. It is driven almost entirely by metric cardinality, which engineers increase by adding a dimension without noticing the multiplication. Second is cross-AZ and egress traffic, which a chatty service architecture pays continuously and which never appears in a capacity plan.
07Multi-cloud for resilience?
No. Running the money path across two providers roughly doubles the operational surface and divides the team's expertise, which usually reduces reliability rather than improving it. Core on AWS with the warehouse on BigQuery is a split I would defend, because the data crosses once in batch with a knowable egress cost.
08Can you run this in Nigeria?
Neither provider has a Nigerian region. For CBN-regulated data that is a hard constraint, and the honest options are a local deployment for the regulated workload with cloud for what is permitted, or an arrangement the regulator has approved. I would not design as though a Lagos region exists.
09You are on a PCI-certified cloud. Are you PCI compliant?
No. The provider certification covers security of the cloud; the configuration, the application and the data are mine. Part 8 minimised scope by keeping the ledger free of card data, and on the cloud that becomes a separate account with its own VPC, keys and HSM. The certification is evidence for an auditor, not a conclusion.
10How do you stop a cost spike becoming a surprise?
Treat it as an SRE concern. Anomaly alerts on daily spend per service against the same day last week, which is the Part 13 business-metric pattern applied to cost. A cost spike is usually an incident: a retry loop, a scan without a partition filter, a stuck job.
11Kubernetes or not?
Only with the team to run it. For nine services and no platform team, ECS Fargate or Cloud Run delivers the same outcome with a fraction of the operational surface. Choosing Kubernetes because it is standard, without that team, means acquiring a second full-time system to operate.
12How do you migrate without downtime?
Shard by shard, with a 20 to 30 second write pause per batch, queued rather than rejected. The shard map from Part 3 is already the routing layer. I would not dual-write, because dual-writing a ledger is the Part 4 dual-write problem applied to the most critical table in the company.
13How do you know the migration was correct?
The same way I know the ledger is correct. Part 10 reconciliation, pointed at the legacy system as a counterparty: row equivalence, the invariant per shard, balance equivalence, trial balance equivalence, external reconciliation, and serving equivalence. Six gates per batch, all of which must pass.
14What would you refuse to move?
The HSM if the regulator requires physical key custody, scheme connectivity if it is contractually location-bound, anything the regulator has not approved, and data the residency rules forbid. A short, reasoned exception list is what a credible migration plan looks like, and it is what the outsourcing notification needs anyway.
15If you had to choose one platform today, which?
AWS, and not because it is technically superior. Eleven of fourteen layers are equivalent. It wins on ecosystem breadth, vendor integrations, auditor familiarity and hiring, which for a bank with six payment integrations matter more than they should. GCP has the better warehouse and the cleaner compliance controls, and if analytics were the product rather than supporting it, that would flip the decision.
the answer shape, one last time
Direct answer, then the reason, then the cost. And where the honest answer is "it depends on something other than technology", say what it depends on. The interviewer is testing whether you can make a decision under ambiguity and defend it, not whether you memorised a catalogue that will be out of date next year.