Part 8 · 8 chapters · ~45 min

Both boards, side by side

Eight parts of comparison, assembled. Eleven of fourteen layers are equivalent and three are not, which is a more useful conclusion than either platform winning. The closing argument is that almost nothing in the CBA module changed: the cloud decided who patches the database and what the bill looks like, and it decided nothing at all about what makes the money correct.

77

The full AWS architecture

LayerServiceWhy
EdgeALB, API Gateway, CloudFrontStandard. WAF in front of public endpoints
Posting, auth, relayECS FargateWarm pools, warm caches, steady load
Spiky workLambdaWebhooks, stream verifier, scheduled sweeps
LedgerAurora PostgreSQL, 8 shardsFastest failover, which is the RTO
ConnectionsRDS ProxyPooling and failover-aware routing
EventsMSK, 48 partitionsReal partitions, 30-day retention, the Kafka ecosystem
ServingDynamoDBPer-account history at any scale
CacheElastiCache for RedisDerived data only, per Part 3
Lake and archiveS3 + Iceberg + GlueOne copy for compliance and analytics
WarehouseAthena, Redshift if justifiedPer-byte, no cluster to size
CDCDMSInto S3 for the analytics path
KeysKMS, CloudHSM in the PCI accountRegional keys enforce residency
ObservabilityCloudWatch + EMF, X-RayWatch cardinality
GuardrailsOrganizations + SCPsRegion pinning that an admin cannot escape
both boards
the same design, two platforms
swipe the figure sideways, or tap expand for full screen
1/9
edge
The edge: load balancing and a web application firewall. Equivalent on both, and nobody chooses a cloud for this.
78

The full GCP architecture

LayerServiceWhy
EdgeCloud Load Balancing, Cloud ArmorGlobal anycast, integrated WAF
Posting, auth, relayCloud Run, min instancesPer-instance concurrency solves the connection problem
Spiky workCloud Run or FunctionsSame runtime, so one mental model
LedgerAlloyDB, 8 shardsAurora-equivalent failover characteristics
ConnectionsAuth Proxy plus app poolingCloud Run concurrency does much of the work
EventsManaged Kafka, or Pub/SubKafka for the ecosystem; Pub/Sub if retention is designed around
ServingBigtableGlobally sorted keys, per-account prefix scans
CacheMemorystoreEquivalent
Lake and archiveGCS + BigLakeSame one-copy property
WarehouseBigQueryThe strongest single product in this table
CDCDatastreamStraight into BigQuery, very clean
KeysCloud KMS, Cloud HSM or EKMEKM is the best answer to a regulator
ObservabilityCloud Monitoring, Logging, TraceBetter log query language
GuardrailsOrganization PolicyCleaner residency control of the two
the shape of the difference
Eleven of fourteen layers are equivalent. The three that are not: Cloud Run's concurrency model genuinely simplifies the connection story, BigQuery is a better warehouse than anything on the AWS side, and Organization Policy plus EKM are cleaner compliance controls. AWS wins on ecosystem breadth and on how many auditors and engineers have seen it before, which is not a technical argument and is a real one.
79

The service mapping table, complete

ComponentAWSGCPEquivalent?
Relational ledgerAurora PostgreSQLAlloyDB / Cloud SQLYes
Horizontally-scaled SQL(none)SpannerNo GCP-only
Event logMSKManaged KafkaYes
Serverless messagingSNS + SQS, KinesisPub/SubDifferent model
Containers, simpleECS FargateCloud RunCloud Run is simpler
KubernetesEKSGKEGKE is better run
FunctionsLambdaCloud FunctionsYes, Lambda deeper
Wide columnDynamoDBBigtableSimilar, different model
WarehouseRedshift / AthenaBigQueryBigQuery ahead
Object storageS3GCSYes
CacheElastiCacheMemorystoreYes
CDCDMSDatastreamYes
OrchestrationStep FunctionsWorkflowsStep Functions richer
SecretsSecrets ManagerSecret ManagerYes
KeysKMS + CloudHSMCloud KMS + HSM + EKMEKM is GCP-only
IdentityIAM roles, IRSAService accounts, Workload IdentityYes
Org guardrailsSCPsOrganization PolicyGCP cleaner
ObservabilityCloudWatch, X-RayMonitoring, TraceGCP logs better
Data pipelinesGlue, EMRDataflowDataflow stronger
ML platformSageMakerVertex AIComparable
80

The monthly bill, itemised and compared

worked numbers
at 20M customers and 10M transactions a day, one region:

                                      AWS         GCP
  ledger (8 shards + replicas)     $16k        $15k
  event backbone                          $2.8k       $2.4k
  compute                                  $5.2k       $4.6k
  serving layer                            $7k         $8k
  warehouse + lake                      $3.4k       $3.8k
  cache                                      $1.4k       $1.3k
  observability                         $4k         $3.6k
  security, keys, HSM                $2.4k       $2.0k
  network + egress                    $2.6k       $2.2k
                                       ──────────────────
  per region, per month       ~$45k     ~$43k
  × 4 regions                         ~$180k   ~$172k
what the numbers are and are not
These are order-of-magnitude figures for shaping a conversation, not a quote. Real pricing depends on commitments, region, negotiated discounts and instance families, and at this spend both providers negotiate. The useful observations are structural: the two platforms are within about 5% of each other, the ledger is roughly 35% of the bill, and observability and egress together are about 15%, which is more than most teams expect and is almost entirely under engineering control.
81

Where each platform is genuinely better

AWS is betterGCP is better
BreadthMore services, more integrations, more third partiesNarrower, more coherent
Data warehouseBigQuery, clearly
Simple containersCloud Run
KubernetesGKE is better operated
FunctionsLambda is deeper and more integrated
NetworkingMore control, more primitivesSimpler, global by default
Compliance controlsOrganization Policy, EKM
Data pipelinesGlue is adequateDataflow is stronger
Hiring poolSubstantially larger
Auditor familiarityMore auditors have seen it
Payments ecosystemMore fintech vendors integrate first
the pattern in that table
GCP wins on individual product quality in the places where it competes; AWS wins on breadth, ecosystem and familiarity. For a bank, breadth and familiarity are worth more than they sound: six payment integrations, a dozen vendors, auditors, and the ability to hire. That is why the recommendation lands on AWS despite GCP having the better warehouse and the cleaner compliance controls, and being able to hold both of those thoughts at once is the point.
82

The hybrid question, and when it is not madness

when a split is defensible
  1. Core on one, analytics on the other. Ledger and payments on AWS, warehouse on BigQuery. The data crosses once, in one direction, in batch, and the egress is a known cost. This is the most common defensible split.
  2. Regulatory concentration risk. A supervisor concerned that a whole sector runs on one provider may push for critical capability on a second, which is a compliance-driven rather than engineering-driven decision.
  3. Acquisition. You bought a company on the other platform and integrating is cheaper than migrating, at least for now.
  4. A genuinely unique service. Spanner or EKM, where nothing equivalent exists and the requirement is real.
when it is a mistake
  1. "For resilience." Running the money path across two providers roughly doubles the operational surface, and a team stretched across two platforms is worse at both. It usually reduces reliability.
  2. "To avoid lock-in." Building to the intersection of two platforms means using the weaker feature set of both, permanently, to buy an option you will probably never exercise.
  3. Without a platform team. Two clouds is two sets of IAM, networking, observability and runbooks. Part 0 chapter 3 identified operational muscle memory as the deepest lock-in, and a split divides it.
the answer
"Hybrid for a reason, never hybrid on principle. The split I would actually consider is core on AWS and warehouse on BigQuery, because the data crosses once in batch, the egress cost is knowable, and BigQuery is genuinely better. What I would not do is run the posting path across two providers for resilience, because that doubles the operational surface to protect against a failure mode that is rarer than the ones the complexity would introduce."
83

What stays the same on both

every design decision from the CBA module, unchanged
  1. Double-entry with a zero-sum invariant per currency. No cloud has an opinion about this.
  2. Integer minor units with the exponent from a currency table.
  3. Append-only entries, corrections as new linked journals, UPDATE and DELETE revoked.
  4. Idempotency on the intent, enforced by a unique constraint.
  5. Suspense accounts holding value whenever an outcome is uncertain.
  6. The transactional outbox, because the dual-write problem is not a cloud problem.
  7. Available versus ledger balance, holds, liens and the Part 18 entitlement layer.
  8. Continuous invariant checks and external reconciliation.
  9. Designed degradation, with kill switches on the front door only.
worked numbers
what the cloud actually changed:

  who patches the database
  how fast you can add a shard
  whether you run a Kafka cluster yourself
  what the bill looks like

what it did not change:
  every single thing that makes the money correct.

that is the whole lesson of this module.
the framing for an interview
"The architecture is the architecture. Moving it to a cloud changed who operates the components and what the failure modes look like, and it changed nothing about double-entry, idempotency, suspense accounts or reconciliation. An engineer who can only design on one platform learned the catalogue; an engineer who can design the system and then map it learned the system."
84

Defending it: the fifteen hardest questions

01Why not serverless everything? It scales to zero.
Because nothing here scales to zero. Part 3 derived a steady 464 transfers per second at peak, so there is no idle to save. And the posting path is connection-heavy and latency-sensitive, which is exactly the profile functions handle worst. I use Lambda for webhooks and stream verifiers, where the work genuinely is spiky and stateless.
02Why not Spanner? It removes your sharding entirely.
It removes sharding I already have working, and it charges for global external consistency inside a single legal region. Part 7 made regions a legal boundary, so cross-region consistency is not available to use. Paying for a property you cannot exercise is the most common cloud mistake.
03Aurora is proprietary. Is that not lock-in?
It speaks the Postgres wire protocol, so the application is portable. The deeper lock-in is operational, per Part 0: runbooks, IAM, dashboards and what the team knows at 3am. I accept Aurora because the failover duration is the RTO and it is consistently the fastest, and because the escape path to plain Postgres exists.
04Pub/Sub has no partitions. Is that disqualifying?
No. Ordering keys give per-account ordering, which is all Part 4 required. The real tradeoff is retention and replay precision: 31 days maximum and seek-by-timestamp rather than by offset. That is acceptable if the rebuild-from-warehouse path is designed and drilled, and dangerous if discovered during an incident.
05Your RTO is 60 seconds. Prove it.
I cannot from documentation, only from a drill. Part 1 decomposed a real failover into detection, promotion, endpoint propagation, pool recovery and backlog drain, which totals 30 seconds to 4 minutes against a documented 30 seconds. The RTO is what the monthly drill measured, and that number is published.
06What is the single biggest cost surprise?
Observability, at around 9% of the bill, above the warehouse. It is driven almost entirely by metric cardinality, which engineers increase by adding a dimension without noticing the multiplication. Second is cross-AZ and egress traffic, which a chatty service architecture pays continuously and which never appears in a capacity plan.
07Multi-cloud for resilience?
No. Running the money path across two providers roughly doubles the operational surface and divides the team's expertise, which usually reduces reliability rather than improving it. Core on AWS with the warehouse on BigQuery is a split I would defend, because the data crosses once in batch with a knowable egress cost.
08Can you run this in Nigeria?
Neither provider has a Nigerian region. For CBN-regulated data that is a hard constraint, and the honest options are a local deployment for the regulated workload with cloud for what is permitted, or an arrangement the regulator has approved. I would not design as though a Lagos region exists.
09You are on a PCI-certified cloud. Are you PCI compliant?
No. The provider certification covers security of the cloud; the configuration, the application and the data are mine. Part 8 minimised scope by keeping the ledger free of card data, and on the cloud that becomes a separate account with its own VPC, keys and HSM. The certification is evidence for an auditor, not a conclusion.
10How do you stop a cost spike becoming a surprise?
Treat it as an SRE concern. Anomaly alerts on daily spend per service against the same day last week, which is the Part 13 business-metric pattern applied to cost. A cost spike is usually an incident: a retry loop, a scan without a partition filter, a stuck job.
11Kubernetes or not?
Only with the team to run it. For nine services and no platform team, ECS Fargate or Cloud Run delivers the same outcome with a fraction of the operational surface. Choosing Kubernetes because it is standard, without that team, means acquiring a second full-time system to operate.
12How do you migrate without downtime?
Shard by shard, with a 20 to 30 second write pause per batch, queued rather than rejected. The shard map from Part 3 is already the routing layer. I would not dual-write, because dual-writing a ledger is the Part 4 dual-write problem applied to the most critical table in the company.
13How do you know the migration was correct?
The same way I know the ledger is correct. Part 10 reconciliation, pointed at the legacy system as a counterparty: row equivalence, the invariant per shard, balance equivalence, trial balance equivalence, external reconciliation, and serving equivalence. Six gates per batch, all of which must pass.
14What would you refuse to move?
The HSM if the regulator requires physical key custody, scheme connectivity if it is contractually location-bound, anything the regulator has not approved, and data the residency rules forbid. A short, reasoned exception list is what a credible migration plan looks like, and it is what the outsourcing notification needs anyway.
15If you had to choose one platform today, which?
AWS, and not because it is technically superior. Eleven of fourteen layers are equivalent. It wins on ecosystem breadth, vendor integrations, auditor familiarity and hiring, which for a bank with six payment integrations matter more than they should. GCP has the better warehouse and the cleaner compliance controls, and if analytics were the product rather than supporting it, that would flip the decision.
the answer shape, one last time
Direct answer, then the reason, then the cost. And where the honest answer is "it depends on something other than technology", say what it depends on. The interviewer is testing whether you can make a decision under ambiguity and defend it, not whether you memorised a catalogue that will be out of date next year.