Part 0 · 8 chapters · ~50 min

The frame: managed is a tradeoff, not a default

Before any service names, a way of deciding. Most cloud architecture discussions go wrong by starting from the catalogue rather than from the requirement, and by treating "managed" as self-evidently correct. This part sets up the questions the rest of the module asks of every component, and establishes the constraints that a regulated bank brings to the table before cost is even discussed.

1

What you are actually buying from a cloud

Not servers. Servers are cheap and were always available. What a cloud sells is more specific, and naming it precisely is what makes the tradeoffs visible.

the five things you are paying for
  1. Elasticity. Capacity in minutes rather than weeks, and the ability to give it back. Worth a great deal if load is variable, and worth little if it is flat.
  2. Operational labour. Somebody else patches, backs up, fails over and monitors the thing. This is usually the largest real saving and it is rarely the one quoted.
  3. Higher-level primitives. A managed queue, a warehouse, a key store. Things you would otherwise build and then maintain forever.
  4. Global footprint. A region in Frankfurt without a lease in Frankfurt, which for Part 7's residency constraints is decisive.
  5. A compliance floor. Physical security, and certifications you inherit rather than earn. Considerable, and frequently misunderstood, per chapter 55.
worked numbers
          what you are not buying:


            cheaper compute per unit      it is usually dearer at steady state

            someone else's correctness     your invariant is still yours

            freedom from capacity planning  quotas and limits are the new form

            compliance itself                   only the floor beneath yours


for a bank at steady 5,600 writes/s, the elasticity argument is

weak and the operational-labour argument is overwhelming.
the honest framing for an interview
"We are not on the cloud because it is cheaper per CPU hour, because at our steady-state load it is not. We are on it because a team of six cannot run multi-region Postgres, a Kafka cluster, an HSM and a warehouse and also build product. The saving is headcount and attention, and the cost is a bill that grows with success and a set of limits we do not control."
2

The four questions to ask of any managed service

The rest of this module applies these four to every component. They are worth memorising, because they work on services that did not exist when you learned them.

one: what does it guarantee, in writing
  1. Not what the marketing page implies. The SLA, the durability figure, the consistency model.
  2. For our ledger the question is specific: does it offer multi-row atomicity and a conditional write? If not, it cannot be the ledger, whatever else it offers.
  3. And what happens when the SLA is missed. A service credit is not a recovery, and for a bank it is close to worthless.
two: what are the limits, and which one will you hit first
  1. Every managed service has quotas: item size, partition throughput, concurrent executions, payload size.
  2. The one that bites is rarely the one in the headline. DynamoDB's 400 KB item limit matters more than its throughput claim, and Lambda's 15-minute ceiling matters more than its concurrency.
  3. Many are soft and raised on request. Knowing which are hard is the actual skill.
three: how does it fail, and what do you do about it
  1. Failover behaviour, and measured rather than documented duration.
  2. Whether failure is visible: a service that silently degrades is worse than one that errors.
  3. What you can do during an incident, which for a managed service is often very little, and that is the real cost.
four: what does it cost at your shape, not at the example's shape
  1. Pricing pages are written for the common case. Model your own workload.
  2. Watch for the dimension that is cheap at demo scale and dominant at production scale, usually data transfer, requests, or storage-months.
  3. And the cost of leaving, which is egress plus the migration in Part 7.
why this framing beats memorising services
Cloud catalogues change every year and these four questions do not. An interviewer asking about a service you have not used is really asking whether you can evaluate one, and answering "I have not used it; here is what I would need to know before I did, and here is which of those would decide it" is a stronger answer than a half-remembered feature list.
3

Lock-in, honestly assessed

The most frequently invoked and least carefully examined consideration in cloud architecture. Lock-in is real and it is mostly not where people look for it.

LayerSwitching costWhy
ComputeLowA container runs anywhere. This is the layer people worry about and it is the cheapest to move
Open-protocol servicesLow to moderatePostgres, Redis and Kafka wire protocols are portable, even when the managed implementation is not
Object storageModerateThe API is similar everywhere; the egress bill is the real barrier
Proprietary data servicesHighDynamoDB, Spanner and BigQuery have no equivalent elsewhere. Moving means a redesign
IAM and the security modelHighEvery policy, role and boundary is platform-specific and there are thousands
Operational muscle memoryHighestRunbooks, dashboards, alerts, and a team that knows where to look at 3am
the point that reframes the argument
The deepest lock-in is not technical, it is operational. Rewriting a DynamoDB access layer is a quarter of work. Rebuilding an organisation's instincts, its runbooks, its on-call knowledge and its security posture on a different platform is years, and it is the reason large migrations fail rather than the API surface.
a proportionate response
  1. Use open protocols where the cost is low. Postgres over a proprietary database when both fit, because the option is nearly free.
  2. Accept deep lock-in where the service is genuinely differentiated. BigQuery and Spanner do things their competitors do not, and refusing them on principle costs you the capability.
  3. Keep the ledger portable specifically, because it is the thing you can least afford to be unable to move, and Postgres makes that cheap.
  4. Do not build an abstraction layer over two clouds you do not use. It costs real complexity to buy an option you will probably never exercise, and it usually means using the intersection of both platforms, which is worse than either.
4

The regulated-bank constraints that narrow the menu

A bank does not get the same menu as a startup. These constraints apply before cost or preference, and several of them eliminate options entirely.

ConstraintEffect on the architecture
Data residencyRegion availability becomes a hard filter, per Part 7. No Nigerian region on either platform is a live constraint for CBN data
Regulator notificationMaterial outsourcing often requires notifying or obtaining approval. Adding a cloud service can be a filing
Exit plansSeveral regulators require a documented, tested plan for leaving the provider. Lock-in becomes a compliance artefact
Audit rightsThe regulator may require the right to audit the provider, which the contract must grant
Concentration riskSupervisors increasingly care that a whole sector runs on one provider, which pushes toward multi-region or multi-cloud for critical services
Encryption and key controlOften customer-managed keys at minimum, sometimes an HSM you control
PCI DSSNarrows compute and network options inside the cardholder data environment, per Part 8
the constraint most engineers do not know exists
An exit plan is a regulatory requirement in several jurisdictions, and a tested one in some. That converts the lock-in discussion from an architectural preference into a documented obligation with an owner and a review date. If you cannot describe how you would leave, you may not be permitted to be there, and mentioning this in an interview signals genuine familiarity with regulated environments.
worked numbers
          the ordering that actually happens:


            1. can we legally put this data here?    → residency, Part 7

            2. does it meet the guarantee we need?  → chapter 2, question one

            3. can we operate it?                        → chapter 2, question three

            4. what does it cost?                         → chapter 6


most architecture discussions start at 4 and work backwards,

which is how you spend two weeks costing an option that was

never legally available.
5

Landing zones, accounts, and blast radius

The structural decision made before any service is provisioned, and the one that is most painful to change later.

AWS
GCP
Accounts
under Organizations, grouped into OUs. Service Control Policies set guardrails an account cannot escape. The account is the primary isolation boundary and is free, so use many.
Projects under folders in a Resource Hierarchy. Organization Policies act as guardrails. The project is the primary boundary, also free, same advice.
Control Tower or Landing Zone Accelerator to scaffold it. aws-nuke for teardown of sandboxes.
Cloud Foundation Toolkit or the Fabric modules. Projects are genuinely easy to delete, which makes sandboxes cheap.
the separation that matters for a bank
  1. Production is its own account or project, with a different access path and different people.
  2. Each region is separate, because Part 7 said the boundary is legal. A shared account across regions makes a residency breach a configuration error away.
  3. The PCI environment is separate, so its audit scope does not swallow everything else, per Part 8 chapter 97.
  4. Security tooling and audit logs live in an account nobody else can write to, so a compromise of production cannot erase its own traces.
  5. Sandboxes are disposable and have hard spending caps, because they are where the surprise bills come from.
the reason to over-separate early
Accounts and projects are free; splitting them later is not. Moving a running workload between accounts means new resource identifiers, new IAM, new network paths and usually downtime. Creating the boundary before there is anything in it costs nothing, and the most common regret in cloud foundations is having used too few.
6

Reading a cloud bill: the shapes that surprise you

Cost is a non-functional requirement in cloud architecture the way latency is in system design, and the surprises are structural rather than arithmetic.

the five shapes that catch people
  1. Egress. Data leaving the provider, and often between regions or even zones. Cross-AZ traffic inside a cluster is a real line item, and a chatty microservice architecture pays it continuously.
  2. Per-request pricing. Cheap at demo scale, dominant at 10M transactions a day. A service costing a fraction of a cent per thousand requests becomes a serious number at 5,600 per second.
  3. Provisioned versus consumed. Provisioned capacity is cheaper per unit and charges while idle. A ledger at steady load wants provisioned; a batch job wants consumed.
  4. Storage-months, plus the copies. Part 15 counted 400 TB across every copy. Backups, replicas, warehouse and archive are each billed.
  5. Observability. Logs and custom metrics at 10M transactions a day are genuinely expensive, which is why Part 13's sampling and EMF matter commercially and not only technically.
worked numbers
          the shape for our bank, roughly, per month:


            ledger cluster (provisioned, multi-AZ)        ~38%

            event backbone                                    ~14%

            compute                                         ~12%

            serving layer                                    ~11%

            observability                                 ~9%

            warehouse and lake                            ~8%

            network and egress                         ~6%

            everything else                                ~2%


observability costing more than the warehouse surprises everyone,

and it is why Part 13 sampled traces rather than keeping them all.
the sentence to have ready
"I would treat cost as an SLO with an owner, reported per service and per transaction, exactly like latency. The two line items I would watch hardest are observability and egress, because both scale with traffic, neither appears in a capacity plan, and both are usually discovered in a quarterly review rather than designed for."
7

Regions, zones, and what Part 7 residency forces

worked numbers
availability zone  one or more datacentres, independent power

                                and cooling, < 2 ms apart

region             a group of zones, tens to hundreds of ms

                                from each other


          synchronous replication across zones:  yes, ~1-2 ms

          synchronous replication across regions: no, not for a money path


so Part 1's RPO of zero is achieved WITHIN a region, across

zones, and cross-region is asynchronous with a real RPO.
what Part 7's residency rule does to this
  1. Each regulated market needs a complete deployment in a permitted region, not a cache of somewhere else.
  2. Cross-region failover is unavailable where residency forbids it, so resilience is built across zones inside the region.
  3. That is more expensive than a global active-active design, and the extra cost is a legal consequence rather than an engineering choice.
  4. Neither AWS nor GCP has a Nigerian region today, which is a genuine constraint for CBN data residency and is usually resolved through a local colocation or an approved arrangement. Say this rather than pretending the region exists.
  5. Region selection for the EU and UK is straightforward on both platforms and is mostly a latency and cost decision.
the answer that shows you have thought about it
"Residency is a hard filter applied before any other consideration. For the EU and UK both platforms have suitable regions and the choice is latency and cost. For Nigerian data neither has an in-country region, so the honest options are a local datacentre for the regulated workload with the cloud for everything else, or an arrangement the regulator has approved. I would not design as though a Lagos region exists."
8

How to answer “AWS or GCP” in an interview

The question is almost never about which is better. It is about whether you can reason about a decision that has no technically correct answer.

the structure that works
  1. Name what actually decides it, which is usually not technology: existing expertise, existing commitments, regional availability, and any regulatory approval already in place.
  2. Identify the components where the platforms genuinely differ, and there are only a handful. Everything else is equivalent.
  3. State a default and the conditions that would change it.
  4. Refuse the tribal version of the question. "AWS is better" is not an answer and neither is the reverse.
ComponentGenuinely different?Which, and why
Relational databaseSomewhatAurora and AlloyDB are comparable. Spanner is unique and expensive
Event backboneYesMSK is Kafka. Pub/Sub is not partitioned, which matters for ordering
WarehouseYesBigQuery is the stronger product, and the serverless model is genuinely different
Wide-column servingSomewhatBigtable and DynamoDB both work; the data models differ meaningfully
ContainersSlightlyCloud Run is the nicest of its kind. EKS and GKE are both Kubernetes
FunctionsSlightlyLambda has the richer ecosystem and the longer track record
Everything elseNoObject storage, queues, secrets, keys and IAM are equivalent in capability
the answer, in full
"For this system either platform works, and I would pick based on what the team already runs, because operational familiarity is worth more than any feature difference here. If I had a free choice: AWS, because the ledger wants Postgres with strong operational tooling, the event backbone wants real Kafka partitions for the per-account ordering from Part 4, and the breadth of the ecosystem matters for the six payment integrations. The thing that would change my mind is analytics: if the warehouse were central to the product rather than supporting it, BigQuery is strong enough to lead the decision, and I would then consider running analytics on GCP against a core on AWS, accepting the egress cost."
what not to say
  1. "Multi-cloud for resilience." Running critical paths across two providers roughly doubles the operational surface and usually reduces reliability. Say it only if you can defend the cost.
  2. "It doesn't matter." It matters; it is just not decided by technology alone, and saying so precisely is the better version.
  3. A feature comparison. Feature lists date within a year and reciting one signals you learned the catalogue rather than the tradeoffs.