The frame: managed is a tradeoff, not a default
Before any service names, a way of deciding. Most cloud architecture discussions go wrong by starting from the catalogue rather than from the requirement, and by treating "managed" as self-evidently correct. This part sets up the questions the rest of the module asks of every component, and establishes the constraints that a regulated bank brings to the table before cost is even discussed.
What you are actually buying from a cloud
Not servers. Servers are cheap and were always available. What a cloud sells is more specific, and naming it precisely is what makes the tradeoffs visible.
- Elasticity. Capacity in minutes rather than weeks, and the ability to give it back. Worth a great deal if load is variable, and worth little if it is flat.
- Operational labour. Somebody else patches, backs up, fails over and monitors the thing. This is usually the largest real saving and it is rarely the one quoted.
- Higher-level primitives. A managed queue, a warehouse, a key store. Things you would otherwise build and then maintain forever.
- Global footprint. A region in Frankfurt without a lease in Frankfurt, which for Part 7's residency constraints is decisive.
- A compliance floor. Physical security, and certifications you inherit rather than earn. Considerable, and frequently misunderstood, per chapter 55.
what you are not buying:
cheaper compute per unit it is usually dearer at steady state
someone else's correctness your invariant is still yours
freedom from capacity planning quotas and limits are the new form
compliance itself only the floor beneath yours
for a bank at steady 5,600 writes/s, the elasticity argument is
weak and the operational-labour argument is overwhelming.The four questions to ask of any managed service
The rest of this module applies these four to every component. They are worth memorising, because they work on services that did not exist when you learned them.
- Not what the marketing page implies. The SLA, the durability figure, the consistency model.
- For our ledger the question is specific: does it offer multi-row atomicity and a conditional write? If not, it cannot be the ledger, whatever else it offers.
- And what happens when the SLA is missed. A service credit is not a recovery, and for a bank it is close to worthless.
- Every managed service has quotas: item size, partition throughput, concurrent executions, payload size.
- The one that bites is rarely the one in the headline. DynamoDB's 400 KB item limit matters more than its throughput claim, and Lambda's 15-minute ceiling matters more than its concurrency.
- Many are soft and raised on request. Knowing which are hard is the actual skill.
- Failover behaviour, and measured rather than documented duration.
- Whether failure is visible: a service that silently degrades is worse than one that errors.
- What you can do during an incident, which for a managed service is often very little, and that is the real cost.
- Pricing pages are written for the common case. Model your own workload.
- Watch for the dimension that is cheap at demo scale and dominant at production scale, usually data transfer, requests, or storage-months.
- And the cost of leaving, which is egress plus the migration in Part 7.
Lock-in, honestly assessed
The most frequently invoked and least carefully examined consideration in cloud architecture. Lock-in is real and it is mostly not where people look for it.
| Layer | Switching cost | Why |
|---|---|---|
| Compute | Low | A container runs anywhere. This is the layer people worry about and it is the cheapest to move |
| Open-protocol services | Low to moderate | Postgres, Redis and Kafka wire protocols are portable, even when the managed implementation is not |
| Object storage | Moderate | The API is similar everywhere; the egress bill is the real barrier |
| Proprietary data services | High | DynamoDB, Spanner and BigQuery have no equivalent elsewhere. Moving means a redesign |
| IAM and the security model | High | Every policy, role and boundary is platform-specific and there are thousands |
| Operational muscle memory | Highest | Runbooks, dashboards, alerts, and a team that knows where to look at 3am |
- Use open protocols where the cost is low. Postgres over a proprietary database when both fit, because the option is nearly free.
- Accept deep lock-in where the service is genuinely differentiated. BigQuery and Spanner do things their competitors do not, and refusing them on principle costs you the capability.
- Keep the ledger portable specifically, because it is the thing you can least afford to be unable to move, and Postgres makes that cheap.
- Do not build an abstraction layer over two clouds you do not use. It costs real complexity to buy an option you will probably never exercise, and it usually means using the intersection of both platforms, which is worse than either.
The regulated-bank constraints that narrow the menu
A bank does not get the same menu as a startup. These constraints apply before cost or preference, and several of them eliminate options entirely.
| Constraint | Effect on the architecture |
|---|---|
| Data residency | Region availability becomes a hard filter, per Part 7. No Nigerian region on either platform is a live constraint for CBN data |
| Regulator notification | Material outsourcing often requires notifying or obtaining approval. Adding a cloud service can be a filing |
| Exit plans | Several regulators require a documented, tested plan for leaving the provider. Lock-in becomes a compliance artefact |
| Audit rights | The regulator may require the right to audit the provider, which the contract must grant |
| Concentration risk | Supervisors increasingly care that a whole sector runs on one provider, which pushes toward multi-region or multi-cloud for critical services |
| Encryption and key control | Often customer-managed keys at minimum, sometimes an HSM you control |
| PCI DSS | Narrows compute and network options inside the cardholder data environment, per Part 8 |
the ordering that actually happens:
1. can we legally put this data here? → residency, Part 7
2. does it meet the guarantee we need? → chapter 2, question one
3. can we operate it? → chapter 2, question three
4. what does it cost? → chapter 6
most architecture discussions start at 4 and work backwards,
which is how you spend two weeks costing an option that was
never legally available.Landing zones, accounts, and blast radius
The structural decision made before any service is provisioned, and the one that is most painful to change later.
aws-nuke for teardown of sandboxes.- Production is its own account or project, with a different access path and different people.
- Each region is separate, because Part 7 said the boundary is legal. A shared account across regions makes a residency breach a configuration error away.
- The PCI environment is separate, so its audit scope does not swallow everything else, per Part 8 chapter 97.
- Security tooling and audit logs live in an account nobody else can write to, so a compromise of production cannot erase its own traces.
- Sandboxes are disposable and have hard spending caps, because they are where the surprise bills come from.
Reading a cloud bill: the shapes that surprise you
Cost is a non-functional requirement in cloud architecture the way latency is in system design, and the surprises are structural rather than arithmetic.
- Egress. Data leaving the provider, and often between regions or even zones. Cross-AZ traffic inside a cluster is a real line item, and a chatty microservice architecture pays it continuously.
- Per-request pricing. Cheap at demo scale, dominant at 10M transactions a day. A service costing a fraction of a cent per thousand requests becomes a serious number at 5,600 per second.
- Provisioned versus consumed. Provisioned capacity is cheaper per unit and charges while idle. A ledger at steady load wants provisioned; a batch job wants consumed.
- Storage-months, plus the copies. Part 15 counted 400 TB across every copy. Backups, replicas, warehouse and archive are each billed.
- Observability. Logs and custom metrics at 10M transactions a day are genuinely expensive, which is why Part 13's sampling and EMF matter commercially and not only technically.
the shape for our bank, roughly, per month:
ledger cluster (provisioned, multi-AZ) ~38%
event backbone ~14%
compute ~12%
serving layer ~11%
observability ~9%
warehouse and lake ~8%
network and egress ~6%
everything else ~2%
observability costing more than the warehouse surprises everyone,
and it is why Part 13 sampled traces rather than keeping them all.Regions, zones, and what Part 7 residency forces
availability zone one or more datacentres, independent power
and cooling, < 2 ms apart
region a group of zones, tens to hundreds of ms
from each other
synchronous replication across zones: yes, ~1-2 ms
synchronous replication across regions: no, not for a money path
so Part 1's RPO of zero is achieved WITHIN a region, across
zones, and cross-region is asynchronous with a real RPO.- Each regulated market needs a complete deployment in a permitted region, not a cache of somewhere else.
- Cross-region failover is unavailable where residency forbids it, so resilience is built across zones inside the region.
- That is more expensive than a global active-active design, and the extra cost is a legal consequence rather than an engineering choice.
- Neither AWS nor GCP has a Nigerian region today, which is a genuine constraint for CBN data residency and is usually resolved through a local colocation or an approved arrangement. Say this rather than pretending the region exists.
- Region selection for the EU and UK is straightforward on both platforms and is mostly a latency and cost decision.
How to answer “AWS or GCP” in an interview
The question is almost never about which is better. It is about whether you can reason about a decision that has no technically correct answer.
- Name what actually decides it, which is usually not technology: existing expertise, existing commitments, regional availability, and any regulatory approval already in place.
- Identify the components where the platforms genuinely differ, and there are only a handful. Everything else is equivalent.
- State a default and the conditions that would change it.
- Refuse the tribal version of the question. "AWS is better" is not an answer and neither is the reverse.
| Component | Genuinely different? | Which, and why |
|---|---|---|
| Relational database | Somewhat | Aurora and AlloyDB are comparable. Spanner is unique and expensive |
| Event backbone | Yes | MSK is Kafka. Pub/Sub is not partitioned, which matters for ordering |
| Warehouse | Yes | BigQuery is the stronger product, and the serverless model is genuinely different |
| Wide-column serving | Somewhat | Bigtable and DynamoDB both work; the data models differ meaningfully |
| Containers | Slightly | Cloud Run is the nicest of its kind. EKS and GKE are both Kubernetes |
| Functions | Slightly | Lambda has the richer ecosystem and the longer track record |
| Everything else | No | Object storage, queues, secrets, keys and IAM are equivalent in capability |
- "Multi-cloud for resilience." Running critical paths across two providers roughly doubles the operational surface and usually reduces reliability. Say it only if you can defend the cost.
- "It doesn't matter." It matters; it is just not decided by technology alone, and saying so precisely is the better version.
- A feature comparison. Feature lists date within a year and reciting one signals you learned the catalogue rather than the tradeoffs.