Part 5 · 9 chapters · ~50 min

Identity, secrets, and the compliance surface

Part 14 designed service identity and Part 8 minimised PCI scope; this part implements both on each platform. The recurring theme is that the cloud provides mechanisms and not decisions: it will give you cryptographic workload identity, immutable audit logs and regional key isolation, and it will not tell you to scope a role by account kind, keep data access logging on, or make a residency violation physically impossible.

48

Workload identity: IAM roles and service accounts

Part 14 chapter 169 established the principle: cryptographic service identity, short-lived, scoped by what it may touch rather than by which method it may call. Both clouds implement it, differently.

AWS
GCP
IAM roles
assumed by workloads. On ECS, a task role. On EKS, IRSA binds a Kubernetes service account to an IAM role via OIDC. On Lambda, an execution role.

Credentials are rotated automatically and never stored.
Service accounts attached to workloads. On GKE, Workload Identity binds a Kubernetes SA to a Google SA. On Cloud Run, a service identity.

Same property: short-lived tokens, no stored keys.
the rules that apply on both
  1. One identity per service, never a shared one. A shared identity makes the audit log useless and the blast radius total.
  2. No long-lived keys. A static access key in an environment variable is the single most common cloud breach vector.
  3. Scope by resource, not only by action. Part 14 made this point: restricting methods is weak because every product service needs Post. Restricting which account kinds it may touch is what bounds a compromise.
  4. Deny by default, with permissions added deliberately and reviewed.
the mapping to Part 14
Our design wanted spiffe://bank/ns/prod/sa/loans scoped to loan and wallet account kinds. On AWS that is an IAM role whose policy conditions on a resource tag or a path prefix; on GCP a service account with a fine-grained IAM binding. The cloud provides the identity; the account-kind scoping is still enforced by the posting API, because no cloud IAM understands what a ledger account kind is.
49

The Part 14 mTLS design, on each platform

AWS
GCP
App Mesh
(Envoy-based) or a self-run Istio or Linkerd on EKS. ACM Private CA issues short-lived certificates.

Many teams instead use IAM plus VPC isolation and skip the mesh, which is a defensible simplification.
Anthos Service Mesh on GKE, Istio-based, with automatic mTLS between workloads and certificate rotation handled for you.

Cloud Run service-to-service uses IAM-authenticated invocation rather than mTLS.
the simplification worth considering
A service mesh is a substantial operational commitment. For nine services in a private network with IAM-authenticated calls, the marginal security gain over IAM plus network isolation is smaller than the operational cost for most teams. The honest answer in an interview is that mTLS everywhere is the right end state and that IAM plus VPC isolation is an acceptable intermediate, and naming that tradeoff is better than asserting a mesh is mandatory.
50

Secrets: the lifecycle, not the store

AWS
GCP
Secrets Manager
for rotatable secrets with built-in rotation Lambdas, Parameter Store for configuration and cheaper static values.

Secrets Manager charges per secret per month plus API calls; Parameter Store standard tier is free.
Secret Manager, versioned, with IAM per secret and per version. Rotation is yours to schedule via Cloud Scheduler and a function.

Simpler surface, slightly less built-in rotation.
the lifecycle, which is the actual problem
  1. Creation: generated, never chosen by a human, never in a ticket.
  2. Distribution: fetched at runtime by workload identity, never baked into an image or a repository.
  3. Rotation: scheduled, automated, and tested. A rotation that has never run will fail the first time it does.
  4. Revocation: possible within minutes, which requires knowing every consumer of every secret.
  5. Audit: every access logged, because unusual secret access is an early compromise signal.
the rotation point
Most teams store secrets well and rotate them never. A secret that has not rotated in three years has been seen by everyone who has ever worked there. Rotation is the control that actually limits exposure, and it is the one that requires the application to handle a credential changing underneath it, which is an application design property rather than a secret-store feature.
51

Key management, HSMs, and PCI scope

AWS
GCP
KMS
for envelope encryption, customer-managed keys, and per-key IAM. CloudHSM for FIPS 140-2 Level 3 single-tenant hardware when the regulator requires it.

KMS keys are regional, which interacts directly with Part 7 residency.
Cloud KMS, with a Cloud HSM protection level for hardware-backed keys, and External Key Manager for keys you hold outside Google entirely.

EKM is the strongest answer to a regulator who wants keys outside the provider.
RequirementAWSGCP
Customer-managed keysKMS CMKCloud KMS
Hardware-backedCloudHSMCloud HSM protection level
Keys outside the providerCloudHSM, partiallyExternal Key Manager
PIN and cryptogram operationsCloudHSM or a payment HSMPayment HSM or third party
Per-subject keys for crypto-shreddingKMS, many keysCloud KMS, many keys
the PCI point from Part 8, on the cloud
Part 8 chapter 97 minimised PCI scope by keeping card data in an isolated vault. On the cloud that becomes a separate account or project, its own VPC, its own KMS keys and its own HSM, with no network path from the main estate except the authorisation API. The compliance boundary and the account boundary should be the same boundary, because an auditor can then reason about it as one thing.
52

Network isolation: VPCs, endpoints, and egress

the shape for a bank
  1. Private subnets for everything stateful. No database, cache or broker has a public address, ever.
  2. Private endpoints for managed services. VPC endpoints on AWS, Private Service Connect on GCP, so traffic to S3 or BigQuery never traverses the public internet.
  3. Controlled egress. A NAT with a stable address, an allowlist of destinations, and logging. Part 14 noted card schemes allowlist source IPs, so this is functional as well as protective.
  4. Separate VPCs per environment and per region, matching the Part 7 legal boundary and the Part 0 chapter 5 account structure.
  5. No peering between production regions where residency forbids data movement, so the network makes the violation impossible rather than merely forbidden.
worked numbers
the private-endpoint argument is also a cost argument:

  traffic to S3 via NAT         → NAT processing + data charges
  traffic to S3 via VPC endpoint → no NAT, no data charge

at 400 TB of lake traffic, that is not a rounding error.
security and cost point the same way, which is unusual and worth using.
the residency enforcement point
Part 7 said to make a violation unrepresentable rather than forbidden. The network layer is where that becomes physical: if the EU VPC has no route to the Nigerian VPC, no misconfigured service can move data between them. Policy documents get violated; missing network routes do not.
53

Audit logging that satisfies an auditor

AWS
GCP
CloudTrail
for control-plane API calls, CloudTrail data events for S3 and DynamoDB object access, VPC Flow Logs for network.

Delivered to a separate account with S3 Object Lock, so production cannot erase its own traces.
Cloud Audit Logs: admin activity always on, data access opt-in and billable.

Exported to a bucket in a separate project with retention locks, same reasoning.
what an auditor actually wants
  1. Immutable. Write-once storage, and the account that writes cannot delete.
  2. Complete. Including read access to customer data, which is the expensive part and the part teams disable to save money.
  3. Attributable. To a human, not to a shared role. This is why Part 14 wanted per-service identities and why human access should be federated with individual identity.
  4. Retained for the statutory period, which Part 13 put at 7 years for money-related audit logs.
  5. Searchable within a reasonable time, because "we have the logs" and "we can answer your question" are different claims.
the one to not economise on
Data access logging is billable and is the log an auditor asks for. Control-plane logging tells you who changed infrastructure; data access logging tells you who read customer records, which is the Part 9 insider-threat question. Turning it off to save money is a decision to be unable to answer the most important question after an incident.
54

Data residency controls, per Part 7

AWS
GCP
Service Control Policies
at the OU level denying any region outside the permitted set. This is a guardrail an account cannot escape, including by an administrator.

Plus S3 bucket policies conditioning on region, and KMS keys that exist only in-region.
Organization Policy constraint gcp.resourceLocations, applied at the folder level, which prevents creating resources outside the allowed locations.

Arguably the cleaner mechanism of the two.
worked numbers
defence in depth for residency:

  1. organisation policy        cannot create the resource
  2. no network route          cannot reach the other region
  3. regional KMS keys       data is unreadable elsewhere
  4. schema with no field    Part 7: unrepresentable
  5. audit alert                 if any of the above is changed

five independent layers, because the consequence of failure
is a regulatory breach rather than an outage.
the Nigerian reality, again
Part 0 chapter 7 said it and it applies here: neither platform has a Nigerian region. The residency controls above are what you apply for EU and UK data. For CBN-regulated data the architecture is a local deployment with cloud used for what is permitted, and the honest answer in an interview is to say so rather than to apply a control to a region that does not exist.
55

Compliance programmes, and what they do not cover

worked numbers
the shared responsibility model, stated precisely:

  the provider: security OF the cloud
    physical, hypervisor, managed service internals

  you: security IN the cloud
    IAM, network, encryption config, application, data

"we are on AWS so we are PCI compliant" is false.
AWS being PCI-certified means the infrastructure can be part of
a compliant environment. your configuration decides whether it is.
what you inherit versus what you must do
  1. Inherited: datacentre physical security, hardware disposal, hypervisor isolation, and the provider’s own certifications as evidence for your auditor.
  2. Yours: every IAM policy, every network rule, every encryption setting, key management, logging, and the application itself.
  3. Yours and often missed: proving the configuration is correct continuously, which is what Config, Security Hub or Security Command Centre are for.
  4. Yours entirely: the Part 10 reconciliation, the Part 13 correctness signal, and everything that makes the ledger right. No cloud certification says anything about whether your money balances.
56

The decision, with numbers

AWS
GCP
Secrets Manager
IAM roles with resource-scoped policies, with scheduled rotation, KMS CMKs per region with CloudHSM in the PCI account, private subnets with VPC endpoints, CloudTrail including data events to a locked separate account, SCPs pinning regions.

Roughly $1,800 to $2,800 per month, dominated by CloudHSM and data-event logging.
Service accounts with Workload Identity, Secret Manager, Cloud KMS with HSM protection or EKM, Private Service Connect, Audit Logs with data access enabled to a locked project, Organization Policy pinning locations.

Roughly $1,500 to $2,500 per month, similar shape.
the answer
"Both platforms give you everything needed, and the difference is ergonomic rather than capability. GCP's Organization Policy for resource locations is the cleaner residency control and External Key Manager is the stronger answer to a regulator who wants keys outside the provider. AWS has the deeper ecosystem and more auditors have seen it before, which matters more than it should. The thing I would not economise on is data access logging, because that is the log that answers the insider-threat question from Part 9, and it is the one teams disable to save money."