Part 3 · 10 chapters · ~55 min

Compute: where the posting path runs

The layer with the most tribal opinion and the clearest technical answer. The posting path is steady, connection-heavy and latency-sensitive, which is precisely the profile function platforms handle worst, and webhook receivers are the opposite. This part works through it per workload rather than per company, and ends on the autoscaling configuration that turns a slow dependency into an outage.

29

The posting path is not a good Lambda

Starting with the conclusion, because it is the one people argue with and the reasons are concrete rather than aesthetic.

five reasons, in order of weight
  1. Connection churn. Part 1 chapter 16: a thousand concurrent functions means a thousand database connections, short-lived, against a database comfortable with a few hundred.
  2. Cold starts on the money path. A VPC-attached function adds latency at exactly the moment a customer is waiting, and provisioned concurrency to avoid it is just a container with extra steps.
  3. The work is not bursty. Part 3 derived a steady 464 transfers per second at peak. Functions price elasticity, and we have none to exploit.
  4. Warm caches matter. The shard map, tariff tables and account-type policy are all read constantly and cached per process. A function discards that cache.
  5. Long-lived connections to the scheme. Part 14 chapter 161 needs persistent ISO 8583 sockets, which a function fundamentally cannot hold.
where Lambda is genuinely right here
Event-driven, spiky, short work with no connection state. Webhook receivers, the Part 11 invariant verifier consuming a stream, image processing for KYC documents, scheduled cleanup. The rule is not "functions are bad"; it is that the posting path has exactly the profile functions are worst at.
30

Lambda: cold starts, concurrency, and the VPC tax

worked numbers
cold start, decomposed:
  download and unpack the package        50 to 300 ms
  start the runtime                             50 to 200 ms
  your init code, pools, config             yours to control
  VPC ENI attachment (historically)   now ~sub-second, was ~10 s

concurrency model: one request per environment
  1,000 concurrent requests = 1,000 environments
  = 1,000 database connections unless proxied
KnobWhat it doesWhen to use it
Provisioned concurrencyKeeps N environments warmPredictable latency. At which point you are paying for idle containers
Reserved concurrencyCaps how many can runProtecting a downstream database from being overwhelmed
SnapStartSnapshots an initialised runtimeJVM cold starts specifically
Memory settingAlso scales CPUOften cheaper to raise: faster execution can cost less
the reserved-concurrency point worth making
Reserved concurrency is a load-shedding mechanism. Capping a function at 50 concurrent executions protects the database behind it, and Part 13 called that graceful degradation. It is one of the few places where a serverless limit is a feature rather than a constraint.
31

Cloud Functions and Cloud Run

AWS
GCP
a real architectural boundary
Lambda plus API Gateway, or a Function URL. Containers on ECS or EKS as the alternative.

The gap between "function" and "container" is on AWS.
Cloud Run blurs it: a container image, scaled to zero or held warm, one URL, concurrency per instance configurable.

Cloud Functions gen 2 is Cloud Run underneath.
LambdaCloud Run
UnitFunction package or containerContainer, always
Concurrency per instance1Up to 1,000, configurable
Connection poolingHard, needs a proxyNatural. One instance serves many requests
Scale to zeroYesYes, or set a minimum
Max duration15 minutes60 minutes
Local developmentEmulationIt is just a container
the concurrency difference is the important one
Cloud Run instances handle many concurrent requests, so a pool of ten database connections serves hundreds of requests. That single difference removes the connection problem that dominates the Lambda discussion. Cloud Run is closer to a managed container service than to a function platform, and comparing it to Lambda directly misses what it is.
32

Containers: ECS, EKS, GKE, and Cloud Run again

ECS FargateEKSGKE AutopilotCloud Run
AbstractionTaskPodPodRequest-driven container
You manageTask definitionsA Kubernetes clusterWorkloads onlyAlmost nothing
Learning curveLowHighModerateLowest
PortabilityAWS-specificPortablePortableKnative-based, mostly portable
Cost at steady loadModerateLowest at scale with reserved capacityModeratePer-request, good when spiky
Fit for the posting pathGoodGood, if you already run itGoodVery good
the Kubernetes question, answered honestly
Kubernetes is worth it when you have enough services and enough engineers that the abstraction pays for its own complexity. For a bank with fifteen services and a platform team, yes. For nine services and no platform team, ECS Fargate or Cloud Run delivers the same outcome with a fraction of the operational surface. Choosing Kubernetes because it is standard, without the team to run it, is how you acquire a second full-time system to operate.
33

Choosing per workload, not per company

WorkloadProfileAWSGCP
Posting serviceSteady, connection-heavy, latency-sensitiveECS Fargate or EKSCloud Run (min instances) or GKE
Card authoriserHard latency budget, warm cachesECS, always warmCloud Run, min instances
ISO 8583 gatewayStateful, persistent socketsECS or EC2, not autoscaled on CPUGKE StatefulSet
Outbox relayContinuous poll loopECS FargateCloud Run, min 1
Webhook receiverSpiky, stateless, shortLambdaCloud Run or Functions
Accrual batchScheduled, parallel, boundedECS scheduled tasks or BatchCloud Run jobs
Stream verifierEvent-driven, smallLambda on the streamCloud Run with Pub/Sub push
ReconciliationLong-running, data-heavyECS or GlueCloud Run jobs or Dataflow
the answer to give
"Per workload, and the profile decides. The posting path is steady, connection-heavy and latency-sensitive, so it gets a long-lived container. Webhook receivers are spiky and stateless, so they get functions. The card gateway holds persistent sockets, so it gets a pinned instance that does not autoscale on CPU. A company-wide rule that everything is serverless or everything is Kubernetes is how you end up fighting the platform on half your workloads."
34

The card authoriser: a latency-bound service

Part 8 gave this a 1,000 ms internal budget against a 2-second scheme timeout, with p99 around 414 ms. That constraint dictates the deployment shape.

what the budget forbids
  1. No cold starts. A 300 ms cold start consumes a third of the budget, so instances are always warm with a floor above expected peak.
  2. No cross-AZ database hops where avoidable. Part 6 noted cross-AZ traffic is billed; here it is also latency.
  3. No autoscaling reaction time on the critical path. Scaling takes tens of seconds and the budget is milliseconds, so capacity is provisioned for peak rather than scaled into it.
  4. No shared connection pool with batch workloads, because a batch job exhausting the pool becomes a card decline.
AWS
GCP
fixed task count sized for peak
ECS service, , spread across AZs, with its own RDS Proxy target group and its own ElastiCache cluster.

Autoscaling is set with a high floor and a slow scale-in, so it never scales down into a peak.
Cloud Run with minimum instances at peak level and concurrency tuned so one instance is never saturated.

Or a GKE deployment with a PodDisruptionBudget, if the rest of the estate is already Kubernetes.
the stand-in consequence
Part 8 chapter 91 built stand-in processing for when the ledger is unreachable. That component must be deployed independently of the ledger and of the authoriser’s primary path, with its own replicated snapshot store. Deploying it in the same task, sharing the same failure domain, defeats the entire point of having it.
35

Batch and accrual jobs

Part 5 chapter 60 designed the accrual batch: sharded, checkpointed, idempotent per day, rate-limited so it does not starve interactive traffic. The cloud question is where it runs.

AWS
GCP
ECS scheduled tasks
via EventBridge, one task per shard range, or AWS Batch for large fan-out with queueing.

Step Functions for orchestration when the job has stages and needs retries per stage.
Cloud Run jobs, which are built for exactly this: run to completion, parallel task index, automatic retry.

Workflows for orchestration, or Dataflow when the work is genuinely a data pipeline.
worked numbers
2M active loans ÷ 4,096 shards ≈ 490 per shard
64 parallel workers → ~2.6 minutes for the whole bank

on Cloud Run jobs: parallelism: 64, one line of config
on ECS: 64 tasks, each given a shard range via env

the batch design from Part 5 maps directly. what the cloud adds
is that 64 workers cost the same as 1 for 1/64th of the time.
the rate-limiting requirement, restated
Part 5 said the batch must not starve interactive traffic. In the cloud that means a separate connection pool with a hard cap, and ideally a separate reader endpoint. Sixty-four parallel workers hammering the same primary that is serving card authorisations is how a nightly job becomes a customer-facing incident, and the cloud makes it easier to accidentally provision that parallelism.
36

Stateful components: the ISO 8583 gateway

Part 14 chapter 161 described it: persistent TCP sockets to the scheme, correlation by STAN rather than by connection, echo tests every 30 seconds. It is the one genuinely stateful service in the estate.

what that rules out
  1. Scaling on CPU. Connections are the resource, not CPU, and adding an instance does not add connections unless the scheme provisions them.
  2. Rolling deploys without care. Draining means letting in-flight authorisations finish while refusing new ones, which needs an explicit drain period.
  3. Ephemeral IPs. The scheme allowlists source addresses, so egress must be through a stable NAT address.
  4. Scale to zero. Ever. The link must stay up or the echo tests fail and the scheme marks us down.
AWS
GCP
fixed desired count
ECS service with a , a NAT Gateway with an Elastic IP for stable egress, and a long deregistration delay for draining.

Inside its own PCI account, per Part 8.
GKE StatefulSet with a fixed replica count and Cloud NAT with a reserved static IP.

Cloud Run is a poor fit: it is request-driven and this is connection-driven.
the general lesson
Every estate has one or two components that do not fit the platform’s preferred shape, and trying to force them usually produces a fragile result. Isolate them, give them their own deployment lifecycle, wrap them in an interface the rest of the system understands, and stop trying to make them cloud-native. Part 6 did exactly this with payment connectors, and it is the same move.
37

Autoscaling that does not amplify an incident

the four rules
  1. Scale on the right signal. CPU is wrong for a connection-bound or latency-bound service. Queue depth or concurrent requests is usually right.
  2. Bound the maximum. An unbounded maximum turns a dependency slowdown into a dependency outage, because every new instance adds load to the thing that is already struggling.
  3. Scale out fast, scale in slow. Scaling in during a lull immediately before a peak is a self-inflicted incident.
  4. Never autoscale into a shared bottleneck. If all instances share one database, adding instances adds connections and contention, not capacity.
the sentence to say
"I would cap the maximum instance count explicitly, because autoscaling on CPU into a slow dependency is a positive feedback loop: more instances, more connections, slower dependency, higher CPU, more instances. The cap converts an outage into a queue, and a queue is something Part 13’s degradation plan can handle."
autoscaling
when scaling out makes it worse
swipe the figure sideways, or tap expand for full screen
1/7
healthy
Two instances of a service, talking to a healthy dependency. Everything is fine.
38

The decision, with numbers

AWS
GCP
ECS Fargate
Posting, authorisation, relay and gateway on , sized for peak with modest headroom. Batch on ECS scheduled tasks. Webhooks and the stream verifier on Lambda.

Roughly $4,500 to $6,000 per month for the always-on services.
The same split on Cloud Run with minimum instances for the always-on services, Cloud Run jobs for batch, and the gateway on GKE.

Roughly $4,000 to $5,500 per month, with Cloud Run’s concurrency model doing real work on the connection side.
the answer
"Long-lived containers for the money path, functions for spiky stateless work, and a pinned instance for the card gateway. The deciding property is not cost, it is connection behaviour and cold starts: the posting path holds a warm pool and warm caches, and functions discard both. On GCP I would lean to Cloud Run, because per-instance concurrency means one instance serves many requests from one pool, which removes the connection problem that dominates the Lambda version of this conversation. And I would cap autoscaling maximums everywhere, because scaling into a slow dependency is how a degradation becomes an outage."