Compute: where the posting path runs
The layer with the most tribal opinion and the clearest technical answer. The posting path is steady, connection-heavy and latency-sensitive, which is precisely the profile function platforms handle worst, and webhook receivers are the opposite. This part works through it per workload rather than per company, and ends on the autoscaling configuration that turns a slow dependency into an outage.
The posting path is not a good Lambda
Starting with the conclusion, because it is the one people argue with and the reasons are concrete rather than aesthetic.
- Connection churn. Part 1 chapter 16: a thousand concurrent functions means a thousand database connections, short-lived, against a database comfortable with a few hundred.
- Cold starts on the money path. A VPC-attached function adds latency at exactly the moment a customer is waiting, and provisioned concurrency to avoid it is just a container with extra steps.
- The work is not bursty. Part 3 derived a steady 464 transfers per second at peak. Functions price elasticity, and we have none to exploit.
- Warm caches matter. The shard map, tariff tables and account-type policy are all read constantly and cached per process. A function discards that cache.
- Long-lived connections to the scheme. Part 14 chapter 161 needs persistent ISO 8583 sockets, which a function fundamentally cannot hold.
Lambda: cold starts, concurrency, and the VPC tax
cold start, decomposed: download and unpack the package 50 to 300 ms start the runtime 50 to 200 ms your init code, pools, config yours to control VPC ENI attachment (historically) now ~sub-second, was ~10 s concurrency model: one request per environment 1,000 concurrent requests = 1,000 environments = 1,000 database connections unless proxied
| Knob | What it does | When to use it |
|---|---|---|
| Provisioned concurrency | Keeps N environments warm | Predictable latency. At which point you are paying for idle containers |
| Reserved concurrency | Caps how many can run | Protecting a downstream database from being overwhelmed |
| SnapStart | Snapshots an initialised runtime | JVM cold starts specifically |
| Memory setting | Also scales CPU | Often cheaper to raise: faster execution can cost less |
Cloud Functions and Cloud Run
The gap between "function" and "container" is on AWS.
Cloud Functions gen 2 is Cloud Run underneath.
| Lambda | Cloud Run | |
|---|---|---|
| Unit | Function package or container | Container, always |
| Concurrency per instance | 1 | Up to 1,000, configurable |
| Connection pooling | Hard, needs a proxy | Natural. One instance serves many requests |
| Scale to zero | Yes | Yes, or set a minimum |
| Max duration | 15 minutes | 60 minutes |
| Local development | Emulation | It is just a container |
Containers: ECS, EKS, GKE, and Cloud Run again
| ECS Fargate | EKS | GKE Autopilot | Cloud Run | |
|---|---|---|---|---|
| Abstraction | Task | Pod | Pod | Request-driven container |
| You manage | Task definitions | A Kubernetes cluster | Workloads only | Almost nothing |
| Learning curve | Low | High | Moderate | Lowest |
| Portability | AWS-specific | Portable | Portable | Knative-based, mostly portable |
| Cost at steady load | Moderate | Lowest at scale with reserved capacity | Moderate | Per-request, good when spiky |
| Fit for the posting path | Good | Good, if you already run it | Good | Very good |
Choosing per workload, not per company
| Workload | Profile | AWS | GCP |
|---|---|---|---|
| Posting service | Steady, connection-heavy, latency-sensitive | ECS Fargate or EKS | Cloud Run (min instances) or GKE |
| Card authoriser | Hard latency budget, warm caches | ECS, always warm | Cloud Run, min instances |
| ISO 8583 gateway | Stateful, persistent sockets | ECS or EC2, not autoscaled on CPU | GKE StatefulSet |
| Outbox relay | Continuous poll loop | ECS Fargate | Cloud Run, min 1 |
| Webhook receiver | Spiky, stateless, short | Lambda | Cloud Run or Functions |
| Accrual batch | Scheduled, parallel, bounded | ECS scheduled tasks or Batch | Cloud Run jobs |
| Stream verifier | Event-driven, small | Lambda on the stream | Cloud Run with Pub/Sub push |
| Reconciliation | Long-running, data-heavy | ECS or Glue | Cloud Run jobs or Dataflow |
The card authoriser: a latency-bound service
Part 8 gave this a 1,000 ms internal budget against a 2-second scheme timeout, with p99 around 414 ms. That constraint dictates the deployment shape.
- No cold starts. A 300 ms cold start consumes a third of the budget, so instances are always warm with a floor above expected peak.
- No cross-AZ database hops where avoidable. Part 6 noted cross-AZ traffic is billed; here it is also latency.
- No autoscaling reaction time on the critical path. Scaling takes tens of seconds and the budget is milliseconds, so capacity is provisioned for peak rather than scaled into it.
- No shared connection pool with batch workloads, because a batch job exhausting the pool becomes a card decline.
Autoscaling is set with a high floor and a slow scale-in, so it never scales down into a peak.
Or a GKE deployment with a PodDisruptionBudget, if the rest of the estate is already Kubernetes.
Batch and accrual jobs
Part 5 chapter 60 designed the accrual batch: sharded, checkpointed, idempotent per day, rate-limited so it does not starve interactive traffic. The cloud question is where it runs.
Step Functions for orchestration when the job has stages and needs retries per stage.
Workflows for orchestration, or Dataflow when the work is genuinely a data pipeline.
2M active loans ÷ 4,096 shards ≈ 490 per shard 64 parallel workers → ~2.6 minutes for the whole bank on Cloud Run jobs: parallelism: 64, one line of config on ECS: 64 tasks, each given a shard range via env the batch design from Part 5 maps directly. what the cloud adds is that 64 workers cost the same as 1 for 1/64th of the time.
Stateful components: the ISO 8583 gateway
Part 14 chapter 161 described it: persistent TCP sockets to the scheme, correlation by STAN rather than by connection, echo tests every 30 seconds. It is the one genuinely stateful service in the estate.
- Scaling on CPU. Connections are the resource, not CPU, and adding an instance does not add connections unless the scheme provisions them.
- Rolling deploys without care. Draining means letting in-flight authorisations finish while refusing new ones, which needs an explicit drain period.
- Ephemeral IPs. The scheme allowlists source addresses, so egress must be through a stable NAT address.
- Scale to zero. Ever. The link must stay up or the echo tests fail and the scheme marks us down.
Inside its own PCI account, per Part 8.
Cloud Run is a poor fit: it is request-driven and this is connection-driven.
Autoscaling that does not amplify an incident
- Scale on the right signal. CPU is wrong for a connection-bound or latency-bound service. Queue depth or concurrent requests is usually right.
- Bound the maximum. An unbounded maximum turns a dependency slowdown into a dependency outage, because every new instance adds load to the thing that is already struggling.
- Scale out fast, scale in slow. Scaling in during a lull immediately before a peak is a self-inflicted incident.
- Never autoscale into a shared bottleneck. If all instances share one database, adding instances adds connections and contention, not capacity.
The decision, with numbers
Roughly $4,500 to $6,000 per month for the always-on services.
Roughly $4,000 to $5,500 per month, with Cloud Run’s concurrency model doing real work on the connection side.