Part 3 · 2 chapters · ~18 min
Compute
The compute spectrum from VMs through orchestrated and serverless containers to functions, what moves at each step and how to choose by the shape of the load; then scaling as a control loop: signals, speed, scheduled peaks, and the stateful limits that cap it.
10
VMs, containers and serverless
control against operations
- VMs: your kernel and your patches. A minute or more to boot, and paid for whether busy or not.
- Orchestrated containers on nodes you own: bin-packing for utilisation, but the nodes are yours to run.
- Serverless containers: you supply an image and the provider supplies the capacity, billed per second. Cloud Run scales to zero.
- Functions: a handler billed per request, with hard limits and cold starts.
- What moves along the spectrum: who patches, how fast it scales, the billing unit, the steady-load price, and how much network control you keep.
- Choose by the shape of the load. Most systems use three of these at once.
| workload | good fit | why |
|---|---|---|
| public API, steady 2k rps | containers on ECS/GKE, committed capacity | cheapest per request at steady load; full network control |
| admin tool, 50 users | Cloud Run / Fargate, min instances 0 or 1 | near-zero idle cost; no nodes to patch |
| resize uploaded images | Lambda / Cloud Functions on the storage event | event-driven, bursty, short |
| nightly reconciliation batch | container job on spot capacity | interruptible, retryable, cheap |
| Postgres | RDS / Cloud SQL / AlloyDB | backups, failover and patches are the hard part |
| Next.js frontend | edge/CDN for static, serverless containers or functions for SSR | spiky, cacheable, latency-sensitive |
the Lambda bill trap
A function at a steady 500 rps with 200 ms at 1 GB costs several times what the same work costs on two committed containers. Functions are cheap for spiky and idle workloads and expensive for busy, steady ones. Re-run the numbers when traffic changes shape.
THE COMPUTE SPECTRUM
from a VM you patch to a function you do not see, and what moves at each step
swipe the figure sideways, or tap expand for full screen
1/6
VMs
Virtual machines (EC2, Compute Engine): a slice of a physical host with its own kernel. You choose the instance type (vCPU, memory, network, GPU), the image, the disk; you patch the OS, install the runtime, run the process supervisor, and handle scaling with instance groups. Full control, slowest to scale (a minute or more to boot), and you pay for the VM whether it is busy or not.
11
Scaling models
a control loop with a signal, a speed and a ceiling
- Vertical or horizontal: horizontal scaling needs stateless services.
- CPU targets lag for I/O-bound services. Prefer request-count or concurrency targets.
- Concurrency scaling follows load closely and pays for it in cold starts.
- Queue-depth scaling absorbs bursts without dropping work.
- Scheduled and predictive scaling gets capacity in place before known peaks.
- The limits: warm-up time, database connections, downstream rate limits and quotas.
code
# Cloud Run: concurrency-based scaling with a floor for latency and a ceiling for the DB
gcloud run deploy payments-api \
--image=europe-docker.pkg.dev/acme/apps/payments:4.18.0 \
--concurrency=40 \
--min-instances=2 \ # no cold start for the first requests of the morning
--max-instances=30 \ # 30 × pool of 5 = 150 connections: under Postgres max_connections
--cpu=1 --memory=1Gi \
--cpu-boost # extra CPU during startup to shorten cold starts
# Kubernetes HPA on a custom metric (requests in flight per pod), not CPU
# metrics: - type: Pods pods: { metric: { name: http_inflight }, target: { averageValue: "30" } }salary day
For a Nigerian fintech, payday (the 25th to the end of the month) and the first days of the month are predictable peaks of 3 to 10× normal transfers. Schedule capacity up on the 24th, and load-test the database and the payment providers' rate limits, not only the API tier.
SCALING MODELS
what triggers more capacity, how fast it arrives, and why the trigger matters as much as the limit
swipe the figure sideways, or tap expand for full screen
1/6
vertical vs horizontal
Vertical versus horizontal: vertical scaling gives one instance more CPU and memory (simple, usually needs a restart, has a ceiling at the biggest instance); horizontal scaling adds more instances behind a load balancer (no ceiling in principle, needs the service to be stateless so any instance can serve any request). Stateless services scale horizontally; state goes to a database, cache or object store.