Part 10 · 3 chapters · ~18 min
Running Node in Production
Containers done right (base images, distroless, signals and PID 1), memory limits against --max-old-space-size, process managers, deployment topologies from VMs to serverless and the edge, cold starts and connection reuse, health checks and load shedding, zero-downtime deploys, and the production stack Node is paired with at scale.
32
Containers, signals and memory limits
code
# Dockerfile: small, non-root, signals handled FROM node:24-slim AS deps WORKDIR /app COPY package.json package-lock.json ./ RUN npm ci --omit=dev --ignore-scripts FROM gcr.io/distroless/nodejs24-debian12 WORKDIR /app COPY --from=deps /app/node_modules ./node_modules COPY dist ./dist USER nonroot ENV NODE_ENV=production NODE_OPTIONS="--max-old-space-size=768 --enable-source-maps" CMD ["dist/server.js"] # exec form: node is PID 1 and receives SIGTERM directly
| problem | cause | fix |
|---|---|---|
| SIGTERM ignored, killed after 30 s | CMD npm start: npm (or a shell) is PID 1 and does not forward signals | exec form CMD ["node", "server.js"], or docker run --init / tini |
| OOMKilled with a "healthy" heap | container limit 1 GiB, heap allowed to grow to its default, plus buffers and native memory outside the heap | --max-old-space-size at ~70-75% of the limit; watch RSS, not heapUsed |
| zombie child processes | PID 1 not reaping children | tini or --init |
| slow image pulls | full Debian image with build tools | multi-stage build, slim or distroless runtime |
Recent Node versions read cgroup memory limits when sizing the default heap, but setting the flag explicitly makes the budget visible in code review.
33
Topologies, serverless, edge and zero-downtime deploys
| topology | what breaks | what to do |
|---|---|---|
| VM with PM2 or systemd | manual scaling, config drift | cluster mode, immutable images, health checks |
| containers on Kubernetes or ECS | memory limits, signals, readiness during deploys | the shutdown sequence (part 7), readiness probes, resource requests |
| serverless (Lambda, Cloud Run) | cold starts; the loop is frozen between invocations; connection storms to the DB | small bundles, create clients outside the handler and reuse, RDS Proxy or a pooler, do not leave work running after the response |
| edge runtimes (Workers, Vercel Edge) | a Web-API subset: no fs, no native addons, limited Node APIs | keep edge code small, push heavy logic to origin services |
health checks, honestly
- Liveness: "is this process stuck?" Cheap, no dependencies. Failing it restarts the pod, so never include the database.
- Readiness: "should I receive traffic now?" False during startup, shutdown and overload.
- Load shedding: when event loop delay or in-flight requests pass a threshold, answer 503 with
Retry-Afterearly (for example with@fastify/under-pressure) instead of queueing until everything times out. - Zero-downtime deploys: rolling updates with readiness gates plus the graceful shutdown from part 7, so in-flight requests drain before a pod dies.
34
The production stack around Node
At scale, a Node service is one box in a larger picture. The pairings below are what most companies add, roughly in this order, as they grow (course 9 compares them across languages).
| stage | added around Node | why |
|---|---|---|
| first deploy | a PaaS or one VM with PM2 behind Nginx or Caddy; Postgres; Sentry | ship, restart on crash, see errors |
| first scale | containers, a load balancer, Redis (cache, sessions), BullMQ for jobs | more instances, slow work off the request path |
| more teams | Kubernetes or ECS, PgBouncer, an API gateway, OpenTelemetry + a Collector, secrets manager (Vault or cloud) | bounded connections, shared auth, tracing across services, no secrets in env files |
| many services | service mesh sidecars (Envoy via Istio or Linkerd) for mTLS and retries, Kafka for events, feature flags, canary deploys | consistent security and traffic control without changing every codebase |
WHAT NODE IS PAIRED WITH AT SCALE
the usual production stack around a Node service
swipe the figure sideways, or tap expand for full screen
1/6
the edge
Static assets and cacheable GETs stop at the CDN. TLS terminates at the edge or the load balancer, not in Node, which then speaks plain HTTP (or mTLS via a mesh) inside the network.
TLS and caching before NodeNode rarely terminates public TLS at scale