Part 10 · 3 chapters · ~18 min

Running Node in Production

Containers done right (base images, distroless, signals and PID 1), memory limits against --max-old-space-size, process managers, deployment topologies from VMs to serverless and the edge, cold starts and connection reuse, health checks and load shedding, zero-downtime deploys, and the production stack Node is paired with at scale.

32

Containers, signals and memory limits

code
# Dockerfile: small, non-root, signals handled
FROM node:24-slim AS deps
WORKDIR /app
COPY package.json package-lock.json ./
RUN npm ci --omit=dev --ignore-scripts

FROM gcr.io/distroless/nodejs24-debian12
WORKDIR /app
COPY --from=deps /app/node_modules ./node_modules
COPY dist ./dist
USER nonroot
ENV NODE_ENV=production NODE_OPTIONS="--max-old-space-size=768 --enable-source-maps"
CMD ["dist/server.js"]          # exec form: node is PID 1 and receives SIGTERM directly
problemcausefix
SIGTERM ignored, killed after 30 sCMD npm start: npm (or a shell) is PID 1 and does not forward signalsexec form CMD ["node", "server.js"], or docker run --init / tini
OOMKilled with a "healthy" heapcontainer limit 1 GiB, heap allowed to grow to its default, plus buffers and native memory outside the heap--max-old-space-size at ~70-75% of the limit; watch RSS, not heapUsed
zombie child processesPID 1 not reaping childrentini or --init
slow image pullsfull Debian image with build toolsmulti-stage build, slim or distroless runtime

Recent Node versions read cgroup memory limits when sizing the default heap, but setting the flag explicitly makes the budget visible in code review.

33

Topologies, serverless, edge and zero-downtime deploys

topologywhat breakswhat to do
VM with PM2 or systemdmanual scaling, config driftcluster mode, immutable images, health checks
containers on Kubernetes or ECSmemory limits, signals, readiness during deploysthe shutdown sequence (part 7), readiness probes, resource requests
serverless (Lambda, Cloud Run)cold starts; the loop is frozen between invocations; connection storms to the DBsmall bundles, create clients outside the handler and reuse, RDS Proxy or a pooler, do not leave work running after the response
edge runtimes (Workers, Vercel Edge)a Web-API subset: no fs, no native addons, limited Node APIskeep edge code small, push heavy logic to origin services
health checks, honestly
  1. Liveness: "is this process stuck?" Cheap, no dependencies. Failing it restarts the pod, so never include the database.
  2. Readiness: "should I receive traffic now?" False during startup, shutdown and overload.
  3. Load shedding: when event loop delay or in-flight requests pass a threshold, answer 503 with Retry-After early (for example with @fastify/under-pressure) instead of queueing until everything times out.
  4. Zero-downtime deploys: rolling updates with readiness gates plus the graceful shutdown from part 7, so in-flight requests drain before a pod dies.
34

The production stack around Node

At scale, a Node service is one box in a larger picture. The pairings below are what most companies add, roughly in this order, as they grow (course 9 compares them across languages).

stageadded around Nodewhy
first deploya PaaS or one VM with PM2 behind Nginx or Caddy; Postgres; Sentryship, restart on crash, see errors
first scalecontainers, a load balancer, Redis (cache, sessions), BullMQ for jobsmore instances, slow work off the request path
more teamsKubernetes or ECS, PgBouncer, an API gateway, OpenTelemetry + a Collector, secrets manager (Vault or cloud)bounded connections, shared auth, tracing across services, no secrets in env files
many servicesservice mesh sidecars (Envoy via Istio or Linkerd) for mTLS and retries, Kafka for events, feature flags, canary deploysconsistent security and traffic control without changing every codebase
WHAT NODE IS PAIRED WITH AT SCALE
the usual production stack around a Node service
CDN / edgecache, TLSload balancerALB, Envoy, NginxAPI gatewayauth, rate limitsNode podsFastify/Nest, 1 proc eachBullMQ workersRedis-backed jobsOTel Collectortraces, metricsPgBouncerpoolingRediscache, queues, sessionsPostgresprimary + replicas
swipe the figure sideways, or tap expand for full screen
1/6
the edge
Static assets and cacheable GETs stop at the CDN. TLS terminates at the edge or the load balancer, not in Node, which then speaks plain HTTP (or mTLS via a mesh) inside the network.
TLS and caching before NodeNode rarely terminates public TLS at scale