Part 4 · 3 chapters · ~22 min

Docker

What a container actually is: a process with namespaces, cgroups and a layered root filesystem on a shared kernel; images, the layer cache, multi-stage builds, bases, digests and secrets; and container networking, volumes and a Compose stack that mirrors production.

12

What a container actually is

a process with a fence
  1. One kernel. A container makes system calls straight into the host kernel. There is no guest OS to boot.
  2. Namespaces give the process its own view of PIDs, network, mounts, hostname and users.
  3. cgroups cap CPU (it gets throttled) and memory (it gets OOM-killed, exit code 137).
  4. The root filesystem is read-only layers plus a writable layer that disappears with the container.
  5. Hardening: drop capabilities, use seccomp, make the root filesystem read-only and run as a non-root user.
  6. Containers versus VMs: containers are fast and dense with weaker isolation. Untrusted code runs in microVMs.
code
# see it for yourself on any Linux host with Docker
docker run -d --name demo --memory=256m --cpus=0.5 nginx:alpine
ps -ef | grep "nginx: master"                     # an ordinary host process
PID=$(docker inspect -f '{{.State.Pid}}' demo)
ls -l /proc/$PID/ns                                 # its namespaces: pid, net, mnt, uts, ipc, user, cgroup
cat /sys/fs/cgroup/system.slice/docker-*.scope/memory.max   # 268435456
cat /sys/fs/cgroup/system.slice/docker-*.scope/cpu.max      # 50000 100000

# build a "container" with no Docker at all
sudo unshare --pid --fork --mount-proc --uts --net bash
hostname inside && ps -ef                           # PID 1 is bash; the host's processes are gone
exit code 137
137 is 128 + 9: the process received SIGKILL, almost always from the OOM killer because it went over memory.max. Node needs --max-old-space-size set below the container limit, or V8 will grow the heap until the kernel kills the process.
WHAT A CONTAINER ACTUALLY IS
a normal Linux process with a restricted view (namespaces), restricted resources (cgroups) and a different root filesystem
swipe the figure sideways, or tap expand for full screen
1/6
one kernel
One kernel, many processes: on the host, ps shows the container's process like any other (with a host PID). There is no guest OS and no hypervisor; the container's process makes system calls straight into the host kernel. That is the whole performance story: nothing to boot, no virtualised hardware.
13

Images, layers and the build

small, cached, reproducible, safe
  1. Layers are content-addressed and shared between images.
  2. The cache: every layer after the first miss is rebuilt, so copy the lockfile and install before copying the source.
  3. Multi-stage builds: build with everything, ship only the output.
  4. Base images: slim, distroless or scratch. Watch for musl on alpine.
  5. Deploy by digest, scan and sign in CI.
  6. .dockerignore keeps node_modules, .git and .env out of the build context.
code
# Dockerfile: a Node API, multi-stage, cached, non-root
FROM node:22-slim AS deps
WORKDIR /app
COPY package.json package-lock.json ./
RUN --mount=type=cache,target=/root/.npm npm ci             # cached until the lockfile changes

FROM deps AS build
COPY . .
RUN npm run build && npm prune --omit=dev

FROM gcr.io/distroless/nodejs22-debian12 AS runtime
WORKDIR /app
COPY --from=build /app/dist ./dist
COPY --from=build /app/node_modules ./node_modules
ENV NODE_ENV=production NODE_OPTIONS=--max-old-space-size=384
USER nonroot
EXPOSE 3000
CMD ["dist/server.js"]
code
# .dockerignore
node_modules
.git
.env*
coverage
dist
the secret in a layer
A secret written in one layer and deleted in the next is still in the image: each layer is a separate tarball, and docker history plus docker save recover it. Pass build secrets with RUN --mount=type=secret, which never writes them to a layer.
IMAGES, LAYERS AND THE BUILD
a Dockerfile as a stack of cached layers, and the ordering that makes rebuilds take seconds instead of minutes
swipe the figure sideways, or tap expand for full screen
1/6
layers
Layers: FROM node:22-slim brings the base layers; COPY package.json and RUN npm ci add a dependencies layer; COPY . adds the source layer; RUN npm run build adds the output. Each layer stores only the files that changed, and identical layers are shared between images on a host and in a registry.
14

Networking, volumes and Compose

plumbing in, state out
  1. The bridge: a veth pair per container, with outbound traffic NATed by the host.
  2. Publishing maps a host port to a container port. Inside the container, bind to 0.0.0.0.
  3. User-defined networks give DNS by service name.
  4. Orchestrators give each pod or task its own routable IP and a stable service name.
  5. Volumes outlive the container. The writable layer does not.
  6. In the cloud, state lives in managed services and containers are disposable.
code
# compose.yaml: the local stack that mirrors production's shape
services:
  api:
    build: .
    ports: ["8080:3000"]                 # host 8080 → container 3000
    environment:
      DATABASE_URL: postgres://app:app@db:5432/app   # "db" resolves via Compose DNS
    depends_on:
      db: { condition: service_healthy }
  db:
    image: postgres:17
    environment: { POSTGRES_USER: app, POSTGRES_PASSWORD: app, POSTGRES_DB: app }
    volumes: ["pgdata:/var/lib/postgresql/data"]      # survives `docker compose down`
    healthcheck: { test: ["CMD", "pg_isready", "-U", "app"], interval: 2s, retries: 15 }
volumes:
  pgdata: {}
the exercise
Run docker compose up, then docker compose down followed by up again. The data survives. Now run down -v: the volume is gone and so is the data. That difference is the whole volume story.
CONTAINER NETWORKING AND VOLUMES
how a container gets an IP and a port, and where its data lives when it dies
swipe the figure sideways, or tap expand for full screen
1/6
bridge
The bridge network (Docker's default): each container gets a veth pair, one end inside as eth0 with an IP like 172.17.0.5, the other end attached to the docker0 bridge on the host. Containers on the same bridge reach each other by IP; outbound traffic is NATed through the host's IP by iptables rules.