Part 7 · 1 chapters · ~8 min

cgroups and Namespaces

The seven namespace types and their system calls (clone, unshare, setns), cgroups v2 controllers and files, PSI pressure metrics, how Docker and Kubernetes map requests and limits to cgroups, systemd slices, and why containers share a kernel.

8

The kernel half of containers

code
sudo unshare --pid --fork --mount-proc --uts --net bash    # a shell in new namespaces
  hostname sandbox; ps aux                                  # sees only itself; hostname changed for it alone
lsns                                                        # list namespaces on the host
cat /proc/$$/cgroup                                         # which cgroup this shell is in
ls /sys/fs/cgroup/system.slice/docker-<id>.scope/          # cpu.max memory.max memory.current io.stat pids.max
cat /sys/fs/cgroup/<group>/memory.events                    # oom, oom_kill counts
cat /proc/pressure/cpu /proc/pressure/memory /proc/pressure/io   # PSI: % of time tasks stalled on each resource
Kubernetescgroup v2 file
resources.requests.cpucpu.weight (relative share under contention)
resources.limits.cpucpu.max (quota per period: throttling)
resources.limits.memorymemory.max (OOM kill when exceeded)
NAMESPACES AND CGROUPS
what a process sees, and what it may use
pid namespaceOwn process tree: the container'sfirst process is PID 1.mount namespaceOwn mount table: its own root filesystem.net namespaceOwn interfaces, routes, ports,iptables.user, uts, ipc, cgroup, timeOwn user ids, hostname, IPCobjects, cgroup view, clocks.cgroups v2One hierarchy limiting cpu,memory, io and pids per group.togetherA container = a process +namespaces + cgroups + a root fs.
swipe the figure sideways, or tap expand for full screen
1/4
visibility
Namespaces change what a process can see: its own PIDs, mounts, network stack and hostname. The kernel is still shared.
namespaces isolate viewsone shared kernel