Part 7 · 2 chapters · ~12 min
Security and Isolation
CPU privilege levels and system calls, memory protection, users, groups and permissions, setuid, capabilities, SELinux and AppArmor, namespaces, cgroups and seccomp, sandboxes and virtual machines, side channels (Spectre, Meltdown), and least privilege in practice.
14
Layers of isolation
code
ls -l /usr/bin/passwd # -rwsr-xr-x: the s bit = setuid root (runs with the owner's privileges)
id # your uid, gid and groups
getcap /usr/bin/ping # capabilities granted to a binary (cap_net_raw instead of setuid root)
# a container that keeps only what it needs (Kubernetes securityContext)
securityContext: { runAsNonRoot: true, readOnlyRootFilesystem: true, allowPrivilegeEscalation: false,
capabilities: { drop: ["ALL"] }, seccompProfile: { type: RuntimeDefault } }LAYERS OF ISOLATION
from hardware privilege levels up to sandboxes
swipe the figure sideways, or tap expand for full screen
1/5
privilege levels
The CPU runs in kernel or user mode. User mode cannot execute privileged instructions or touch device registers; the only way in is a system call or a fault.
hardware enforces kernel vs usersystem calls are the only door
15
Side channels
Spectre and Meltdown (2018) showed that speculative execution (Computers part 1) could leak data across isolation boundaries through cache timing. Fixes (kernel page table isolation, retpolines, microcode) cost some performance, and the lesson persists: isolation is only as strong as the hardware it relies on, which is why cloud providers dedicate cores or hosts for sensitive tenants.