Part 5 · 1 chapters · ~8 min
seccomp and Capabilities
Linux capabilities and the default set granted to containers, dropping capabilities, seccomp-bpf filters and the runtime default profile, AppArmor and SELinux, non-root users, read-only root file systems, no-new-privileges, user namespaces and rootless containers, and Pod Security Standards in Kubernetes.
6
Least privilege, enforced
code
# Kubernetes securityContext matching the Pod Security "restricted" standard
securityContext:
runAsNonRoot: true
runAsUser: 10001
allowPrivilegeEscalation: false
readOnlyRootFilesystem: true
capabilities: { drop: ["ALL"] }
seccompProfile: { type: RuntimeDefault }
# see a process's capabilities and seccomp mode (Linux)
grep -E 'Cap(Eff|Bnd)|Seccomp' /proc/$(pgrep -o node)/status
capsh --decode=00000000a80425fb # Docker's default effective set, decoded
docker run --rm --cap-drop=ALL --security-opt no-new-privileges --read-only -u 10001 api:1.42Privileged containers (--privileged, or hostPID, hostNetwork, hostPath mounts of / or the Docker socket) remove most isolation; treat them as root on the node. Enforce the restricted standard with Pod Security Admission and allow exceptions per namespace with reasons.
HARDENING A CONTAINER
remove privileges until it barely works, then stop
swipe the figure sideways, or tap expand for full screen
1/5
root by default
Without a USER instruction, the process runs as root inside the container, which is root on the host kernel's terms for many operations. A container escape then starts with root.
default root is riskyset USER