Part 3 · 1 chapters · ~8 min
Memory Management
Physical pages, zones and the buddy allocator, slab allocators for kernel objects, virtual memory areas and page tables, page faults (minor and major), copy-on-write, transparent huge pages, reclaim with LRU lists and kswapd, overcommit, the OOM killer, cgroup memory limits, and reading /proc/meminfo.
4
Pages, faults, reclaim and the OOM killer
code
grep -E 'MemTotal|MemFree|MemAvailable|Cached|Dirty|AnonPages' /proc/meminfo # MemFree is small on a healthy server: free memory becomes page cache. MemAvailable is the honest number. cat /proc/$(pgrep -o node)/status | grep -E 'VmRSS|VmSwap' cat /proc/$(pgrep -o node)/smaps_rollup # Rss, Pss, Anonymous, shared vs private /usr/bin/time -v node app.js 2>&1 | grep -E 'Maximum resident|page faults' # major vs minor faults cat /proc/sys/vm/overcommit_memory # 0 heuristic, 1 always, 2 strict dmesg | grep -i 'out of memory' # OOM killer decisions, with the victim's oom_score
| allocator | job |
|---|---|
| buddy allocator | hands out physical pages in power-of-two blocks |
| slab (SLUB) | caches of same-sized kernel objects (inodes, dentries, sk_buffs): /proc/slabinfo |
| page cache | file data in memory, reclaimable |
| reclaim (kswapd, direct reclaim) | frees clean cache, writes back dirty pages, swaps anonymous pages |
Containers: a cgroup memory limit (memory.max) triggers reclaim and then an OOM kill inside that cgroup only, which is why a pod is "OOMKilled" while the node has free memory. Page cache from the container's own files counts toward its limit.
A PAGE FAULT, STEP BY STEP
virtual memory made real on first touch
swipe the figure sideways, or tap expand for full screen
1/4
mappings first
mmap and malloc create virtual memory areas (VMAs) but no physical pages. Memory is committed lazily, on first touch.
VMAs without pageslazy allocation