Part 2 · 1 chapters · ~8 min
Scheduling: CFS and EEVDF
Tasks and the run queue, scheduling classes (stop, deadline, real-time, fair, idle), virtual runtime and nice weights, CFS and its red-black tree, EEVDF since Linux 6.6, preemption and context switches, CPU affinity and load balancing, cgroup CPU controls and throttling, and sched_ext.
3
Who runs next
| class | policies | use |
|---|---|---|
| deadline | SCHED_DEADLINE | tasks with runtime, period and deadline guarantees |
| real-time | SCHED_FIFO, SCHED_RR | audio, control loops; can starve everything below |
| fair | SCHED_NORMAL, SCHED_BATCH | almost every process (CFS, now EEVDF) |
| idle | SCHED_IDLE | runs only when nothing else wants the CPU |
code
chrt -p $(pgrep -o postgres) # current policy and priority taskset -cp 0-3 $(pgrep -o node) # pin to CPUs 0-3 cat /proc/$(pgrep -o node)/sched | head # se.vruntime, nr_switches, voluntary vs involuntary cat /sys/fs/cgroup/<group>/cpu.stat # nr_throttled, throttled_usec: CPU limits biting
CPU limits and latency: a container with a CPU limit uses the CFS bandwidth controller (cpu.max: quota per period). A multi-threaded runtime can spend its quota early in a 100 ms period and be throttled for the rest, adding tail latency even at low average CPU. Watch nr_throttled. sched_ext (merged in 6.12) lets schedulers be written as eBPF programs.
SCHEDULING: FROM CFS TO EEVDF
fair shares of CPU, and who runs next
swipe the figure sideways, or tap expand for full screen
1/4
per-CPU queues
Each CPU has its own run queue, so scheduling decisions are local and cheap; load balancing moves tasks between CPUs periodically.
a queue per CPUbalanced periodically