Part 1 · 2 chapters · ~12 min

CPU Scheduling

Goals (throughput, latency, fairness), FIFO, shortest job first and round robin, multi-level feedback queues, Linux CFS and EEVDF, priorities and nice values, real-time scheduling, multicore scheduling and affinity, and what containers' CPU limits mean (CFS quota throttling).

3

Policies and what they optimise

code
nice -n 10 ./batch-job          # lower priority (higher nice) for background work
chrt -f 50 ./latency-critical   # real-time FIFO priority (Linux; use with great care)
taskset -c 2,3 ./service        # pin to CPUs 2 and 3 (affinity)
SCHEDULING POLICIES
from simple queues to fair share, and what each optimises
FIFOrun to completionSJF / SRTFshortest firstround robintime slicesMLFQpriorities that adaptCFSfair share by virtual runtimeEEVDFfair + latency deadlines
swipe the figure sideways, or tap expand for full screen
1/5
FIFO
First in, first out: simple, but one long job delays every short one behind it (the convoy effect).
simple; convoys behind long jobsbad for interactive work
4

CPU limits in containers

Kubernetes CPU limits are enforced with CFS bandwidth control: a container with a limit of 1 CPU gets 100 ms of CPU time per 100 ms period. A multi-threaded runtime using 4 threads can burn the quota in 25 ms and then be throttled for 75 ms, adding latency spikes even though average CPU looks fine. Watch container_cpu_cfs_throttled_periods_total; many teams set requests but no CPU limits for latency-sensitive services, and size runtime thread counts (GOMAXPROCS, JVM active processor count, UV_THREADPOOL_SIZE) to the quota.