9 parts · 9 chapters

Systems Performance and eBPF

How to find out why a machine is slow instead of guessing: a method, the resources to check in order, and the tools that answer each question on a live Linux system. Built on Brendan Gregg's methods and the Kernel Internals course.

Nine parts: methodology (the USE method, workload characterisation, the 60-second checklist); CPU; memory; file systems; disks; network; perf and flame graphs; eBPF and bpftrace; and tuning with a change-one-thing discipline.

methodology (USE, workload characterisation) · CPU · memory · file systems · disks · network · perf and flame graphs · eBPF and bpftrace · tuningsenior → staff · backend, infrastructure and SRE engineers
methodUSE for every resource; RED for every service.
cpuUtilisation, saturation, run queues, profiling.
memoryWorking set, faults, swapping, reclaim.
storageFile system latency first, then the disk.
networkThroughput, retransmits, drops, queues.
toolsperf, flame graphs, bpftrace one-liners.
Built on Kernel Internals and DiagnosisUses Kernel Internals for the paths being measured and Backend Diagnosis for the application-level counterpart.