9 parts · 9 chapters
Systems Performance and eBPF
How to find out why a machine is slow instead of guessing: a method, the resources to check in order, and the tools that answer each question on a live Linux system. Built on Brendan Gregg's methods and the Kernel Internals course.
Nine parts: methodology (the USE method, workload characterisation, the 60-second checklist); CPU; memory; file systems; disks; network; perf and flame graphs; eBPF and bpftrace; and tuning with a change-one-thing discipline.
methodUSE for every resource; RED for every service.
cpuUtilisation, saturation, run queues, profiling.
memoryWorking set, faults, swapping, reclaim.
storageFile system latency first, then the disk.
networkThroughput, retransmits, drops, queues.
toolsperf, flame graphs, bpftrace one-liners.
00
Methodology: USE, Workload Characterisation, the 60-Second Checklist
A method before tools
1 ch · ~8 min01CPU
Utilisation, saturation, efficiency
1 ch · ~8 min02Memory
Pressure, reclaim, faults
1 ch · ~8 min03File Systems
Latency the application feels
1 ch · ~8 min04Disks
Operations, latency, queues
1 ch · ~8 min05Network
Throughput, loss, latency
1 ch · ~8 min06perf and Flame Graphs
Profile, then optimise
1 ch · ~8 min07eBPF and bpftrace
One-liners that answer real questions
1 ch · ~8 min08Tuning
Settings with reasons
1 ch · ~8 minBuilt on Kernel Internals and DiagnosisUses Kernel Internals for the paths being measured and Backend Diagnosis for the application-level counterpart.