Part 6 · 1 chapters · ~8 min
perf and Flame Graphs
perf stat for counters, perf record and report for profiles, sampling rates, call graphs with frame pointers, DWARF or LBR, on-CPU and off-CPU flame graphs, differential flame graphs, profiling JIT runtimes (Node, Java, Python 3.12+ perf support), and continuous profiling in production.
7
Profile, then optimise
code
perf record -F 99 -a -g -- sleep 30 # sample all CPUs with stacks for 30 s perf script | ./stackcollapse-perf.pl | ./flamegraph.pl > cpu.svg # Brendan Gregg's FlameGraph scripts node --perf-basic-prof server.js # Node writes /tmp/perf-<pid>.map so JS frames have names python3.12 -X perf app.py # CPython 3.12+ perf trampoline support ./asprof -d 30 -f cpu.html <java-pid> # async-profiler for the JVM # off-CPU: where threads wait (locks, I/O, sleeps): the other half of latency offcputime-bpfcc -df -p $(pgrep -o node) 30 | ./flamegraph.pl --color=io > offcpu.svg
Off-CPU analysis matters because a slow request often spends most of its time waiting, not computing; an on-CPU flame graph cannot show that. Continuous profiling (Parca, Pyroscope, Grafana) keeps low-rate profiles always on, so you can compare before and after a deploy.
READING A FLAME GRAPH
stack samples merged into one picture
swipe the figure sideways, or tap expand for full screen
1/4
sample
perf samples every CPU's stack 99 times a second (an odd rate avoids lockstep with timers). Thirty seconds gives thousands of samples.
99 Hz stack sampleslow overhead