Part 6 · 1 chapters · ~8 min

perf and Flame Graphs

perf stat for counters, perf record and report for profiles, sampling rates, call graphs with frame pointers, DWARF or LBR, on-CPU and off-CPU flame graphs, differential flame graphs, profiling JIT runtimes (Node, Java, Python 3.12+ perf support), and continuous profiling in production.

7

Profile, then optimise

code
perf record -F 99 -a -g -- sleep 30            # sample all CPUs with stacks for 30 s
perf script | ./stackcollapse-perf.pl | ./flamegraph.pl > cpu.svg     # Brendan Gregg's FlameGraph scripts

node --perf-basic-prof server.js               # Node writes /tmp/perf-<pid>.map so JS frames have names
python3.12 -X perf app.py                      # CPython 3.12+ perf trampoline support
./asprof -d 30 -f cpu.html <java-pid>          # async-profiler for the JVM

# off-CPU: where threads wait (locks, I/O, sleeps): the other half of latency
offcputime-bpfcc -df -p $(pgrep -o node) 30 | ./flamegraph.pl --color=io > offcpu.svg

Off-CPU analysis matters because a slow request often spends most of its time waiting, not computing; an on-CPU flame graph cannot show that. Continuous profiling (Parca, Pyroscope, Grafana) keeps low-rate profiles always on, so you can compare before and after a deploy.

READING A FLAME GRAPH
stack samples merged into one picture
perf record -F 99 -gstack samplesstackcollapsemerge identical stacksflamegraph.plSVGwide boxon CPU oftentop edgerunning the CPU itself
swipe the figure sideways, or tap expand for full screen
1/4
sample
perf samples every CPU's stack 99 times a second (an odd rate avoids lockstep with timers). Thirty seconds gives thousands of samples.
99 Hz stack sampleslow overhead