Part 2 · 2 chapters · ~12 min

CPU Profiling Across Runtimes

Sampling versus instrumenting profilers, Linux perf and perf maps for JIT runtimes, Go pprof, async-profiler and JFR for the JVM, py-spy for Python, Node's profilers, rbspy and stackprof for Ruby, flame graphs and differential flame graphs, and on-CPU versus off-CPU time.

5

Profilers by runtime

code
# Go
curl -s localhost:6060/debug/pprof/profile?seconds=30 > cpu.pb.gz && go tool pprof -http=:8081 cpu.pb.gz
# JVM
asprof -d 30 -e cpu -f flame.html <pid>          # async-profiler; -e alloc, -e lock, -e wall
# Python
py-spy record -o flame.svg --pid <pid> --duration 30
# any native or JIT process on Linux
perf record -F 99 -g -p <pid> -- sleep 30 && perf script | stackcollapse-perf.pl | flamegraph.pl > flame.svg
CPU PROFILERS, BY RUNTIME
sampling profilers that are safe enough for production
Linux perfKernel-level sampling of anyprocess, native frames; JITruntimes need perf maps.Go pprofnet/http/pprof endpoints: CPU,heap, goroutines, mutex and blockprofiles.async-profilerJVM: CPU, allocations and lockswithout safepoint bias; flamegraph output.py-spyPython: attach to a runningprocess, no code changes, top-likeand flame graph modes.Node --cpu-profV8 sampling profiles, clinic.jsflame and doctor, 0x.rbspy / stackprofRuby: rbspy attaches externally;stackprof in-process sampling.
swipe the figure sideways, or tap expand for full screen
1/6
perf
perf samples stacks at a frequency (99 Hz avoids lockstep) for any process, including the kernel. For JIT languages, enable perf maps (--perf-basic-prof in Node, -XX:+PreserveFramePointer and perf-map-agent for the JVM) to name JIT frames.
samples anything, including the kernelJIT runtimes need perf maps
6

On-CPU, off-CPU and differential profiles

A CPU profile shows where time is spent running. Requests that are slow because they wait (on locks, IO, the database, the threadpool) do not appear in it. For those, use off-CPU or wall-clock profiles (async-profiler -e wall, Go block and mutex profiles, offcputime from bcc), or traces.

Differential flame graphs compare two profiles (before and after a deploy): frames that grew are coloured red. They answer "what got slower in this release?" faster than reading two graphs side by side. Continuous profilers (Pyroscope, Parca, cloud profilers) keep a profile history so the comparison is always available.