CPU Profiling Across Runtimes
Sampling versus instrumenting profilers, Linux perf and perf maps for JIT runtimes, Go pprof, async-profiler and JFR for the JVM, py-spy for Python, Node's profilers, rbspy and stackprof for Ruby, flame graphs and differential flame graphs, and on-CPU versus off-CPU time.
Profilers by runtime
# Go curl -s localhost:6060/debug/pprof/profile?seconds=30 > cpu.pb.gz && go tool pprof -http=:8081 cpu.pb.gz # JVM asprof -d 30 -e cpu -f flame.html <pid> # async-profiler; -e alloc, -e lock, -e wall # Python py-spy record -o flame.svg --pid <pid> --duration 30 # any native or JIT process on Linux perf record -F 99 -g -p <pid> -- sleep 30 && perf script | stackcollapse-perf.pl | flamegraph.pl > flame.svg
On-CPU, off-CPU and differential profiles
A CPU profile shows where time is spent running. Requests that are slow because they wait (on locks, IO, the database, the threadpool) do not appear in it. For those, use off-CPU or wall-clock profiles (async-profiler -e wall, Go block and mutex profiles, offcputime from bcc), or traces.
Differential flame graphs compare two profiles (before and after a deploy): frames that grew are coloured red. They answer "what got slower in this release?" faster than reading two graphs side by side. Continuous profilers (Pyroscope, Parca, cloud profilers) keep a profile history so the comparison is always available.