Part 8 · 1 chapters · ~8 min

Performance and Profiling

pprof for CPU, heap, goroutines, mutexes and blocking, benchmarks with allocation counts, reducing allocations (escape analysis, sync.Pool, preallocation), strings and bytes, JSON encoding costs and faster alternatives, profile-guided optimisation, and tracing with go tool trace.

14

Profiles and allocations

code
import _ "net/http/pprof"     // exposes /debug/pprof on the default mux (serve it on an internal port only)
go func() { log.Println(http.ListenAndServe("127.0.0.1:6060", nil)) }()

go tool pprof -http=:8081 http://127.0.0.1:6060/debug/pprof/profile?seconds=30   # CPU flame graph
go tool pprof http://127.0.0.1:6060/debug/pprof/heap                            # in-use memory
curl -s http://127.0.0.1:6060/debug/pprof/goroutine?debug=2 | head              # every goroutine's stack
go test -bench=. -benchmem -cpuprofile cpu.out ./ledger                          # profile a benchmark
go build -pgo=default.pgo                                                        # profile-guided optimisation (1.21+)
common costfix
many small allocations in hot loopspreallocate slices (make([]T, 0, n)), reuse buffers, sync.Pool
string concatenation in loopsstrings.Builder
encoding/json reflection costfewer fields, json/v2 or code-generated encoders for hot paths
lock contentionmutex profile; shard or remove the lock
too many goroutines blockedgoroutine profile; bounded worker pools