Part 8 · 1 chapters · ~8 min
Performance and Profiling
pprof for CPU, heap, goroutines, mutexes and blocking, benchmarks with allocation counts, reducing allocations (escape analysis, sync.Pool, preallocation), strings and bytes, JSON encoding costs and faster alternatives, profile-guided optimisation, and tracing with go tool trace.
14
Profiles and allocations
code
import _ "net/http/pprof" // exposes /debug/pprof on the default mux (serve it on an internal port only)
go func() { log.Println(http.ListenAndServe("127.0.0.1:6060", nil)) }()
go tool pprof -http=:8081 http://127.0.0.1:6060/debug/pprof/profile?seconds=30 # CPU flame graph
go tool pprof http://127.0.0.1:6060/debug/pprof/heap # in-use memory
curl -s http://127.0.0.1:6060/debug/pprof/goroutine?debug=2 | head # every goroutine's stack
go test -bench=. -benchmem -cpuprofile cpu.out ./ledger # profile a benchmark
go build -pgo=default.pgo # profile-guided optimisation (1.21+)| common cost | fix |
|---|---|
| many small allocations in hot loops | preallocate slices (make([]T, 0, n)), reuse buffers, sync.Pool |
| string concatenation in loops | strings.Builder |
| encoding/json reflection cost | fewer fields, json/v2 or code-generated encoders for hot paths |
| lock contention | mutex profile; shard or remove the lock |
| too many goroutines blocked | goroutine profile; bounded worker pools |