Part 0 · 2 chapters · ~12 min
The Diagnostic Method
Symptom, scope, hypothesis, evidence, fix and verify; the USE and RED methods as checklists; correlating with changes; avoiding the common traps (fixing the loudest error, trusting averages, changing several things at once); and keeping an investigation log.
1
Six steps
Debugging production is a search. The method keeps it from becoming a random walk: every step should shrink the space of possible causes, and every claim should be backed by a measurement someone else could repeat.
THE DIAGNOSTIC METHOD
from a symptom to a proven cause, without guessing
swipe the figure sideways, or tap expand for full screen
1/5
symptom
Start from what users feel and a number: p99 latency on POST /transfers went from 300 ms to 4 s at 14:05. Not "the database is slow".
a user-visible symptom with a number and a timenot a guessed cause
2
Checklists and traps
| trap | instead |
|---|---|
| fixing the loudest error in the logs | check whether it was there before the incident started |
| trusting averages | percentiles and per-host breakdowns |
| changing three things at once | one change, one measurement |
| debugging in production by restarting | capture evidence first (a heap dump, a profile, a thread dump), then restart |
| assuming the newest dependency is guilty | let the timeline and the scope decide |
Use USE (utilisation, saturation, errors) for every resource and RED (rate, errors, duration) for every service as checklists (SRE part 8), and keep a timestamped investigation log in the incident channel: it becomes the postmortem timeline.