Part 0 · 2 chapters · ~12 min

The Diagnostic Method

Symptom, scope, hypothesis, evidence, fix and verify; the USE and RED methods as checklists; correlating with changes; avoiding the common traps (fixing the loudest error, trusting averages, changing several things at once); and keeping an investigation log.

1

Six steps

Debugging production is a search. The method keeps it from becoming a random walk: every step should shrink the space of possible causes, and every claim should be backed by a measurement someone else could repeat.

THE DIAGNOSTIC METHOD
from a symptom to a proven cause, without guessing
symptomp99 latency 4 s on /transfersscopewhich hosts, routes, tenants, since when?hypothesispool exhaustion?evidencemetrics → traces → profilesfixsmallest changeverifysame measurement, before/after
swipe the figure sideways, or tap expand for full screen
1/5
symptom
Start from what users feel and a number: p99 latency on POST /transfers went from 300 ms to 4 s at 14:05. Not "the database is slow".
a user-visible symptom with a number and a timenot a guessed cause
2

Checklists and traps

trapinstead
fixing the loudest error in the logscheck whether it was there before the incident started
trusting averagespercentiles and per-host breakdowns
changing three things at onceone change, one measurement
debugging in production by restartingcapture evidence first (a heap dump, a profile, a thread dump), then restart
assuming the newest dependency is guiltylet the timeline and the scope decide

Use USE (utilisation, saturation, errors) for every resource and RED (rate, errors, duration) for every service as checklists (SRE part 8), and keep a timestamped investigation log in the incident channel: it becomes the postmortem timeline.