Part 0 · 1 chapters · ~8 min
Methodology: USE, Workload Characterisation, the 60-Second Checklist
Anti-methods (streetlight, random change, blame someone else), the USE method for resources, the RED method for services, workload characterisation (who, why, what, how), latency analysis by drilling down, and Netflix's 60-second checklist on a fresh Linux box.
1
A method before tools
code
# the first 60 seconds on a slow Linux host (after Brendan Gregg / Netflix) uptime # load averages: rising or falling? dmesg -T | tail # OOM kills, TCP drops, disk errors vmstat 1 # r (run queue), free, si/so (swap), us/sy/wa/st CPU mpstat -P ALL 1 # one hot CPU? (a single-threaded bottleneck) pidstat 1 # which processes use CPU iostat -xz 1 # disk r/s, w/s, await, aqu-sz free -m # available memory, page cache sar -n DEV 1 # NIC throughput sar -n TCP,ETCP 1 # new connections, retransmits top # a final overview
| workload characterisation question | example answer |
|---|---|
| who is causing the load? | one tenant's export job (PID, user, client IP, API key) |
| why is it called? | a dashboard refreshing every 5 seconds |
| what is the load? | large range scans, 40 MB responses |
| how is it changing? | doubled since the release on Tuesday |
Often the biggest win is eliminating unnecessary work found by characterisation, not tuning the system to do it faster.
THE USE METHOD
for every resource: utilisation, saturation, errors
swipe the figure sideways, or tap expand for full screen
1/4
list resources
Start from a checklist of every resource: CPUs, memory, each disk and NIC, interconnects, and software resources such as thread pools, connection pools and locks.
every resource, in turnhardware and software