Part 0 · 1 chapters · ~8 min

Methodology: USE, Workload Characterisation, the 60-Second Checklist

Anti-methods (streetlight, random change, blame someone else), the USE method for resources, the RED method for services, workload characterisation (who, why, what, how), latency analysis by drilling down, and Netflix's 60-second checklist on a fresh Linux box.

1

A method before tools

code
# the first 60 seconds on a slow Linux host (after Brendan Gregg / Netflix)
uptime                    # load averages: rising or falling?
dmesg -T | tail           # OOM kills, TCP drops, disk errors
vmstat 1                  # r (run queue), free, si/so (swap), us/sy/wa/st CPU
mpstat -P ALL 1           # one hot CPU? (a single-threaded bottleneck)
pidstat 1                 # which processes use CPU
iostat -xz 1              # disk r/s, w/s, await, aqu-sz
free -m                   # available memory, page cache
sar -n DEV 1              # NIC throughput
sar -n TCP,ETCP 1         # new connections, retransmits
top                       # a final overview
workload characterisation questionexample answer
who is causing the load?one tenant's export job (PID, user, client IP, API key)
why is it called?a dashboard refreshing every 5 seconds
what is the load?large range scans, 40 MB responses
how is it changing?doubled since the release on Tuesday

Often the biggest win is eliminating unnecessary work found by characterisation, not tuning the system to do it faster.

THE USE METHOD
for every resource: utilisation, saturation, errors
list resourcesCPU, memory, disks, NICs, locks, poolsutilisation% busysaturationqueued workerrorserror countsnext resourceor drill down
swipe the figure sideways, or tap expand for full screen
1/4
list resources
Start from a checklist of every resource: CPUs, memory, each disk and NIC, interconnects, and software resources such as thread pools, connection pools and locks.
every resource, in turnhardware and software