Part 8 · 1 chapters · ~11 min
Observability for Operators
Metrics, logs and traces and the questions each answers, the RED method for services and the USE method for resources, structured correlated logs, traces propagated from the browser, and why high-cardinality labels belong in logs and traces rather than metrics.
13
Three signals, two methods, one bill
code
# RED for a service, in PromQL
sum(rate(http_requests_total{service="transfers"}[5m])) # Rate
sum(rate(http_requests_total{service="transfers", status=~"5.."}[5m])) # Errors
histogram_quantile(0.99, sum by (le) (rate(http_request_duration_seconds_bucket{service="transfers"}[5m]))) # Duration p99
# USE for a connection pool
db_pool_in_use / db_pool_size # Utilisation
db_pool_waiting # Saturation: alert when > 0 for 5m
rate(db_pool_acquire_errors_total[5m]) # Errors| question | signal | why |
|---|---|---|
| Is the service healthy right now? | metrics (RED, SLO burn) | cheap, fast, aggregated |
| Why did this customer's transfer fail? | logs, by request id | the specific event and its fields |
| Why are some requests slow? | traces, filtered to slow ones | which span took the time |
| Is the database pool the bottleneck? | metrics (USE) | saturation before failure |
| What happened in the browser before the error? | RUM session events plus the trace id | the Frontend's share, part 6 |
the frontend's link
Start the trace in the browser: send a
traceparent header from the client, and log the trace id with client errors. One id then follows a tap through the BFF and the ledger, and a user's complaint becomes a single search.OBSERVABILITY FOR OPERATORS
metrics, logs and traces, the RED and USE methods, and the cardinality bill
swipe the figure sideways, or tap expand for full screen
1/6
metrics
Metrics: counters, gauges and histograms with labels, scraped or pushed every few seconds and kept for months. Cheap to store and query because they are pre-aggregated. Latency must be a histogram, not an average: averages hide the slow tail that users feel.