Part 1 · 2 chapters · ~18 min

SLIs, SLOs and Error Budgets

Choosing SLIs from user journeys: availability, latency as a proportion, freshness and correctness, where to measure and how to write them down; setting targets from data; then error budgets over the window, burn rates, time to exhaustion, fast and slow burns, and a policy enforced by tools.

2

Choosing SLIs and setting SLOs

good events over valid events, for journeys users feel
  1. Start from journeys, not from services.
  2. Availability: good over valid. Client-caused 4xx responses count as good.
  3. Latency: the proportion of requests under a threshold.
  4. Freshness and correctness for asynchronous work and data pipelines.
  5. Measure at the load balancer, and check against the client and synthetic probes.
  6. Write it down precisely: numerator, denominator, source, exclusions.
code
# an SLO as code (OpenSLO-style), generated into recording rules and alerts
apiVersion: openslo/v1
kind: SLO
metadata: { name: transfer-availability }
spec:
  service: payments
  description: Customers can send money to another bank
  budgetingMethod: Occurrences
  timeWindow: [{ duration: 28d, isRolling: true }]
  objectives:
    - target: 0.999
      ratioMetric:
        good:  { metricSource: { type: Prometheus, spec: { query: 'sum(rate(lb_requests_total{route="/v1/transfers",code!~"5..",synthetic="false"}[{{window}}]))' } } }
        total: { metricSource: { type: Prometheus, spec: { query: 'sum(rate(lb_requests_total{route="/v1/transfers",synthetic="false"}[{{window}}]))' } } }
setting the target
Start from what the service achieves today (from the last 90 days of data), set the SLO a little below it, and tighten it as reliability improves. An SLO set above current performance is exhausted on day one and teaches everyone to ignore it. Review targets quarterly with product.
CHOOSING SLIS
indicators that measure what users feel, measured where users feel it, as good events over valid events
swipe the figure sideways, or tap expand for full screen
1/6
journeys
Start from the journey, not the service: "a customer sends money to another bank" touches the app, the API gateway, payments, the ledger and a payment provider. Users feel the journey. List the critical journeys first (sign in, fund wallet, transfer, pay a bill), then the services each depends on.
3

Error budgets and burn rates

the rate matters more than the total
  1. The budget over the window must not go below zero.
  2. Burn rate is the observed error rate divided by the allowed error rate.
  3. Time to exhaustion is the window divided by the burn rate.
  4. A fast burn needs a human now.
  5. A slow burn needs a ticket this week.
  6. The policy is enforced by tools, not argued about.
code
// burn-rate arithmetic, the way you'd check it in a review
const slo = 0.999, allowed = 1 - slo;               // 0.001
const burn = (errorRate: number) => errorRate / allowed;
const hoursToExhaust = (b: number, windowDays = 28) => (windowDays * 24) / b;
const budgetSpentInHours = (b: number, h: number, windowDays = 28) => (b * h) / (windowDays * 24);

burn(0.05);                       // 50
hoursToExhaust(50);               // 13.4 h
budgetSpentInHours(50, 1);        // 0.074: 7.4% of the month in one hour
budgetSpentInHours(14.4, 1);      // 0.021: the classic 2%-per-hour page threshold
ERROR BUDGETS AND BURN RATES
how fast the budget is being spent, and why the rate matters more than the total
swipe the figure sideways, or tap expand for full screen
1/6
the budget
The budget over the window: 99.9% over 28 days gives a budget of 0.1% of requests. Plot the remaining budget over time: a steady background error rate slopes it down gently; an incident drops it sharply; the line must stay above zero at the end of the window.