Part 0 · 1 chapters · ~10 min

What SRE Is

The contract between reliability and velocity: why 100% is the wrong target, the SLO as a user-visible promise, the error budget as permission to ship, the policy agreed in advance, and what SRE adds around it.

1

What SRE is: the contract

reliability as a number both sides agree to
  1. The old conflict: features against uptime, with no number to settle it.
  2. Why not 100%: users cannot tell the difference, and the target freezes all change.
  3. The SLO: a user-visible indicator, a target and a window.
  4. The error budget: 100% minus the SLO, spent by any cause.
  5. The policy: ship while budget remains, fix reliability once it is spent. Agree this in advance.
  6. Around it: automation, capped toil, shared on-call and blameless learning.
SLObudget per 28 days (time)per 1 M requestswhat it usually needs
99%~6 h 43 min10,000 failuresone zone, manual recovery
99.5%~3 h 22 min5,000health checks, automatic restarts
99.9%~40 min1,000multi-zone, automated rollback, on-call
99.95%~20 min500canaries, fast detection, tested failover
99.99%~4 min100cells, multi-region, no single manual step in recovery
the conversation it replaces
"Can we ship the new checkout on Friday?" stops being about nerves and becomes a lookup. Is there budget? Is it burning? What does the canary say? The SRE book's phrase is that the error budget aligns incentives: developers now want reliability, because it buys them the right to ship.
THE CONTRACT BETWEEN RELIABILITY AND VELOCITY
why 100% is wrong, what the gap buys, and how a number ends the argument between shipping and stability
swipe the figure sideways, or tap expand for full screen
1/6
the old conflict
The old conflict: developers are rewarded for features, operators for uptime. Every release is a risk to the operator and a deliverable to the developer. The result is release committees, change freezes, and blame after every incident, with nobody able to say how reliable is reliable enough.