Part 5 · 1 chapters · ~8 min

Site Reliability Engineering

The Google SRE book and workbook reduced to the ideas that spread: software engineering for operations with a toil cap, embracing risk, SLOs and error budgets, eliminating toil, symptom-based monitoring and alerting, and incident command with blameless postmortems.

7

Site Reliability Engineering

The book is free to read online at sre.google, along with the workbook. The SRE course in Learna is these two books applied to a fintech, frontend included.

book chaptersread forLearna
3 Embracing risk, 4 SLOsthe contract and the budgetSRE P0-P1
5 Eliminating toilthe definition and the 50% capSRE P6
6 Monitoring distributed systemsgolden signals, symptoms versus causesSRE P2, Cloud P8
14 Managing incidents, 15 Postmortem cultureroles, blamelessnessSRE P3-P4
21-22 Overload and cascading failuresload shedding, retries, deadlinesDistributed Systems P0, this course P9
Workbook 5: Alerting on SLOsmulti-window burn ratesSRE P2
where it has aged
Some practices assume Google's scale and staffing, such as a dedicated SRE team per service and handing services back to developers. Most companies adopt the practices (SLOs, budgets, blameless reviews) without the org structure, and the workbook is the more practical of the two volumes.
SITE RELIABILITY ENGINEERING (THE GOOGLE BOOK)
Beyer, Jones, Petoff and Murphy (eds.), 2016; plus The Site Reliability Workbook, 2018
swipe the figure sideways, or tap expand for full screen
1/6
SRE is software engineering for operations
SRE as software engineering: Google's answer to the developer-operations split was to staff operations with software engineers, cap their operational load at half their time, and spend the rest building automation. The cap forces engineering investment when load grows.