12 parts · 16 chapters

Site Reliability Engineering

Site reliability engineering is what happens when you treat operations as a software problem and reliability as a number. Its central idea is that 100% is the wrong target: users cannot tell 99.99% from 100%, every extra nine costs more than the last, and the gap between the target and perfect is a budget the team can spend on shipping. SLOs make that budget explicit, error budgets make it enforceable, and everything else in this course (alerts, incidents, postmortems, release safety, toil) is how a team lives within it.

Twelve parts. What SRE is and the contract it makes; SLIs, SLOs and error budgets with burn-rate maths; alerting on symptoms with a sane paging policy; incident response roles, severity and communication; blameless postmortems with action items that get done; capacity planning, load testing and release safety; measuring and removing toil; and the frontend's share: client SLOs, crash-free rate, and what should page a frontend engineer. Four deeper parts follow: observability for operators, dependencies and graceful degradation, disaster recovery and game days, and production readiness with on-call health.

the contract · SLOs and error budgets · alerting · incidents · postmortems · capacity and release safety · toil · the frontend share · observability · degradation · disaster recovery · readinesssenior → staff · every engineer who carries a pager or decides what ships
the contractReliability versus velocity, settled with a number both sides agree to.
SLIs, SLOs, budgetsChoosing indicators users feel, setting objectives, burn rates and budget policies.
alertingSymptom-based alerts, multi-window burn-rate pages, and an end to alert fatigue.
incident responseCommander, operations, communications; severity levels; status pages; runbooks.
postmortemsBlameless practice, contributing factors over root causes, and action items that close.
capacity and releasesForecasting, load testing, canaries, rollbacks, and freezes tied to the budget.
toilWhat counts as toil, measuring it, and what to automate first.
the frontend shareClient SLOs, crash-free sessions, and the pages a frontend engineer should get.
Built on Cloud, Infra and Distributed SystemsThis course assumes the earlier infrastructure courses for the mechanics (observability, deployment, failure domains) and gives them a purpose: a reliability target users feel and a team can afford. The Big-company FE course part 8 is the frontend on-call view; part 7 here connects them.