Site Reliability Engineering
Site reliability engineering is what happens when you treat operations as a software problem and reliability as a number. Its central idea is that 100% is the wrong target: users cannot tell 99.99% from 100%, every extra nine costs more than the last, and the gap between the target and perfect is a budget the team can spend on shipping. SLOs make that budget explicit, error budgets make it enforceable, and everything else in this course (alerts, incidents, postmortems, release safety, toil) is how a team lives within it.
Twelve parts. What SRE is and the contract it makes; SLIs, SLOs and error budgets with burn-rate maths; alerting on symptoms with a sane paging policy; incident response roles, severity and communication; blameless postmortems with action items that get done; capacity planning, load testing and release safety; measuring and removing toil; and the frontend's share: client SLOs, crash-free rate, and what should page a frontend engineer. Four deeper parts follow: observability for operators, dependencies and graceful degradation, disaster recovery and game days, and production readiness with on-call health.