🛠️ Site Reliability Engineering Foundations: Sub-Curriculum Index
Welcome to the SRE Foundations & Operations curriculum. This track covers the core disciplines of Site Reliability Engineering: quantitative reliability targeting (SLIs/SLOs/Error Budgets), structured incident command, toil reduction, capacity planning, and chaos engineering.
Every guide in this series strictly follows a two-part learning format:
- ⚡ Quick Dive: Architecture cheat sheets, severity classification matrices, and SLO calculations.
- 📖 Extended Guide: Deep operational frameworks, postmortem templates, load testing scripts, and fault-injection manifests.
📚 Curriculum Roadmap
| # | Guide | Primary Topics Covered |
|---|---|---|
| 01 | SRE Principles & DevOps Alignment | SRE core tenets, DevOps alignment, error budgeting, automation mindset, and risk management. |
| 02 | SLIs, SLOs & Error Budgets | Quantitative reliability metrics, the "Nines" of uptime, multi-window burn-rate alerting, and freeze policies. |
| 03 | Incident Management & Postmortems | SEV-1/2/3 severity tiers, Incident Commander protocols, blameless postmortem culture, and the 5-Whys RCA. |
| 04 | Toil Reduction & Automation | Identifying operational toil, the SRE 50% engineering cap, and transitioning runbooks into self-healing loops. |
| 05 | Capacity Planning & Load Testing | Resource forecasting, stress/spike testing, and distributed performance testing using k6 and Locust. |
| 06 | Chaos Engineering & GameDays | Principles of Chaos, blast-radius mitigation, Chaos Mesh Kubernetes manifests, and organizing GameDays. |