🛠️ Site Reliability Engineering Foundations: Sub-Curriculum Index

Welcome to the SRE Foundations & Operations curriculum. This track covers the core disciplines of Site Reliability Engineering: quantitative reliability targeting (SLIs/SLOs/Error Budgets), structured incident command, toil reduction, capacity planning, and chaos engineering.

Every guide in this series strictly follows a two-part learning format:

  • ⚡ Quick Dive: Architecture cheat sheets, severity classification matrices, and SLO calculations.
  • 📖 Extended Guide: Deep operational frameworks, postmortem templates, load testing scripts, and fault-injection manifests.

📚 Curriculum Roadmap

# Guide Primary Topics Covered
01 SRE Principles & DevOps Alignment SRE core tenets, DevOps alignment, error budgeting, automation mindset, and risk management.
02 SLIs, SLOs & Error Budgets Quantitative reliability metrics, the "Nines" of uptime, multi-window burn-rate alerting, and freeze policies.
03 Incident Management & Postmortems SEV-1/2/3 severity tiers, Incident Commander protocols, blameless postmortem culture, and the 5-Whys RCA.
04 Toil Reduction & Automation Identifying operational toil, the SRE 50% engineering cap, and transitioning runbooks into self-healing loops.
05 Capacity Planning & Load Testing Resource forecasting, stress/spike testing, and distributed performance testing using k6 and Locust.
06 Chaos Engineering & GameDays Principles of Chaos, blast-radius mitigation, Chaos Mesh Kubernetes manifests, and organizing GameDays.