Reliability & Observability
Keeping production systems healthy: SLOs, the three pillars of observability, incident response, and chaos engineering.
Subject: Operations & Reliability · Roles: DevOps Engineer, Site Reliability Engineer
Concepts
SLIs, SLOs & Error Budgets
The vocabulary reliability is measured in, and how an error budget turns 'be reliable' into a concrete number.
Monitoring, Logging & Tracing
The three pillars of observability, and what question each one is actually good at answering.
Incident Management & Postmortems
Running an incident from detection to resolution, and writing a blameless postmortem that actually prevents a repeat.
Capacity Planning & Load Testing
Knowing how much traffic a system can take before it does, through modeling and deliberate load testing.
Chaos Engineering
Deliberately injecting failure into a system to verify it actually survives what you assume it survives.
On-Call & Alerting Design
Designing alerts that page for what actually needs a human, and structuring on-call sustainably.
