Overview
Curated: · Written: · Reviewed:
Reliability objectives are product contracts implemented as control loops
A service level indicator is a measured representation of behavior users care about. A service level objective is a target for that indicator over a declared population and time window. An error budget is the allowed amount of nonconforming service implied by the target, together with an agreed policy for what happens as it is consumed. These are different from an SLA: an SLA is a business or legal agreement and may define remedies, exclusions and measurement rules that do not match an internal engineering SLO.
Interviewers open here almost every time: "What's the difference between an SLI, an SLO, and an SLA?" The weak answer recites the acronyms and stops. The strong answer distinguishes them by what they commit you to — an SLI is a measurement, an SLO is an internal target with consequences, an SLA is a contract with penalties — and then says the part that matters: an SLO without a policy attached is just a dashboard gauge, not an objective.
Start from the user journey, not the available metric
Begin with a user journey or decision, not an available metric. Define the service boundary, consumer class, valid event population, good-event condition, measurement point, exclusions, target and window. For request availability, an SLI can be good eligible requests divided by valid eligible requests. For latency, define a threshold and count events below it; an average hides the tail. For pipelines, freshness or correctness may be more useful than HTTP availability. Separate interactive and batch workloads when their expectations differ.
The follow-up you should expect: "Why is an SLI a ratio rather than a raw count or a gauge?" Because a raw count of errors is meaningless without the denominator — 500 failures is catastrophic at 10k requests and invisible at 50M. The ratio of good events to valid events is what makes the indicator comparable across windows and traffic levels. A weak answer names a metric ("we track 500s in Grafana") without ever defining the population it is a fraction of.
On latency specifically, expect "why not average latency?" The answer to give: a thresholded SLI — the fraction of requests faster than, say, 300 ms — or a percentile target like p99, because averages collapse the tail. If 1% of requests take 4 seconds and the rest take 50 ms, the average looks fine and the p99 does not. Users who abandon are in the tail, not the mean.
The measurement point changes truth
Server-side success can disagree with load balancer, client or synthetic experience because requests can fail before reaching the service, after its response, or within a dependency. Choose the closest trustworthy proxy for user experience, document its blind spots, and compare independent viewpoints. Telemetry loss is unknown, not good. Protect the denominator from invalid traffic without excluding real failures merely because they are inconvenient.
A classic interview probe: "Your server-side availability is 99.98% but support tickets are up. What's going on?" Candidates who only know the definitions stall. Candidates who have operated SLOs enumerate the gaps: clients timing out before the server responds, requests failing at the LB, 200 responses with garbage bodies, and telemetry dropouts counted as success. The measurement point is part of the SLI specification, and moving it changes the number even when the service is unchanged.
Choosing targets and windows
Targets require product judgment about user harm, dependency capabilities, cost, staffing, architecture and innovation. Do not copy a fashionable number or set the target equal to current performance. A 100% objective provides no budget for ordinary change and can demand disproportionate cost; extremely high objectives may exhaust their budget faster than detection and mitigation are physically possible. Use separate objectives for distinct critical journeys, but keep the set small enough to understand and operate.
Two follow-ups to be ready for:
- "Why isn't 99.99% strictly better than 99.9%?" Because a target tighter than users can perceive is reliability spend with no user-visible return, and each additional nine roughly multiplies the required engineering. If users notice nothing until the error rate hits 0.5%, a 99.99% objective buys nothing except a budget too small to release through.
- "Rolling or calendar window?" Rolling windows continually expire old events and represent recent experience; calendar windows align to business reporting but reset abruptly. Both have boundary effects. State which model is used, because the same incident consumes different amounts of each.
Error budget arithmetic
For a good-event ratio target S over a window, the allowed bad fraction is 1 − S. A 99.9% target therefore permits 0.1% bad eligible events, not universally forty-three minutes: time-based downtime is only equivalent when traffic and measurement semantics support that interpretation. Request-based budgets weight busy periods more heavily; time-based availability weights intervals. State which model is used. Compute consumed and remaining budget from the same versioned numerator, denominator, target, and window.
The "43 minutes" trap is a standard interview filter. The weak answer converts 99.9% to 43.2 minutes of downtime per month by reflex. The strong answer says: that conversion assumes uniform traffic and a time-based SLI. Under a request-based SLI, the budget is a number of bad events, and an outage at peak burns it orders of magnitude faster than the same wall-clock outage overnight.
Burn rates and alerting
Burn rate is observed bad-event rate divided by the allowed bad-event rate. A burn rate of one would consume the budget exactly over the objective window; faster rates exhaust it sooner. Multiwindow, multi-burn-rate alerting combines a fast window with a longer confirmation window for rapid, sustained incidents, and slower windows for gradual budget threats. Page only when a human must act now to protect users or budget. Use tickets for slower actionable work and dashboards/logs for investigation. Test alert precision, recall, detection time and reset behavior.
The error budget policy is what makes it real
An error-budget policy turns measurement into a shared decision rule. It should define owners, thresholds, release or rollout constraints, reliability work, incident review, exception classes, approval, expiration and escalation. It is not punishment, an outage allowance to spend deliberately, or permission to ignore individual catastrophic incidents while the aggregate remains green. Security fixes and urgent mitigations may need explicit exception handling. Freeze risky change based on agreed evidence and restore normal delivery using declared exit criteria.
This is where the interview usually lands, because it is the part that separates reading the Google SRE book from running the loop: "What actually happens when the budget is exhausted?" The weak answer is "we monitor it closely." The strong answer names the mechanism: feature freezes or rollout slowdowns, reprioritization of reliability work over features, an exception process with a named approver, and declared exit criteria for resuming normal delivery. The budget is the explicit trade of release velocity against reliability — when it's spent, velocity is what gets cut. An SLO with no agreed consequence is theatre: a number on a dashboard that degrades quietly because nothing in the organisation is arranged to notice.
Dependencies and segmentation
Budgets across dependencies do not simply add or multiply without a traffic and failure model. Map critical paths, fanout, fallback, correlated failures and dependency objectives; otherwise a service can promise more than its chain can deliver. Internal SLOs can be tighter than an external commitment to preserve margin, but reliability cannot be created by arithmetic alone.
Segment enough to reveal a harmed region, tenant or operation, while also retaining the aggregate contract. A healthy global SLO must not hide a severe low-volume cohort, and a tiny cohort should not automatically create an unactionable global page. Pair the primary objective with diagnostic slices and explicit minimum-volume or synthetic strategies.
Operating the loop
Treat SLOs as versioned production configuration. Record rationale, owner, applicable version, data query, labels, threshold, window, exclusions, missing-data policy and change history. Backtest candidates against incidents and normal periods, shadow new alerts, reconcile raw events to aggregates, test late data and telemetry outages, and audit exclusion growth. Review objectives when users, architecture, volume or business consequences change — without relaxing them merely to erase a miss.
Operate the complete loop: measure, compare, decide, act and learn. Dashboards should show current compliance, remaining budget, burn rates, affected journeys and data freshness. Incidents and escaped user harm should update instrumentation, definitions or targets when evidence shows they are poor proxies. A green SLO is evidence about the declared events and window, not proof that the service is healthy in every dimension — expect the interviewer to probe exactly that: "Your SLO is green. Why are users still complaining?" The answer is the population, the window, the measurement point, or the slice — one of them is hiding the harm.
Where objectives usually go wrong
The most common failure is an objective that nobody would act on. An availability target chosen because it looked responsible, with no agreed consequence when the budget is exhausted, is a number on a dashboard rather than a control loop. The test for a real objective is whether someone can say what changes when the budget runs out — which deployments pause, which work is reprioritised, and who decides — and whether that has ever actually happened.
The second failure is measuring the wrong thing. An indicator computed from server-side success codes misses the requests that never arrived, the clients that timed out before the server replied, and the responses that were technically successful and useless. Defining the indicator from the user's perspective — including what counts as an eligible event, and what happens to requests the system never saw — is what makes the objective correspond to experience rather than to internal health.
The third is arithmetic that hides the tail. Averaging availability across a long window lets a total outage disappear into a month of good hours, and averaging latency across endpoints lets a slow critical path hide behind a fast static one. Objectives should be stated per user journey, evaluated over windows short enough to detect a sustained problem, and paired with burn-rate alerting so that a fast burn pages immediately while a slow burn creates work rather than an interruption.
Worked example: 99.9% is 43,200 bad events, not 43 minutes
30-day rolling, request-based, 43.2M eligible events. Target 99.9% → budget 43,200 bad.
| incident | bad events | remaining budget | page? |
|---|---|---|---|
| 2 h overnight at 2 rps | 14,400 | 28,800 | ticket (slow burn) |
| 12 min peak at 500 rps | 360,000 | exhausted 8× | page now |
The conversion to wall-clock minutes assumes uniform traffic. Peak is where the budget actually lives.
