Top 100 Site Reliability Engineer Interview Questions and Answers
The questions most likely to actually come up in your Site Reliability Engineer interview, ranked by likelihood — with detailed, senior-level answers covering what an interviewer is really listening for.
Curated: · Written: · Reviewed:
QA-1A product team wants to SLO on CPU idle and HTTP 200 rate. Which measurements belong in the SLO, and which do not?(show answer)
The first thing to settle about choosing user-journey SLIs is what decision it actually changes.
Bind every SLO to a completed user journey, because a green host metric can hide a checkout that never finishes.
Concretely, name the journeys that lose money or trust, instrument success at the last hop the user waits on, and keep resource gauges as diagnostic signals rather than as the SLO itself.
The reason for that specificity is a failure I have seen: Paging on CPU idle while 18% of payment confirmations stall at the bank callback produced a week of quiet dashboards and a 12% drop in completed orders.
Map each candidate SLI to the journey it does or does not observe.
| Candidate | User sentence it measures | Keep as SLO? |
|---|---|---|
| CPU idle > 40% | none | no, diagnostic only |
| HTTP 200 on /healthz | probe can reach the process | no |
| checkout_ok / checkout_started | buyer sees an order id | yes |
| p99 render of search results | shopper sees a result list | yes |
I would not consider it settled without evidence: For each proposed SLI, write the user sentence it measures and reject any candidate that still holds when that sentence is false.
An SLI that a user cannot feel is a dashboard hobby, not a reliability contract.
Curated: · Written: · Reviewed:
QA-2A 99.9% availability SLO is green while shoppers abandon a 2.4-second search page. How do you split the objectives?(show answer)
I would start by making the assumption behind availability versus latency SLOs explicit.
Give availability and latency separate SLOs on the same journey, because a late success is a different product failure from an error.
Concretely, define availability as the proportion of valid attempts that succeed, and latency as the proportion of valid attempts completed within a product-defined threshold; maintain separate error budgets for each.
The reason for that specificity is a failure I have seen: Counting every response after 8 seconds as available hid a p99 of 2,410 ms, and conversion on search fell 9% while the availability tile stayed at 99.94%.
Two SLOs on checkout, 28-day rolling window.
| Objective | Indicator | Target | 28-day result |
|---|---|---|---|
| Availability | successful checkouts / attempts | 99.90% | 99.94% |
| Latency | successful checkouts <= 300 ms / attempts | 99.90% | 91.2% |
| Combined "good" events | success AND <= 300 ms | not used | 91.2% |
I would not consider it settled without evidence: Plot completion rate and p99 on the same journey for 28 days and show that either series can burn budget while the other stays inside target.
One number cannot describe both "did it work" and "was it fast enough to use".
Curated: · Written: · Reviewed:
QA-3Your 30-day 99.9% SLO has 43.2 minutes of remaining budget and a risky migration is scheduled. What policy do you apply?(show answer)
This is one of those areas where the default answer and the right answer for error-budget policy differ.
Use remaining error budget as a primary release-risk signal, alongside security, compliance, and business constraints, and slow or halt risky changes when the published policy requires it.
Concretely, publish a written policy that maps remaining budget bands to freeze, canary-only, or unrestricted change, and require the service owner and SRE on-call to sign the band before a risky migration.
The reason for that specificity is a failure I have seen: Shipping a schema rewrite with 4 minutes remaining caused a 6-hour outage—8.33 times the entire 30-day 99.9% budget—and kept the rolling SLO exhausted until the outage aged out of the window.
30-day 99.9% budget is 43.2 minutes; policy bands decide the launch.
| Remaining budget | Band | Allowed change |
|---|---|---|
| > 21.6 min (50%) | green | normal deploys |
| 8.6–21.6 min | amber | canary + 2h soak |
| < 8.6 min (20%) | red | reliability work only |
| 4 min at review | — | migration blocked |
I would not consider it settled without evidence: Show the last four budget-band decisions in the change log, with the remaining minutes at approval and the actual burn in the following 7 days.
A budget that never constrains a launch is a slide, not a control.
Curated: · Written: · Reviewed:
QA-4A 30-day SLO pages only after the month is already lost. How do you alert on burn rate instead?(show answer)
The useful framing here is to ask what evidence would change my mind about multi-window burn-rate alerts.
Page on a fast burn window and ticket on a slow burn window so a short outage and a quiet leak both get a human before the budget is gone.
Concretely, for a 30-day SLO, page when both 1h and 5m windows exceed 14.4× burn, or both 6h and 30m windows exceed 6×; open a ticket when both 3d and 6h windows exceed 1×, using the same bad-event ratio as the SLO.
The reason for that specificity is a failure I have seen: A 2% error leak for 11 days never crossed a 5-minute error-rate page, then the monthly report showed 0 minutes of budget left.
Multi-window burn rates for a 30-day 99.9% SLO (43.2 min budget).
# 14.4x for 1h consumes 2% of the budget; sustained 14.4x burn exhausts the full budget in about 50 hours.
(
sum(rate(sli_bad[1h])) / sum(rate(sli_total[1h]))
) / (1 - 0.999) >= 14.4
and
(
sum(rate(sli_bad[5m])) / sum(rate(sli_total[5m]))
) / (1 - 0.999) >= 14.4
I would not consider it settled without evidence: Replay the last three budget-exhausting incidents through the burn-rate rules and record page time versus the minute the budget actually hit zero.
If the page fires after the monthly slide is already red, the alerting window is the wrong length.
Curated: · Written: · Reviewed:
QA-5On-call spent 19 hours last week restarting a batch worker. Is that toil, and what do you do with it?(show answer)
Before choosing an implementation I would establish what success looks like for toil definition and reduction.
Classify work as toil when it is manual, repetitive, automatable, tactical, lacks enduring value, and tends to scale with service growth; it may be scheduled or interrupt-driven.
Concretely, track toil hours per rotation, set a ceiling such as 50% of SRE time, and fund an automation ticket for every recurring runbook step that has been executed more than twice in 30 days.
The reason for that specificity is a failure I have seen: Leaving the restart as a known workaround consumed six hours a week until the team fixed the memory defect and added bounded automatic recovery with alerting.
Classify last week's on-call hours before asking for more headcount.
| Activity | Hours | Toil? | Next action |
|---|---|---|---|
| Restart batch worker on OOM | 19 | yes | cgroup limit + retry budget |
| Design review for new SLO | 4 | no | keep |
| Copy-paste config to 12 shards | 7 | yes | generate from one source |
| Customer incident command | 6 | no | keep |
I would not consider it settled without evidence: Compare toil hours and ticket count for the same task in the four weeks before and after the automation shipped.
A runbook that is executed on a schedule is a job description for a program you have not written yet.
Curated: · Written: · Reviewed:
QA-6Engineering wants a freeze after two Sev-2s, while product wants the launch date. How do you decide?(show answer)
The engineering question underneath error budget vs feature freeze is which failure is acceptable.
Let remaining error budget, not incident count or a calendar date, decide whether feature work continues.
Concretely, compute remaining budget, apply the published policy, and document any executive risk exception explicitly; the exception does not restore budget or make an SLO miss disappear.
The reason for that specificity is a failure I have seen: Freezing after two short Sev-2s that burned 6 minutes, then shipping through a later 0-minute budget, taught the org that freezes are political rather than quantitative.
Same two Sev-2s, opposite freeze calls depending on remaining budget.
| Day | Remaining 30-day budget | Sev-2 count | Policy |
|---|---|---|---|
| 4 | 38.1 min | 2 | no freeze, 6 min burned |
| 22 | 3.4 min | 0 | freeze, budget nearly gone |
| 22 exception | 3.4 min | 0 | allowed only with signed grant |
I would not consider it settled without evidence: Show freeze and unfreeze timestamps next to remaining budget, and confirm no feature deploy occurred below the red band without a signed grant.
Incident folklore is a poor substitute for the minutes the users already paid for.
Curated: · Written: · Reviewed:
QA-7Availability rose to 99.98% after the team started retrying inside the SLI window. What went wrong?(show answer)
My approach to SLI gaming and Goodhart starts from the constraint rather than from the technique.
Measure the user's first unsuccessful attempt, not the server's eventual recovery, or the SLI will be gamed into a vanity number.
Concretely, measure one logical user attempt from start to terminal outcome, including retry latency; count it good only if it succeeds within the user-visible latency threshold, and separately expose first-attempt and retry metrics diagnostically.
The reason for that specificity is a failure I have seen: Client retries turned a 7% payment-gateway failure into a 99.97% SLI, while 1 in 14 buyers still saw a spinner for 11 seconds.
First-attempt success versus eventual success after in-window retries.
| Counting rule | 7-day success | What the buyer saw |
|---|---|---|
| Eventual HTTP 200 | 99.97% | 11 s spinner on 7% of pays |
| First attempt only | 92.80% | matches the spinner |
| First attempt, 300 ms budget | 88.10% | honest latency+availability |
I would not consider it settled without evidence: Compare first-attempt success against eventual success on a 7-day sample and treat a gap above 0.5 points as SLI contamination.
When the metric is the target, the shortest path is to stop counting the failures users already felt.
Curated: · Written: · Reviewed:
QA-8The API returns 200 while 3.1% of invoices contain the wrong tax. Is the availability SLO enough?(show answer)
I would treat correctness SLIs vs request success as a design decision with a stated trade-off rather than a best practice.
Add a correctness SLI wherever a successful response can still be a wrong business result, because HTTP success is not product success.
Concretely, sample completed writes against an independent oracle or invariant, count a mismatch as a bad event even when the status is 200, and give correctness its own budget.
The reason for that specificity is a failure I have seen: A rounding change shipped behind a 99.95% availability SLO and billed 14,200 invoices 1–3 cents high for 9 days before finance noticed.
HTTP success versus invoice-correctness on the same 9-day window.
bad_http = events.filter(status="5xx").count() # 0.04%
bad_tax = invoices.filter(tax != oracle.tax).count() # 3.10%
correctness_sli = 1 - (bad_tax / invoices.count()) # 96.90%
# Availability SLO: 99.95% green. Correctness SLO: 99.99% exhausted on day 1.
I would not consider it settled without evidence: Run a nightly invariant job over 100% of invoices and publish the mismatch rate beside the HTTP success rate.
A 200 that posts the wrong number is an outage that monitoring will not name unless you ask it to.
Curated: · Written: · Reviewed:
QA-9A search index is "available" at 99.99% while new products take 47 minutes to appear. What SLI are you missing?(show answer)
The place where teams go wrong on freshness SLIs for async work is usually the step before the one they are debating.
For asynchronous pipelines, SLO on the age of the newest successfully processed record, not on whether the consumer process is up.
Concretely, for a time-based freshness SLI, measure current time minus the newest completely processed source watermark and count observation intervals over the threshold as bad; alternatively measure per-record publication delay, but do not mix the two denominators.
The reason for that specificity is a failure I have seen: Paging only on consumer CPU left a stuck Kafka partition undetected for 47 minutes, and launch-day inventory for 8,400 SKUs was invisible to search.
Freshness SLO: 99% of documents searchable within 120 seconds.
# Bad minutes: watermark lag exceeds 120s.
avg_over_time((
time() - search_index_watermark_unixtime
)[30d:1m] > 120)
# 47 min stall on one partition: 47 / 43,200 min = 0.109% bad, which already
# burns a 99.9% freshness SLO (budget 43.2 min).
I would not consider it settled without evidence: Inject a tagged event at the source and assert it becomes queryable within the freshness SLO, then leave that probe running.
A worker that is alive but behind is indistinguishable from an outage for anyone waiting on the output.
Curated: · Written: · Reviewed:
QA-10Mean search latency is 42 ms and users still complain that search feels slow. What are you measuring wrong?(show answer)
What matters most about tail latency vs average latency is whether the result can be checked afterwards.
Define latency success as the proportion of valid requests completed within a user-relevant threshold, while publishing p95 and p99 as diagnostic summaries.
Concretely, record latency histograms with fine buckets around the product threshold, publish p95 and p99 for the journey, and ignore the mean except as a cost signal.
The reason for that specificity is a failure I have seen: A mean of 42 ms sat next to a p99 of 1,860 ms caused by a 2% path that waited on a cold cache, and that 2% was the mobile users on a slow radio.
Same week, mean looks healthy while the tail is the complaint.
| Statistic | Value | Product reading |
|---|---|---|
| mean | 42 ms | "fast" |
| p50 | 28 ms | typical shopper |
| p95 | 210 ms | noticeable pause |
| p99 | 1,860 ms | abandon on mobile |
| share > 300 ms | 2.1% | the actual SLO miss |
I would not consider it settled without evidence: Show the histogram, the mean, and the p99 on the same week, and confirm the SLO threshold sits inside a histogram bucket you actually collect.
Nobody experiences the average request; they experience their own, which is why the tail is the product.
Curated: · Written: · Reviewed:
QA-11Should a 99.9% SLO reset on the first of the month or slide every minute?(show answer)
I would answer this by separating what rolling versus calendar SLO windows guarantees from what it merely usually does.
Prefer a rolling window so a disaster on the 2nd cannot hide behind a fresh calendar month, unless a contract forces calendar reporting.
Concretely, compute the SLI over the trailing 28 or 30 days every minute, keep a calendar view only for external reports, and drive freezes from the rolling number.
The reason for that specificity is a failure I have seen: A 14-hour outage on the 1st was "reset" on the 1st of the next month, so feature launches resumed while users were still inside the pain of the previous window.
14-hour outage starting 00:10 on the 1st, 99.9% / 30-day budget = 43.2 min.
| Window | Budget remaining on the 2nd | Freeze? |
|---|---|---|
| Calendar month (resets 00:00 on 1st) | 0 min until the 1st next month | freeze all month |
| Calendar month if you "start over" next month | 43.2 min on the 1st | false green |
| Rolling 30d | 0 min for the next 29 days | freeze until burn ages out |
I would not consider it settled without evidence: Replay a 14-hour outage on day 1 of a calendar month against both windows and show remaining budget on day 2.
Users do not forgive you because a spreadsheet rolled over at midnight.
Curated: · Written: · Reviewed:
QA-12Legal signed 99.9% monthly availability with a customer, while the team tracks 99.95% internally. How should the two relate?(show answer)
The judgement in customer SLO vs internal SLO lies in what you refuse to do as much as in what you build.
Keep the internal SLO strictly tighter than the customer contract so internal burn triggers action before a credit is owed.
Concretely, choose an internal objective with enough quantified buffer to detect and mitigate burn before breaching the SLA, based on incident and response-time modeling rather than an arbitrary extra nine.
The reason for that specificity is a failure I have seen: Matching the contract at 99.9% meant the first internal page was also the first billable credit, and finance learned of the miss from the customer.
Internal 99.95% / 30d versus customer 99.9% / calendar month.
| SLO | Budget | Used as |
|---|---|---|
| Internal 99.95% rolling 30d | 21.6 min | pages, freezes |
| Customer 99.9% calendar month | 43.2 min | credits |
| Incident of 28 min | internal exhausted, contract intact | working as designed |
I would not consider it settled without evidence: For the last four incidents, show internal page time, remaining internal budget, and whether the contractual SLO was still intact.
The SLO you operate to must trip before the SLO you sold, or you have no warning channel.
Curated: · Written: · Reviewed:
QA-13Your service SLO is 99.9% and you call four dependencies each at 99.9%. How do you budget for that?(show answer)
The first thing to settle about dependency error budgets is what decision it actually changes.
Reserve a slice of your own error budget for each critical dependency, and refuse a dependency whose published reliability cannot fit in that slice.
Concretely, list critical dependencies, assign each a maximum contribution to your bad-event rate, and require a fallback or a tighter vendor SLO when the arithmetic does not close.
The reason for that specificity is a failure I have seen: Four serial 99.9% calls without fallbacks produced a 99.6% ceiling, and the team's 99.9% SLO was mathematically unreachable before they wrote a line of code.
Serial dependency math for a 99.9% service (budget 0.10% bad).
| Dependency | Their SLO | If independent and serial | Allocated slice |
|---|---|---|---|
| Auth | 99.90% | 0.10% | 0.03% |
| Payments | 99.90% | 0.10% | 0.04% |
| Catalog | 99.95% | 0.05% | 0.02% |
| Tax API | 99.90% | 0.10% | 0.01% + cache fallback |
| Product if all fail open | — | 0.35% (~99.65%) | does not fit 0.10% |
I would not consider it settled without evidence: Publish the budget table in the service spec and show measured dependency contribution versus the allocated slice over 30 days.
You cannot inherit four nines of blame and still sell three nines of your own without a plan for the remainder.
Curated: · Written: · Reviewed:
QA-14A nightly billing job has no request rate. What does an SLO look like for it?(show answer)
I would start by making the assumption behind SLOs for batch jobs explicit.
SLO a batch job on freshness of its output and on completeness of records processed, not on process uptime.
Concretely, define a due time, count a run late when the watermark is past due, count a run incomplete when processed rows disagree with the source count, and budget those failures over 30 days.
The reason for that specificity is a failure I have seen: A job that "ran" every night with exit 0 skipped 6.4% of accounts after a silent filter change, and invoices went out wrong for 11 days.
Billing batch SLO: complete by 06:00 UTC, 99.5% of accounts billed.
| Run date | Due | Finished | Rows in / out | Good? |
|---|---|---|---|---|
| 2026-08-02 | 06:00 | 05:12 | 1,204,331 / 1,204,331 | yes |
| 2026-08-11 | 06:00 | 05:08 | 1,211,004 / 1,133,492 | no, 6.4% dropped |
| 2026-08-19 | 06:00 | 07:41 | 1,209,880 / 1,209,880 | no, 101 min late |
I would not consider it settled without evidence: Compare source row counts to output row counts every run, and page when either lateness or incompleteness exceeds the SLO.
Exit status 0 is not an SLO; the output either arrived on time and whole, or it did not.
Curated: · Written: · Reviewed:
QA-15A CI/CD pipeline is down 9% of weekday hours. Why is that an SRE problem rather than a developer inconvenience?(show answer)
This is one of those areas where the default answer and the right answer for pipeline availability differ.
Treat the path that produces production artifacts as a user journey for engineers, with an availability SLO, because a dark pipeline is a change freeze you did not declare.
Concretely, sLI on successful pipeline starts and on time-to-green for the default branch, page when the pipeline SLO burns, and give the pipeline a named owner on the SRE rotation.
The reason for that specificity is a failure I have seen: A 4-hour Git-runner outage during a Sev-1 blocked the rollback job itself, stretching MTTR from 22 minutes to 4.6 hours.
Pipeline SLO beside the service it ships.
| Journey | SLI | Target | Last 30d |
|---|---|---|---|
| checkout API | success and p99 <= 300 ms | 99.90% | 99.93% |
| default-branch pipeline | start succeeds and finishes < 18 min | 99.00% | 91.00% |
| rollback job | triggered and complete < 10 min | 99.90% | unmeasured, failed in Sev-1 |
I would not consider it settled without evidence: Report pipeline availability and mean time-to-green weekly next to product SLOs, and show at least one incident where pipeline downtime delayed mitigation.
If you cannot ship a fix, you do not have a reliable production system, you have a museum.
Curated: · Written: · Reviewed:
QA-16Leadership wants MTTR under 15 minutes. What do you split out first?(show answer)
The useful framing here is to ask what evidence would change my mind about MTTD versus MTTR.
Split detect time from repair time, because a 15-minute MTTR target is a lie if 40 minutes of that incident were spent noticing it.
Concretely, stamp detect, diagnose, mitigate, and recover on every Sev-1/2, report MTTD and MTTR separately, and fund detection work when MTTD dominates.
The reason for that specificity is a failure I have seen: A disk-full outage took 14.2 minutes to page on a 5-minute scrape interval with a 3-for-3 alert, 8.4 minutes to diagnose, and 6.1 minutes to mitigate; the postmortem still quoted "MTTR 28.7 minutes" as if repair were the whole problem.
Ten Sev-1s: detection, not repair, owns the clock.
| Phase | Median minutes | What moves it |
|---|---|---|
| MTTD (start to page) | 14.2 | SLI paging, not host paging |
| Diagnose | 8.4 | dashboards, traces |
| Mitigate | 6.1 | rollback, feature flag |
| "MTTR" as commonly quoted | 28.7 | hides that 14.2 min was noticing |
I would not consider it settled without evidence: For the last 10 Sev-1s, show MTTD and MTTR as stacked bars and fund the larger slice.
You cannot repair a failure you have not yet admitted is happening.
Curated: · Written: · Reviewed:
QA-17Three engineers are all debugging during a Sev-1 and nobody is talking to customers. How do you structure the response?(show answer)
Before choosing an implementation I would establish what success looks like for incident command.
Name a single incident commander who does not debug, plus a dedicated comms lead, so diagnosis has a coordinator and customers have a voice.
Concretely, on Sev-1/2, the first responder either becomes commander or hands command within 5 minutes, assigns an ops lead and a comms lead, and records those names in the incident doc.
The reason for that specificity is a failure I have seen: Four people patched four different services for 38 minutes because nobody owned the hypothesis list, and the actual stuck deploy ran the whole time.
Roles on a Sev-1 at T+5 minutes.
| Role | Person | Allowed to SSH/debug? |
|---|---|---|
| Incident commander | SRE-2 | no |
| Ops lead | SWE-4 | yes, one change at a time |
| Comms lead | support-1 | no, owns customer + status page |
| Scribe | SRE-2 (dual-hat if short) | timeline only |
I would not consider it settled without evidence: Audit the last 8 Sev-1s for a named commander within 5 minutes and a comms update on the agreed cadence.
An incident with many helpers and no commander is just a group of people interrupting each other.
Curated: · Written: · Reviewed:
QA-18A dashboard is wrong but checkout works. Is that Sev-1? How do you classify?(show answer)
The engineering question underneath severity classification is which failure is acceptable.
Classify severity by user impact and blast radius against written criteria, not by how embarrassed the team feels.
Concretely, publish a matrix of user-visible impact versus fraction of users, train the on-call to pick from it in the first 2 minutes, and forbid upgrading severity to "get more people" without a new impact fact.
The reason for that specificity is a failure I have seen: Calling a broken internal admin theme Sev-1 paged 22 people at 02:14, while a 14% checkout 5xx the previous afternoon sat as Sev-3 because it "looked like a client bug".
Written severity matrix used at T+2 minutes.
| Impact | < 5% of users | 5–50% | > 50% or revenue path |
|---|---|---|---|
| No user-visible effect | Sev-4 | Sev-4 | Sev-3 |
| Degraded, workaround exists | Sev-3 | Sev-2 | Sev-2 |
| Hard down or wrong money | Sev-2 | Sev-1 | Sev-1 |
I would not consider it settled without evidence: Re-score the last 15 incidents against the matrix and report the mismatch rate; target under 10%.
Severity is a description of harm, not a volume knob for attention.
Curated: · Written: · Reviewed:
QA-19During an outage, Slack fills with hypotheses and the status page stays on "investigating" for 70 minutes. What is the comms contract?(show answer)
My approach to incident communication starts from the constraint rather than from the technique.
Send time-boxed updates on a stated cadence with impact, current action, and next update time, even when the cause is still unknown.
Concretely, comms lead posts internally every 10 minutes and externally every 20 minutes for Sev-1, using a template that forbids silence and forbids guessing a root cause.
The reason for that specificity is a failure I have seen: A 70-minute "investigating" status page next to a 40% error rate produced a 1,200-ticket flood and a journalist story before engineering had a mitigation.
Sev-1 external update template, minute 20.
Status: identified
Impact: 38% of checkouts returning HTTP 503 since 14:12 UTC
What we are doing: rolling back deploy checkout@2026.09.07.1311
Next update: 14:52 UTC (20 min)
Cause: not confirmed; we will not speculate here
I would not consider it settled without evidence: For each Sev-1, check the incident doc for update timestamps at the cadence, and count customer tickets that arrived after the first accurate impact statement.
People will invent a narrative if you do not provide one on a clock.
Curated: · Written: · Reviewed:
QA-20A postmortem draft says "the intern ignored the runbook". How do you rewrite it?(show answer)
I would treat blameless postmortems as a design decision with a stated trade-off rather than a best practice.
Describe the conditions, incentives, and missing guards that made the action reasonable, and never name a person as the cause.
Concretely, require contributing factors in systems language, a timeline of facts, and a review that rejects personal attributions before the doc is marked done.
The reason for that specificity is a failure I have seen: Naming an intern ended the investigation at "training", and the same missing canary check shipped a second bad config six weeks later with a different person at the keyboard.
Rewrite the contributing factor so it survives a staffing change.
| Draft line | Rewrite |
|---|---|
| The intern ignored the runbook. | Production apply was possible without a plan file, and the UI defaulted to the last target. |
| Jane should have paged sooner. | The only page was on CPU, which stayed green while 5xx rose to 18%. |
| Fat-fingered the cluster name. | Two clusters accepted the same context name; kubectl had no confirmation prompt. |
I would not consider it settled without evidence: Scan the last 10 postmortems for personal attributions and for at least one system-level action item per doc.
If the same process with a different human would fail the same way, the human was not the defect.
Curated: · Written: · Reviewed:
QA-21A postmortem has 23 action items and three of them are "be more careful". What makes an action item real?(show answer)
The place where teams go wrong on postmortem action items is usually the step before the one they are debating.
Every action item must remove a contributing factor with an owner, a due date, and a test that would have changed this incident.
Concretely, cap the list at the few items that close the actual factors, require a verification step, and track them in the same system as other reliability work with a 30-day default due date.
The reason for that specificity is a failure I have seen: Twenty-three items aged into a graveyard, including the one that would have added a plan file, and the next apply-without-plan outage was a copy of the first.
Keep three items that would have changed the incident.
| Item | Owner | Due | Verification |
|---|---|---|---|
| Require terraform plan artifact in CI | platforms | 14d | apply job fails without plan |
| Page on 5xx burn, not CPU | sre | 7d | replay of incident pages at T+3 min |
| "Be more careful in prod" | — | — | rejected |
I would not consider it settled without evidence: Report action-item completion at 30 days and re-open any item whose verification step cannot be demonstrated.
An action item without a check is a wish that will lose to the next feature deadline.
Curated: · Written: · Reviewed:
QA-22After a year of outages, leadership asks "what keeps breaking". How do you answer without a slide of anecdotes?(show answer)
What matters most about incident taxonomy is whether the result can be checked afterwards.
Tag every incident with a small, stable set of cause classes so you can spend error budget on the classes that actually recur.
Concretely, adopt a fixed taxonomy at incident close, require one primary class, and review quarterly which class burned the most customer-facing minutes.
The reason for that specificity is a failure I have seen: Free-text causes ("network-ish", "bad luck", "GCP") produced 90 unique labels in a year and no way to fund the 41% of minutes that were change-related.
Customer-facing minutes, last 90 days, one primary class per incident.
| Class | Incidents | Minutes | Share |
|---|---|---|---|
| Change / deploy | 11 | 412 | 41% |
| Capacity / overload | 4 | 188 | 19% |
| Dependency | 6 | 171 | 17% |
| Config | 5 | 140 | 14% |
| Hardware / zone | 2 | 91 | 9% |
I would not consider it settled without evidence: Show a quarterly histogram of customer-minutes by taxonomy class, with the top class mapped to a funded project.
You cannot prioritize a pile of stories; you can prioritize a distribution.
Curated: · Written: · Reviewed:
QA-23The on-call is paged for disk at 70% while users are fine, and not paged when checkout 5xx hits 12%. What do you page on?(show answer)
I would answer this by separating what symptom-based paging guarantees from what it merely usually does.
Page on symptoms that match SLO burn, and demote cause-based resource alerts to tickets unless they are a proven precursor with a tight deadline.
Concretely, bind pages to SLI burn-rate or to user-journey error/latency thresholds, and move host-level warnings to a next-business-day queue.
The reason for that specificity is a failure I have seen: A 70% disk page at 03:11 woke the rotation 14 nights in a row, while a 12% checkout 5xx at noon waited for Twitter because no symptom page existed.
Reclassify the paging set against the checkout SLO.
| Alert | Fires on | Page? |
|---|---|---|
| checkout 5xx burn 14x / 1h | user symptom | yes |
| latency bad-event burn on requests > 300 ms | user symptom measured against latency SLO | yes |
| disk 70% on a log node | possible cause | ticket |
| CPU > 80% on one replica | possible cause | no, HPA should absorb |
I would not consider it settled without evidence: Count pages that correlated with SLO burn versus pages that did not, over 14 days, and drive the uncorrelated ones off the pager.
A page that does not correspond to user harm trains the rotation to ignore the ones that do.
Curated: · Written: · Reviewed:
QA-24The primary pager received 61 pages last week and acknowledged 58 of them in under 30 seconds. What is broken?(show answer)
The judgement in alert fatigue lies in what you refuse to do as much as in what you build.
A page that is acknowledged without investigation is a ticket wearing a siren, and those must leave the pager until the remainder is rare and urgent.
Concretely, measure pages per shift and ack-without-action rate, set a budget such as 2 pages per 12-hour shift, and delete or downgrade every alert above that until the budget holds.
The reason for that specificity is a failure I have seen: Sixty-one weekly pages trained the rotation to auto-ack; the 62nd was a real 18% 5xx event that sat acknowledged for 41 minutes.
One week of primary-pager traffic versus a 2-page/shift budget.
| Metric | Last week | Target |
|---|---|---|
| Pages | 61 | <= 14 (2 x 7 shifts) |
| Ack < 30 s with no ticket | 58 | ~0 |
| Pages that became incidents | 2 | most of them |
| Real 5xx sitting on an acked alert | 41 min | 0 min |
I would not consider it settled without evidence: Report pages/shift, median time-to-ack, and fraction of pages that produced an incident doc; the last should be high, not near zero.
Silence is a feature of a good paging set; volume is a defect.
Curated: · Written: · Reviewed:
QA-25A 7-day primary rotation ends with 19 hours of night work and a 4-line Slack handoff. How do you run load and handoff?(show answer)
The first thing to settle about on-call load and handoff is what decision it actually changes.
Cap interrupt load per shift and replace folklore handoffs with a written state file the next primary can act on without a conversation.
Concretely, track interrupt hours and sleep-disrupting pages, keep primary rotations at most 24 hours where load is high, and require a handoff doc covering open incidents, silenced alerts, and in-flight changes.
The reason for that specificity is a failure I have seen: A Friday 4-line handoff omitted a silenced burn-rate alert; the weekend primary discovered it when the monthly budget hit zero on Sunday night.
Handoff doc the next primary can use at 09:00 without a call.
Primary: 2026-09-06 09:00 -> 2026-09-07 09:00
Open incidents: INC-4412 (Sev-2, waiting on vendor, next update 10:30)
Silences: checkout_burn_1h until 11:00 (owner: sre-lee, reason: load test)
In-flight: canary 12% on payments-v2, rollback command in runbook RB-88
Pages this shift: 3 (2 real, 1 mis-aimed). Interrupt hours: 2.1
Follow-ups: raise HPA max from 30 -> 45 after traffic report
I would not consider it settled without evidence: Audit 8 consecutive handoffs for the required fields and plot interrupt hours per primary; act when either misses the cap.
On-call is a designed workload, not a personality test for who can absorb the most 03:00 pages.
Curated: · Written: · Reviewed:
QA-26A 5xx spike has a red graph, 40 GB of logs, and no trace ids. How do you wire the three signals?(show answer)
I would start by making the assumption behind metrics logs and traces together explicit.
Correlate logs and traces with trace context, and connect aggregate metrics to representative traces through exemplars rather than request-id labels.
Concretely, propagate a trace context on every hop, attach the trace id to structured logs, and link SLO burn alerts to an exemplar trace from the bad bucket.
The reason for that specificity is a failure I have seen: During a 22-minute 5xx event, on-call grepped 40 GB of logs by timestamp and never found the 3% of requests that hit a stale replica.
One checkout request, three signals, one id.
metric: checkout_sli_bad{code="503"} += 1 # 14:12:03.441
trace: 00-9f2c41ab77de0031c0ffee4411aa22bb-7a11e4c0ffee1234-01 # W3C traceparent
log: {"trace":"9f2c41ab77de0031c0ffee4411aa22bb","span":"7a11e4c0ffee1234",
"msg":"payments_upstream_timeout","ms":2104,"replica":"pay-7b"}
# Page -> exemplar -> log: 8 seconds, not 22 minutes of grep.
I would not consider it settled without evidence: From a burn-rate page, open one exemplar and confirm you can see the failing span and the matching log line without a search.
Three observability products that cannot name the same request are three separate hobbies.
Curated: · Written: · Reviewed:
QA-27Someone added user_id as a Prometheus label to "debug faster". What happens, and what do you allow instead?(show answer)
This is one of those areas where the default answer and the right answer for high-cardinality labels differ.
Keep metric label sets bounded to things you would actually alert or aggregate on, and push unbounded identifiers into logs and traces.
Concretely, cap label cardinality per metric, block new labels in CI against a budget such as 100k series per service, and store user_id only on spans and structured logs.
The reason for that specificity is a failure I have seen: user_id on http_requests_total created 12.4 million series in 3 days, Prometheus scrape lagged 90 seconds, and the burn-rate page that needed to fire did not.
Series budget for checkout metrics; user_id does not belong.
| Label | Distinct values / 24h | On metrics? | Where it lives |
|---|---|---|---|
| code | 8 | yes | Prometheus |
| route | 24 | yes | Prometheus |
| region | 3 | yes | Prometheus |
| user_id | 2,100,000 | no | trace attribute + log field |
| series if user_id added | 12.4e6 | — | scrape lag 90 s |
I would not consider it settled without evidence: Graph series count per metric daily and fail the build when a change would push a metric over its series budget.
A metric that cannot be scraped in time is worse than no metric, because it also blinds the ones you still needed.
Curated: · Written: · Reviewed:
QA-28p99 latency is 1.8 s and nobody can find a slow request to inspect. What do you attach to the histogram?(show answer)
The useful framing here is to ask what evidence would change my mind about trace exemplars.
Attach exemplars from the tail buckets of the latency histogram so a p99 tile is one click from a real trace.
Concretely, enable exemplars on the request histogram, store trace ids in the tail buckets, and teach the latency dashboard to open that trace.
The reason for that specificity is a failure I have seen: A week of "p99 is 1.8 s" reviews sampled random traces from the 28 ms majority and concluded the application was fine.
Histogram bucket 1.0–2.5 s carries an exemplar into the trace backend.
histogram_quantile(0.99, sum by (le) (rate(checkout_latency_seconds_bucket[5m])))
# 1.86 s
# exemplar in le="2.5": trace_id=9f2c41ab77de0031c0ffee4411aa22bb duration=1.81s span=payments.wait
I would not consider it settled without evidence: Click the p99 tile, land on a trace whose duration is in the p99 bucket, and confirm that path matches a known slow dependency.
A percentile without an example is a number you cannot debug.
Curated: · Written: · Reviewed:
QA-29You cannot keep 8,000 traces per second. Which traces do you keep?(show answer)
Before choosing an implementation I would establish what success looks like for trace sampling.
Sample at the head for cheap traffic and use tail policies to retain errors and latency-SLO misses, so the expensive cases are not the ones you preferentially throw away.
Concretely, use tail sampling or a collector rule configured to retain error and SLO-miss traces, plus a small baseline such as 1% of the rest; monitor late or capacity-dropped traces and record the policy in the service spec.
The reason for that specificity is a failure I have seen: Head sampling at 0.1% dropped 99.9% of the 5xx traces during a 0.8% error incident, leaving on-call with 14 healthy traces and no failing ones.
Collector policy: keep the bad, downsample the fine.
# With no drop policy here, a sample vote from any of these policies retains the trace.
processors:
tail_sampling:
policies:
- name: errors
type: status_code
status_code: {status_codes: [ERROR]}
- name: slo_miss
type: latency
latency: {threshold_ms: 300}
- name: baseline
type: probabilistic
probabilistic: {sampling_percentage: 1.0}
I would not consider it settled without evidence: During a synthetic error test, compare retained error-trace count with an unsampled error counter and confirm nearly all eligible error traces were retained; do not infer production error rate from the biased sampled population.
Sampling that is blind to failure preferentially deletes the evidence.
Curated: · Written: · Reviewed:
QA-30On-call is regexing "timeout" across 14 log formats. What logging contract do you enforce?(show answer)
The engineering question underneath structured logging is which failure is acceptable.
Emit one JSON schema per service with required fields for timestamp, severity, service, trace id, and a stable event name.
Concretely, ship a library that refuses to log free text at warn+, validate the schema in CI on a sample, and index event and trace_id as first-class fields.
The reason for that specificity is a failure I have seen: A format change from "timeout after 2s" to "deadline exceeded (2000ms)" broke the only dashboard that detected payment hangs, for 6 days.
Required fields; the message is optional commentary.
{"ts":"2026-09-07T14:12:03.441Z","sev":"error","svc":"checkout",
"event":"payments_upstream_timeout","trace":"9f2c41ab77de0031c0ffee4411aa22bb",
"span":"7a11e4c0ffee1234","http_code":503,"upstream_ms":2104,"retry":2}
I would not consider it settled without evidence: Query last 24 hours for events missing trace_id or event, and keep that rate under 0.1%.
A log line that cannot be grouped is a note to a future human, not an operational signal.
Curated: · Written: · Reviewed:
QA-31When do you instrument a service with RED, and when do you reach for USE?(show answer)
My approach to RED versus USE methods starts from the constraint rather than from the technique.
Use RED on request-serving APIs to track user harm, and USE on resources to find saturation before that harm appears.
Concretely, for each user-facing service, export rate, errors, and duration; for each scarce resource (CPU, disk, pool, queue), export utilization, saturation, and errors.
The reason for that specificity is a failure I have seen: Watching only USE on a fleet of happy CPUs missed a 9% 400-rate from a bad client header, because the machines were not saturated, the users were.
Checkout API (RED) and its database pool (USE) on the same incident.
| Method | Signal | 14:12 value | Reading |
|---|---|---|---|
| RED | rate | 1,840 rps | load not unusual |
| RED | errors | 9.1% 5xx | user harm |
| RED | duration p99 | 1,920 ms | SLO miss |
| USE | pool utilization | 0.98 | 98 of 100 conns |
| USE | pool saturation | wait queue 240 | the bottle |
I would not consider it settled without evidence: Show a dashboard pair: RED for the journey SLO and USE for the first resource that historically preceded a burn.
RED tells you users are hurting; USE tells you which bottle is about to empty.
Curated: · Written: · Reviewed:
QA-32Error rate is still 0.02% and you want to know you will run out of capacity next Thursday. What do you alert on?(show answer)
I would treat saturation signals as a design decision with a stated trade-off rather than a best practice.
Alert on saturation and time-to-exhaustion of finite pools, not only on the errors that appear after the pool is empty.
Concretely, track utilization and queue wait for connections, threads, disk, and quotas; page or ticket when time-to-limit at the current slope is inside the lead time to add capacity.
The reason for that specificity is a failure I have seen: A connection pool hit 100% at 18:42 Friday; the first 5xx page arrived 11 minutes later, after 8,400 checkout attempts queued and timed out.
Time-to-exhaustion on a 500-connection database limit.
limit, used, slope = 500, 400, 4.2 # connections per minute
minutes_left = (limit - used) / slope # 23.8 minutes
lead_time_to_scale = 15 # HPA + pool bounce
# Alert when minutes_left < lead_time_to_scale * 2 -> 23.8 < 30, page.
I would not consider it settled without evidence: In the last capacity incident, show that the saturation signal crossed its threshold before the SLI did, with enough lead time to scale.
Errors are a lagging confirmation that you ignored a full tank gauge.
Curated: · Written: · Reviewed:
QA-33A load test at 10k rps passed on empty tables, then production fell over at 2.1k rps. What was the test missing?(show answer)
The place where teams go wrong on load testing versus production is usually the step before the one they are debating.
Load-test the production shape of data, caches, and dependencies, because empty-state throughput is a different system.
Concretely, replay anonymized production traffic against production-sized datasets, include cache cold-start and downstream rate limits, and compare saturation signals, not just HTTP 200 counts.
The reason for that specificity is a failure I have seen: A 10k rps test against 200-row fixtures never touched the 48 ms p99 index miss that 2.1k rps of real SKUs produced on a 90-million-row catalog.
Same 10k rps generator, two data shapes.
| Condition | rps | p99 | First saturation |
|---|---|---|---|
| 200-row fixture, warm cache | 10,000 | 18 ms | none |
| 90e6-row catalog, 61% hit rate | 2,100 | 640 ms | DB CPU 96%, pool wait |
| 90e6-row, cold cache after deploy | 740 | 2,400 ms | origin overload |
I would not consider it settled without evidence: Before a peak, show a test report whose dataset size, cache hit rate, and dependency mock limits sit within 20% of last week's production.
A passing test of a toy database is not evidence about the database you actually run.
Curated: · Written: · Reviewed:
QA-34Traffic grows 6% a week. How much unused capacity do you keep, and how do you know it is real?(show answer)
What matters most about capacity headroom is whether the result can be checked afterwards.
Keep quantified headroom against the next peak plus the time to add capacity, and verify it with a production soak, not with a spreadsheet of instance counts.
Concretely, set a headroom SLO such as 30% above the last peak on the bottleneck resource, schedule a quarterly soak to that level, and start procurement when projected days of headroom fall below lead time.
The reason for that specificity is a failure I have seen: Counting 40% extra pods ignored that the database primary was already at 88% CPU, so the "headroom" was a row of idle app containers in front of a full bottle.
Headroom is computed on the bottleneck, not on replica count.
| Resource | Last peak | Limit | Headroom | Lead time |
|---|---|---|---|---|
| app pods | 42 | 80 | 47% | 10 min HPA |
| db primary CPU | 88% | 100% | 12% | 14 days for a larger SKU |
| bottleneck | db CPU | — | 12% < 30% target | order hardware now |
I would not consider it settled without evidence: Show the bottleneck resource at last peak, the soak result at peak+30%, and the date when headroom will cross the lead-time line.
Headroom that is not on the actual bottleneck is decorative.
Curated: · Written: · Reviewed:
QA-35Latency doubled after you raised the worker count. How does Little's law explain that, and what do you change?(show answer)
I would answer this by separating what Little's law and queues guarantees from what it merely usually does.
Bound queue depth, because latency is occupancy divided by throughput, and an unbounded queue turns a slowdown into an arbitrarily long wait.
Concretely, use Little’s law for steady-state planning, and use measured backlog divided by net drain rate for incident recovery estimates; bound queue depth and shed above the cap.
The reason for that specificity is a failure I have seen: A 50,000-item Redis backlog represented about 10.4 minutes of waiting at 80 items per second; continued arrivals prevented recovery and made actual waits much longer.
L = λW: 80 rps arrival, 50,000 in queue, predicted wait.
lambda_rps = 80
L = 50_000
W_sec = L / lambda_rps # 625 s = 10.4 minutes already in queue
# After a 12 min outage at 80 rps, L ~= 57,600, W ~= 12 min more.
# Cap at 800 (10 s at 80 rps) and return 429 above that.
I would not consider it settled without evidence: Measure arrival rate, observed occupancy, and latency, and show they agree within 10%; then show latency at the cap versus unbounded.
A queue that never refuses work is a latency amplifier with a misleadingly green consumer.
Curated: · Written: · Reviewed:
QA-36At 2.4x expected QPS the service thrashes and nobody gets an answer. What do you drop first?(show answer)
The judgement in load shedding lies in what you refuse to do as much as in what you build.
Shed low-priority, expensive, or retry-safe work first, and keep a declared minimum of capacity for the revenue path.
Concretely, classify endpoints by criticality, apply concurrency limits per class, return a cheap 429/503 with Retry-After, and prove the critical class still meets SLO under overload.
The reason for that specificity is a failure I have seen: Fair-share across all routes let search-autocomplete crowd out checkout, and at 2.4x QPS both went to 18% 5xx instead of search degrading alone.
3x overload: shed search, keep checkout inside 99.9%.
| Class | Share of QPS | Limit | At 3x offered load |
|---|---|---|---|
| checkout | 12% | never shed | 99.92% success, p99 280 ms |
| search | 54% | 40% of workers | 41% 429, remainder 200 |
| analytics ingest | 34% | first to drop | 100% 429 |
I would not consider it settled without evidence: Run an overload test at 3x and show checkout success >= SLO while a non-critical class absorbs the shed.
If everything is equally important under overload, nothing important survives.
Curated: · Written: · Reviewed:
QA-37A 4-second payments blip produced 19 minutes of 5xx. Where did the extra load come from?(show answer)
The first thing to settle about retry storms is what decision it actually changes.
Retries without a budget multiply offered load during an outage, so every client must cap attempts and stop when the server says it is overloaded.
Concretely, set a per-request retry budget (for example 2 retries), honor Retry-After and 429, disable retries on non-idempotent calls without a key, and watch outbound retry rate as its own metric.
The reason for that specificity is a failure I have seen: Three layers each retrying 3 times turned a 4-second 20% error blip into a 27x offered-load spike that lasted until circuit breakers were added 19 minutes later.
Retry amplification: 3 layers, 3 retries each, no budget.
layers, retries_each = 3, 3
amplification = (1 + retries_each) ** layers # 4 ** 3 = 64x worst case
# Observed in the 4 s blip: 27x offered QPS, p99 8.1 s, 19 min to recover.
# With budget=1 total retry across the call chain: peak 1.8x, recover 40 s.
I would not consider it settled without evidence: In staging, fail 20% of a dependency for 30 seconds and show offered load stays under 2x, not tens of times.
The outage you remember is often the retry storm, not the original blip.
Curated: · Written: · Reviewed:
QA-38After a 10-second outage, all clients retry at t=10, t=20, t=40 and knock the service over again. What is missing?(show answer)
I would start by making the assumption behind backoff with jitter explicit.
Decorrelate retries with equal-jitter backoff so recovery is not a synchronized herd.
Concretely, implement exponential backoff with equal jitter, cap the delay, and include jitter in every official client; reject clients that retry on a fixed interval.
The reason for that specificity is a failure I have seen: A mobile fleet of 2.1 million clients with a fixed 10-second retry produced a 2.1-million-QPS thump every 10 seconds for 6 minutes after a DNS fix.
Equal jitter: delay in [backoff/2, backoff], cap 32 s.
import random
def sleep_seconds(attempt, cap=32):
base = min(cap, 0.5 * (2 ** attempt)) # 0.5, 1, 2, 4, ... 32
return random.uniform(base / 2, base) # decorrelated
# Fixed 10 s retry: 2.1e6 hits at t=10. Jittered: spread over 5–10 s bins.
I would not consider it settled without evidence: Plot retry timestamps after a 30-second induced outage and show the distribution is spread across the backoff window, not spiked at the powers of two.
A perfectly synchronized retry is just a scheduled distributed denial of service against yourself.
Curated: · Written: · Reviewed:
QA-39When should a client stop calling a sick dependency, and when is that the wrong tool?(show answer)
This is one of those areas where the default answer and the right answer for circuit breakers differ.
Open a breaker on a local error budget for that dependency so you fail fast, and keep it from hiding a dependency that still has spare capacity for this caller.
Concretely, track a rolling error ratio per dependency, open after a threshold such as 50% of 20 calls, half-open with a probe, and size the threshold so a small caller cannot trip on noise.
The reason for that specificity is a failure I have seen: A global breaker opened after 5 errors from a 0.01% client, blocking 100% of checkout while payments was healthy for everyone else.
Per-caller breaker, 20-call window, 50% trip, 10 s open.
| State | Condition | Checkout behaviour |
|---|---|---|
| closed | errors < 10 / 20 | call payments |
| open | errors >= 10 / 20 | fail fast 503, 8 ms |
| half-open | 10 s later | 1 probe, then close or re-open |
| anti-pattern | 5 errors global | 0.01% client black-holes all |
I would not consider it settled without evidence: Fail a dependency at 80% for one caller and show that caller sheds while others continue; then restore and show half-open probes succeed within 30 seconds.
A breaker is a local admission decision, not a cluster-wide rumor about a host being dead.
Curated: · Written: · Reviewed:
QA-40Recommendations are down. Should checkout go down with them? How do you degrade?(show answer)
The useful framing here is to ask what evidence would change my mind about graceful degradation.
Declare which features may go empty or stale so the core journey can continue, and implement those fallbacks as tested paths, not as hope.
Concretely, list dependencies as required or optional, serve a cached or empty optional surface on timeout, and SLO the core journey independently of the optional ones.
The reason for that specificity is a failure I have seen: A 2.2-second recommendations timeout sat on the checkout critical path with no fallback, converting a recs outage into a 100% checkout outage for 16 minutes.
Required versus optional on the checkout page.
| Surface | Dependency | On timeout (250 ms) | Checkout SLO |
|---|---|---|---|
| pay | payments | fail the request | required |
| tax | tax API | last cached rate + flag | required, stale ok 10 min |
| recs | recs svc | empty module | optional, not in SLO |
| before fix | recs | block 2.2 s | 16 min hard down |
I would not consider it settled without evidence: Kill the optional dependency in staging and show checkout success and p99 still inside SLO, with the optional surface marked degraded.
An optional feature that cannot fail open safely is a single point of failure wearing a product name.
Curated: · Written: · Reviewed:
QA-41A team wants to "run chaos" by killing random pods in production on Friday. What makes an experiment real?(show answer)
Before choosing an implementation I would establish what success looks like for chaos experiment design.
A chaos experiment has a hypothesis, a blast radius, an abort signal, and a success criterion tied to an SLO, not a random kill.
Concretely, write the hypothesis against a specific failure, limit scope to one AZ or a percentage of pods, abort if the SLI burns above a set rate, and run it when the error budget and staffing can absorb a miss.
The reason for that specificity is a failure I have seen: Unscoped pod kills on a Friday drained the only ready replicas in two shards, spent 31 minutes of error budget, and taught the org that chaos means unplanned Sev-2s.
Hypothesis-driven experiment, abort on SLO burn.
hypothesis: "killing 1 of 12 checkout pods in az-a keeps p99 < 300ms"
blast_radius: {az: "az-a", max_pods: 1, duration_min: 15}
abort: {checkout_burn_rate_1h: 6} # well below page-at-14x
window: {error_budget_remaining_min: "> 21.6", staffing: "full primary"}
result: {p99_ms: 188, burn_x: 1.1, abort_fired: false}
I would not consider it settled without evidence: The experiment report shows hypothesis, abort threshold, observed SLI, and a go/no-go for the next larger blast radius.
If you cannot say what should stay true while you inject pain, you are not experimenting, you are outaging.
Curated: · Written: · Reviewed:
QA-42How is a game day different from a chaos experiment, and what do you score?(show answer)
The engineering question underneath game days is which failure is acceptable.
A game day rehearses humans and runbooks against a scripted failure, and you score detection, command, comms, and time-to-mitigate, not whether the platform stayed pretty.
Concretely, schedule a 90-minute facilitated session, inject a failure the runbook claims to cover, and record MTTD, commander assignment, update cadence, and wrong turns without blaming individuals.
The reason for that specificity is a failure I have seen: Calling a silent failover test a game day skipped the comms and command practice, so the first real AZ outage had a 28-minute MTTD and no commander for 17 minutes.
AZ-loss game day scorecard, 90 minutes.
| Skill | Target | This game day |
|---|---|---|
| MTTD | < 5 min | 4.2 min |
| Named commander | < 5 min | 11.0 min |
| External update cadence | 20 min | first update at 34 min |
| Mitigate (fail over) | < 15 min | 13.1 min |
| Wrong-turn count | note | 2 (scaled the dead AZ) |
I would not consider it settled without evidence: Publish game-day scores over a year and show MTTD and time-to-commander falling; a single "it worked" checkbox is not a score.
The platform can pass while the organization fails, which is why you time the people.
Curated: · Written: · Reviewed:
QA-43How do you prove a fallback works without waiting for the vendor to actually die?(show answer)
My approach to dependency failure injection starts from the constraint rather than from the technique.
Inject the dependency's failure mode on the call path in a controlled environment, and assert the fallback and the core SLO, not merely that an error was logged.
Concretely, use a fault proxy or test double to return timeouts, 500s, and slow success, run it against staging and a small production canary, and compare core-journey SLIs with and without the fault.
The reason for that specificity is a failure I have seen: A mock that always returned an empty 200 trained the fallback to look green, then the real vendor hung TCP for 30 seconds and checkout threads exhausted.
Fault proxy on the tax API, production canary at 2% traffic.
| Injected fault | Expected fallback | Canary result |
|---|---|---|
| HTTP 500 | cached rate <= 10 min old | 99.94% checkout, stale tax 2.1 min |
| 5,000 ms hang | deadline 250 ms, then cache | p99 271 ms, no pool exhaustion |
| TCP reset | same as 500 | 99.93% |
| mock empty 200 (anti-pattern) | looks fine | hides the hang |
I would not consider it settled without evidence: Show a report: injected timeout of 5 seconds, fallback engaged in < 300 ms, checkout success still >= 99.9%.
A fallback never exercised is a comment in the code, not a reliability control.
Curated: · Written: · Reviewed:
QA-44Autoscaling is still spinning up pods and latency is already 2 s. What should have rejected work earlier?(show answer)
I would treat admission control as a design decision with a stated trade-off rather than a best practice.
Admit only the concurrency the service can finish inside the latency SLO, and reject the rest at the edge before queues grow.
Concretely, set a global and per-instance concurrency limit from Little's law using SLO latency and measured service time, return 429 at the load balancer, and scale on rejected QPS plus utilization.
The reason for that specificity is a failure I have seen: Waiting for HPA to add pods at 2 s p99 admitted 14k waiting requests, then the new pods spent 8 minutes draining a queue that admission control would have refused.
Concurrency cap from SLO: 300 ms, 25 ms service time, 40 pods.
slo_s, service_s, pods = 0.300, 0.025, 40
concurrency_per_pod = slo_s / service_s # 12 in-flight
cluster_cap = concurrency_per_pod * pods # 480
max_admit_rps = cluster_cap / slo_s # 1,600 rps
# Offered 3,200 rps -> admit 1,600, 429 the rest, p99 of admitted 240 ms.
I would not consider it settled without evidence: At 2x offered load, show admitted QPS flat at the cap, p99 of admitted work inside SLO, and a 429 rate that matches the overflow.
Scaling is how you raise the cap; admission control is how you survive until it rises.
Curated: · Written: · Reviewed:
QA-45HPA on CPU sits at 40% while p99 is 1.4 s. Which signal should scale the service?(show answer)
The place where teams go wrong on autoscaling signals is usually the step before the one they are debating.
Scale on a signal that represents the bottleneck users feel, such as concurrency, queue wait, or SLO-adjacent utilization, not on CPU by default.
Concretely, identify the bottleneck with USE, export it as a custom metric, point HPA at that metric with a target that leaves SLO headroom, and keep CPU as a secondary cap.
The reason for that specificity is a failure I have seen: CPU-based HPA never moved during a lock-contention incident because cores were idle in futex waits, while p99 climbed from 180 ms to 1.4 s.
HPA target: in-flight requests per pod, not CPU.
metrics:
- type: Pods
pods:
metric: {name: checkout_inflight}
target: {type: AverageValue, averageValue: "8"} # cap 12, SLO 300 ms
- type: Resource
resource:
name: cpu
target: {type: Utilization, averageUtilization: 70} # secondary
# Incident: CPU 40%, inflight 19, p99 1.4 s — CPU target never fired.
I would not consider it settled without evidence: Show a load ramp where the chosen metric crosses its target before p99 crosses the SLO, and that replica count follows it.
Autoscaling on a comfortable metric scales a fleet that is not the problem.
Curated: · Written: · Reviewed:
QA-46A canary looks "about the same" on a dashboard. How do you decide to proceed or roll back?(show answer)
What matters most about canary analysis is whether the result can be checked afterwards.
Compare canary and control on the same SLIs with a statistical test and a pre-registered abort threshold, not with a glance at two graphs.
Concretely, send a fixed share of traffic to the canary, measure SLI deltas over a soak window, abort if error rate or latency exceeds the registered threshold with confidence, and refuse manual override without a new hypothesis.
The reason for that specificity is a failure I have seen: A 5-minute eyeball of overlapping lines missed a 0.8-point error-rate delta that burned 18 minutes of monthly budget after 100% rollout.
Registered abort: +0.20 points error or +50 ms p99, 15 min soak.
| Slice | n requests | Error % | p99 ms | Decision |
|---|---|---|---|---|
| control | 84,200 | 0.12 | 188 | — |
| canary 5% | 4,410 | 0.94 | 210 | abort, +0.82 pts |
| eyeball at 5 min | — | "looks close" | "looks close" | would have shipped |
I would not consider it settled without evidence: Show the analysis report: sample sizes, p-values or sequential test bounds, and the go/no-go that the pipeline actually obeyed.
If a human can wave a canary through, you do not have canary analysis, you have a waiting period.
Curated: · Written: · Reviewed:
QA-47After a canary passes, how do you raise traffic without turning the rest of the rollout into a big-bang?(show answer)
I would answer this by separating what progressive delivery guarantees from what it merely usually does.
Raise exposure on a schedule of shares and soaks, and make each step automatically reversible from the same SLI gates as the canary.
Concretely, use a pipeline that walks 1% / 5% / 25% / 50% / 100% with a soak at each step, holds on SLO burn, and rolls back to the last good share without a meeting.
The reason for that specificity is a failure I have seen: Jumping 5% canary to 100% after a green 10-minute window hit a region that had 0% canary traffic and reproduced the bug for 31% of users.
Share schedule with soak and auto-rollback.
| Step | Share | Soak | Gate | Result |
|---|---|---|---|---|
| 1 | 1% | 20 min | error +0.20 pts | pass |
| 2 | 5% | 20 min | same | pass |
| 3 | 25% | 30 min | same + regional | hold, eu-west +0.41 pts |
| 4 | 50% | — | not started | rolled back to 0% |
I would not consider it settled without evidence: For the last 20 production changes, show the share schedule, any holds, and that no step skipped a gate.
Progressive delivery is a series of small bets with an automatic fold, not a delayed big-bang.
Curated: · Written: · Reviewed:
QA-48When is rollback the correct mitigation, and when is it the wrong reflex?(show answer)
The judgement in rollback criteria lies in what you refuse to do as much as in what you build.
Roll back when the change is still the leading hypothesis and is reversible inside the RTO; forward-fix only when rollback is impossible or would destroy data.
Concretely, pre-compute rollback for every change type, define SLI abort lines that trigger it, and practice the command in game days so it is faster than a design discussion.
The reason for that specificity is a failure I have seen: Debating a forward-fix for 26 minutes on a reversible feature flag left 18% 5xx in place; the flag off took 40 seconds once someone ran it.
Choose rollback versus forward-fix by reversibility.
| Change | Reversible in < 5 min? | Abort line | Action |
|---|---|---|---|
| app binary | yes, previous replica set | +0.20 pts error | rollback |
| feature flag | yes, 40 s | same | turn flag off |
| data backfill | no, rows already rewritten | — | forward-fix + pause writers |
| 26 min debate | yes | already crossed | wasted 26 min of 18% 5xx |
I would not consider it settled without evidence: Measure time from abort-line crossed to old version serving 99% of traffic, and keep that under the incident RTO.
Rollback is a designed product of the release system, not a confession that the team was wrong.
Curated: · Written: · Reviewed:
QA-49Should a risky behaviour change ship as a binary deploy or as a flag? How do you operate flags in production?(show answer)
The first thing to settle about feature flags versus deploys is what decision it actually changes.
Separate code presence from code execution so you can halt a behaviour without rolling back the artifact, and treat flag changes as production releases.
Concretely, ship dark, enable by cohort with the same SLI gates as a canary, require a named owner and an expiry, and log every flag mutation in the change timeline.
The reason for that specificity is a failure I have seen: A long-lived flag flipped at 02:00 without a ticket, with no soak, and with the binary 19 versions stale, producing a Sev-1 that "was not a deploy".
Flag rollout uses the same gates as the binary canary.
| Event | Artifact | User share | Gate |
|---|---|---|---|
| Tue 11:00 | binary v441 (flag off) | 100% | standard canary |
| Tue 14:00 | flag payments_v2 | 5% | error +0.20 pts, 20 min |
| Tue 14:40 | flag payments_v2 | 25% | same |
| Wed 02:00 (anti-pattern) | flag, no ticket | 100% | none, Sev-1 |
I would not consider it settled without evidence: Show flag changes in the same audit as deploys, with soak metrics, and a weekly list of flags older than 30 days.
A flag that bypasses the release process is an untracked production change with extra steps.
Curated: · Written: · Reviewed:
QA-50A config-map edit took checkout down. Why was that not "just config", and how do you ship it?(show answer)
I would start by making the assumption behind configuration as a release explicit.
Treat production configuration as a versioned release with the same canary, audit, and rollback as binaries, because the process reads it on the hot path.
Concretely, store config in git, render it immutably, roll it to a subset of instances first, and keep a one-command rollback to the previous config hash.
The reason for that specificity is a failure I have seen: An in-cluster edit of a 2-line timeout from 250 ms to 25000 ms bypassed CI, skipped the canary, and pinned every checkout thread on a hung vendor for 14 minutes.
Config change 250 ms -> 25,000 ms, two shipping paths.
| Path | Canary | Rollback | Outcome |
|---|---|---|---|
| kubectl edit | none | unknown previous | 14 min hang, 22% 5xx |
| git hash a1c3, 5% then 100% | 20 min SLI gate | kubectl apply -f config-a1c3.yaml | abort at 5%, 0.3% extra errors |
I would not consider it settled without evidence: Diff prod config against git HEAD every 5 minutes and page on drift; show that the last 10 config changes had canaries.
The runtime does not care that the bytes came from a ConfigMap instead of a container; neither should your release bar.
Curated: · Written: · Reviewed:
QA-51Live cluster state no longer matches the last merged manifest. How do you detect that, and what is the allowed repair?(show answer)
This is one of those areas where the default answer and the right answer for GitOps drift differ.
Make git the only write path, detect live drift continuously, and reconcile by applying git, never by editing the cluster and copying back.
Concretely, run a drift detector every 5 minutes, page when live objects differ from HEAD, disable kubectl mutate for humans except break-glass, and auto-heal from git.
The reason for that specificity is a failure I have seen: A 02:00 kubectl scale survived 11 days, then a GitOps sync reset replicas to 3 during a sale and dropped p99 from 140 ms to 2.6 s.
Drift on checkout replicas, detected at T+5 minutes.
# git HEAD: checkout Deployment spec.replicas: 12
# live: spec.replicas: 4 (kubectl scale at 02:00, no PR)
status:
drift: true
objects: 1
detected_after_s: 300
repair: apply HEAD # not "update git to 4"
sale_impact_if_unhealed: "p99 2.6s at 18:00"
I would not consider it settled without evidence: Count drift events per week, time-to-heal, and confirm no production mutate succeeded outside break-glass.
Two sources of truth is how a cluster acquires a folklore configuration that no review ever saw.
Curated: · Written: · Reviewed:
QA-52The default-branch pipeline is red 18% of the time and people just retry. Why is that a reliability defect?(show answer)
The useful framing here is to ask what evidence would change my mind about flaky pipelines.
A flaky pipeline trains engineers to ignore red builds, so a real regression waits behind retries until it reaches production.
Concretely, quarantine tests that fail without a product change, count retries as pipeline unavailability, and freeze merge when flake rate exceeds a budget such as 2%.
The reason for that specificity is a failure I have seen: An 18% flake rate hid a genuine auth regression for 9 retries over 2 days, then production 401s hit 6% of logins.
Default-branch health, 30 days, 2% flake budget.
| Metric | Value | Budget |
|---|---|---|
| Runs | 1,140 | — |
| Failed first try | 18.0% | — |
| Failed after 3 retries (real) | 1.1% | — |
| Flake rate | 16.9% | 2.0% |
| Auth regression time-to-detect | 2 days / 9 retries | 1 run |
I would not consider it settled without evidence: Publish flake rate, median retries-to-green, and the last regression that landed because a red build was retried rather than read.
Retry is not a test strategy; it is how you launder a red signal into a green one.
Curated: · Written: · Reviewed:
QA-53Retail wants a freeze from Black Friday through Cyber Monday. How do you freeze without blocking incident response?(show answer)
Before choosing an implementation I would establish what success looks like for deploy freezes.
Freeze routine change while keeping a rehearsed exception path for SLO-saving rollbacks and security fixes.
Concretely, close the merge queue to all but a freeze-exception label, require dual approval for that label, and keep rollback jobs exempt and tested on day one of the freeze.
The reason for that specificity is a failure I have seen: A hard freeze that also blocked rollbacks turned a 6-minute bad canary into a 4-hour wait for a VP exception during the highest-revenue window of the year.
Black Friday freeze policy, 4 days, 99.9% SLO still in force.
| Change type | Allowed? | Gate |
|---|---|---|
| Feature merge | no | queue closed |
| Config tweak | no | same |
| Rollback to last green | yes | automated, no extra approval |
| Sev-1 security patch | yes | freeze-exception + 2 approvers |
| Bad canary at 5% | rollback | 6 min, not a VP ticket |
I would not consider it settled without evidence: On freeze day, demonstrate a rollback in staging, list who can approve exceptions, and show the queue rejects a normal feature PR.
A freeze that cannot roll back is a vow to sit in any production fire you already started.
Curated: · Written: · Reviewed:
QA-54When do you cut all traffic to a new environment, and when do you drip it?(show answer)
The engineering question underneath blue-green versus canary is which failure is acceptable.
Use canary when you need production traffic to falsify a hypothesis cheaply; use blue-green when you need instant cutback and the two environments can both carry 100% load.
Concretely, run a canary as a traffic share with SLI gates, or run blue-green as two full stacks with a health-checked green, a balancer flip, and a warm blue sized for the RTO rather than a 100% flip onto an untested green.
The reason for that specificity is a failure I have seen: A blue-green flip of an unsoaked green environment took 100% of checkout to a missing IAM role, and the "instant cutback" was 11 minutes of DNS because blue had been scaled to zero to save money.
Same release, two topologies, two failure costs.
| Topology | First users exposed | Cutback | Last miss |
|---|---|---|---|
| Canary 1% | 1% for 20 min | instant to 0% | would have caught IAM miss |
| Blue-green, blue warm | 100% at flip | 40 s balancer | IAM miss = 100% Sev-1 |
| Blue-green, blue scaled 0 | 100% at flip | 11 min reprovision | freeze-time outage |
I would not consider it settled without evidence: State the RTO of the cutback and prove it in a drill; for canary, show the smallest share that would have caught the last three bugs.
Blue-green without a live blue is just a big-bang deploy with extra infrastructure.
Curated: · Written: · Reviewed:
QA-55How do you change a production table without a long blocking lock on checkout or making old binaries crash?(show answer)
My approach to production schema migrations starts from the constraint rather than from the technique.
Expand-contract in separate releases so every running binary can read both shapes, and keep the expand step's unavoidable lock acquisition short and bounded.
Concretely, use expand-contract across separate releases; verify the database engine’s exact DDL behavior, set short lock timeouts, use online migration tooling where needed, backfill in throttled chunks, test a mixed-version fleet, and contract only after old readers and writers are gone.
The reason for that specificity is a failure I have seen: An in-place ALTER with a volatile default rewrote the checkout table while holding an ACCESS EXCLUSIVE lock, blocking 2.1 million writes over 14 minutes.
Expand-contract on orders.tax_cents, 2.1e6 rows, bounded lock and no table rewrite.
PR1: ALTER ADD COLUMN tax_cents INT NULL; # short ACCESS EXCLUSIVE lock, no rewrite
PR2: backfill 50k rows / 2s, 84 min, lag SLO 120s
PR3: binaries write tax_cents; mixed fleet 24 h
PR4: ALTER DROP COLUMN tax_legacy; # after 0 old binaries
# PostgreSQL 11+: constant DEFAULT is metadata-only; a volatile DEFAULT still rewrites rows.
# Any ALTER can wait on or block behind its ACCESS EXCLUSIVE lock, so set lock_timeout.
I would not consider it settled without evidence: Show the four PRs, the backfill rate, and that old and new binaries passed a mixed-fleet test for 24 hours.
A migration that requires a maintenance window is an admission that the change was not sequenced.
Curated: · Written: · Reviewed:
QA-56A pod is killed in a loop while it is only slow to warm. Which probe is wrong?(show answer)
I would treat liveness versus readiness probes as a design decision with a stated trade-off rather than a best practice.
Readiness removes a pod from traffic until it can serve; liveness restarts a pod that is deadlocked, and using liveness as readiness turns slowness into a crash loop.
Concretely, use a startup probe to protect slow initialization, readiness to report whether this pod can currently accept traffic, and liveness only for local unrecoverable process failure. Include downstream state in readiness only when losing that dependency truly makes the pod unable to serve and cannot cause a fleet-wide readiness collapse.
The reason for that specificity is a failure I have seen: A liveness probe that hit payments timed out during a 4-second vendor blip, Kubernetes restarted 80 pods, and the herd of cold starts extended a 4-second blip to 11 minutes.
Probe split for checkout, warmup 40 s, vendor blip 4 s.
readinessProbe:
httpGet: {path: /ready, port: 8080} # local ability to serve, not shared payments
initialDelaySeconds: 40
timeoutSeconds: 1
livenessProbe:
httpGet: {path: /live, port: 8080} # in-process only, no payments
periodSeconds: 10
failureThreshold: 18 # 3 min before restart
# Wrong: liveness=/ready -> 80 restarts, 11 min herd.
I would not consider it settled without evidence: Fail a shared downstream and verify the service follows its intended fallback without restart or total endpoint removal; separately deadlock the process and verify liveness restarts it.
Restarting a slow pod is how you convert a brownout into a thundering herd of cold starts.
Curated: · Written: · Reviewed:
QA-57Pods OOM at 600 mi while CPU HPA is happy. Do you raise replicas or raise limits?(show answer)
The place where teams go wrong on HPA versus VPA is usually the step before the one they are debating.
Use HPA to add replicas for concurrency, and VPA or a measured request/limit change for per-pod resource shape; they solve different shortages.
Concretely, if latency rises with in-flight work, add pods via HPA; if a single replica OOMs or CPU-throttles at low QPS, fix requests/limits (optionally VPA in recommend mode) before adding more of the same broken shape.
The reason for that specificity is a failure I have seen: HPA scaled an OOM-prone pod from 4 to 30, multiplying restart storms, while the working set was 780 Mi and the limit was 600 Mi.
OOM versus concurrency: pick VPA/limits first.
| Symptom | Working set | Limit | Concurrency | Tool |
|---|---|---|---|---|
| OOMKilled every 11 min | 780 Mi | 600 Mi | 4 rps/pod | raise limit, not HPA |
| p99 1.4 s, CPU 85% | 410 Mi | 1 Gi | 22 rps/pod | HPA |
| VPA recommend | 900 Mi request | — | — | apply in a release, not live surprise |
I would not consider it settled without evidence: Show working-set histograms versus limits, and a load test where HPA alone does or does not keep p99 inside SLO.
Thirty copies of a pod that cannot hold its heap is not capacity, it is a coordinated crash.
Curated: · Written: · Reviewed:
QA-58A node upgrade drained 6 of 6 checkout pods at once. What should have blocked that?(show answer)
What matters most about pod disruption budgets is whether the result can be checked afterwards.
Set a PodDisruptionBudget that preserves the minimum healthy replicas the SLO needs, so voluntary drains cannot take the service below that floor.
Concretely, calculate the required healthy replica floor, enforce it with a PDB for voluntary evictions, spread replicas across failure domains, and separately provision capacity for involuntary node or AZ loss.
The reason for that specificity is a failure I have seen: No PDB allowed a rolling node AMI update to evict all 6 replicas in 40 seconds, producing a 7-minute complete checkout outage during a "non-event" upgrade.
PDB for 6 checkout pods, peak 1,200 rps, 250 rps/pod.
apiVersion: policy/v1
kind: PodDisruptionBudget
metadata: {name: checkout}
spec:
minAvailable: 5 # 5 * 250 rps = 1,250 > 1,200 peak
selector: {matchLabels: {app: checkout}}
# Without PDB: 6 evicted in 40 s, 0 rps served for 7 min.
I would not consider it settled without evidence: Run a drain of one node in staging and show evictions pause when the PDB would be violated; count production drains that waited rather than evicted.
A drain is a voluntary outage unless a budget says how many pods must remain.
Curated: · Written: · Reviewed:
QA-59Pods with no CPU request schedule densely, then p99 explodes under load. What do you set, and why not unlimited?(show answer)
I would answer this by separating what CPU requests and limits guarantees from what it merely usually does.
Set requests to the measured working CPU so the scheduler packs honestly, and set limits only if you accept throttle; choose Burstable or Guaranteed QoS deliberately.
Concretely, set CPU requests from measured sustained demand so scheduling and HPA are meaningful; set CPU limits only after testing their tail-latency effect, and use Guaranteed QoS only when every container has equal CPU and memory requests and limits.
The reason for that specificity is a failure I have seen: Unset requests packed 18 latency-critical pods onto a 4-core node; under load they stole each other's cycles and p99 went from 160 ms to 2.1 s with a green HPA.
Checkout QoS: Burstable to preserve CPU bursts, measured 280 millicores p95.
| Field | Value | If omitted |
|---|---|---|
| request.cpu | 300m | packed 18-on-4 cores, p99 2.1 s |
| limit.cpu | omit or size above tested bursts | setting 300m can throttle latency-sensitive bursts |
| throttle_seconds[peak] | 0.4 / min | 40 / min when limit 100m |
| HPA CPU target | 70% of request | meaningless without a request |
I would not consider it settled without evidence: Show a node that cannot overcommit checkout requests, and a graph of throttle seconds remaining near zero at peak.
The scheduler can only protect a pod whose request is a true reservation, not a polite hint.
Curated: · Written: · Reviewed:
QA-60How do you drain a node in production without a user-visible blip?(show answer)
The judgement in node drains lies in what you refuse to do as much as in what you build.
Drain only after replacements are Ready and the PDB is satisfied, and fail the drain if it would violate the budget or the latency SLO.
Concretely, cordon the node, ensure spare schedulable capacity or pre-scale the workload, drain through the Eviction API so the PDB is honored, let the workload controller replace evicted pods, and abort rather than bypassing the PDB or termination grace period.
The reason for that specificity is a failure I have seen: A 15-minute drain timeout plus force deleted 9 pods that still had in-flight checkouts, cutting 1,200 in-flight payments mid-TCP.
Safe drain sequence, 12 checkout pods, PDB minAvailable 10.
t=0 cordon node-7
t=90s 2 replacement pods Ready in other nodes (HPA)
t=95s kubectl drain node-7 --timeout=8m # waits on PDB
t=4.1m 2 pods evicted, 10 remain serving
forced_evictions: 0
checkout_error_delta: +0.01 pts (inside noise)
# Anti-pattern: --force --grace-period=0 at t=15m -> 1,200 cut payments.
I would not consider it settled without evidence: Record drain duration, forced evictions (target 0), and SLI burn during each production drain.
Force-deleting a pod is an unplanned SIGKILL of user work, not a maintenance trick.
Curated: · Written: · Reviewed:
QA-61A database password lives for 18 months. How do you rotate it without downtime?(show answer)
The first thing to settle about secrets rotation is what decision it actually changes.
Overlap two valid secrets for a window, roll consumers to the new one, then retire the old, because a simultaneous change is an outage with extra ceremony.
Concretely, issue secret N+1, deploy consumers that try N+1 then N, confirm zero authentications on N, disable N, and alert if age exceeds the policy such as 30 days.
The reason for that specificity is a failure I have seen: A single-shot password change at 03:00 raced 40 pods; 12 still used the old password and 5xx'd for 19 minutes until the next kubelet sync.
30-day DB password overlap, two ids live for 40 minutes.
| Minute | Valid ids | Pods on N+1 | Auth failures |
|---|---|---|---|
| 0 | N | 0 / 40 | 0 |
| 5 | N and N+1 | 0 / 40 | 0 |
| 25 | N and N+1 | 40 / 40 | 0 |
| 40 | N+1 only | 40 / 40 | 0 |
| single-shot (anti-pattern) | N+1 only at t=0 | 28 / 40 | 12 pods x 19 min |
I would not consider it settled without evidence: Show auth metrics for both secret ids during the overlap, then a zero on the old id before disable.
Rotation that requires a global restart is a planned outage, not a security control.
Curated: · Written: · Reviewed:
QA-62Developers want to mount all secrets as environment variables. What do you inject, and how?(show answer)
I would start by making the assumption behind secret injection explicit.
Inject the minimum secret at process start from a runtime identity, never from git or a shared env file, and prefer tmpfs files over env because env leaks to child processes and crash dumps.
Concretely, use workload identity to fetch from a manager, render onto a memory-backed volume, grant the container only those keys, and block env-based injection in the admission policy except for a documented exception list.
The reason for that specificity is a failure I have seen: A secret injected into the application environment was exposed through /proc, inherited by a spawned diagnostic child process, and copied from a crash dump into a ticket for six hours.
Admission policy: files on tmpfs, not env, identity-bound.
volumeMounts:
- name: payments-token
mountPath: /var/run/secrets/payments
readOnly: true
volumes:
- name: payments-token
csi:
driver: secrets-store.csi.k8s.io
readOnly: true
# env PAYMENT_TOKEN: denied by policy (leaked in crash dump, ticket T-9182)
I would not consider it settled without evidence: Scan running pods for secrets in env, scan git for secret material, and show the last deploy fetched a short-lived credential bound to the workload identity.
If a secret can appear in ps, a crash dump, or a CI log, the injection path is the incident.
Curated: · Written: · Reviewed:
QA-63A Terraform apply in the "shared" stack took down billing and search. How do you bound blast radius?(show answer)
This is one of those areas where the default answer and the right answer for IaC blast radius differ.
Split state and apply scope so one plan cannot mutate unrelated production surfaces, and treat a complete plan as the normal change rather than a targeted apply.
Concretely, split state by bounded ownership and failure domain, review complete plans for each state, and reserve targeted applies for documented break-glass recovery rather than normal deployment.
The reason for that specificity is a failure I have seen: A one-character CIDR typo in a shared networking stack applied to 14 services; 9 lost database routes for 23 minutes.
Split stacks; the CIDR typo can no longer reach billing.
| Stack | Resources | Who applies | Typo blast |
|---|---|---|---|
| net-prod (old) | VPC, all routes, all SGs | anyone with prod | 14 services, 23 min |
| net-prod-search | search subnets + SGs | search owners | search only |
| net-prod-billing | billing subnets + SGs | billing owners | billing only |
| checkout-prod | checkout IAM + deploy | checkout | checkout only |
I would not consider it settled without evidence: Show that a plan for service A cannot include resources of service B, and that a networking change lists the exact resource addresses.
Shared state is a blast radius you reaffirm on every apply.
Curated: · Written: · Reviewed:
QA-64Is a 02:00-04:00 change window safer than shipping at 14:00 with a full staff?(show answer)
The useful framing here is to ask what evidence would change my mind about change windows.
Prefer changes when detection and rollback staffing are highest, and use windows only to avoid known peak traffic, not to hide risk in the night.
Concretely, schedule risky changes in staffed hours below the traffic peak, require on-call plus a second engineer, and reserve night windows for the rare change that cannot overlap peak.
The reason for that specificity is a failure I have seen: A 03:10 schema apply with one sleepy primary took 47 minutes to notice a lock; the same change at 14:00 with two people rolled back in 6 minutes in rehearsal.
Same schema apply, two clocks.
| When | Staff | Traffic | MTTD | Rollback |
|---|---|---|---|---|
| 03:10 | 1 primary | 12% of peak | 47 min | 19 min |
| 14:00 | primary + SWE | 70% of peak | 3 min | 6 min |
| 19:00 peak | full | 100% | — | forbidden by policy |
I would not consider it settled without evidence: Compare MTTD and rollback time for staffed-hour versus night changes over a quarter.
A quiet graph at 03:00 is not safety; it is a missing audience for the alarm.
Curated: · Written: · Reviewed:
QA-65Engineers SSH to production boxes daily to "just check". What access model do you put in its place?(show answer)
Before choosing an implementation I would establish what success looks like for production access.
Default production to no standing access; grant time-bounded, ticket-linked roles for a named task, and do the rest through debug tooling that is audited.
Concretely, remove persistent SSH keys, issue short-lived certificates or SSM sessions bound to a change ticket, record session video or command logs, and review weekly.
The reason for that specificity is a failure I have seen: A standing SSH key on a laptop that was stolen at a conference stayed valid for 16 days and had write access to the payments bastion.
Session grant: 60 minutes, ticket INC-4412, commands logged.
who: sre-lee
role: checkout-debug-ro
ttl: 60m
ticket: INC-4412
path: ssm start-session --target i-0ab1c --reason INC-4412
standing_ssh_keys: 0
stolen-laptop window under old model: 16 days write on payments bastion
I would not consider it settled without evidence: Count standing credentials (target 0), median grant TTL, and the share of sessions with a ticket id.
Standing production access is a permanent incident waiting for a laptop to leave a bag.
Curated: · Written: · Reviewed:
QA-66When every control is locked down, how does on-call still mitigate a Sev-1?(show answer)
The engineering question underneath break-glass access is which failure is acceptable.
Provide a break-glass path that is fast, fully audited, and painful enough afterwards that it is not the daily workflow.
Concretely, a sealed role that opens in < 2 minutes, pages security when assumed, expires in 60 minutes, and automatically opens a review ticket that cannot be closed without a diff of actions.
The reason for that specificity is a failure I have seen: Hiding break-glass behind a 45-minute security-manager approval stretched a Sev-1 from 8 minutes of mitigation to 53 minutes of waiting.
Sev-1 break-glass drill, target < 2 minutes to credentials.
| Step | Target | Drill |
|---|---|---|
| Assume checkout-break-glass | < 2 min | 71 s |
| Security page | immediate | 12 s |
| TTL | 60 min | 60 min |
| Review ticket | auto | 48 API calls listed |
| Manager approval before assume | — | rejected, that was 45 min |
I would not consider it settled without evidence: Time a drill from decision to usable credentials, confirm the security page fired, and show the review ticket lists every API call.
Break-glass that is slower than the outage is decoration; break-glass that is easier than the default path is the default path.
Curated: · Written: · Reviewed:
QA-67Which checks belong in CI because they prevent production incidents, not because they are style?(show answer)
My approach to CI as a reliability control starts from the constraint rather than from the technique.
Put in CI the gates that would have stopped your last year of change-caused outages: tests of SLI-critical paths, migration safety, IAM diffs, and secret scanning.
Concretely, map each change-caused incident to a missing gate, add that gate as blocking, and measure the share of production incidents that still had a green pipeline.
The reason for that specificity is a failure I have seen: A green pipeline that ran unit tests but not a migration expander shipped the locking ALTER, 14 minutes of blocked writes, and a "CI was green" postmortem line.
Gates mapped to last year's change-caused minutes.
| Gate | Incident class caught | Minutes last year |
|---|---|---|
| expand-contract linter | locking ALTERs | 210 |
| IAM plan diff required | overly broad policy | 88 |
| secret scan | tokens in repo | 40 |
| checkout SLI e2e | auth header regression | 160 |
| unit tests only (old CI) | none of the above | 0 caught |
I would not consider it settled without evidence: For 12 months of change-caused Sev-1/2s, show which would now fail CI, and keep that catch rate rising.
A pipeline that cannot fail the class of change that pages you is a compiler, not a reliability control.
Curated: · Written: · Reviewed:
QA-68A "latest" tag on the registry was overwritten. How do you know what actually runs in production?(show answer)
I would treat artifact provenance as a design decision with a stated trade-off rather than a best practice.
Deploy by digest from a signed provenance record that names git SHA, builder, and tests, and refuse tags that can move.
Concretely, build once, sign the digest, store SLSA-style provenance, and have the cluster admit only signatures from the CI identity; pin Deployments to sha256.
The reason for that specificity is a failure I have seen: Someone retagged latest to an untested laptop build; 30% of pods pulled it on the next restart and served a debug backdoor for 3.5 hours.
Admission: digest + signature, not :latest.
image: registry.example/checkout@sha256:9f2c41ab77de0031c0ffee4411aa22bb
# provenance:
# git: 7a11e4c
# builder: github-oidc:repo:acme/checkout
# tests: slI-e2e passed
# :latest retag at 11:04 -> 30% pods, 3.5 h debug binary, unsigned
I would not consider it settled without evidence: For a running pod, show digest, signature, git SHA, and that the same digest is in the provenance log.
A movable tag is an unsigned instruction to production to trust a stranger later.
Curated: · Written: · Reviewed:
QA-69A feature flag and a JSON config both change behaviour. How do you roll config to 5% of users rather than 5% of pods?(show answer)
The place where teams go wrong on progressive config rollout is usually the step before the one they are debating.
Target config by user or request cohort, not only by replica, or you pin unlucky users to a sticky bad pod and call it a canary.
Concretely, evaluate config in the request path from a central store with cohort hashing, soak on SLI by cohort, and keep pod-level rollout for binary-incompatible config only.
The reason for that specificity is a failure I have seen: Pushing a bad timeout to 5% of pods stuck a subset of sticky-session users on those pods at 100% error, while the fleet-wide error rate looked like 5% and passed the gate.
Cohort hash versus pod percent, sticky sessions on.
| Rollout unit | What 5% means | User experience | Gate saw |
|---|---|---|---|
| 5% of pods | unlucky stickies 100% bad | some users fully down | +5 pts mixed |
| 5% of user ids | each treated user 100% on new | 5% of users | cohort SLI +0.9 pts, abort |
| rollback cohort | hash off | instant | no restart |
I would not consider it settled without evidence: Show cohort-level SLIs for the config version, and that a user hashed into treatment can be moved back without a pod restart.
A 5% pod canary is not a 5% user canary when sessions stick.
Curated: · Written: · Reviewed:
QA-70Cluster A is on fire. How do you fail closed or over to cluster B without a split brain on writes?(show answer)
What matters most about multi-cluster failover is whether the result can be checked afterwards.
Promote a new writer only after it has acquired exclusive authority through quorum or an external fencing mechanism, and ensure the former writer can no longer commit writes.
Concretely, run B as a warm replica, replicate state with a measured lag SLI, shift reads first, promote writes with a fencing token, and block A from writing after promotion.
The reason for that specificity is a failure I have seen: A DNS flip sent writes to both clusters for 8 minutes; inventory decremented twice and 1,140 orders oversold.
Write fencing on promotion, RTO 15 min, RPO 60 s.
t=0 A control plane dead, lag on B = 22 s (inside 60 s RPO)
t=2 min reads -> B (read-only)
t=6 min fencing token epoch=14 on B, A tokens revoked
t=7 min writes -> B
t=8 min (anti-pattern without fence): both accept writes, 1,140 oversold
drill RTO: 7 min (inside 15)
I would not consider it settled without evidence: In a drill, show lag at promotion, fencing of A, and order-id uniqueness holding; time the whole path against RTO.
Failover without fencing is two primaries, which is a data incident with a networking story.
Curated: · Written: · Reviewed:
QA-71Hit rate is 97% and the origin is sized for 3%. Why is that a reliability problem?(show answer)
I would answer this by separating what cache as a reliability hazard guarantees from what it merely usually does.
Treat cache hit rate as a dependency: size the origin for a miss storm, and make a cache failure degrade rather than multiply load by 1/(1-hit rate).
Concretely, provision origin for a stated miss-rate floor, add request coalescing, and pre-warm a new key generation before a staggered cutover so a deploy does not create a fleet-wide cold miss; page on origin QPS, not only cache process health.
The reason for that specificity is a failure I have seen: A redis restart dropped a 97% hit rate to 0%; origin received 33x traffic and melted in 40 seconds, taking search with it.
Origin load at 97% hit versus a flush.
| State | Hit rate | Origin QPS | User p99 |
|---|---|---|---|
| steady | 97% | 300 | 40 ms |
| sized for | 70% floor | 3,000 | 90 ms |
| redis restart | 0% | 10,000 (33x) | timeout |
| coalesced + shed | 0% | 3,000 cap | 429 on overflow |
I would not consider it settled without evidence: Fail the cache in staging and show origin QPS, shed behaviour, and user SLI; the origin must survive the floor miss rate.
A 97% hit rate is a 33x landmine unless the origin was bought for the other 3% becoming 100%.
Curated: · Written: · Reviewed:
QA-72A popular key expires and 4,000 requests miss together. How do you stop the stampede?(show answer)
The judgement in cache stampedes lies in what you refuse to do as much as in what you build.
Coalesce misses on a key and serve stale while one winner refills, so expiry is a background refresh rather than a synchronized origin flood.
Concretely, use singleflight or a lock per key, set stale-while-revalidate, add TTL jitter so keys planted together do not die together, and bound origin concurrency.
The reason for that specificity is a failure I have seen: A homepage key with TTL 60 s and no jitter expired on the minute; 4,000 requests missed at once and the origin 5xx'd for 90 seconds every minute.
Homepage key: 4,000 waiters, one origin call.
# singleflight: 4,000 concurrent Get("home") -> 1 origin fetch
waiters, origin_calls = 4000, 1
stale_age_s = 8
# TTL jitter: 60s * uniform(0.8, 1.2) -> expiries spread over 24 s
# Without: 4,000 misses at t=60, 90 s of origin 5xx, every minute.
I would not consider it settled without evidence: Expire a hot key under load and count origin requests (target 1) while clients still receive a 200.
The popular key is the one whose expiry you cannot afford to treat as a miss.
Curated: · Written: · Reviewed:
QA-73Producers keep publishing while consumers lag by 2.4 million messages. What backpressure do you implement?(show answer)
The first thing to settle about queue backpressure is what decision it actually changes.
Push back to the producer when lag exceeds a time budget, because an unbounded queue only postpones overload into a longer recovery.
Concretely, measure lag in time not only depth, reject or slow producers at a threshold such as 30 seconds of lag, scale consumers on lag, and alert on time-to-drain.
The reason for that specificity is a failure I have seen: A 2.4e6 backlog at 80 rps drain needed 8.3 hours; producers added 200 rps the whole time, so lag never fell and freshness SLO burned continuously.
Lag in time: 2.4e6 messages, 80 rps drain, producers still at 200 rps.
depth, drain, produce = 2_400_000, 80, 200
seconds_to_drain_if_paused = depth / drain # 30,000 s = 8.3 h
net_rps = produce - drain # +120, lag grows
# Policy: if depth/drain > 30 s, return 429 to producers.
I would not consider it settled without evidence: In a consumer-pause test, show producers receiving 429/backpressure before lag crosses the freshness SLO.
A queue that always accepts is polite to producers and cruel to every user waiting on the consumer.
Curated: · Written: · Reviewed:
QA-74One bad payload retries forever and blocks the partition. How do you isolate it?(show answer)
I would start by making the assumption behind poison messages explicit.
After a bounded retry budget, park the message on a dead-letter path and continue the stream, then page on dead-letter depth.
Concretely, set max attempts, capture the payload and trace id on the dead-letter topic, keep the main consumer moving, and require an owner to replay or drop.
The reason for that specificity is a failure I have seen: A null-price order retried 14,000 times on a single Kafka partition, stalling 40,000 good orders behind it for 3.1 hours.
Retry budget 5, then dead-letter, partition unblocked.
| Step | Main partition lag | DLQ depth | Page? |
|---|---|---|---|
| poison arrives | 0 | 0 | no |
| attempts 1–5 | +1 | 0 | no |
| attempt 6 | 0 (skipped) | 1 | yes |
| 40,000 good orders | 0 | 1 | — |
| infinite retry (old) | 40,000 | 0 | 3.1 h stall |
I would not consider it settled without evidence: Inject a permanently failing payload and show the main lag stays flat while dead-letter depth goes to 1 and a page fires.
Infinite retry of a poison payload is how one bad byte becomes a freshness outage.
Curated: · Written: · Reviewed:
QA-75Two regions cannot talk for 9 minutes. What does the service still promise, and what does it refuse?(show answer)
This is one of those areas where the default answer and the right answer for consistency under partition differ.
Choose explicit behavior per journey under partition: preserve invariants by rejecting or fencing writes, or accept divergence only with a defined reconciliation rule; separately document the normal-operation latency/consistency trade-off.
Concretely, for money and inventory, fail writes that cannot reach a quorum; for feeds, serve stale with a freshness SLI; never let both sides accept conflicting unique writes.
The reason for that specificity is a failure I have seen: Both regions accepted "create account" during a 9-minute partition and minted 640 duplicate user ids, then merge was a 6-hour manual dedupe.
9-minute inter-region partition, per-journey choice.
| Journey | During partition | SLO |
|---|---|---|
| checkout pay | 503 unless quorum | availability drops, correctness holds |
| inventory decrement | 503 | no oversell |
| homepage feed | stale, 4 min old | freshness SLI burns, page still 200 |
| create account (old) | both accept | 640 duplicate ids |
I would not consider it settled without evidence: Run a partition test and show which APIs return 503, which serve stale, and that unique ids stay unique.
Hoping both sides stay consistent during a partition is a third, unnamed consistency model, and it loses.
Curated: · Written: · Reviewed:
QA-76A payment event is delivered twice. How does the consumer avoid charging twice?(show answer)
The useful framing here is to ask what evidence would change my mind about idempotent consumers.
Make each external side effect idempotent at the system that performs it, and use a durable inbox or state machine locally so crash recovery resumes the same operation.
Concretely, persist an inbox record and state transition under a unique event ID, call the payment provider with a stable provider-supported idempotency key, and persist the returned result; retries must query or repeat the same key rather than issue a new charge.
The reason for that specificity is a failure I have seen: A consumer that charged then wrote the id crashed between the two; the retry charged again and 220 customers saw a duplicate $19.00 line.
Reserve the event locally, then reuse one provider idempotency key across every recovery path.
def consume(event):
with tx():
op = inbox.reserve(event.id, args_hash=hash(event.cents)) # unique
if op.args_hash != hash(event.cents): raise KeyReuseConflict()
if op.status == "done": return op.result
result = payments.charge(key=event.id, cents=event.cents) # provider dedupes
with tx(): inbox.finish(event.id, result)
# Crash test: 2 deliveries, 1 charge of 1900 cents, 0 duplicates.
I would not consider it settled without evidence: Kill the process between charge and record in a test, replay the event, and assert one charge and one ledger row.
At-least-once delivery is a promise of duplicates; idempotency is the only reason that is survivable.
Curated: · Written: · Reviewed:
QA-77A partner retries failed webhooks every second with no cap. How do you protect the service and still accept the event once?(show answer)
Before choosing an implementation I would establish what success looks like for webhook retry storms.
Authenticate, accept quickly with an idempotency key, and rate-limit per partner, or their retry policy becomes your outage.
Concretely, return 2xx only after the event is durably queued, dedupe on partner event id, apply a per-partner QPS cap, and publish a Retry-After when shedding.
The reason for that specificity is a failure I have seen: A 4-second 500 from our side caused a partner to retry at 1 Hz from 60 workers; 60 QPS of the same order id saturated the payments pool for 17 minutes.
Accept in < 50 ms, dedupe, cap 20 rps per partner.
| Step | Cost | Duplicate 10,000 | Over cap |
|---|---|---|---|
| auth + enqueue | 18 ms | 1 job | 429 + Retry-After: 5 |
| charge inline then 200 (old) | 210 ms | 60 QPS storms | pool 100% 17 min |
| partner workers | 60 at 1 Hz | same event id | our incident |
I would not consider it settled without evidence: Replay 10,000 duplicate webhooks of one id and show one queued job; then exceed the partner cap and show 429 with Retry-After.
A webhook endpoint that does expensive work before 200 is volunteering to run the partner's retry loop on your CPU.
Curated: · Written: · Reviewed:
QA-78The payments call has no deadline and checkout threads pile up. How do you set timeouts end to end?(show answer)
The engineering question underneath API timeouts is which failure is acceptable.
Give every outbound call a deadline smaller than the caller's remaining budget, and fail when it expires rather than waiting forever.
Concretely, start with the user SLO, subtract a margin, budget each hop, set client timeouts and server-side deadlines to that, and propagate remaining time.
The reason for that specificity is a failure I have seen: A 30-second default HTTP client timeout on a 300 ms SLO held 900 threads during a payments hang, then the load balancer 504'd everyone else.
300 ms checkout SLO, hop budgets, 50 ms margin.
user SLO p99: 300 ms
margin: 50 ms
auth: 40 ms deadline
catalog: 60 ms
payments: 120 ms # was 30,000 ms default
tax (cached): 30 ms
# Hang payments: caller 120 ms -> 503, threads +4% not +900.
I would not consider it settled without evidence: Hang the dependency in staging and show the caller returns inside the budget with a 503/504, and that thread count stays bounded.
An unset timeout is a promise to wait longer than the user will.
Curated: · Written: · Reviewed:
QA-79A single API key is 40% of QPS. How do you rate-limit without becoming a Sev-1 for everyone else?(show answer)
My approach to rate limiting starts from the constraint rather than from the technique.
Enforce per-principal limits at the edge, isolate noisy neighbors, and keep a reserved slice for anonymous or new keys so one key cannot consume the admission cap.
Concretely, token-bucket per API key and per IP, a global reserved band for the rest, 429 with Retry-After, and dashboards that show shed per principal.
The reason for that specificity is a failure I have seen: A global 1,600 rps cap with no per-key split let one integration eat 1,520 rps of retries and 429 everyone else, including checkout.
Edge buckets: per-key 200 rps, reserved 400 rps for others.
| Principal | Offered | Admitted | 429 |
|---|---|---|---|
| key_9f2c (broken retries) | 1,520 rps | 200 | 1,320 |
| other keys combined | 900 rps | 900 | 0 |
| global cap only (old) | 2,420 rps | 1,600 mixed | checkout in the 429s |
I would not consider it settled without evidence: Run one key at 10x its contract and show that key 429'd while other keys' success rates stay inside SLO.
A single shared bucket is how one broken client becomes a site-wide outage.
Curated: · Written: · Reviewed:
QA-80You want two regions serving writes. What has to be true for that not to split inventory?(show answer)
I would treat multi-region active-active as a design decision with a stated trade-off rather than a best practice.
Active-active writes require either a conflict-free datatype, a globally serialized invariant, or pinning each write-set to one region; otherwise you will double-spend.
Concretely, classify every write as CRDT-mergeable, sticky-primary, or globally coordinated; put inventory and payments in the sticky or coordinated class; run a partition drill before going live.
The reason for that specificity is a failure I have seen: Active-active increment of remaining_stock in two regions oversold 1,140 SKUs during a 9-minute link cut because both read 4 remaining and both wrote 3.
Write classes before flipping a second primary.
| Write | Class | During partition |
|---|---|---|
| page view counter | CRDT | both accept, merge |
| remaining_stock | sticky-primary | only primary region writes |
| capture payment | coordinated / sticky | never both |
| remaining_stock as naive +1/+1 | race | 1,140 oversold |
I would not consider it settled without evidence: Under a 10-minute partition, show stock never goes negative and payments are not captured twice; document the class of each write.
Two writers on a number that must stay non-negative is not active-active, it is a race with extra latency.
Curated: · Written: · Reviewed:
QA-81Product says "we cannot lose more than 5 minutes of orders and we must be up in 15". What do you actually build?(show answer)
The place where teams go wrong on RPO and RTO is usually the step before the one they are debating.
RPO is the maximum tolerable data-loss interval, while RTO is the maximum tolerable restoration time; replication, backups, and spare capacity must be designed and drilled to meet both.
Concretely, replicate orders off-AZ within 5 minutes worst lag, keep a warm spare that can pass health checks in 15 minutes, monitor lag as an SLI, and refuse a backup-only design that snapshots hourly.
The reason for that specificity is a failure I have seen: Hourly snapshots were sold as 5-minute RPO; a 14:40 primary death lost 41 minutes of orders and 3.2 hours of restore from a cold AMI.
5-minute RPO / 15-minute RTO versus what the old design delivered.
| Design | Worst data loss | Time to first 200 | Meets? |
|---|---|---|---|
| hourly snapshot + cold AMI | 41 min | 3.2 h | no / no |
| async replica, lag SLI 90 s | 90 s | 12 min promote | yes / yes |
| sync cross-region | ~0 | 8 min | yes, + write latency |
I would not consider it settled without evidence: Show max replication lag over 30 days versus 5 minutes, and a drill clock from declared disaster to first successful checkout versus 15 minutes.
An untested RTO is a wish; an RPO shorter than the worst recoverable backup or replication interval is unmet by design.
Curated: · Written: · Reviewed:
QA-82Backups are green every night. How do you know they restore?(show answer)
What matters most about backup restore drills is whether the result can be checked afterwards.
A backup is only real if a restore drill meets RPO and RTO on a schedule, from the same artifacts the alarm thinks are healthy.
Concretely, once a month, restore into an isolated environment, run a checksum against a known watermark, time the clock, and page if the drill misses RTO or the data is incomplete.
The reason for that specificity is a failure I have seen: 14 months of green backup jobs hid a missing --single-transaction flag; the first real restore applied a torn dump and failed foreign keys after 2.1 hours.
Monthly restore drill against a 15-minute RTO, 5-minute RPO.
| Check | Target | Last drill |
|---|---|---|
| restore clock | <= 15 min | 11.4 min |
| watermark vs disaster time | <= 5 min | 3.1 min |
| orders row count vs source | 0 diff | 0 |
| torn dump (old, no drill) | — | FK errors at 2.1 h |
I would not consider it settled without evidence: Publish the last drill's duration, row-count diff, and watermark time versus RPO.
The backup job succeeding is a test of the job, not of the restore.
Curated: · Written: · Reviewed:
QA-83You flip a DNS record to a spare region. What else has to be true for users to actually move?(show answer)
I would answer this by separating what DNS failover guarantees from what it merely usually does.
Failover only as fast as TTL, resolver caches, and health-check truth allow, so a 60-second TTL with 300-second resolvers is a 5-minute RTO lie.
Concretely, keep TTL low on the failover name, use health-checked records, flush or bypass sticky resolvers where you can, and measure time-to-shift from the users' resolvers, not from dig on your laptop.
The reason for that specificity is a failure I have seen: A 60 s TTL flip left 38% of mobile users on the dead IP for 11 minutes because carrier resolvers clamped TTL to 300 s.
TTL 60 s versus observed shift from 12 resolver vantage points.
| Vantage | Clamped TTL | 50% shifted | 95% shifted |
|---|---|---|---|
| your laptop dig | 60 s | 60 s | 90 s |
| major public DNS | 60 s | 70 s | 3 min |
| carrier resolvers | 300 s | 6 min | 11 min |
| declared RTO | — | — | 5 min (missed) |
I would not consider it settled without evidence: From synthetic probes in major resolvers, graph the share of traffic on the new IP after a flip and compare to RTO.
DNS failover is a cache-invalidation problem wearing an operations badge.
Curated: · Written: · Reviewed:
QA-84The load balancer thinks a pod is healthy while /ready is false. What is the check wrong about?(show answer)
The judgement in load-balancer health checks lies in what you refuse to do as much as in what you build.
Health checks must exercise the same readiness condition you use to serve users, including critical dependencies you are not willing to serve without.
Concretely, point the load balancer at a shallow readiness endpoint that proves this instance can accept work; include only dependencies whose failure makes this particular instance uniquely unable to serve, and test shared-dependency outages against the intended fallback.
The reason for that specificity is a failure I have seen: A TCP health check kept 12 deadlocked pods in rotation; they accepted connections and never answered, adding 8 seconds of tail latency for 20% of requests.
Check interval 5 s, unhealthy threshold 3 = 15 s MTTD budget.
| Check | Deadlocked pod in pool? | Extra tail |
|---|---|---|
| TCP :8080 | yes, 12 pods | p99 +8.0 s |
| HTTP /live (shallow local handler) | maybe; can miss a stuck worker pool | p99 +2.1 s |
| HTTP /ready (local serving path) | removed in 15 s | p99 +0.2 s |
I would not consider it settled without evidence: Deadlock a worker in staging and show it leaves the balancer before the SLO burns; time that against the check interval times threshold.
A check that only proves a port is open will happily send users to a process that will never speak HTTP.
Curated: · Written: · Reviewed:
QA-85A pod dies and its sticky users error until their cookie expires. How should affinity behave on failure?(show answer)
The first thing to settle about sticky sessions after failure is what decision it actually changes.
Stickiness must die with the backend; failing pods must drop affinity so users are hashed onto a live replica immediately.
Concretely, use balancer affinity that is ignored when the target is unhealthy, set a short cookie TTL as a backstop, and prefer stateless tokens over pod-sticky sessions.
The reason for that specificity is a failure I have seen: A 12-hour cookie pinned 8% of users to a terminating pod; they 502'd for 11 minutes while the other 92% were fine.
Cookie TTL 12 h versus health-aware affinity.
| Policy | Pod killed | Stuck users | Recover |
|---|---|---|---|
| cookie 12 h, ignore health | 8% of traffic | 11 min 502 | cookie expiry |
| cookie 12 h, drop if unhealthy | 8% | 0 | next request, 80 ms |
| stateless JWT, no stickiness | 0% stuck | 0 | any replica |
I would not consider it settled without evidence: Kill a pod under sticky load and show those users succeed on another replica within one retry, not at cookie expiry.
Affinity to a corpse is not session continuity, it is a private outage.
Curated: · Written: · Reviewed:
QA-86After a fleet restart, every pod stamps the origin at once. How do you stagger the return to traffic?(show answer)
I would start by making the assumption behind thundering herd after restart explicit.
Admit pods to the balancer on a jittered ramp and coalesce the cold-cache fills, so restart is not a synchronized miss storm.
Concretely, roll restart in small batches, pre-warm before reporting ready, coalesce cache fills, and implement randomized warmup in the application or rollout controller; use a real readiness gate only if a controller sets the corresponding Pod condition.
The reason for that specificity is a failure I have seen: A rolling restart with maxUnavailable 50% still aligned 40 pods' first requests on the same 20 hot keys; origin QPS jumped 18x for 70 seconds.
Restart 80 pods, hot key set 20, origin cap during warmup.
maxUnavailable: 10% # 8 pods at a time, not 50%
prewarm: {keys: 20, originConcurrency: 4, jitterSec: [5, 25]}
# Readiness gate only if a controller sets the matching conditionType.
# Old 50% restart: 40 pods * 20 keys = 800 misses, 18x origin, 70 s
# New: 8 pods, coalesced, origin <= 2x, p99 210 ms
I would not consider it settled without evidence: Restart 50% of the fleet in staging and show origin QPS stay under 2x steady, with p99 inside SLO.
The most dangerous moment for a cache-backed service is not failure, it is recovery.
Curated: · Written: · Reviewed:
QA-87Certificates look expired on one node and tokens fail randomly. What clock discipline does production need?(show answer)
This is one of those areas where the default answer and the right answer for clock skew differ.
Bound clock skew across the fleet and treat NTP failure as a paging condition, because auth, TLS, and fencing tokens all assume time is roughly shared.
Concretely, run a chrony/NTP daemon, page when offset exceeds 100 ms, prefer monotonic clocks for timeouts, and avoid wall-clock for fencing where a logical epoch exists.
The reason for that specificity is a failure I have seen: A VM with 12 minutes of drift rejected TLS as not-yet-valid and served 5xx to 6% of users hashed onto that node for 2 hours.
Offset page at 100 ms; 12-minute drift is a Sev-2, not a curiosity.
| Node | NTP offset | TLS | Action |
|---|---|---|---|
| ip-a | 4 ms | ok | — |
| ip-b | 12 min | not-yet-valid on new certs | page, drain |
| timeouts using wall clock | jump -2 s | leaked 2 s deadline | use monotonic |
I would not consider it settled without evidence: Dashboard max |offset| across nodes, with a page at 100 ms, and a game-day where you skew a node and watch it leave the pool.
A node that does not know what time it is cannot speak TLS or tokens honestly.
Curated: · Written: · Reviewed:
QA-88App replicas autoscale from 10 to 40 and the database starts refusing connections. What was not scaled?(show answer)
The useful framing here is to ask what evidence would change my mind about connection pool exhaustion.
Cap total client connections at the database max minus admin reserve, and scale pools down per replica as replica count rises.
Concretely, set per-pod pool size so pods × pool ≤ db_max − reserve, emit pool wait metrics, and fail fast when the pool is empty instead of growing threads.
The reason for that specificity is a failure I have seen: Each of 40 pods opened 50 connections against a max of 500; 1,500 attempted sockets left 0 for migrations and on-call, and new checkouts 500'd.
db_max 500, reserve 20, HPA max 40 pods.
db_max, reserve, pods = 500, 20, 40
pool_per_pod = (db_max - reserve) // pods # 12, not 50
# Old: 40 * 50 = 2,000 attempted, 500 accepted, migrations locked out
# Alert: pool_wait_ms p99 > 20 or db_used > 400
I would not consider it settled without evidence: At max HPA replicas, show used connections ≤ 80% of db_max and a remaining admin reserve.
Autoscale without a connection budget moves the outage from CPU to the one resource that does not autoscale.
Curated: · Written: · Reviewed:
QA-89p99 climbed and CPU is idle. How do you tell a full disk path from a full NIC, and why does it matter?(show answer)
Before choosing an implementation I would establish what success looks like for disk versus network saturation.
Saturation of disk and saturation of network need different mitigations, so measure queue depth, await, and NIC utilization separately before you add replicas.
Concretely, export disk await and utilization, NIC bytes and drops, and application wait; add disks or reduce chatty queries for storage, and add bandwidth or cut payload for network.
The reason for that specificity is a failure I have seen: HPA added 20 pods because CPU was idle during a disk storm; they amplified random IOPS and await went from 18 ms to 210 ms.
Two idle-CPU brownouts, two bottlenecks.
| Signal | Incident A | Incident B |
|---|---|---|
| CPU | 11% | 9% |
| disk await | 210 ms | 1.2 ms |
| NIC utilization | 12% | 95%, 2.1% drops |
| Fix | bigger gp3 IOPS | compress + extra ENI |
| HPA on CPU | made A worse | irrelevant |
I would not consider it settled without evidence: Show a pair of incidents where the same p99 had disk await 210 ms versus NIC at 95% and dropped packets, with different fixes.
Replicas cannot create IOPS that the volume does not have, and they cannot create NIC credits either.
Curated: · Written: · Reviewed:
QA-90All private-subnet egress shares one NAT. What fails when that NAT or its AZ fails?(show answer)
The engineering question underneath NAT gateway as a SPOF is which failure is acceptable.
Treat a single NAT as a hard dependency for every outbound call, and spread egress across AZs and optionally across gateways with per-AZ routes.
Concretely, use a NAT gateway per AZ with local routes, monitor ErrorPortAllocation, active connections, packets dropped, and bandwidth, and add secondary public IP addresses or distribute destinations when connection capacity is approached.
The reason for that specificity is a failure I have seen: A single AZ-local NAT gateway exhausted its concurrent connections to the payments destination and blocked egress from every AZ because all private-subnet routes pointed through it.
About 55,000 concurrent connections per destination per NAT IP; one NAT versus per-AZ NAT.
| Design | AZ-a NAT down | Peak connections to payments |
|---|---|---|
| one NAT in AZ-a | 100% egress dead | ErrorPortAllocation at one destination |
| NAT per AZ, local routes | AZ-a egress dead, others live | connections split across AZs |
| extra NAT IPs | — | raise per-destination connection cap |
I would not consider it settled without evidence: Fail NAT in AZ-a in a drill and show AZ-b and AZ-c egress still succeeding; graph port allocation under 60% at peak.
One NAT is a single network control plane for every dependency you do not host.
Curated: · Written: · Reviewed:
QA-91A whole availability zone disappears. What must already be true for checkout to stay inside SLO?(show answer)
My approach to AZ outage design starts from the constraint rather than from the technique.
Run N-1 capacity across AZs, keep stateful dependencies multi-AZ, and automatically drain the dead AZ rather than waiting for DNS to guess.
Concretely, provision enough per-AZ capacity that the surviving zones carry peak traffic at the chosen headroom target; at 250 rps per pod and a 1,200-rps peak with a 70% utilization ceiling, at least seven pods must survive, so an even three-AZ layout needs four pods per AZ.
The reason for that specificity is a failure I have seen: Two-AZ deploy at 50/50 left 100% of remaining capacity at 100% utilization when AZ-b died; p99 went to 3.4 s and the SLO burned in 9 minutes.
Peak 1,200 rps, 250 rps/pod, three AZs.
| Layout | Pods per AZ | After one AZ loss | p99 |
|---|---|---|---|
| 2 AZ, 3+3 | 3 | 3 pods = 750 rps cap | 3.4 s |
| 3 AZ, 4+4+4 | 4 | 8 pods = 2,000 rps cap, 60% utilized | inside target |
| 3 AZ, 2+2+2 | 2 | 4 pods = 1,000 rps cap | SLO miss, undersized |
I would not consider it settled without evidence: In an AZ-loss game day, show remaining utilization ≤ 70%, SLI inside SLO, and no manual DNS edit.
Multi-AZ that is not N-1 is a hope that the surviving zone was bored.
Curated: · Written: · Reviewed:
QA-92A deploy bot uses a long-lived access key in CI. What identity should automation use instead?(show answer)
I would treat IAM for automation identities as a design decision with a stated trade-off rather than a best practice.
Give automation a workload identity that mints short-lived credentials from a signed job claim, scoped to one repo, one environment, and one task.
Concretely, federate CI via OIDC, map subject claims to a role, set a session TTL of minutes, deny iam:* and unused services, and alert on any long-lived key still in the account.
The reason for that specificity is a failure I have seen: A 2-year-old static key in a forked PR's secret store was used from an unexpected IP to list and copy the production bucket.
OIDC trust pinned to repo and environment, TTL 15 minutes.
trust:
federated: token.actions.githubusercontent.com
sub: repo:acme/checkout:environment:prod
ttl_minutes: 15
deny: ["iam:*", "s3:*"] # deploy role deploys, it does not list buckets
static_keys_for_this_bot: 0
# Stolen 2-year key: production bucket copy from unknown IP
I would not consider it settled without evidence: Show the last deploy's credentials expired within 60 minutes, the role's trust policy pins the repo, and key count for bots is 0.
A robot with a permanent key is a human attacker who never sleeps and never rotates.
Curated: · Written: · Reviewed:
QA-93A CI role can read every secret and attach any policy. How do you shrink that?(show answer)
The place where teams go wrong on IAM blast radius is usually the step before the one they are debating.
Scope each automation role to the least actions and resources that job needs, and split jobs so a compromised frontend deploy cannot touch payments data.
Concretely, one role per pipeline job, resource ARNs not wildcards, deny-by-default SCP on destructive APIs, and a quarterly access review that diffs unused actions.
The reason for that specificity is a failure I have seen: A docs-site deploy role with AdministratorAccess was stolen via a poisoned action; it disabled CloudTrail and snapshotted the payments DB in 8 minutes.
Split roles; docs deploy cannot see payments.
| Role | Actions | Resource | Stolen docs job |
|---|---|---|---|
| checkout-deploy | ecs:UpdateService | checkout cluster | n/a |
| docs-deploy | s3:PutObject | docs bucket | cannot DisableTrail |
| old shared admin | * | * | CloudTrail off, DB snapshot in 8 min |
I would not consider it settled without evidence: For each role, show allowed actions, last used actions, and that a dry-run of the docs job cannot DescribeDBInstances.
The blast radius of a stolen deploy is exactly the IAM you were too busy to split.
Curated: · Written: · Reviewed:
QA-94An engineer runs terraform apply because "the plan looked fine yesterday". What rule do you enforce?(show answer)
What matters most about apply without a plan is whether the result can be checked afterwards.
Apply only a plan file produced in CI for this commit, because a live apply against drifted state is a different change than the one that was reviewed.
Concretely, cI writes a plan artifact, production apply consumes that file, reject apply without it, and treat any lock-step local apply as break-glass.
The reason for that specificity is a failure I have seen: A local apply after 40 unrelated merges destroyed a queue "not in yesterday's plan" and dropped 18 minutes of events.
CI plan artifact is the only apply input.
ci: terraform plan -out=plan.bin # commit 7a11e4c
prod: terraform apply plan.bin # rejects if state != planned
break-glass: logged, pages security
# Local apply on drifted state: destroyed sqs.checkout_events, 18 min gap
I would not consider it settled without evidence: The apply job fails in a drill when the plan artifact is missing, and state versions show apply == planned commit.
A plan you are not applying is a review of a hypothetical; production will apply reality.
Curated: · Written: · Reviewed:
QA-95A debug echo printed an API token into a public CI log. What controls stop that class of leak?(show answer)
I would answer this by separating what secrets in CI logs guarantees from what it merely usually does.
Treat CI logs as public, mask known secret shapes, forbid echo of env, and rotate any credential that could have been printed.
Concretely, enable runner masking, block set -x on secret-bearing steps, scan logs for high-entropy strings, and auto-rotate on a hit.
The reason for that specificity is a failure I have seen: A token sat in a public log for 11 hours, was scraped, and minted 4,200 fraudulent API calls before the nightly rotate.
Synthetic leak drill: mask, ticket, rotate.
| Control | Synthetic token in echo | Result |
|---|---|---|
| runner mask | *** | never visible |
| set -x blocked | step failed | no dump of env |
| entropy scanner | ticket SEC-441 | 2 min |
| auto-rotate | old token dead | 6 min |
| no controls (old) | 11 h public | 4,200 fraud calls |
I would not consider it settled without evidence: Push a synthetic secret in a PR and show the log is masked, the scanner tickets, and the real secret is rotated within minutes.
Anything a CI log can print is already on the internet; design for that, do not hope nobody looks.
Curated: · Written: · Reviewed:
QA-96A TLS cert expired at 00:01 and took login down. How do you make expiry a non-event?(show answer)
The judgement in certificate expiry lies in what you refuse to do as much as in what you build.
Automate issuance and renewal well inside the lifetime, page on days-left, and probe the served cert from outside the cluster, not the one on disk.
Concretely, use an ACME controller or managed certs, alert at 21 and 7 days, and run an external synthetic that checks the presented notAfter.
The reason for that specificity is a failure I have seen: A disk cert was renewed but the load balancer still presented the old one; dashboards showed 40 days left while users got expiry errors at 00:01.
Public probe of notAfter versus on-disk expiry.
disk: notAfter = 2026-10-19 (40 days) # renewed
served: notAfter = 2026-09-07T00:01:00Z # LB still old
probe: days_left = 0 -> page
policy: page if served days_left < 21
uptime at 00:01 without probe: login 100% down
I would not consider it settled without evidence: Show last renewal, days-left on the served cert from a public probe, and a game-day where you serve a short-lived cert and the page fires.
The certificate that matters is the one the client sees, not the one the disk wants you to believe.
Curated: · Written: · Reviewed:
QA-97Payments is a vendor with no public SLO. How do you still run a reliability program on that hop?(show answer)
The first thing to settle about third-party dependency SLOs is what decision it actually changes.
Measure the vendor as a dependency SLI from your callers, give it a budget slice, and negotiate or fall back when the measured reliability cannot fit.
Concretely, if this dependency alone receives a 0.04% bad-event slice, its caller-observed success target must be at least 99.96%, unless measured fallbacks prevent dependency failures from becoming service bad events.
The reason for that specificity is a failure I have seen: Treating the vendor as "their problem" left a 2.8% timeout rate unowned for 6 weeks, burning 80% of the checkout error budget.
Vendor hop allocated 0.04% of the service bad-event budget: required effective success ≥99.96%.
| Signal | 30d | Slice | Action |
|---|---|---|---|
| payments success | 97.20% | 99.96% | overrun |
| payments p99 | 1,840 ms | 250 ms | overrun |
| checkout budget burned by this hop | 80% | 40% max | enable cached auth + queue |
I would not consider it settled without evidence: A 30-day vendor SLI next to the allocated slice, plus a dated conversation or fallback when it overruns.
A dependency without an SLO is still in your SLO; you are just not looking at it.
Curated: · Written: · Reviewed:
QA-98The status page says "degraded performance". What does on-call actually do in the first 15 minutes?(show answer)
I would start by making the assumption behind vendor outage runbooks explicit.
A vendor runbook names detection, user impact, fallback, comms, and the condition to revert, so on-call does not rediscover the vendor on a Sev-1.
Concretely, write a 1-page runbook per critical vendor, link it from the burn-rate page, rehearse the fallback in a game day, and time the first 15 minutes.
The reason for that specificity is a failure I have seen: Without a runbook, on-call spent 22 minutes hunting Slack for a vendor ticket while a documented cache-fallback would have restored checkout in 90 seconds.
Payments vendor runbook, first 15 minutes.
detect: checkout_payments_success < 99% for 5 min OR vendor RSS
impact: % of checkouts in 5xx, start comms 10 min cadence
fallback: PAYMENTS_CACHE=1 flag, stale 10 min tax/auth (rehearsed)
do not: retry harder (amplifies)
revert when: vendor success > 99.5% for 20 min
game day: fallback 90 s; without runbook: 22 min Slack archaeology
I would not consider it settled without evidence: Game-day the vendor down: time to fallback, comms sent, and checkout SLI; keep that path inside the RTO.
If the runbook is "check their status page", you have a news feed, not a mitigation.
Curated: · Written: · Reviewed:
QA-99Synthetics ping /healthz from inside the VPC and stay green during a DNS outage. Where should probes run?(show answer)
This is one of those areas where the default answer and the right answer for synthetic monitoring placement differ.
Place synthetics on the user path — public DNS, TLS, edge, and a full journey — or they will certify a system users cannot reach.
Concretely, run journey probes from outside the network in each market, assert on checkout completion not /healthz, and keep one internal probe only as a splitter for "is it us or the path".
The reason for that specificity is a failure I have seen: Internal /healthz probes stayed 200 while a DNS TTL trap sent users to a dead VIP for 11 minutes; MTTD was a social-media report.
DNS outage: internal health versus public journey.
| Probe | Path | During DNS trap | Page? |
|---|---|---|---|
| /healthz in VPC | skip DNS + edge | 200 | no |
| TCP to VIP | skip DNS | 200 | no |
| public journey checkout | DNS + TLS + edge + pay | fail at 30 s | yes, MTTD 30 s |
| users | 11 min | that was MTTD before |
I would not consider it settled without evidence: Fail public DNS in a drill and show the external journey probe pages while the internal /healthz does not, and that on-call follows the external page.
A synthetic that shares fate with the service is a unit test, not a user.
Curated: · Written: · Reviewed:
QA-100The budget is half gone at day 12. Do you spend the rest on features or on deleting toil?(show answer)
The useful framing here is to ask what evidence would change my mind about error budget spend on toil vs features.
Use budget burn and incident evidence to shift engineering effort toward reliability when policy thresholds are crossed, while subjecting reliability changes to the same staged rollout and SLI gates as feature changes.
Concretely, at 50% budget consumed, shift the team's next sprint capacity to the top toil and the top incident class; allow feature launches only with a budget grant and a tighter canary.
The reason for that specificity is a failure I have seen: Shipping a new recommendation module at 48% budget remaining added 0.9 points of errors and a 14-hour on-call week, then the freeze arrived anyway with the toil untouched.
Day 12, 50% of 43.2 min spent; sprint mix changes.
| Band | Remaining | Feature launches | Toil / reliability |
|---|---|---|---|
| green (> 21.6 min) | 30 min | normal | 20% of SRE time |
| amber (day 12) | 21.6 min | grant + canary only | 70% of SRE time |
| spend on recs module | then 6 min | 1 launch | toil unchanged, 14 h nights |
| red (< 8.6 min) | freeze | 0 | 100% reliability |
I would not consider it settled without evidence: Show budget remaining, toil hours, and feature-launch count by band for a quarter, and that amber weeks reduce toil hours rather than add launches.
Burning the last minutes on features is how you buy a freeze with nothing to show for the reliability you skipped.
Curated: · Written: · Reviewed:
