Skip to content
Tech Interview Prep home

Site Reliability Engineer Interview Prep

Overview

Builds and operates the production side of services: SLOs and error budgets, incident command, observability, and automation that removes whole classes of repeat failure.

Curated: · Written: · Reviewed:

View Site Reliability Engineer leaderboard →

Top 100 Site Reliability Engineer Interview Questions and Answers

The questions most likely to actually be asked, ranked by likelihood, with pro-level model answers.

Top 100 Site Reliability Engineer Practice MCQs

Quick multiple-choice self-checks covering the same high-value ground, with an explanation for every answer.

What Site Reliability Engineer interviews evaluate

SRE interviews evaluate whether you can translate user impact into reliability targets you will defend under questioning, reason about a failure while it is still unfolding, and stop a release on budget arithmetic—not whether you can name observability tooling or recite an on-call checklist.

  • Translate user journeys into SLIs, keep availability and latency SLOs separate, compute error-budget burn against a stated window, and turn that arithmetic into an explicit go/no-go on the next release.
  • Design symptom-based alerts and run incident command, then rebuild the timeline—detection time, mitigation time, contributing factors—into corrective work that demonstrably lowers recurrence risk.
  • Reason about change failure as a distribution: canary and rollback thresholds, load-test evidence at real traffic shape, and the failure modes hiding in configuration, IAM, capacity headroom, and third-party dependency SLOs.

How to prepare: Practise each Top 100 scenario aloud in one unbroken pass—user impact, SLI and SLO, the evidence you would gather, immediate mitigation, the release or rollback call—and close on the specific thing you would refuse to ship and the concept that justifies the refusal.

Site Reliability Engineer preparation roadmap

Follow these concepts in order. Each opens its guide, interview QA, and practice MCQs while keeping this role as your study context.

  1. Scalability Fundamentals

    Production scalability fundamentals for technical interviews: bottlenecks, scaling, load balancing, autoscaling, capacity, overload control, and failure behavior.

  2. Caching Strategies

    Production caching for technical interviews: placement, read/write patterns, freshness, stampedes, HTTP caching, observability, failure recovery, and decision tradeoffs.

  3. Database Scaling (Sharding & Replication)

    Splitting data across machines (sharding) and copying it across machines (replication) — solving two different scaling problems.

  4. Message Queues & Async Processing

    Decoupling a slow or unreliable step from the request path by handing it to a queue and processing it separately.

  5. CAP Theorem & Consistency Models

    Why a distributed system can't have perfect consistency, availability, and partition tolerance all at once — and what real systems trade off.

  6. API Design & REST Fundamentals

    Designing HTTP APIs that are predictable to call and safe to retry — resource modeling, status codes, versioning, and idempotency.

  7. API Authentication & Authorization

    Verifying who's calling an API (authentication) and what they're allowed to do (authorization) — API keys, OAuth, and JWTs.

  8. Webhooks & Asynchronous API Integration

    Handling work that can't complete within a single request/response cycle — inbound webhooks and long-running async job APIs.

  9. URL Shortener Design

    Designing a URL shortener: unique keys, redirect semantics, cache TTLs, click accounting off the GET path, and open-redirect abuse.

  10. Cloud Networking Fundamentals

    VPCs, subnets, and security groups — the building blocks every other cloud topic assumes.

  11. IAM & Security Fundamentals

    The principle of least privilege, and how roles/policies enforce it instead of relying on long-lived credentials.

  12. Infrastructure as Code

    Defining infrastructure in version-controlled configuration instead of clicking through a console — reproducible, reviewable, and diffable.

  13. High Availability & Disaster Recovery

    Designing for component failure as the expected case, and the RTO/RPO trade-off that shapes disaster-recovery strategy.

  14. SLIs, SLOs & Error Budgets

    The vocabulary reliability is measured in, and how an error budget turns 'be reliable' into a concrete number.

  15. Monitoring, Logging & Tracing

    The three pillars of observability, and what question each one is actually good at answering.

  16. Incident Management & Postmortems

    Running an incident from detection to resolution, and writing a blameless postmortem that actually prevents a repeat.

  17. Capacity Planning & Load Testing

    Knowing how much traffic a system can take before it does, through modeling and deliberate load testing.

  18. Chaos Engineering

    Deliberately injecting failure into a system to verify it actually survives what you assume it survives.

  19. On-Call & Alerting Design

    Designing alerts that page for what actually needs a human, and structuring on-call sustainably.

  20. CI/CD Pipeline Design

    Continuous integration and continuous delivery — automating the path from commit to a shippable build.

  21. Containerization & Orchestration

    Packaging an app with its dependencies via containers, and how Kubernetes schedules and manages them at scale.

  22. Deployment Strategies (Blue-Green, Canary, Rolling)

    Different ways to roll a new version out safely, trading off speed, blast radius, and infrastructure cost.

  23. Configuration Management

    Keeping infrastructure and application configuration consistent, versioned, and reproducible across environments.

  24. GitOps

    Using a Git repository as the single source of truth for infrastructure and deployment state.

  25. Secrets Management in Pipelines

    Keeping credentials and keys out of source control and pipeline logs, while still letting automation use them.