Skip to content
Tech Interview Prep home
Technical interview guide

High Availability & Disaster Recovery

Designing for component failure as the expected case, and the RTO/RPO trade-off that shapes disaster-recovery strategy.

Read
26 min
Practice MCQs
25
Interview QA
25
Edition
v6
Editorial status
Reviewed

Scope: AWS Well-Architected reliability and DR guidance, Google Cloud DR architecture, and Azure reliability guidance accessed 2026-08-30..

Overview

Curated: · Written: · Reviewed:

Key takeaways

  • High availability (HA) keeps an agreed function operating through expected failures; disaster recovery (DR) restores it after disruption exceeds the active design. They overlap, but redundant serving capacity is not a substitute for recoverable data.
  • Define failure scope, service-level indicators, business impact, recovery time objective (RTO), recovery point objective (RPO), and degraded-mode expectations before selecting architecture.
  • Remove correlated failure: spread capacity across independent failure domains, keep dependencies and deployment waves from failing together, and preserve enough headroom to lose a domain.
  • Replication improves continuity but can replicate deletion, corruption, or malicious change. Maintain isolated, versioned backups and prove restore integrity and duration.
  • Failover is a consistency and control decision, not just a traffic switch. Establish authority, fencing, data readiness, dependency health, and abort criteria before promotion.
  • A runbook or standby is unproven until exercised under representative constraints. Measure achieved RTO/RPO, data correctness, operator load, and failback—not merely whether automation started.

1. Define reliability in business terms

Availability is the proportion of time a workload performs its agreed function when required. Measure the user journey or business operation, not merely whether a virtual machine responds. A service can return HTTP 200 while serving stale or incorrect data and still be unavailable for its purpose. Define service-level indicators, targets, error budgets, dependencies, and which degraded functions remain acceptable.

RTO is the maximum acceptable delay from disruption to restored service. RPO is the maximum acceptable gap between the latest recoverable data point and disruption. These are business risk decisions tied to a declared scenario: a process crash, zone loss, regional isolation, destructive credential compromise, or data corruption can demand different objectives. Zero RTO or RPO is an expensive architectural claim that must be demonstrated end to end; simply enabling replication does not establish it.

2. Failure domains and high availability

Design for the failure domain the objective names. Multiple instances on one host, rack, zone, region, account, identity provider, DNS control plane, or deployment wave may share the same failure. Spread stateless serving capacity across independent zones and route with health signals that reflect real readiness. Keep capacity to absorb the loss of a domain, or prove how fast and reliably safe scale-out occurs without depending on the failed control plane.

Redundancy only helps when copies fail independently. Common configuration, a global dependency, shared credentials, bad deployment, exhausted quota, or overloaded database can remove all replicas together. Use timeouts, bounded retries with jitter, load shedding, circuit breaking, queues, graceful degradation, and bulkheads to stop local failures becoming cascading failures. Test the complete dependency graph and deploy progressively across domains.

3. Data durability, replication, and backup

Choose replication topology from consistency, latency, write authority, and failure objectives. Synchronous replication can reduce acknowledged-data loss but adds latency and can reduce availability during partition. Asynchronous replication preserves distance and availability but creates measurable lag and a nonzero loss window. Monitor durable replication position, not just process health, and define the point that is safe to promote.

Replication and backup solve different threats. A replica may quickly copy an accidental deletion, ransomware, schema bug, or logical corruption. Retain isolated, immutable or protected, versioned recovery points across the failure scope, including keys, configuration, code, and dependency metadata. A backup is evidence only after a restore has produced an application-consistent dataset whose integrity and recovery duration were checked.

4. Disaster-recovery strategies

Backup and restore has low steady cost but must provision infrastructure and restore data during recovery. A pilot light keeps core data and foundational components ready while application capacity is created or activated. Warm standby runs a complete but reduced-capacity environment that can serve at limited scale before scaling up. Multi-site active/active serves from several sites continuously, reducing failover delay but increasing data-conflict, routing, operational, and cost complexity. Strategy labels do not prove objectives; measure the actual end-to-end path.

Keep IaC, images, configuration, secrets/key recovery, quotas, certificates, DNS or global routing, observability, and external integrations ready in the recovery location. Prefer recovery actions that use already-operational data planes; depending on many control-plane mutations during a broad incident can extend RTO. Protect the standby from configuration drift and from the same compromised identity or automation that damaged primary.

5. Failover, split brain, and failback

Detection must distinguish a genuine site failure from a local observer, dependency, or network partition. Fast automatic failover can amplify a false positive; slow committee approval can violate RTO. Define corroborating signals, decision authority, automation boundaries, and an emergency manual path. Before promotion, verify recovery data, replication position, dependency availability, capacity, and whether the old writer is fenced. Without fencing, two primaries can accept conflicting writes and create split brain.

Failback is another high-risk migration. Stabilize the recovery site, understand changes accumulated there, restore replication in the correct direction, reconcile conflicts, test the original site, and shift traffic progressively. Do not rush back simply to restore the architectural diagram. Preserve evidence and complete the incident review.

6. Exercises and continuous readiness

Exercise component, zone, region, control-plane, dependency, credential, and corruption scenarios at a cadence tied to criticality. Include unavailable people, expired credentials, quota pressure, stale runbooks, and realistic data volume. Start with contained game days, then increase scope with explicit abort conditions and business approval. Measure detection time, decision time, infrastructure readiness, restore time, replication lag/data loss, integrity, traffic recovery, backlog drain, operator steps, and failback.

After each exercise or incident, assign gaps to owners and deadlines, update automation and runbooks, and repeat the failed step. Continuously detect standby drift, test backups, monitor replication and capacity, and verify that organizational changes have not invalidated access. Reliability is an operated capability, not a one-time topology.

7. Worked example: what an availability target costs, and what a failover actually takes

An availability target is a downtime budget, and the budget is what makes the target arguable:

TargetDowntime per 365-day yearPer 30-day monthWhat it rules out
99.0%3d 15h 36m7h 12mnothing much
99.9%8h 46m43m 12sunattended manual failover
99.95%4h 23m21m 36sa single-AZ database
99.99%52m 34s4m 19sany human in the recovery path
99.999%5m 15s26sanything that reboots for a deploy

State the basis or the column is meaningless: a 365-day year and a 30-day month do not divide into each other, and quoting one row on one basis and the next on the other is a common way for these tables to be quietly wrong.

The jump from 99.9% to 99.99% is where the architecture changes rather than the effort: 4 minutes 19 seconds a month cannot absorb a human deciding to fail over, so detection and promotion must be automatic, which means split-brain becomes a design problem you now own.

Then price a real regional failover against an RTO of 15 minutes:

  t+0:00   Region A stops serving. Health checks are 10s interval, 3 failures to trip.
  t+0:30   Detection fires (3 x 10s). Two independent clocks start here, in parallel:
           the DNS record is updated, and replica promotion begins.
  t+1:15   Promotion completes (45s). The async replica was 2.3s behind, so that
           much write history is gone: this is the RPO, and it is not zero.
  t+1:30   Every resolver honouring the 60s TTL has now re-resolved. Most moved earlier,
           since a cached record is usually already part-expired when the update lands.
           ~5% ignore the TTL and keep sending to the dead region for longer.
  t+1:30   Application tier goes from warm 20% to 100%. The 8 instances start in
           parallel, so this costs one ~90s cold start, not eight.
  t+3:00   Instances healthy. Connection pools re-establish over the next ~1m45s.
  t+4:45   Serving, degraded: the cache is cold, so p99 runs 3-4x for about 12 minutes.
  t+16:45  Cache warm, p99 normal.

RTO is met at under 5 minutes if "serving" is the definition and missed at nearly 17 minutes if "serving normally" is. Deciding which one the target means is the work.

Two details in that timeline are worth arguing with, because both are usually drawn wrong. Promotion does not have to wait for DNS: they are independent clocks, and serializing them pads the RTO by a full TTL for no reason. And a TTL is not a switch that flips at t+TTL — a resolver's cached copy is typically already part-expired when the update lands, so clients trickle over across the window and t+1:30 is the last honouring client, not the first. An RPO of "near zero" under async replication is likewise not a design property; it is whatever the replication lag graph says at the instant you fail over, which is why that percentile belongs on the same dashboard as the availability number.