Skip to content
Tech Interview Prep home
Technical interview guide

High Availability & Disaster Recovery

Designing for component failure as the expected case, and the RTO/RPO trade-off that shapes disaster-recovery strategy.

Read
26 min
Practice MCQs
25
Interview QA
25
Edition
v6
Editorial status
Reviewed

Scope: AWS Well-Architected reliability and DR guidance, Google Cloud DR architecture, and Azure reliability guidance accessed 2026-08-30..

Interview QA

Treat each question like a live interview question: answer out loud first (structure, assumptions, tradeoffs), then open the model answer to spot gaps and rehearse a tighter follow-up.

Curated: · Written: · Reviewed:

QA-1

Explain the difference between high availability and disaster recovery.

QA-2

How do you derive RTO and RPO for a business service?

QA-3

Design a multi-zone highly available web service.

QA-4

How do you identify hidden correlated failures in a redundant architecture?

QA-5

Compare synchronous and asynchronous replication for DR.

QA-6

Design a backup strategy that can survive logical corruption or ransomware.

QA-7

How would you choose among backup/restore, pilot light, warm standby, and active-active?

QA-8

Design a safe regional failover decision process.

QA-9

Explain split brain and how you prevent it during disaster recovery.

QA-10

How do you design health checks for HA and failover?

QA-11

How do you determine and enforce service dependency ordering during full-region disaster recovery failover?

QA-12

How do you test DR without creating unacceptable production risk?

QA-13

Which measurements prove a DR exercise met its goals?

QA-14

How would you design and test regional DNS failover?

QA-15

How do you plan safe failback after a regional disaster?

QA-16

How do you prevent a bad deployment from defeating multi-zone redundancy?

QA-17

How would you design graceful degradation for a dependency outage?

QA-18

How do you size capacity for failure and recovery scenarios?

QA-19

What is the cloud provider versus customer responsibility for reliability?

QA-20

How would you protect the DR environment from the same security incident as primary?

QA-21

How do you keep a passive recovery environment from drifting?

QA-22

How do queues change HA and DR design?

QA-23

How would you recover from regional data corruption rather than infrastructure loss?

QA-24

How do compliance and data residency constrain DR architecture?

QA-25

How do you improve reliability after a failed DR exercise?