Skip to content
Tech Interview Prep home
Technical interview guide

Chaos Engineering

Deliberately injecting failure into a system to verify it actually survives what you assume it survives.

Read
28 min
Practice MCQs
25
Interview QA
25
Edition
v5
Editorial status
Reviewed

Scope: Chaos Engineering principles and AWS, Azure, Google Cloud, Kubernetes, LitmusChaos, Chaos Mesh, and OpenTelemetry guidance current 2026-08-31.

Overview

Curated: · Written: · Reviewed:

Chaos engineering tests a specific resilience hypothesis with controlled failure

Chaos engineering is disciplined experimentation on a system to learn whether it preserves an explicitly defined steady state when a realistic adverse event occurs. It is not random breakage, a dramatic demo, or permission to discover safety boundaries at users' expense. Start from architecture, incident history, threat and dependency analysis: identify a credible failure, the assumed defense, the user-visible outcome that must remain acceptable, and evidence that can falsify the assumption. If the hypothesis cannot fail or the observation cannot distinguish success from failure, the experiment does not produce useful confidence.

Define steady state through critical user journeys and correctness, not merely host health. Include availability/error, latency, data integrity, security, queue age, freshness, durability and business or safety signals as relevant. Specify measurement boundaries, populations, intervals and telemetry health. A green component metric cannot prove a user journey survived, and missing telemetry is unknown. Establish a clean baseline and account for normal variance before injecting a fault.

Choose the smallest experiment that can answer the question. A fault may terminate a process or instance, remove a zone, delay or drop network traffic, exhaust CPU/memory/disk/connections, throttle a dependency, corrupt a non-production copy, expire credentials, skew clocks, deny an API, overload a queue, or make observability/control tooling unavailable. Model the mechanism closely enough to exercise the intended defense. Killing a stateless pod does not validate regional failover; adding latency does not establish data-recovery correctness. State exclusions and residual uncertainty.

Safety is designed before execution. Name an accountable experiment owner and independent abort authority; scope exact accounts, clusters, regions, services, tenants and identities; cap magnitude and duration; define stop conditions from user/data/security and dependency signals; verify rollback or natural fault expiry; protect credentials; coordinate on-call, product, support, security, privacy, vendors and change calendars; and retain an incident path. Use least-privilege experiment identities and audited templates. Stop conditions should be automatic when tooling supports them, but automation supplements rather than removes human judgment.

Progress from tabletop and deterministic tests through development/staging to canary production experiments when the remaining question requires production realism. Production often contains traffic shape, scale, state, dependency, configuration and human-response behavior absent elsewhere, but that does not make production the default starting point. Expand blast radius only after evidence at the prior scope, and never run during an unrelated incident or risky change unless the exercise explicitly and safely models concurrency. Consider customer agreements, regulated data, physical safety and irreversible side effects.

Observe the entire response. Record experiment version, hypothesis, target selection, baseline, injected action, timestamps and clock basis, expected signal, actual user and component outcomes, alerts, autoscaling/failover, retries, queues, data reconciliation, responder decisions, abort/rollback and recovery. Verify that the fault was actually applied and that the generator or control plane did not fail first. Separate detection, mitigation and recovery. Continuing after steady state breaches just to reach planned duration is not rigor; stop safely and preserve the evidence.

The recovery phase is part of the experiment. Removing the fault may unleash retry storms, backlog, rebalancing, cache fill, leader elections or scale-down oscillation. Confirm critical journeys, data and security state, drain or quarantine outstanding work, remove emergency access and temporary controls, and return ownership explicitly. A system that remains healthy during injection but cannot recover cleanly did not pass the resilience hypothesis.

Treat findings blamelessly and turn them into owned risk work. Record what held, what failed, causal limits, severity, exposure and residual risk. Actions need an owner, priority, due date, interim control and verification experiment. Rerun the same hypothesis after remediation and periodically as architecture changes. Do not use a pass as permanent certification or count experiments as a vanity target; track coverage of material risks, discovered user harm, action closure, recurrence and whether controls remain effective.

Build a governed program from reusable versioned experiment templates, target allowlists, policy checks, audit logs, scheduling controls, safety reviews and searchable results. Prevent selection drift and overly broad wildcards. Test the chaos platform itself: authorization denial, stop-condition failure, partial injection, controller outage, stale templates and cleanup. The platform is a privileged production control surface and must have stronger boundaries than an ordinary test runner.

Mature programs combine chaos experiments with conventional unit, integration, load, security, backup/restore, disaster-recovery and incident drills. Each method answers different questions. Chaos engineering is most valuable where defenses exist but their behavior under realistic failure is uncertain; it supplies bounded evidence and learning, never proof that the system is resilient under every untested event.

What an interviewer is testing

The question behind most chaos-engineering interview prompts is whether the candidate can distinguish an experiment from a stunt. An experiment names a defense that is believed to exist, states the user-visible outcome that must survive, and specifies in advance the observation that would show the belief was wrong. A stunt terminates something and reports that the system stayed up, which establishes only that the particular thing terminated was not load-bearing at that moment. Being able to say what result would have falsified the hypothesis is the fastest way to demonstrate the difference.

The second thing being tested is judgment about blast radius and consent. Production realism is a reason to run in production, not an entitlement to do so: the argument has to be that the remaining uncertainty cannot be resolved anywhere else, and it has to be paired with a scoped identity, a capped magnitude and duration, automatic stop conditions tied to user-visible signals, and a named person who can abort without seeking approval. A candidate who reaches for production first, or who cannot describe how the experiment is stopped when it starts harming users, is describing an outage with a change ticket attached.

The third is what happens after the fault is removed. Recovery is where retry storms, backlog drain, cache refill, leader elections and scale-down oscillation appear, and a system that behaved well under injection can still fail here. Treating recovery as part of the experiment — with its own expected signals and its own stop conditions — is what separates a resilience programme from a demonstration.