Skip to content
Tech Interview Prep home
Technical interview guide

Infrastructure as Code

Defining infrastructure in version-controlled configuration instead of clicking through a console — reproducible, reviewable, and diffable.

Read
25 min
Practice MCQs
25
Interview QA
25
Edition
v6
Editorial status
Reviewed

Scope: Terraform language and CLI documentation accessed 2026-08-30; OpenTofu 1.13 state encryption; current AWS CloudFormation change-set and drift behavior..

Overview

Curated: · Written: · Reviewed:

Key takeaways

  • Infrastructure as Code (IaC) makes desired infrastructure versioned, reviewable, repeatable, and testable; it does not make every change safe or every provider operation reversible.
  • A preview is evidence, not a guarantee. Review replacements, deletes, identity-policy changes, data migration, cost, quota, and runtime preconditions before applying a saved artifact.
  • State is a sensitive coordination database that maps configuration addresses to real objects. Protect it with restricted remote storage, encryption, versioning, locking, audit, and tested recovery.
  • Separate environments by credentials, state, and blast radius. Naming conventions or CLI workspaces alone are weak security boundaries for complex production estates.
  • Build small, opinionated modules with explicit inputs, outputs, versions, ownership, upgrade notes, and policy tests. Avoid a single module that hides an entire platform.
  • Treat drift, imports, moves, replacements, and emergency console changes as controlled reconciliation workflows, not reasons to edit state JSON or overwrite reality blindly.

1. Desired state, providers, and dependency graphs

Declarative IaC describes resources and relationships rather than a sequence of console clicks. A provider translates the configuration into API operations, reads existing objects, and computes a dependency graph. References normally create implicit dependencies; explicit dependencies are for real ordering constraints the tool cannot infer. An apply can still partially succeed because remote APIs, quotas, permissions, and eventual consistency exist outside the declarative model. Idempotence means converging repeated runs tend toward the same desired state—not that every provider action is intrinsically harmless.

Pin tool and provider versions within an intentionally supported range, commit dependency lock data where the tool recommends it, and upgrade through reviewed pull requests. Provider schema changes can alter a plan even when configuration did not change. Validate formatting and syntax, then add policy, security, cost, and behavior checks appropriate to the resource rather than treating successful parsing as production proof.

2. State is production data

State associates a resource address with a real object and records attributes needed to plan later changes. It may contain resource IDs, network details, generated values, and secrets returned by providers. A sensitive display marker can redact CLI output but does not necessarily remove the value from state. Never commit state to source control or edit its JSON directly.

For teams, use a remote backend with least-privilege access, encryption in transit and at rest, locking or an equivalent serialization control, immutable/versioned recovery copies, audit logs, and a rehearsed restore path. Back up before backend migration or high-risk state operations. Encryption keys are part of recovery: losing a required key can make encrypted state unrecoverable. Split state by lifecycle and blast radius so one plan does not require broad access or refresh thousands of unrelated resources.

3. Plan, review, apply, and rollback

A speculative plan answers “what would this configuration do against state and observed remote objects now?” A saved plan binds the reviewed proposal to an apply, but freshness still matters: credentials, quotas, data, remote resources, or provider behavior can change. CI should produce the plan from an immutable revision, expose adds/updates/replacements/deletes and policy findings, require approval for protected environments, then apply that exact artifact with a narrowly scoped workload identity.

Rollback usually means a new forward change. Reapplying old configuration may not restore deleted data, an old database engine, or an externally changed dependency. Protect stateful resources with backups, retention or deletion controls, preconditions, staged migrations, and tested restore procedures. Use create-before-destroy only when names, quotas, capacity, and dependencies permit overlap; use ignore-change rules narrowly because they can conceal material drift.

4. Modules, environments, and delivery architecture

A useful module owns a coherent capability, exposes a small typed interface, publishes outputs consumers need, and encodes safe organizational defaults. Version modules, document compatibility and breaking changes, and test representative examples. Keep provider configuration and environment-specific identity at the composition root when practical. Excessively deep or universal modules create hidden coupling and risky upgrades.

Separate production from non-production with distinct accounts/projects or subscriptions where feasible, credentials, policy guardrails, and state. Reuse modules, not the same state. Promote a reviewed module or configuration version through environments while supplying environment-specific data. Avoid long-lived cloud keys in CI; prefer workload federation with audience and subject restrictions.

5. Drift, import, refactoring, and incidents

Drift is a difference between expected configuration and observed infrastructure. Detection coverage is not universal, and computed or omitted properties may not be comparable. Classify drift: revert unauthorized change, encode an intentional emergency change, or import an unmanaged resource. Preview reconciliation before applying; an automatic “fix” can erase a valid incident response.

Import associates an existing object with a configuration address; it does not prove the authored configuration matches every important remote property. Refactoring addresses should use declarative move mechanisms where available so objects are not destroyed and recreated. State commands are surgical tools: authenticate strongly, lock and back up state, peer-review the exact addresses, and verify afterward. During an IaC incident, stop competing writers, preserve plan/state/run evidence, assess actual cloud state, restore coordination safely, and reconcile through a reviewed forward plan.

6. Worked example: reading a plan before approving it

A plan is the artifact review acts on, and the verbs matter more than the count. This one carries ten resource actions -- 7 to add, 2 to change, 1 to destroy -- and exactly one line decides whether it is safe. Two of those actions are shown:

# terraform plan -out=tf.plan  ->  7 to add, 2 to change, 1 to destroy
  ~ aws_db_instance.primary
      ~ allocated_storage       = 200 -> 500
      ~ instance_class          = "db.r6g.xlarge" -> "db.r6g.2xlarge"
      ~ backup_retention_period = 7 -> 30
- /+ aws_instance.bastion                            # forces replacement
      ~ subnet_id = "subnet-0a1" -> "subnet-0c3"

The -/+ is the whole review. subnet_id on aws_instance is a force-new attribute: an instance cannot move between subnets, so Terraform destroys this one and builds another. The replacement arrives with a new instance id, a new private address, and nothing on its instance-store volumes. Whether that is routine or an incident depends entirely on what the bastion is holding, and the plan cannot tell you.

Do not read the three ~ lines as therefore safe. In-place is a statement about identity, not about impact: allocated_storage really is close to free, but instance_class resizes the instance, which reboots it — on a multi-AZ deployment that is a failover of roughly 60 to 120 seconds, applied whenever the maintenance setting says rather than when the reviewer expects. A plan that reads as "2 to change" contains a scheduled outage.

The provider decides which properties force replacement, and the answer is frequently not the intuitive one: changing the members of an aws_db_subnet_group, for instance, updates in place through ModifyDBSubnetGroup rather than replacing the group. Guessing is how a reviewer approves a destroy. Read the verb the plan printed, not the verb the change sounds like.

A reviewer who counts resources sees "7 to add" and approves; a reviewer who reads verbs sees one destructive line and one rebooting line, and asks for both to be split out.

The verbs, and what each is worth checking for:

SymbolMeaningWhat to check
+createDoes anything already exist at this address, unmanaged?
~update in placeIs the provider's in-place claim true for this property?
-/+destroy then createWhat is attached, and what is the downtime?
+/-create then destroyDo both versions coexist safely, e.g. name collisions?
-destroyIs the data recoverable if this is wrong?

Two properties of the plan bound how much it is worth: it is computed against state, so it cannot see drift the provider does not report, and it expires. A plan reviewed 3 hours ago and applied against 3 hours of other changes is not the plan that was approved — which is why terraform apply tf.plan on the saved file, inside a short window, is the reviewable unit rather than a bare apply.