Skip to content
Tech Interview Prep home
Technical interview guide

Configuration Management

Keeping infrastructure and application configuration consistent, versioned, and reproducible across environments.

Read
30 min
Practice MCQs
25
Interview QA
25
Edition
v4
Editorial status
Reviewed

Scope: Ansible Core, Puppet Core, Chef Infra Client, Salt, Kubernetes, NIST SP 800-128, and CIS Controls guidance current 2026-08-31.

Overview

Curated: · Written: · Reviewed:

Configuration management makes intended state reviewable, reproducible, and recoverable

Configuration management is one of those interview topics where senior candidates lose points by talking tools. The interviewer usually starts with "how do you keep a hundred servers consistent?" and is really asking whether you think in terms of intended state: a reviewed, versioned description of how things should be, plus a mechanism that converges reality toward it and tells you when it can't. A weak answer names a tool and lists its features. A strong answer describes the control loop — represent, review, apply, verify, detect divergence, recover — and then argues about the edge cases: imperative escape hatches, drift that should not be overwritten, and rollouts that can be safely paused.

The rest of this guide walks that arc.

Desired state, convergence, and idempotency

Declarative resources describe properties of an end state — a package installed, a file containing specific content, a service running and enabled. Imperative tasks prescribe operations — run this command, in this order. Real systems combine both: a package, user, or registry value is naturally declarative, while a vendor migration binary is unavoidably imperative.

What "declarative" buys you is convergence: the provider checks observed state against declared state and applies only the difference. That check is what makes re-running the same automation a no-op, and it's the property interviewers probe hardest. Idempotency is a property of the resource implementation and the environment, not a label attached to a tool. A command resource that runs on every invocation is not idempotent; a package resource backed by the system package manager is, because the manager itself answers "is this already installed at this version?" When an imperative escape hatch is unavoidable, you owe it preconditions, postconditions, timeouts, exit-code semantics, durable markers tied to the desired version, and a compensating recovery path — because a zero exit code is not proof that the desired state exists. Verify the resulting service and security behavior, not just the exit code. This is the single most common weak-answer tell: candidates assert "Ansible modules are idempotent" without being able to say which module argument busts idempotency (command/shell without creates:/when: guards) or how they'd verify convergence after an imperative step.

Also expect the follow-up: "how does the tool know a resource is already compliant?" The honest answer is that each provider has its own detection logic — comparing file checksums and permissions, querying package databases, asking the service manager — and each detection method has blind spots. A file whose content matches but whose containing directory has wrong ownership looks compliant to a naive check. Test the resource's actual convergence behavior, including partial failure and retry.

Execution models: agentless push vs pull

Agentless push (Ansible) opens SSH (or WinRM) to targets and ships modules over the wire each run. No agent to install or upgrade, so bootstrapping is trivially short — you need reachable credentials and that's it. The cost is that the control node must be able to reach every target, targets do nothing between runs, and every run needs the network path open. Intermittently connected or firewalled hosts fall behind.

Pull agents (Chef, Puppet, Salt) run daemons on each host that fetch their configuration from a server on a schedule. This handles hosts behind NAT, lets you converge on a cadence (say every 30 minutes) without a central push, and keeps local facts current. The cost is agent sprawl: upgrade the agent across the fleet, keep the agent from being a lateral-movement foothold, and accept configuration lag — a node's state reflects the last pull, not the present.

The trade-offs the interviewer wants named: bootstrapping burden, network reach direction, agent lifecycle management, convergence cadence, and behavior of disconnected hosts. If a candidate says "push is better" or "pull is better" without naming the constraint that decided it, that's the weak answer. The deciding constraints are usually fleet size, network topology, and how quickly an emergency change must reach every host.

Dependencies, ordering, and parallelism

Model ordering explicitly. Package installation precedes file rendering; validated configuration should precede reload; a reload should happen only when relevant content actually changes. Incidental file ordering — resource A happening to be evaluated before resource B — is a fragile dependency that breaks the first time someone reorders the file or enables parallelism. Build a real resource graph, reject cycles, understand how notifications coalesce (ten changed files → one service reload, not ten), and distinguish a config reload from a full restart — a distinction with real availability impact. Parallelism speeds up runs only across genuinely independent resources. Test dependency failure paths so a downstream resource doesn't advertise success after its prerequisite failed, and make interrupted runs converge on resume.

Drift: detection, classification, and reconciliation

Drift is a difference between authorized intended state and observed relevant state. Every word in that definition is load-bearing, and interviewers will test each one: an approved difference between dev and prod is not drift; a difference on an attribute your policy doesn't claim is out of scope. Weak answers treat drift as "anything that differs, must overwrite" — and that policy destroys evidence after an incident and can re-apply an incorrect baseline over an emergency mitigation.

Instead: detect (scheduled audit runs, continuous reconciliation, or independent sampling of reported state — because a broken or compromised agent can lie or go silent), then classify. Drift may be harmful, benign, or an emergency repair. Identify the actor and source before remediating; if a legitimate emergency change happened, reconcile it back into the reviewed source rather than letting the next run erase it. And define precedence explicitly, because images, bootstrap/user-data, the CM agent, orchestration manifests, application defaults, feature-flag systems, and human operators can all write the same setting. If you can't say who wins when Terraform and Ansible both claim a security group, you haven't finished the design.

Git as source of truth, promotion, and the provisioning boundary

Represent intended state as reviewed artifacts in version control: pull requests with diffs, approvals for high-risk changes, tags pinning module and dependency versions, and a way to reproduce last week's configuration exactly. Promote environments by promoting a reference — a branch ref, tag, or variable value — not by maintaining per-environment branches of whole configurations, which decay independently and hide production differences. Keep prod deltas explicit and minimal in variable files, and validate schema, allowed values, cross-field invariants, and rendered output before apply. A syntactically valid template can still be unsafe for the installed daemon, OS, or dependency version.

Know where provisioning ends and configuration begins. Terraform and cloud APIs own resource lifecycle — create the database instance, the subnet, the managed service. CM owns what runs and is configured on top. The classic interview question is the overlap: provisioned resource attributes flow into configuration (an instance's ID or hostname is an input to the node's role), which forces you to answer how state crosses that boundary — Terraform state → CM inventory, or a shared fact store — and what happens when each side disagrees.

Secrets: inputs, not content

Secrets are inputs to configuration but must not become ordinary configuration content. Keep plaintext out of source, rendered diffs, logs, agent facts, caches, command lines, process listings, backups, and reusable images. The tools differ in mechanism but share the shape: Ansible Vault (encrypted values in your repo, key distribution still on you), SOPS (encrypted fields inside otherwise-readable YAML), or an external store like HashiCorp Vault with runtime injection via sidecar or template — short-lived credentials, scoped retrieval, audited access, rotation, and revocation you've actually tested. Encryption at rest protects stored files; it does not solve decryption-key distribution or runtime exposure. Kubernetes ConfigMaps are documented for non-confidential data, and base64 encoding is encoding, not encryption — saying otherwise is an instant weak-answer flag.

Prefer injection at runtime over templating secrets into files where the platform allows, because a rendered file with mode 0644 on a shared host outlives the run.

Rollout, exceptions, and recovery

Roll out progressively by representative cohort or failure domain: preflight inventory and compatibility checks, render and validate, preview the plan while acknowledging preview limitations, canary a small set, observe service/security/data guardrails, then expand in bounded batches. Serialize conflicting changes. Define who can pause or abort, and account for disconnected hosts that will converge late. A successful agent run proves what the agent measured — not user journeys, data correctness, or absence of hidden drift.

Recovery depends on change semantics. Restoring last week's file is usually safe; downgrading a package, database schema, identity policy, or certificate usually is not — prefer forward repair when rollback breaks compatibility. Back up the right state, retain last-known-good configuration and dependency versions, and rehearse restore, restart, credential recovery, and break-glass access, including the case where the management plane itself is impaired (DNS down, identity down, the CM server unreachable).

Treat configuration as a security boundary throughout: least-privilege execution identities, protected control planes, signed artifacts where appropriate, immutable audit, and exceptions with an owner, rationale, compensating controls, scope, expiry, and revalidation — not an undocumented override.

What to measure, and what interviewers probe

Measure inventoried vs managed assets, last successful convergence, drift age and recurrence, exception age, rollout duration, restart impact, baseline compliance, secret exposure events, and recovery success. Don't reward raw enforcement frequency — an agent that runs hourly and converges to a wrong baseline scores perfectly on that metric.

The likely follow-ups, and what a weak answer sounds like on each:

  • "Re-run this playbook twice — what happens?" Weak: "it's idempotent because Ansible is." Strong: names which resources report changed on the second run and which cannot be re-checked, and what marker or guard fixes it.
  • "An operator hand-edited a prod file to mitigate an incident. What does the next run do?" Weak: reverts it. Strong: classifies the drift, preserves the evidence, reconciles the mitigation into review if it was correct.
  • "Where do secrets live in this design?" Weak: "in the vault." Strong: traces the path from store to process environment, who can decrypt, what appears in logs and diffs, and how rotation reaches the workload.
  • "Roll back the last configuration change." Weak: git revert and re-run. Strong: asks what the change's reversibility semantics are before answering, and prefers forward repair when they aren't safe.

Good configuration management creates controlled convergence, not automatic certainty: automation can consistently distribute an unsafe value, continuous enforcement can erase an emergency mitigation, and an idempotent run can repeatedly converge to the wrong baseline. Keep authority explicit, changes small, evidence scoped, failure modes rehearsed, and humans accountable for the intended state.