Overview
Curated: · Written: · Reviewed:
GitOps is a controlled reconciliation system, not merely deployment from Git
The loop you should be able to draw on a whiteboard
Every GitOps question reduces to one mechanism. If you can draw it and trace a change through it, you can derive the rest of your answers instead of reciting them:
git (desired state) cluster (observed state)
│ ▲
│ pull, on interval + trigger │
▼ │
┌─────────── reconciler agent ──────────┐
│ fetch → render → diff → apply → report │
└───────────────────────────────────────┘
An agent inside or near the cluster pulls a declarative desired state from versioned, immutable revisions, compares it with what it observes in the cluster, acts to reduce relevant divergence, and reports evidence. The loop runs continuously — on an interval (Flux defaults to a short poll of the source repository; Argo CD compares roughly every three minutes by default, both configurable) and on webhook triggers for faster reaction. Self-healing is nothing more than this loop doing its job again: someone edits a replica count with kubectl, the next diff sees it, the apply restores the declared value.
The interview version: "GitOps = the desired state of my system lives in Git, and a controller continuously reconciles the cluster toward it." Then stop and let the interviewer pick the thread, because each thread has a hard follow-up.
What a weak answer sounds like: "We deploy from Git." That describes a CI job that runs kubectl apply on merge — push-based, no loop, no drift correction, no evidence. Repository use alone is not GitOps, and commit history alone is not proof of production state. If your answer has no agent, no comparison, and no repeated reconciliation, it isn't GitOps.
The four principles, and what each one buys you
The OpenGitOps principles (CNCF) state four properties: declarative, versioned and immutable, pulled automatically, and continuously reconciled. Interviewers ask for them, but the senior-level answer is what each buys over a push pipeline:
- Declarative — the repo states what, not how. The reconciler owns the imperative work, so the same intent applies the same way every time and the diff is meaningful.
- Versioned and immutable — every change to production is a reviewed commit. You get history, blame, rollback-as-revert, and audit for free from Git's mechanics.
- Pulled automatically — the cluster fetches its own intent. No external system needs inbound cluster credentials (more below).
- Continuously reconciled — drift is corrected, not just noticed. A push pipeline is correct at the moment it pushes and blind afterward.
Division of labor, which interviewers probe to see if you've actually run this: CI still owns build and test — it compiles, runs tests, scans, and publishes an immutable artifact (an image digest, not a mutable tag) with provenance. The GitOps layer owns promotion and deployment — a reviewed change selects that digest for an environment, and the reconciler deploys the verified artifact by immutable identity. Rebuilding from the same source can produce different bytes; a mutable tag like :latest can change with no manifest diff at all. If your pipeline rebuilds the image at deploy time, you've lost the artifact identity that makes rollback and audit trustworthy.
Drift: the operational problem the loop creates
Drift is where GitOps stops being a definition and becomes an on-call decision. Causes, all of which you should be able to name from experience:
- Manual
kubectl edit/kubectl scaleby a human debugging at 2am - Out-of-band controllers (an operator or HPA writing fields the reconciler also manages)
- Autoscaling changing replica counts the manifest also declares
- Emergency mitigations that were correct but never flowed back to Git
Detection is the diff step: the agent compares observed state against the rendered desired state and reports divergence (Argo CD surfaces this as OutOfSync; Flux reports it per-resource). The correction policy is the real question, and it's a policy, not a default:
- Auto-remediate (self-heal on) when the declared state is trusted and the divergence is unauthorized and reversible.
- Alert-only when an out-of-band writer legitimately owns the field — HPA-owned replica counts are the canonical example; you ignore those fields deliberately rather than fight the autoscaler.
The follow-up that separates staff answers: "Self-heal just reverted my emergency fix." Self-healing can erase a live mitigation, fight another legitimate controller, destroy forensic evidence, or repeatedly reapply a bad baseline. So: establish per-field ownership, ignore externally owned or defaulted fields explicitly, alert on unknown writers, and — critically — re-flow the emergency change back through Git. An emergency fix that isn't reconciled into source will be silently reverted by the next sync; the runbook is "fix in Git, or pause automation with a time bound." Pause reconciliation during an investigation when enforcement would amplify harm.
The pull model's credential topology — and its flip side
Pull-based operation exists for a security reason, and interviewers want both halves.
The benefit: a central CI job no longer holds broad inbound credentials to every cluster. The agent runs in or near the cluster and pulls, so no firewall hole, no NAT traversal problem, and air-gapped or customer-owned clusters can fetch their own intent over an outbound connection. This is why GitOps is the standard answer for fleet and edge management — you cannot kubectl apply into a cluster behind NAT that you can't reach.
The flip side, which the interviewer is waiting for:
- The agent still needs privileged access — it creates, updates, and deletes resources in its cluster, so a compromised reconciler is a cluster-wide or fleet-wide attack path.
- A compromised repo is a supply-chain risk: a repo writer must not automatically gain unrestricted cluster-admin. Scope source repositories, paths, namespaces, and resource kinds per application; separate tenants and high-risk platform controls; protect the controller's service account, tokens, network reachability, extensions, and logs.
- Git availability becomes a runtime dependency. If Git is down, new deploys stop — running workloads keep running, but you've coupled your deploy path to your repo host's uptime. Know your answer for that.
- Validate more than the commit: templates, Helm dependencies, plugins, remote bases, registries, and image automation all change the effective desired state. Verify source authenticity and the rendered manifest bundle, not just the Git revision.
"Single source of truth" — qualify it or lose the point
The phrase is bait. Git holds desired state; the cluster holds observed state; and the gap between them is where every hard GitOps problem lives. Say this explicitly and you'll sound like someone who has operated it:
- Secrets don't belong in Git. You need a companion story — sealed secrets, external secret operators, vault injection — and the reconciler has to treat the rendered secret as desired state without the plaintext being versioned.
- Runtime-only state (HPA-managed replicas, PVCs, node state) exists in the cluster and never in Git; your diff policy has to know which fields it owns.
- Git shows no live health. A
Syncedcondition means the controller's comparison found the expected managed state — nothing more.Healthyis tool-specific. Neither proves critical user journeys, authorization, data integrity, or downstream compatibility. Observe source fetch, render, comparison, apply, dependency readiness, health, and post-deployment outcomes as separate signals. - Human approvals don't fit naturally in a pull loop — they live in the PR review that produces the commit, or in a promotion gate between environments, not inside reconciliation.
Deletion, ordering, and promotion — where the loop gets dangerous
Pruning turns absence from desired state into deletion, so it deserves its own safeguards: scope the inventory precisely, preview the deletion set, protect namespaces, persistent data, CRs and shared infrastructure, and require confirmation for high-consequence objects. Empty or failed rendering must never be interpreted as an instruction to erase an environment. Test rename, move, generator failure, source outage, stuck finalizers, partial deletion, and restoration before you trust prune in production.
Ordering: CRDs before custom resources, foundational services before consumers, migrations obeying data compatibility. Sync waves and hooks sequence actions, but imperative hooks need idempotency, timeouts, retries, cleanup, and durable evidence — and selective sync can bypass them. A permanently unhealthy early wave blocks all later progress, so expose and rehearse that recovery path.
Promotion: promote revisions across environments with reviewed changes that identify the exact artifact digest, not implicit branch copying. Use progressive cohorts, observation windows, and guardrails; avoid one commit instantaneously reconciling every production target. Rollback is a new reviewed desired revision — and be ready for the follow-up: reverting Git cannot reverse a database migration, issued credentials, or consumed messages. Sometimes roll-forward or data repair is the safe move, and saying so is the senior answer.
Bootstrap, recovery, and what to measure
Design bootstrap and disaster recovery without circular dependencies: controller installation, source definitions, trust roots, credentials, and last-known-good revisions must be recoverable under a tested procedure. Decide what happens when Git, DNS, identity, the registry, admission, the cluster API, or the reconciler itself is unavailable — reconciliation must resume without applying stale or unverified intent. Break-glass changes are time-bound, audited, visible as drift, and reconciled back afterward.
Measure the loop, not just the deploys: source age, desired vs. observed revision, render failures, sync duration, retry loops, drift count and age, prune candidates, health guardrails, controller saturation, and recovery success. Distinguish source, render, authorization, apply, and workload failure so one root cause doesn't produce an alert storm. And because a compromised controller can report success falsely, independently verify a sample of effective state.
Likely follow-ups and the traps in each
- "Why not just have CI run kubectl apply?" — no loop, no drift correction, CI holds cluster credentials, no evidence after push time.
- "What happens when Git is down?" — running state persists; new promotions stop; you need a last-known-good story and a break-glass procedure.
- "How do you handle secrets?" — never plaintext in Git; companion tooling; the reconciler reconciles a reference.
- "Self-heal reverted my hotfix — what went wrong?" — the fix wasn't reconciled into source; the policy question is field ownership, not the self-heal flag.
- "How do you roll back a schema migration?" — you often can't; roll forward. Git revert reverses intent, not external state.
- "Is GitOps right for everything?" — no: rapid-iteration dev environments sometimes want push; non-declarative systems need adapters; the loop's value shows at environment count and audit requirements.
The one-sentence version to carry into the room: GitOps buys traceability and repeatability only when the entire control loop is governed — authority least-privileged, artifacts immutable and verified, deletions explicit, rollouts bounded, evidence scoped, and recovery rehearsed.
