Overview
Curated: · Written: · Reviewed:
Operators are API products backed by convergent control loops
This is the concept interviewers reach for when they want to know whether you can build Kubernetes extensions that survive production, not just pass a demo. The weak answer sounds like a feature list: "an operator is a controller plus a CRD, it automates day-2 operations." The strong answer explains the division of labour, why reconciliation is level-triggered, and where the sharp edges are: finalizers, external resources, versioning, and the security model. Expect follow-ups on each of those, and expect to be asked to sketch a reconcile function on a whiteboard.
Mental model: what Kubernetes gives you and what you own
A CRD registers a new resource kind with the Kubernetes API server. Once installed, you get for free:
- Structured storage in etcd behind a schema you publish (structural OpenAPI v3, plus CEL rules for object-local validation).
- API discovery, so
kubectl get yourthingsworks with no client code. - RBAC verbs on the new resource, watch/streaming endpoints, and standard machinery like
metadata.resourceVersion-based optimistic concurrency. - Garbage collection of owned objects via
ownerReferences.
What Kubernetes does not give you: anything that makes the spec true. A CRD with no controller is a typed database row. The controller is code you own — a process that watches your resource kind and repeatedly reduces the gap between spec and observed reality. An operator is that controller plus the domain knowledge encoded in it: provisioning, backup, failover, upgrade, deletion.
The interview version: "Kubernetes gives me an API surface, storage, and lifecycle plumbing; I own the control loop that converges state." If you can't say which half a piece of work belongs to, you'll design the wrong boundary — usually by putting imperative workflow steps into the API instead of a cohesive domain object.
When to use one, and when not to
Use a CRD plus controller when users genuinely benefit from declarative, Kubernetes-native state: discovery, RBAC, watch, kubectl, and reconciliation of long-lived domain objects (a PostgreSQLCluster, not a RunMigration).
Don't use one when:
- Ordinary config or built-in resources already express the need. A Deployment plus an HPA beats a custom
AutoScalerCRD. - The need is a transient imperative command. Modeling "trigger a job" as a durable resource creates objects that exist only to be deleted, and the API server becomes a message queue it was never designed to be.
- You need custom storage, protocols, or API semantics CRDs can't express — that's the aggregated API server path, a much heavier commitment.
A common misconception worth naming in an interview: an operator is not "a way to run Helm from Kubernetes." If the CRD's spec is just a rendered template and the controller is a one-shot apply, you've hidden an imperative workflow behind declarative plumbing and added a privileged control-plane component to run it.
How reconciliation actually works
The loop is level-triggered, not edge-triggered. Every invocation means "re-read authoritative state and move it toward spec," never "perform this event exactly once." Watches can close, coalesce, or redeliver; informer caches lag; the process can crash mid-reconcile. None of that may corrupt the outcome, so:
- Reconcile from a full re-read of spec, status, and owned/external state — never from the event payload.
- Make every side effect idempotent, keyed on stable external identifiers (not names, which can be reused after deletion).
- Compare before writing, and retry conflicts with a fresh read.
- Requeue on transient failure with exponential backoff and jitter; classify terminal errors separately so you don't retry the unretryable.
- Periodic resync or explicit drift checks repair missed notifications.
A "state machine" in the spec is not imperative steps. It maps to status transitions: the controller observes which phase reality is in, records it in status, and each reconcile pass decides the next idempotent action from current state. A successful reconcile means "no useful action is currently required," not "never touch this again."
Here is the shape of a reconcile function (Go, controller-runtime, Kubernetes 1.29-era APIs) with a worked trace:
func (r *Reconciler) Reconcile(ctx context.Context, req ctrl.Request) (ctrl.Result, error) {
var db v1.Database
if err := r.Get(ctx, req.NamespacedName, &db); err != nil {
return ctrl.Result{}, client.IgnoreNotFound(err) // deleted: nothing to do
}
if !db.DeletionTimestamp.IsZero() {
return r.finalize(ctx, &db)
}
if db.Generation != db.Status.ObservedGeneration {
meta.SetStatusCondition(&db.Status.Conditions, metav1.Condition{
Type: "Ready", Status: metav1.ConditionFalse,
Reason: "Reconciling", Message: "spec changed",
ObservedGeneration: db.Generation})
if err := r.Status().Update(ctx, &db); err != nil {
return ctrl.Result{}, err // conflict: re-reconcile with fresh read
}
}
id := string(db.UID) // stable across name reuse
inst, err := r.Provider.EnsureInstance(ctx, id, db.Spec.Size)
if err != nil {
return ctrl.Result{}, err // requeue with backoff
}
db.Status.Endpoint = inst.Endpoint
db.Status.ObservedGeneration = db.Generation
return ctrl.Result{}, r.Status().Update(ctx, &db)
}
Trace it for a Database with Generation: 3, Status.ObservedGeneration: 2, and no finalizer:
Getsucceeds;DeletionTimestampis zero, so we skip finalization.Generation != ObservedGeneration, so we write aReady=False, Reason=Reconcilingcondition. If that write conflicts (someone else updated first), we return the error and the next pass re-reads — no stale write lands.EnsureInstanceis keyed ondb.UID, so if this pass runs twice — say the process crashed after step 3 but before step 5 — the second run adopts the existing instance instead of creating a duplicate.- Status is updated with the endpoint and
ObservedGeneration: 3. A client watching status can now trust it reflects spec generation 3.
If the process had crashed between creating the instance and writing status, the next reconcile finds ObservedGeneration still 2, re-runs EnsureInstance (idempotent), and converges. That is the whole design test: any step can run twice or never complete, and the system still moves toward spec without leaking or duplicating.
API design: spec, status, and the subresources
Design the API before the controller. Name a cohesive domain object, not provider implementation steps.
- User-owned desired state goes in
spec; controller-observed state goes instatus; identity and coordination stay inmetadata. - The API server bumps
metadata.generationon spec changes; the controller recordsstatus.observedGeneration. Clients compare the two to know whether status reflects the latest spec — the single most useful status field, and a common thing interviewers probe for. - Use
metav1.Conditionconventions (type, status, reason, message,lastTransitionTime,observedGeneration) rather than ad-hoc string fields. - Enable the status subresource. It splits the object into two write endpoints:
/statuswrites can only change status, and main-resource writes can only change spec and metadata. That gives you two things: separate RBAC (ordinary spec writers can't forge status), and no clobbering — a status write can't overwrite a user's concurrent spec edit, and a spec write can't wipe the controller's status. Note that any write, either endpoint, still bumpsresourceVersion; the subresource changes what a write may touch, not the optimistic-concurrency mechanism. The scale subresource does the same for replica-count fields. - Events are transient diagnostics, not durable state. Anything a client must rely on belongs in status.
Schema design details that come up in follow-ups:
- Structural schemas are mandatory for practical purposes (defaulting, pruning of unknown fields in some API versions, CEL). Publish CEL rules for object-local validation — e.g.
spec.replicas >= 1— before reaching for admission webhooks. - Defaulting and immutability are API contracts, not conveniences. A defaulted field (
x-kubernetes-list-type, ordefault: 3onspec.replicas) is persisted into stored objects, so changing the default later only affects newly created objects — old objects keep the old value until touched, which surprises people during upgrades. Mark fields immutable withx-kubernetes-validationsCEL (e.g.old.self == self) or an admission rule when changing them would be destructive — astorageClassNameor an engine version you can't switch in place. Decide both at schema-design time; retrofitting either is a breaking change. - Decide namespaced vs cluster-scoped deliberately; cluster-scoped widens blast radius.
- Add printer columns, short names, and categories so
kubectl getis usable. - Objects are capped at ~1.5 MB in etcd; large specs or status blobs are a design smell and can pressure the API server.
- Version the CRD as a public API from day one. Multiple served versions force conversion (a no-op only if versions are trivially compatible); exactly one version is the storage version. Changing the storage version does not rewrite existing objects — you must migrate stored objects and check
status.storedVersionsbefore dropping an old schema. Conversion webhooks are control-plane dependencies needing availability, auth, and rollback plans. This is real engineering cost; interviewers like hearing that multiple versions are not free.
Ownership, deletion, and finalizers
Child Kubernetes objects get a label selector plus a valid ownerReference with controller: true. Garbage collection then cascades deletion — but only within ownership rules: a namespaced object cannot own a cluster-scoped object, and cross-namespace ownership is invalid. Observe child readiness; creating a Deployment is not the same as a ready service.
External resources (cloud databases, DNS records, S3 buckets) get no automatic GC, so:
- Persist a stable binding between the CR's UID and the provider resource, and tag external objects for discovery.
- Make create/delete idempotent and handle "create succeeded but the response was lost" by looking up before creating.
- Have an explicit authority policy for out-of-band drift: does the operator win, or does it surface the conflict? An operator that fights a human or another automation oscillates forever.
Finalizers turn deletion into a two-phase protocol. When deletionTimestamp appears: stop provisioning, run bounded idempotent cleanup, report progress in status, and remove only your own finalizer once the obligation is satisfied or an authorized abandonment decision is recorded. A stuck finalizer blocks deletion indefinitely — that's the failure mode interviewers ask about — and blindly removing one leaks infrastructure or data. The runbook answer: alert on stuck deletion, retry with backoff, and have a controlled force-removal process that records impact evidence.
Security and operations at scale
The operator is a privileged control-plane component:
- Grant narrowly scoped verbs and resources; separate namespace-scoped from cluster-scoped permissions; isolate its service account and network; validate external endpoints; protect the build and dependency chain.
- The tenant authorization boundary belongs at the API/admission layer. A controller that turns "permission to create a CR" into unrestricted infrastructure authority is a confused deputy — a classic senior-level follow-up.
- Run leader election (or another single-active mechanism) where concurrent work is unsafe, but keep reconciliation idempotent anyway: leases reduce duplicates, they don't eliminate them (a partitioned old leader can still act).
- Bound work queues and reconcile concurrency; use backoff with jitter; monitor reconcile latency, error rates, queue depth, stale generations, stuck deletions, and external rate limits. Hot loops and full relists can take down an API server — this is the "what changes at scale" answer.
- Prefer schema validation and CEL over admission webhooks; webhooks sit on the synchronous write path, so slow or broad ones disrupt the cluster. If you need one, minimize scope and latency, make mutation idempotent, support dry-run, and plan failure behavior explicitly. Admission validates and defaults; the controller does asynchronous provisioning.
Testing and what interviewers probe
Test more than the happy path, because the design contract is convergence under interruption: schema rejection and defaulting, status permission failures, repeated reconciliation, crash between side effect and status write, stale caches, update conflicts, deleted-then-recreated names, partial child readiness, external drift, deletion during provisioning, finalizer failure, leader failover, API throttling, conversion round trips, storage migration. Use envtest for API behavior and integration environments for external effects.
The follow-up ladder you should expect:
- "What does a CRD give you for free?" — storage, discovery, RBAC, watch, GC. Not automation.
- "Your controller crashed mid-reconcile — what happens?" — the level-triggered re-read makes it safe; walk the trace above.
- "How do you clean up a cloud resource when the CR is deleted?" — finalizers, UID binding, idempotent delete.
- "You added a v2 of your CRD — now what?" — conversion, storage version,
storedVersionsmigration, old-client testing. - "Who's allowed to create your CR, and what does that actually grant?" — the confused-deputy question.
The weak answer to every one of these is a feature name with no mechanism. The strong answer names the mechanism and its failure mode. An operator earns its keep when it reliably packages operational knowledge behind a well-designed API; it's a liability when it hides an imperative workflow behind a CRD while adding an under-observed, over-privileged control plane.
