Overview
Curated: · Written: · Reviewed:
Deployment strategies: bounded exposure while evidence accumulates
A deployment strategy is a risk-control policy for moving a verified release into service. Every strategy is a point on four axes: delivery speed, blast radius, infrastructure cost, and rollback speed. All-at-once is fast and cheap but exposes everyone and rolls back slowly. Blue-green is slow to provision and costs double capacity but flips traffic in seconds and reverts the same way. Interviewers grade you on whether you reason along these axes or just name a favorite.
What interviewers actually probe
The question "how do you deploy?" is rarely about Kubernetes flags. The probes that follow a first answer are:
- "What happens to in-flight requests during the switch?" — testing whether you know cutover is not atomic for connections, sessions, and queues.
- "How does this work with a schema migration?" — the question that separates people who have run rollouts from people who have read about them. A traffic rollback cannot undo a destructive migration.
- "How do you know the canary is good enough to promote?" — testing guardrail design, sample size, and whether you know a global average can hide one burning tenant (Simpson's paradox).
- "What runs during the rollout that isn't versioned the same way?" — background jobs, cron, consumers, caches.
A weak answer sounds like: "We use canary deployments because they're best practice," followed by silence when asked what fraction, for how long, measured against what baseline, and what happens to data written by the canary if you abort. Another weak pattern is describing only the happy path — no abort authority, no recovery observation, no compatibility window.
A strong answer names the strategy, states the four-axis trade-off explicitly, then volunteers the hard constraint unprompted: "canary at 1/5/25/100 with error-budget-burn guardrails, expand-contract schema so both versions can run, and abort is a weight revert plus a write-pause while we reconcile."
The mechanics, precisely
All-at-once. Terminate old, start new. No old version exists after cutover; rollback is a full redeploy of the previous artifact. Blast radius is 100% of capacity at the moment of switch. It's the right answer for stateless internal tools and the wrong answer for anything with users.
Rolling. Replace instances incrementally. In Kubernetes, maxSurge controls how many extra replicas may exist above the desired count during rollout, and maxUnavailable controls how much of the desired count may be not-ready at once. Old and new versions coexist for the whole rollout — that's the defining property, not a bug. Rollback is a slow reverse-roll: you re-run the same mechanism against the previous artifact, so recovery time is proportional to rollout time. A bad release spreads while evidence lags, so you need backward/forward protocol compatibility, correct readiness probes, a progress deadline, and surge capacity.
Blue-green. Stand up a complete replacement environment, validate it on a preview or test path, then switch production traffic — a pointer flip at the load balancer. The old environment stays warm, so rollback is switching the pointer back, seconds not minutes. The costs: duplicate capacity, and cutover does not remove state risks — cache warming, long-lived connections, DNS TTLs, sticky sessions, in-flight queue messages, and external callbacks all cross the boundary on their own schedules. Both environments must not perform conflicting singleton or background work. Keep the old environment only for a bounded, audited window.
Canary. Expose a small population or traffic fraction, increase through explicit steps (e.g., 1% → 5% → 25% → 100%), compare against a concurrent control cohort, and abort by reverting the traffic weight. The old version remains fully deployed the entire time — that's why canary rollback is the fastest of any strategy. The cohort must be representative enough to reveal the target risks without concentrating harm; route by stable random sampling, tenant, or region, and prevent session/version inconsistency where it breaks semantics.
Linear. A canary with fixed-size, evenly spaced steps and fixed dwell time per step — the App Engine / Argo-style model. It trades the judgment calls of ad-hoc canary steps for predictability, and it's what most progressive-delivery controllers actually implement.
| Strategy | Old version during rollout | Traffic shift | Rollback | Extra capacity |
|---|---|---|---|---|
| All-at-once | None after switch | Instant, total | Redeploy (slow) | None |
| Rolling | Mixed, per-instance | Gradual with topology | Reverse-roll (slow) | Surge only |
| Blue-green | Full parallel environment | Pointer flip | Pointer flip back (fast) | ~2× |
| Canary / linear | Full parallel deployment | Weighted, stepped | Weight revert (fast) | ~2× at full canary |
Choosing between them
The decision inputs: statelessness, traffic volume, risk tolerance, compliance and audit needs, cost of double infrastructure, and — the one people forget — how fast you can detect failure. A canary with a 24-hour observation window is worthless if your guardrail metrics have 6 hours of lag; you'll promote on noise.
- Stateless, low-traffic internal service → rolling is fine; the blast radius is small and the team is the victim.
- High-traffic user-facing service with good observability → canary or linear; you can afford the double capacity and your metrics arrive fast enough to act on.
- Cutover must be auditable and instant (compliance, financial) → blue-green; the switch is one logged, reversible action.
- Cannot afford double capacity → rolling with strict compatibility requirements, accepting slower rollback.
What each buys you: rolling buys economy; blue-green buys cutover and rollback speed; canary buys evidence before exposure; all-at-once buys only simplicity. None of them buys data safety — that's the next section.
Schema changes are the hard constraint
Blue-green and rolling both break on non-backward-compatible migrations, for the same reason: both require old and new code to run against the same data store simultaneously. If the migration renames a column the old version still reads, your traffic rollback now routes old code to an incompatible schema — you've converted a recoverable failure into an unrecoverable one.
The fix is expand-contract (parallel change), and it's the answer to the schema follow-up every time:
- Expand: add the new column/table/field, nullable or defaulted. Both versions can write it; old code ignores it.
- Migrate: backfill with a tolerant, idempotent migration while both versions run. New code reads new, falls back to old, or dual-writes per your read strategy.
- Contract: only after the last old-version reader is gone, remove the old shape. This is a separate, independently gated deployment — never bundled with the release that introduced the change.
Separate application exposure from destructive state contraction. A traffic rollback cannot reverse corrupted, duplicated, or transformed data. Define the trusted source, repair/replay procedures, and control totals before you need them.
Version skew: N and N+1 at once
During any rollout except all-at-once, version N and N+1 serve traffic simultaneously. This is where mixed-version semantics bite:
- Queues: a message produced by N+1 may be consumed by N. The contract — schema, headers, semantics — must be forward- and backward-compatible, or consumers must tolerate and defer what they can't parse.
- Caches: values written by one version and read by another need versioned keys or tolerant readers, or you get cross-version corruption that outlives the rollout.
- Clients: long-lived mobile clients and WebSockets don't upgrade on your schedule; your API surface has to keep old clients working for the whole compatibility window, not just the rollout duration.
Traffic splitting must match the actual request graph. Weighting an edge proxy does not ensure the same fraction reaches every async worker, WebSocket, callback, batch job, or downstream service. Retries distort observed exposure; mirroring duplicates requests and must be restricted to safe, idempotent, read-only paths with protected data. Verify the achieved weight, not the configured one, and test routing-controller outage and fallback.
Guardrails, promotion, and abort
Guardrails cover more than error rate: critical-journey success and latency, error-budget burn, data invariants and reconciliation, security and authorization outcomes, queue depth and freshness, dependency health, and monitoring health itself. Predeclare thresholds, minimum sample sizes, evaluation delay, allowed missing data, and abort authority. Low sample counts, delayed outcomes, and multiple comparisons can make automatic promotion unsafe — a 1% canary on a low-traffic service may see ten requests in an hour, which proves nothing.
Promotion and abort are state machines, not buttons. Serialize changes to one target, reject stale controllers, make steps idempotent, and record artifact, exposure, analysis, decision, and owner. On failure, choose from traffic rollback, feature disablement, isolation, write pause, roll-forward, or data repair — based on compatibility and consequence, not habit. Then observe recovery: connection drains, retry storms, queue backlog, cache refill, and autoscaling can extend harm after the abort itself.
Feature flags decouple code deployment from feature exposure and are the cheapest abort lever of all — but they're runtime configuration and authorization-adjacent control, not automatically safe. Define owner, default/fail behavior, targeting privacy, audit, and expiry. Prevent client-controlled context from granting protected functionality, test provider outage and kill-switch behavior, and delete both flag and dead branches after full rollout.
What a completed rollout does and does not prove
A successful canary proves only the tested cohort, time window, workload, metrics, and state. A blue-green switch does not guarantee data rollback. A rolling update completing does not prove mixed-version compatibility — it proves only that the orchestrator saw ready replicas. Validate the running version distribution and user journeys, not the available-replica count.
Rehearse the strategy before depending on it: rollout, pause, abort, switchback, delayed and missing metrics, controller outage, node loss, insufficient capacity, sticky sessions, and partial state change. Measure lead time, guardrail detection time, change failure and user harm, abort and recovery success, and stale releases and flags — not raw deployment count.
Likely follow-ups and how to answer them
- "Canary or blue-green?" — Answer with the axes: detection speed, cost of double infra, state compatibility. Never pick without stating the workload's properties.
- "How do you roll back a canary that already wrote data in the new format?" — Expand-contract means old code still reads the data; abort is a weight revert plus reconciliation. If you can't answer this, the interviewer has found the gap they were looking for.
- "What's your minimum canary cohort and dwell time?" — Give real figures from your system and justify them from metric latency and sample-size needs, even if hypothetical and labelled as such.
- "What breaks during the rollout that isn't the new code?" — Capacity, connection draining, background jobs, caches, external callbacks. Naming these unprompted is what a staff-level answer sounds like.
