Skip to content
Tech Interview Prep home

Top 100 DevOps Engineer Interview Questions and Answers

The questions most likely to actually come up in your DevOps Engineer interview, ranked by likelihood — with detailed, senior-level answers covering what an interviewer is really listening for.

Curated: · Written: · Reviewed:

Reviewed 50Review pending 50
QA-1Staging ran image checkout:4.12 and production rebuilt checkout:4.12 from the same git SHA. Why is that still two different releases?(show answer)

The first thing I would pin down about immutable artifact promotion is which merge or promote it is allowed to block.

Promote a content-addressed artifact; never rebuild for a later environment, because a second build is a second binary even when the commit is the same.

Concretely, build once, record the digest, and have staging, canary, and production apply that digest. Refuse any job whose Dockerfile, base image, or build-arg set can differ by environment. Record the digest, the gate that ran, and the rollback command for immutable artifact promotion on the same change, and store that record beside the promote job so an incident later can name the bits that shipped.

The reason for that specificity is a failure I have seen: Production rebuilt checkout:4.12 after a Debian mirror rotated; the digest diverged and a glibc patch that staging never saw caused a 14-minute crash loop on 12% of pods.

Same git SHA, two builds, two productions.

EnvironmentGit SHAImageDigest
staging9f3c1a2checkout:4.12sha256:aa11…
production rebuild9f3c1a2checkout:4.12sha256:bb22…
required9f3c1a2checkout@sha256:aa11…sha256:aa11…

I would not consider it settled without evidence: Show the same sha256 in the staging deploy record, the production GitOps manifest, and the cluster's running image, with a pipeline that fails if any of the three differ.

A tag that moved is not a release; a digest that was promoted is. That record is part of the change, not a wiki page written afterwards.

Curated: · Written: · Reviewed:

QA-2The prod workflow runs docker build so "production uses production credentials". What do you change?(show answer)

I would start rebuild-in-prod as a defect from the artifact that actually ships, not from the job that happened to go green.

Credentials belong in the runtime identity, not in a second compile; rebuilding in production is a supply-chain and reproducibility defect.

Concretely, build in the isolated CI identity, sign the digest, and have production only pull and run. Inject runtime secrets via the workload identity, not via build-args that change the image. Record the digest, the gate that ran, and the rollback command for rebuild-in-prod as a defect on the same change, and store that record beside the promote job so an incident later can name the bits that shipped.

The reason for that specificity is a failure I have seen: A prod-only build-arg baked a database URL into a layer; the image scanned clean in CI and dumped the URL in docker history after a support copy.

Prod workflow before and after.

# reject
jobs.prod: docker build --build-arg DB_URL=...
# accept
jobs.prod: cosign verify $DIGEST && helm upgrade --set image=$DIGEST
runtime: IRSA/workload identity, not a build-arg

I would not consider it settled without evidence: A policy that production workflows contain no build or compile steps, plus one failing check on a PR that adds docker build under the prod job.

If production still compiles, you do not have a release; you have a second factory. That record is part of the change, not a wiki page written afterwards.

Curated: · Written: · Reviewed:

QA-3Every manifest says image: checkout:stable. A node reschedule pulled a different binary. What should GitOps store?(show answer)

This is a place where a passing pipeline and a safe production change for digest pinning versus mutable tags are not the same event.

Store the digest in the desired state; a floating tag is a late-binding rebuild that Git cannot explain.

Concretely, write image@sha256:… (or a tag plus digest) in the environment overlay. Automation may open a PR that updates the digest; humans review that PR, not a moving tag. Record the digest, the gate that ran, and the rollback command for digest pinning versus mutable tags on the same change, and store that record beside the promote job so an incident later can name the bits that shipped.

The reason for that specificity is a failure I have seen: checkout:stable was retagged during a registry GC; 30 pods on a new node ran 4.11 while 270 still ran 4.12, and canary analysis compared the wrong binary.

Floating tag versus pinned digest after a node recycle.

SpecAfter recycleRunning digest
checkout:stablepull :stablesha256:old or sha256:new
checkout@sha256:aa11pull aa11sha256:aa11
mismatch30 of 300 podsundeclared 4.11

I would not consider it settled without evidence: kubectl get pod -o jsonpath of the running image ID matches the overlay digest on 100% of replicas after a 24-hour node recycle.

If Git does not name the bits, the cluster will pick some. That record is part of the change, not a wiki page written afterwards.

Curated: · Written: · Reviewed:

QA-4integration is required, but a push with [skip ci] still landed on main. How did the gate fail, and what makes it real?(show answer)

My answer to required checks that can be skipped begins at the promotion boundary, because that is where delivery either has a gate or has a story.

A required check that never ran stays pending and should block merge; a merge that skipped CI means bypass, a direct push, or a path filter that posted success without running the suite.

Concretely, inventory bypass actors and force-push rights. Use a merge queue. Path filters must fail closed: a no-match reporter still posts a status, and skipped-path success is an explicit empty-suite policy, not a missing job. Record the digest, the gate that ran, and the rollback command for required checks that can be skipped on the same change, and store that record beside the promote job so an incident later can name the bits that shipped.

The reason for that specificity is a failure I have seen: A docs-only path filter treated a Helm values change under charts/ as documentation and posted success; production shipped without the integration job for 11 commits.

200 merges to main, status presence.

Path filterIntegration status postedMerges without bypass
docs/** fail-open successsuccess, suite not run11 bad
docs/** always-run reportersuccess or fail200
missing required check, no bypasspending0

I would not consider it settled without evidence: For the last 200 merges to main, every SHA has a success or failure (not missing) status from the gate jobs, and the bypass-actor list is empty except named break-glass.

A missing status is not a pass; if it merged anyway, someone had a bypass. That record is part of the change, not a wiki page written afterwards.

Curated: · Written: · Reviewed:

QA-5PR #441 was green on Monday and merged on Thursday after main moved 40 commits. What did you fail to test?(show answer)

I would treat merge queue versus a stale green as a control in the path to production rather than as a dashboard tile after the fact.

Test the merge result, not the branch that was green against an old main, or you ship an unbuilt combination.

Concretely, require a merge queue (or rebase-and-retest) so the required jobs run on main + this PR. Delete the ability to merge a green check older than the queue's freshness window. Record the digest, the gate that ran, and the rollback command for merge queue versus a stale green on the same change, and store that record beside the promote job so an incident later can name the bits that shipped.

The reason for that specificity is a failure I have seen: #441 was green against main@A; main had migrated a proto by Thursday; the merge compiled in isolation and broke checkout for 22 minutes.

Stale green versus merge-queue SHA.

EventSHA testedResult
Monday CIpr441 vs main@Agreen
Thursday mergepr441 vs main@A+40 (untested)proto break, 22 min
merge queuemerge(main@NOW, pr441)red, not merged

I would not consider it settled without evidence: Queue logs showing the tested SHA is a merge commit of current main, and a rejected merge when the PR's last check is older than the freshness window (for example 60 minutes).

Green on an ancestor is not green on the thing you are about to ship. That record is part of the change, not a wiki page written afterwards.

Curated: · Written: · Reviewed:

QA-6The team cuts a release branch every two weeks with 180 commits. How do you shrink what production has to absorb?(show answer)

The useful question for trunk-based delivery batch size is what a developer is still able to merge if the control is skipped.

Keep the batch that can reach production small; a fortnight of commits is an incident waiting for a name.

Concretely, trunk-based development with feature flags, multiple promotes per day, and a batch-size SLO (for example ≤15 production-bound commits per promote). Long-lived release branches are an exception with an expiry. Record the digest, the gate that ran, and the rollback command for trunk-based delivery batch size on the same change, and store that record beside the promote job so an incident later can name the bits that shipped.

The reason for that specificity is a failure I have seen: Release/2026-08-22 contained 180 commits; the rollback reverted 179 good changes to undo one bad cache key, and recovery took 3 hours 40 minutes.

Batch size versus rollback blast radius, 28 days.

ModeCommits / promoteRollback time
fortnightly branch1803 h 40 m
daily train2241 m
trunk + flags, 6 / day89 m

I would not consider it settled without evidence: Median commits per production promote over 28 days, with a cap, and a count of release branches older than 48 hours.

You cannot roll back what you cannot name; a 180-commit train has no name. That record is part of the change, not a wiki page written afterwards.

Curated: · Written: · Reviewed:

QA-7CI runs unit tests only when src/** changes. A change to terraform/modules/network shipped without the integration job. What is the rule?(show answer)

I would settle path filters that hide tests against a reversible promote, so a bad change has a measured way back.

Path filters must fail closed for anything that can change runtime behaviour, including IaC, charts, policies, and workflow files.

Concretely, map path families to jobs explicitly. Workflow, chart, terraform, and policy paths always run the integration and policy jobs. A docs-only skip is a named allowlist, not the default. Record the digest, the gate that ran, and the rollback command for path filters that hide tests on the same change, and store that record beside the promote job so an incident later can name the bits that shipped.

The reason for that specificity is a failure I have seen: A one-line change to terraform/modules/network matched no src/** filter; the apply from main took production DNS private and CI still reported green.

Fail-closed path map.

Pathunitintegrationpolicy
src/**yesyesyes
terraform/**noyesyes
charts/**noyesyes
docs/**nonono

I would not consider it settled without evidence: A table in the platform repo: path glob → jobs, with a test that adding a file under terraform/ fails the PR unless integration ran.

If the filter can ignore the thing that ships, the filter is the outage. That record is part of the change, not a wiki page written afterwards.

Curated: · Written: · Reviewed:

QA-8A flaky checkout spec has been skipped with xfail for 11 weeks. What policy makes quarantine something other than a junk drawer?(show answer)

The judgement in flaky test quarantine with expiry is which check is required to ship, not which check looks impressive on a pull request.

Quarantine is a time-boxed ticket with an owner and a kill date, not a skip that silently lowers the suite.

Concretely, a skip requires a ticket, an owner, and an expiry ≤14 days. CI fails if a quarantine annotation is past expiry. Track quarantine count as a DORA input, not as a badge. Record the digest, the gate that ran, and the rollback command for flaky test quarantine with expiry on the same change, and store that record beside the promote job so an incident later can name the bits that shipped.

The reason for that specificity is a failure I have seen: 11 quarantined specs hid a real payment regression for 9 days; the suite was "green" at 94% of the tests the team thought they had.

Quarantine ledger.

SpecOwnerExpiryAge
checkout_spec:88payments2026-06-0211 weeks, expired
policyany≤14 daysCI fails if overdue
current allowed3all < 8 days2000 specs

I would not consider it settled without evidence: A report of open quarantines with ages, and a pipeline check that fails the build when any quarantine date is in the past.

A skip without an expiry is a deleted test with extra ceremony. That record is part of the change, not a wiki page written afterwards.

Curated: · Written: · Reviewed:

QA-9coverage is required and the job prints 82%. The PR still merges at 61%. Where is the gate?(show answer)

Where teams lose a week on coverage gates that never fail the build is usually a bypass that was left in for one hotfix and never removed.

A coverage number that is not a failing status check is decoration; the gate is the exit code, not the badge.

Concretely, fail the job when coverage drops below the branch floor or when diff coverage on changed lines is below the policy (for example 80%). Upload the report, but the required check is the failure. Record the digest, the gate that ran, and the rollback command for coverage gates that never fail the build on the same change, and store that record beside the promote job so an incident later can name the bits that shipped.

The reason for that specificity is a failure I have seen: The coverage job always exited 0 and posted 82% from main's last default branch; 40 PRs merged at 61% on the changed lines.

Badge versus failing check.

SetupPostedExitMerge
report only82% (main)0yes at 61% diff
diff coverage 80%61%1blocked
floor 75% on main82%0ok if diff also holds

I would not consider it settled without evidence: A PR that deletes a tested branch without a replacement is red, and the required check name matches the job that can exit non-zero.

If the badge cannot stop a merge, you have a screenshot, not a gate. That record is part of the change, not a wiki page written afterwards.

Curated: · Written: · Reviewed:

QA-10Checkout and payments are different repos. Checkout merged a field rename; payments produced 19% 400s after the next promote. What should have blocked the first merge?(show answer)

I would answer contract tests as a promote gate by separating what the pipeline proved from what production has not yet seen.

Consumer-driven contracts run in CI of both sides before either artifact is promotable.

Concretely, publish the consumer contract from checkout's pipeline; payments verifies it on every PR. Breaking changes require a versioned field and a dual-read window, not a rename in one deploy. Record the digest, the gate that ran, and the rollback command for contract tests as a promote gate on the same change, and store that record beside the promote job so an incident later can name the bits that shipped.

The reason for that specificity is a failure I have seen: The rename shipped in checkout 4.12 while payments still required card_last4; 19% of payments were 400 for 47 minutes.

Contract gate on a field rename.

Changecheckout CIpayments CIPromote
rename card_last4green (unit)not run19% 400s
pact + dual-readred until v2 fieldgreen on v1+v2allowed
after 14d drop v1greengreenallowed

I would not consider it settled without evidence: A payments PR that drops a field used by the published contract is red, and a checkout PR that removes a field without a deprecation version is red.

If the first time two services meet is production, the pipeline did not integrate them. That record is part of the change, not a wiki page written afterwards.

Curated: · Written: · Reviewed:

QA-11A 40-package monorepo runs every test on every PR (41 minutes). Skipping by git diff broke a shared library. How do you select work without lying?(show answer)

The engineering content of monorepo affected-package CI is the lead time and the change-fail cost, not the number of stages.

Affected-package selection must follow the dependency graph, not the list of dirty files, or a library change ships untested consumers.

Concretely, use a graph-aware tool (Bazel, Nx, Gradle, or equivalent) so a change to lib/payments marks every dependent package. Unknown or graph-generation failure runs the full suite, fail closed. Record the digest, the gate that ran, and the rollback command for monorepo affected-package CI on the same change, and store that record beside the promote job so an incident later can name the bits that shipped.

The reason for that specificity is a failure I have seen: git-diff selection ran only lib/payments tests; checkout was not marked; a tax-rounding change passed CI and overcharged 4,200 orders.

Selection methods on a library change.

SelectorPackages testedResult
git difflib/paymentstax bug in checkout
Nx affectedlib/payments, checkout, bffcaught
graph staleall 40slow, safe

I would not consider it settled without evidence: A fixture where only lib/payments changes and checkout's tests still run, plus a fail-closed run when the graph file is stale.

Dirty files are not dependents; the graph is. That record is part of the change, not a wiki page written afterwards.

Curated: · Written: · Reviewed:

QA-12The remote cache cut CI from 24 minutes to 6. An untrusted fork PR wrote a poisoned object. What is the trust model?(show answer)

Before adding another job I would write what a good delivery of remote build cache authentication looks like in DORA terms.

Write access to a shared build cache is equivalent to write access to every downstream binary; forks and untrusted events get read-only or no cache.

Concretely, authenticate cache writes with the trusted default-branch identity. pull_request from forks uses a read-only token or an isolated cache namespace. Sign or hash cache keys that include the command line. Record the digest, the gate that ran, and the rollback command for remote build cache authentication on the same change, and store that record beside the promote job so an incident later can name the bits that shipped.

The reason for that specificity is a failure I have seen: A fork PR uploaded a cache entry for //checkout:test that always passed; main reused it for 6 hours and skipped a failing assertion.

Cache trust.

EventReadWriteNamespace
push to mainyesyesprod-cache
PR from forkoptionalnonone or fork-*
poisoned put from fork—blockedprod-cache

I would not consider it settled without evidence: Fork workflows cannot cache put on the main namespace; a canary job on main that tampers a key is rejected.

A cache that anyone can fill is a remote code execution with extra steps. That record is part of the change, not a wiki page written afterwards.

Curated: · Written: · Reviewed:

QA-13GitHub Actions still has AWS_SECRET_ACCESS_KEY in repository secrets. What do you replace it with?(show answer)

The first thing I would pin down about OIDC to cloud instead of long-lived keys is which merge or promote it is allowed to block.

The pipeline assumes a cloud role through a short-lived OIDC token whose aud and sub match that job; a static key in GitHub is a standing credential.

Concretely, create an identity provider for the CI issuer and a role whose condition matches the exact aud plus the exact sub GitHub will issue for that job. Jobs that use a GitHub Environment get repo:ORG/REPO:environment:NAME; jobs that do not get a ref: subject. Bind production to the environment subject and restrict which branches may use that environment; do not write a trust policy that claims both ref and environment in one sub. Record the digest, the gate that ran, and the rollback command for OIDC to cloud instead of long-lived keys on the same change, and store that record beside the promote job so an incident later can name the bits that shipped.

The reason for that specificity is a failure I have seen: A leaked Actions secret from a forked workflow dump had AdministratorAccess for 18 days; the key was still in GitHub after the person who created it left.

OIDC role condition.

{"StringEquals":{
  "token.actions.githubusercontent.com:aud":"sts.amazonaws.com",
  "token.actions.githubusercontent.com:sub":"repo:acme/checkout:environment:production"
}}

PR jobs do not use that environment, so they cannot assume this role. Allowed branches are set on the GitHub environment, not stuffed into the same sub.

I would not consider it settled without evidence: CloudTrail (or equivalent) shows AssumeRoleWithWebIdentity from the CI issuer and zero AccessKeySignin for the old user, plus a secret scan that finds no AWS_SECRET in the repo.

If the key still works next quarter, it is not CI; it is a standing admin. That record is part of the change, not a wiki page written afterwards.

Curated: · Written: · Reviewed:

QA-14Anyone with write on the repo can click "Deploy to production". What belongs on the production environment?(show answer)

I would start environment protection rules from the artifact that actually ships, not from the job that happened to go green.

Production is a protected environment: required reviewers, restricted secrets, and a deployment branch policy, not a job anyone with push can dispatch.

Concretely, map the production job to a GitHub (or equivalent) environment with required reviewers, wait timer if policy asks, and secrets that do not exist on pull_request. Limit which refs may deploy. Record the digest, the gate that ran, and the rollback command for environment protection rules on the same change, and store that record beside the promote job so an incident later can name the bits that shipped.

The reason for that specificity is a failure I have seen: A contractor with write pushed a workflow_dispatch to production at 01:12 and rolled a debug image; there was no reviewer and the secret was the same as staging.

Production environment policy.

RuleStagingProduction
required reviewers01 (not the author)
allowed refsanymain
secretsSTAGING_*PROD_* only on this env
wait timer05 min

I would not consider it settled without evidence: An attempt to deploy from a feature branch is rejected, and an attempt without the second reviewer stays waiting, not running.

Write access to git is not authority to run production. That record is part of the change, not a wiki page written afterwards.

Curated: · Written: · Reviewed:

QA-15The on-call is also the author of the release. Policy says four eyes. GitHub Environments need one reviewer. How do you actually get two?(show answer)

This is a place where a passing pipeline and a safe production change for two-person production promote are not the same event.

GitHub environment required reviewers typically need one approval from the reviewer list and can block the deploying actor from self-approving; they do not by themselves prove a second distinct human besides author and initiator.

Concretely, keep the environment reviewer rule, then add a check that the GitHub actor who queued, the committer, and a third recorded approver are three different people, or use a deploy system that stores those identities. Break-glass is a separate environment with a recorded reason and a 60-minute token. Record the digest, the gate that ran, and the rollback command for two-person production promote on the same change, and store that record beside the promote job so an incident later can name the bits that shipped.

The reason for that specificity is a failure I have seen: Self-approvals were assumed to be impossible; 7 of 12 production promotes in a month had the same person as committer and sole environment approver because they queued from another account.

Approver identity, 90 days.

PromotesAuthor=approverValid two-personBreak-glass
12 (before)750 recorded
12 (after)0111, 38 min token

I would not consider it settled without evidence: Audit log: for 90 days, production approval user ≠ committer and ≠ queueing actor, except on named break-glass records.

If you can approve your own promote, you do not have four eyes; you have a checkbox. That record is part of the change, not a wiki page written afterwards.

Curated: · Written: · Reviewed:

QA-16Sev-1, payments is down, the integration job is stuck 24 minutes. How do you skip without making skip the culture?(show answer)

My answer to emergency skip of gates begins at the promotion boundary, because that is where delivery either has a gate or has a story.

A skip is a named break-glass workflow with a ticket, a time limit, and a mandatory follow-up that re-runs the skipped gates on the same digest.

Concretely, one prod-break-glass environment, two-person, reason required, 60-minute credentials. The next pipeline on that digest must run the skipped jobs or the release stays marked skipped. Record the digest, the gate that ran, and the rollback command for emergency skip of gates on the same change, and store that record beside the promote job so an incident later can name the bits that shipped.

The reason for that specificity is a failure I have seen: A sticky continue-on-error: true left on the integration job after a Sev-1; 16 more promotes skipped it for 11 days.

Break-glass ledger.

DigestJobs skippedWindowRe-run
sha256:cc33integration60 mingreen +18 min
sticky continue-on-errorall integration11 daysnever, 16 promotes

I would not consider it settled without evidence: A ledger of skips with digest, jobs skipped, restoral time, and a failing check if a skipped digest is still the production pointer without a later green run.

An undocumented skip is just how you deploy now. That record is part of the change, not a wiki page written afterwards.

Curated: · Written: · Reviewed:

QA-17The dashboard says lead time is 2.1 days because that is PR open to merge. Production happens the next afternoon. What should you measure?(show answer)

I would treat DORA lead time from commit to production as a control in the path to production rather than as a dashboard tile after the fact.

Lead time for changes is first production-bound commit to production traffic, not merge time, or you will optimise the pull request and ignore the train.

Concretely, stamp commit time, merge time, and first-healthy-in-production time on the digest. Report p50/p90 of commit→prod. A merge that sits in staging overnight still counts. Record the digest, the gate that ran, and the rollback command for DORA lead time from commit to production on the same change, and store that record beside the promote job so an incident later can name the bits that shipped.

The reason for that specificity is a failure I have seen: The team got PRs merging in 4 hours and still took 2 days to production; a Friday merge sat until Monday's train and a Monday incident blamed "CI".

Two clocks, 28 days, p50.

Clockp50p90
PR open → merge4.2 h1.1 d
merge → production18 h2.4 d
commit → production22 h2.8 d

I would not consider it settled without evidence: A 28-day plot of commit→prod and merge→prod; if they diverge by more than a few hours, the train is the bottleneck, not the review.

A fast merge of a change that is not in production is inventory, not delivery. That record is part of the change, not a wiki page written afterwards.

Curated: · Written: · Reviewed:

QA-18Change fail rate is 4% because only Sev-1s count. Hotfixes the same afternoon are "operational". What do you count?(show answer)

The useful question for change fail rate as a pipeline metric is what a developer is still able to merge if the control is skipped.

A failed change is a production deployment that required remediation causally tied to that deployment; a same-day hotfix is a candidate, not automatic membership.

Concretely, link each production digest to incidents, rollbacks, and hotfixes. Include an item in CFR only when the remediation addresses a defect introduced by that digest (revert, rollback, or forward-fix of that change). Review the definition quarterly so renaming Sev-1 cannot hide it. Record the digest, the gate that ran, and the rollback command for change fail rate as a pipeline metric on the same change, and store that record beside the promote job so an incident later can name the bits that shipped.

The reason for that specificity is a failure I have seen: CFR looked like 4% while 19% of digests had a same-day hotfix, several of which were unrelated Friday-traffic scaling; leadership could not tell signal from coincidence.

90 production promotes.

DefinitionFailedCFR
Sev-1 only44.4%
every same-day hotfix (over-count)1718.9%
causally attributed rollback/fix910.0%

I would not consider it settled without evidence: A 90-day table of production digests with rollback/hotfix flags, and a CFR that moves when those flags move.

If only Sev-1 counts, you will have fewer Sev-1 names; if every hotfix counts, you will over-accuse. That record is part of the change, not a wiki page written afterwards.

Curated: · Written: · Reviewed:

QA-19Legal requires a human on every production release. The team wants "CD". Which CD do they mean?(show answer)

I would settle continuous delivery versus continuous deployment against a reversible promote, so a bad change has a measured way back.

Continuous delivery keeps every main commit a releasable digest waiting on a business trigger; continuous deployment lets every green digest take production without that trigger.

Concretely, build, test, and stage automatically. Production is a protected promote of the same digest. Do not weaken tests to simulate deployment if the constraint is the human trigger. Record the digest, the gate that ran, and the rollback command for continuous delivery versus continuous deployment on the same change, and store that record beside the promote job so an incident later can name the bits that shipped.

The reason for that specificity is a failure I have seen: They enabled auto-deploy to production to "be CD" and spent a quarter adding after-the-fact CAB screenshots; two bad digests reached users before the screenshot.

Two CD shapes.

ShapeStagingProduction
deliveryauto on digestapproved promote
deploymentautoauto on green
fake CDweekly rebuildCAB after deploy

I would not consider it settled without evidence: Main is always releasable (staging green on the digest) and production history shows an explicit approve per digest, not a rebuild.

A human on the last mile is still CD if the artifact was already proven; auto-prod is a different product. That record is part of the change, not a wiki page written afterwards.

Curated: · Written: · Reviewed:

QA-20Helm returned 0 and the pipeline went green. Checkout was 500 for 8 minutes because the new config missed a key. Where does the smoke live?(show answer)

The judgement in smoke test after promote is which check is required to ship, not which check looks impressive on a pull request.

A promote is not done when the orchestrator exits 0; it is done when a post-deploy smoke against the live version passes, or the job rolls back.

Concretely, after rollout status is complete, hit version-aware smoke (readiness of the new ReplicaSet plus one synthetic checkout). Non-zero smoke triggers the documented rollback of that digest. Record the digest, the gate that ran, and the rollback command for smoke test after promote on the same change, and store that record beside the promote job so an incident later can name the bits that shipped.

The reason for that specificity is a failure I have seen: helm --wait saw pods Ready on a handler that returned 200 to /healthz and 500 to /checkout; users saw 8 minutes of errors while CI was already on the next PR.

Promote job sequence.

1 helm upgrade --wait --timeout 5m
2 kubectl rollout status deploy/checkout
3 curl -f /checkout/version == $DIGEST
4 curl -f POST /checkout/smoke (synthetic)
fail => helm rollback $PREV  && exit 1

I would not consider it settled without evidence: A deliberate bad config in staging: helm succeeds, smoke fails, rollback runs, and the pipeline is red.

Orchestrator success is not user success; smoke is the last job, not a dashboard. That record is part of the change, not a wiki page written afterwards.

Curated: · Written: · Reviewed:

QA-21Blue-green flipped the load balancer immediately. In-flight POSTs died on blue. What is the cutover sequence?(show answer)

Where teams lose a week on blue-green cutover with drain is usually a bypass that was left in for one hotfix and never removed.

Cutover stops new traffic to blue, drains in-flight work, then terminates blue; flipping first drops connections.

Concretely, start green and smoke it, deregister blue from new traffic, drain existing blue connections until the load balancer reports zero active (idle timeout plus max request), then terminate blue. Record the digest, the gate that ran, and the rollback command for blue-green cutover with drain on the same change, and store that record beside the promote job so an incident later can name the bits that shipped.

The reason for that specificity is a failure I have seen: An instant target-group swap dropped 1,104 in-flight checkouts (p99 8.2s). The pipeline was green because both colours had been "healthy".

Cutover timing.

SteptBlue activeGreen
smoke green02,4000 traffic
start drain (no new to blue)10s2,400 → 0100% new
flip complete40s02,410
instant swap (bad)0RST 1,1045xx

I would not consider it settled without evidence: LB metrics: new traffic leaves blue at deregister, active connections then reach 0, then terminate; a 200 rps cutover shows zero 5xx from RST.

Two healthy colours are not a cutover; drained blue with no new traffic is. That record is part of the change, not a wiki page written afterwards.

Curated: · Written: · Reviewed:

QA-22Canary is 5% for 20 minutes, then 100%, with no abort. Error rate on the canary was 9%. What was missing?(show answer)

I would answer canary traffic split with abort by separating what the pipeline proved from what production has not yet seen.

A canary without an abort condition is a delayed 100% rollout; analysis must halt promotion when the canary is worse than baseline.

Concretely, compare canary vs baseline on the same SLI (for example success ratio and p99). Abort if canary success drops by more than the agreed delta (for example 0.5 points) over the window. Rollback to the previous digest automatically. Record the digest, the gate that ran, and the rollback command for canary traffic split with abort on the same change, and store that record beside the promote job so an incident later can name the bits that shipped.

The reason for that specificity is a failure I have seen: 5% ran 20 minutes at 9% errors (baseline 0.3%); then 100% copied the error to everyone for 11 more minutes until a human stopped it.

Abort math, 20-minute window.

SliceSuccessp99Action
baseline99.70%180 ms—
canary 5%91.00%240 msabort, delta 8.7 pts
policyhalt if success −0.5 pts3 min minauto rollback

I would not consider it settled without evidence: A staging drill where canary is forced to 8% errors: abort at minute 3, production pointer unchanged, alert fired.

If nothing can stop the 100%, the 5% was theatre. That record is part of the change, not a wiki page written afterwards.

Curated: · Written: · Reviewed:

QA-23The dangerous code is behind a flag, so the team ships the binary on Friday and flips Monday. The flag default was on. What was the release?(show answer)

The engineering content of feature flags as a separate release from the binary is the lead time and the change-fail cost, not the number of stages.

A flag is a second release with its own default, audience, and rollback; the binary ship is not safe if the default exposes the change.

Concretely, default new flags off in production. Change the flag through the same change process as a promote (audit, owner, soak). Cleaning the flag is a dated ticket, not a hope. Record the digest, the gate that ran, and the rollback command for feature flags as a separate release from the binary on the same change, and store that record beside the promote job so an incident later can name the bits that shipped.

The reason for that specificity is a failure I have seen: NEW_TAX_ENGINE defaulted true in the Helm values that shipped Friday; Monday's "flip" was a no-op and Saturday's orders used the new engine at 100%.

Two releases.

EventBinaryFlag defaultUsers
Friday ship (bad)4.12on100%
Friday ship (good)4.12off0%
Monday flip4.12on for 5%5%

I would not consider it settled without evidence: A checklist: default off in prod values, flag change is a separate audited event, and flags older than 30 days after full rollout are deleted in CI.

A flag whose default is the new behaviour is just a deploy with extra YAML. That record is part of the change, not a wiki page written afterwards.

Curated: · Written: · Reviewed:

QA-24Deploy 4.12 ran a column drop, then 4.12 misbehaved. Helm rollback restored 4.11. Why is the app still dead?(show answer)

Before adding another job I would write what a good delivery of rollback of code when a migration already ran looks like in DORA terms.

You cannot roll back a binary past a destructive migration; expand/contract and forward-fix are the delivery design, not hoping helm rollback undoes SQL.

Concretely, migrations are expand (additive) in one promote, code that uses both shapes, then contract (drop) only after the old binary is gone. Rollback targets the previous compatible pair. Record the digest, the gate that ran, and the rollback command for rollback of code when a migration already ran on the same change, and store that record beside the promote job so an incident later can name the bits that shipped.

The reason for that specificity is a failure I have seen: 4.12 dropped orders.legacy_tax; rollback to 4.11 selected a missing column; checkout stayed 500 for 51 minutes until a forward migration.

Compatible pairs.

StepSchemaBinaryRollback of binary
expandadd new_tax4.11ok
dual-writeboth4.12ok to 4.11
contractdrop legacy4.13not to 4.11

I would not consider it settled without evidence: A staging drill: expand, deploy new, rollback binary, still green; a separate drill that contract-then-rollback is red and forbidden in prod scripts.

Helm does not restore columns. If the migration is irreversible, the rollback plan cannot include it. That record is part of the change, not a wiki page written afterwards.

Curated: · Written: · Reviewed:

QA-25The migration job is in the same Helm chart as the app and runs on every promote. How do you stop a drop from riding with a new feature?(show answer)

The first thing I would pin down about expand/contract schema in the pipeline is which merge or promote it is allowed to block.

Schema changes are their own artifacts and promotes, ordered expand → dual-running → contract, not a side effect of the app chart.

Concretely, a migrate job image is versioned and promoted separately. CI fails if a contract migration and an app binary that still reads the old column are in the same change set. Record the digest, the gate that ran, and the rollback command for expand/contract schema in the pipeline on the same change, and store that record beside the promote job so an incident later can name the bits that shipped.

The reason for that specificity is a failure I have seen: One chart applied DROP COLUMN and the new frontend in one revision; a partial rollout mixed 4.11 pods with a contracted schema.

Split artifacts.

PromoteArtifactAllowed with
1migrate:expand-19app 4.11
2app 4.12expand-19 present
3migrate:contract-19no 4.11 in cluster

I would not consider it settled without evidence: Policy as code on the migration directory: contract files cannot share a PR with app code that references the dropped column, proven by a fixture PR that is red.

If SQL and pods share a revision, mixed versions are a certainty. That record is part of the change, not a wiki page written afterwards.

Curated: · Written: · Reviewed:

QA-26helm upgrade exits 0 in 8 seconds because --wait was omitted. Pods crash 40 seconds later. What does the job owe?(show answer)

I would start helm wait as a deploy gate from the artifact that actually ships, not from the job that happened to go green.

The deploy job waits until the new revision is ready or the timeout fires, and a failed wait rolls back; --wait alone leaves a half-applied release.

Concretely, use helm upgrade --wait --timeout --atomic (or kubectl rollout status plus an explicit rollback). Size the timeout to the worst Ready time plus slack. --wait-for-jobs if the release includes Jobs that must finish. Record the digest, the gate that ran, and the rollback command for helm wait as a deploy gate on the same change, and store that record beside the promote job so an incident later can name the bits that shipped.

The reason for that specificity is a failure I have seen: A ConfigMap typo shipped in 8 seconds of green CI; the new ReplicaSet never became Ready and the old pods were scaled down, leaving 0/3 for 6 minutes because --wait was omitted and nothing rolled back.

Job duration versus cluster state.

helm upgrade --wait --timeout 5m --atomic checkout charts/checkout
# --wait without --atomic: timeout fails, cluster may keep the bad revision
# no --wait: job 8s green, 0/3 Ready

I would not consider it settled without evidence: A staging apply with a broken probe: the job is red, revision is rolled back automatically, and wall time ≥ probe failure, not 8 seconds.

Submitted is not ready. The timeout is the gate. That record is part of the change, not a wiki page written afterwards.

Curated: · Written: · Reviewed:

QA-27Deployment sets maxUnavailable: 50% on a 4-pod checkout to "go faster". What did you trade?(show answer)

This is a place where a passing pipeline and a safe production change for maxUnavailable during a rolling update are not the same event.

maxUnavailable is capacity you have accepted to lose during the rollout; size it from the SLO, not from impatience.

Concretely, for a 4-pod service that needs 3 to hold p99, maxUnavailable is 1 (25%) and maxSurge is 1. Faster rollouts use surge, not double-unavailable. Record the digest, the gate that ran, and the rollback command for maxUnavailable during a rolling update on the same change, and store that record beside the promote job so an incident later can name the bits that shipped.

The reason for that specificity is a failure I have seen: 50% unavailable on 4 pods left 2 running; p99 went from 180 ms to 1.9s for 4 minutes and the canary analysis blamed the new binary.

4 replicas, need 3 to hold p99.

maxUnavailableRunning during rollp99
50% (2)21.9 s
25% (1)3210 ms
0 + maxSurge 14190 ms

I would not consider it settled without evidence: A load test during rollout: in-flight success stays inside the SLO at the chosen maxUnavailable, and a 50% fixture fails that test.

A fast rollout that burns the SLO is just an outage with a progress bar. That record is part of the change, not a wiki page written afterwards.

Curated: · Written: · Reviewed:

QA-28Readiness is tcpSocket:8080. The new version accepts TCP before loading config and serves 500 for 12 seconds. Why did CI still pass?(show answer)

My answer to readiness versus pipeline success begins at the promotion boundary, because that is where delivery either has a gate or has a story.

Readiness must mean the process can serve the user journey, or the pipeline will cut traffic to a listener that is not yet an application.

Concretely, probe an endpoint that exercises config load and a cheap dependency (for example GET /ready that checks the tax table). initialDelay and timeout must cover worst-case boot. Record the digest, the gate that ran, and the rollback command for readiness versus pipeline success on the same change, and store that record beside the promote job so an incident later can name the bits that shipped.

The reason for that specificity is a failure I have seen: tcpSocket flipped Ready in 0.4s; kube-proxy sent checkout to a process still reading a 40 MB config; 12 seconds of 500s per pod during a 6-pod roll.

Probe versus first good request.

ProbeReady atFirst /checkout 200
tcp 80800.4 s12.4 s
HTTP /healthz (liveness)0.5 s12.4 s
HTTP /ready (config+db ping)12.1 s12.2 s

I would not consider it settled without evidence: Boot a pod with a delayed config: Ready stays false until /ready is 200, and a tcp-only probe fixture is rejected in the chart tests.

A port that is open is not a service that is ready. That record is part of the change, not a wiki page written afterwards.

Curated: · Written: · Reviewed:

QA-29Dockerfile says FROM node:20-alpine. Monday's CI and Thursday's CI produced different musl/ICU layers. What do you pin?(show answer)

I would treat image FROM digest as a control in the path to production rather than as a dashboard tile after the fact.

Pin base images by digest in the Dockerfile that produces the release artifact, and rebuild on a controlled cadence with a PR.

Concretely, fROM node:20-alpine@sha256:…. A dependabot/renovate job opens a PR when the tag moves; that PR is the rebuild, not an accidental pull. Record the digest, the gate that ran, and the rollback command for image FROM digest on the same change, and store that record beside the promote job so an incident later can name the bits that shipped.

The reason for that specificity is a failure I have seen: Thursday's rebuild pulled a base with a broken ICU package; production (rebuilt, not promoted) failed locale parsing for 3.1% of EU checkouts.

Base pin.

FROM node:20-alpine@sha256:9c0e1e2d3c4b5a6978877665544332211
# renovate: pin update via PR, not on every build

I would not consider it settled without evidence: Two CI runs 48 hours apart on the same Dockerfile digest produce the same base layer id, unless a pin-update PR changed it.

latest-of-a-tag is a supply chain with a weekly personality. Alpine images change musl and apk, not glibc. That record is part of the change, not a wiki page written afterwards.

Curated: · Written: · Reviewed:

QA-30A multi-stage build copied .env into the builder and the final image still has it in layer 4. How do you prove it is gone?(show answer)

The useful question for secrets leftover in image layers is what a developer is still able to merge if the control is skipped.

Secrets must never enter a layer that ships; multi-stage copies only the runtime bits, and CI scans history and the filesystem.

Concretely, build with BuildKit secrets mounts (--secret) or a builder stage that is not copied. CI runs docker history and a trufflehog/gitleaks scan on the image filesystem. Record the digest, the gate that ran, and the rollback command for secrets leftover in image layers on the same change, and store that record beside the promote job so an incident later can name the bits that shipped.

The reason for that specificity is a failure I have seen: COPY . . in the builder included .env; a later COPY --from=builder /app still contained .env because /app was the whole tree; the digest was public on the registry.

Layer audit.

InstructionFinal image
COPY . . then COPY --from=builder /app.env present
RUN --mount=type=secret,id=npm npm cinot in layers
COPY --from=builder /app/distdist only

I would not consider it settled without evidence: A fixture that COPY .env fails the image scan, and docker history on the release digest shows no secret-sized layer in the final image.

Deleted in the Dockerfile is not gone from the layer that copied it. That record is part of the change, not a wiki page written afterwards.

Curated: · Written: · Reviewed:

QA-31The chart runs as root because "the process binds 80". What do you change in the pipeline?(show answer)

I would settle non-root runtime images against a reversible promote, so a bad change has a measured way back.

The shipped image and the admission policy both require a non-root user; privilege is not a convenience for a port number.

Concretely, uSER in the Dockerfile, runAsNonRoot in the pod spec, and a Gatekeeper/Kyverno check in CI and at admit. Bind 8080 and let the Service/mesh map 80. Record the digest, the gate that ran, and the rollback command for non-root runtime images on the same change, and store that record beside the promote job so an incident later can name the bits that shipped.

The reason for that specificity is a failure I have seen: A CVE in the app became host compromise on a misconfigured node because the container was uid 0; CI had only linted YAML indentation.

Runtime identity.

securityContext:
  runAsNonRoot: true
  runAsUser: 65532
  allowPrivilegeEscalation: false
# Service: 80 -> 8080, not CAP_NET_BIND_SERVICE

I would not consider it settled without evidence: kubectl exec id on the running pod is non-zero, and a PR that removes runAsNonRoot is red in policy CI.

Root in production is a pipeline that did not care who the process is. That record is part of the change, not a wiki page written afterwards.

Curated: · Written: · Reviewed:

QA-32On-call asked for a debug image "just for this incident". It stayed the production tag for 9 days. What is the rule?(show answer)

The judgement in distroless versus debug tags in production is which check is required to ship, not which check looks impressive on a pull request.

Production runs the distroless (or equivalent) digest; debug images are a separate time-boxed promote with an expiry, not a tag overwrite.

Concretely, two artifacts: app and app-debug. Break-glass deploys app-debug with a 4-hour TTL and a pipeline that restores app. Image policy denies *:debug in the prod overlay. Record the digest, the gate that ran, and the rollback command for distroless versus debug tags in production on the same change, and store that record beside the promote job so an incident later can name the bits that shipped.

The reason for that specificity is a failure I have seen: checkout:debug included a shell and curl; it remained after the incident and was the image that later leaked a token from a copied heap dump.

TTL of debug.

ImageProd overlayTTL
checkout@sha256:aayesnone
checkout:debugdenied—
break-glass debugtemporary overlay4 h then restore

I would not consider it settled without evidence: Policy test: prod kustomize overlay mentioning :debug is red; a break-glass record shows restore within TTL.

A debug tag that can stay is the production image. That record is part of the change, not a wiki page written afterwards.

Curated: · Written: · Reviewed:

QA-33GC deleted untagged images older than 7 days. Monday's rollback needed Friday's digest. What do you retain?(show answer)

Where teams lose a week on registry retention of the promoted digest is usually a bypass that was left in for one hotfix and never removed.

Every digest that was production in the retention window (and its rollback predecessor) is excluded from GC, by tag or by a retain list.

Concretely, tag promoted digests prod-YYYYMMDD or put them on an immutable retain list the GC reads. Keep N previous production digests (for example 20) regardless of age. Record the digest, the gate that ran, and the rollback command for registry retention of the promoted digest on the same change, and store that record beside the promote job so an incident later can name the bits that shipped.

The reason for that specificity is a failure I have seen: Friday's sha256 was untagged after a Saturday retag of :prod; Monday rollback pulled 404 and the team rebuilt from git, which was a different binary.

Retain set.

DigestLast prodGC
aa11 (Fri)Frikeep, tagged prod-20260904
bb22 (Sat retag)Satkeep
untagged old featurenever prodcollect after 7d

I would not consider it settled without evidence: A 30-day list of production digests all docker pull successfully, and a GC dry-run that does not list them.

If you cannot pull last week's production, you do not have a rollback; you have a scavenger hunt. That record is part of the change, not a wiki page written afterwards.

Curated: · Written: · Reviewed:

QA-34The SBOM job ran on the source repo and filed a CSV. Production was compromised via a base-image CVE that was not in the CSV. What should the SBOM describe?(show answer)

I would answer SBOM attached to the image that shipped by separating what the pipeline proved from what production has not yet seen.

The SBOM is generated from the image digest that is promoted, not from the lockfile of a different build, and it is stored with that digest.

Concretely, syft (or equivalent) on the release image, attach or attest to the digest, and gate on that attestation existing before promote. Record the digest, the gate that ran, and the rollback command for SBOM attached to the image that shipped on the same change, and store that record beside the promote job so an incident later can name the bits that shipped.

The reason for that specificity is a failure I have seen: The source SBOM listed app npm packages and omitted the Debian openssl in the base; the CVE scanner on the image found it 6 hours after production was live.

What the SBOM covered.

Sourcenpm pkgsbase opensslAttached to digest
repo lockfileyesnono
image syftyesyesyes
CVE found—CVE-2026-…blocked if gated

I would not consider it settled without evidence: For the production digest, an attestation exists and includes the base OS packages; a promote without the attestation is red.

An SBOM of the repo is not an SBOM of the bits that ran. That record is part of the change, not a wiki page written afterwards.

Curated: · Written: · Reviewed:

QA-35CVE-2024-1234 is ignored in .trivyignore with no date. It is still ignored 14 months later. What is the control?(show answer)

The engineering content of vulnerability gate with ignore expiry is the lead time and the change-fail cost, not the number of stages.

Every ignore is a risk acceptance with an owner and an expiry; a permanent ignore is a gate you turned off.

Concretely, ignore entries require an until date no more than 90 days out plus an owner, CI fails expired ignores, and critical CVEs cannot be ignored without a break-glass record that names residual risk and the compensating control in the same ticket. Record the digest, the gate that ran, and the rollback command for vulnerability gate with ignore expiry on the same change, and store that record beside the promote job so an incident later can name the bits that shipped.

The reason for that specificity is a failure I have seen: A 14-month ignore covered a CVE that became remotely exploitable; scanners stayed green while production was listed in a CISA advisory.

Ignore file policy.

CVEuntilOwnerCI
CVE-2024-1234(none)—fail lint
CVE-2024-12342026-09-30paymentspass until then
critical, no glassany—fail

I would not consider it settled without evidence: A linter on the ignore file for dates and owners, and a fixture where yesterday's date fails the build.

An ignore without a clock is a policy of never. That record is part of the change, not a wiki page written afterwards.

Curated: · Written: · Reviewed:

QA-36uses: actions/checkout@v4. A compromised tag moved v4. What do you pin?(show answer)

Before adding another job I would write what a good delivery of pinning GitHub Actions by commit SHA looks like in DORA terms.

Reusable actions are pinned to a full commit SHA; tags are mutable pointers and are not a supply-chain control.

Concretely, uses: actions/checkout@<40-char-sha> with the version in a comment. Renovate opens PRs to move the SHA. Org rules block unpinned uses. Record the digest, the gate that ran, and the rollback command for pinning GitHub Actions by commit SHA on the same change, and store that record beside the promote job so an incident later can name the bits that shipped.

The reason for that specificity is a failure I have seen: v4 was retagged to a commit that exfiltrated secrets: GITHUB_TOKEN walked private repos for 7 hours until the pin was restored.

Pin form.

- uses: actions/checkout@9a87ee2b4d8c3e1f0a1b2c3d4e5f678901234567 # v4.2.2
# forbidden: actions/checkout@v4 and actions/checkout@main

I would not consider it settled without evidence: A CI grep that fails on any third-party uses: whose ref is not a 40-character lowercase hex SHA, including @v4, @main, and @master.

A moving tag is someone else's release, not yours. That record is part of the change, not a wiki page written afterwards.

Curated: · Written: · Reviewed:

QA-3740 repos call acme/platform/.github/workflows/deploy.yml@main. Platform main broke production in all 40. How should they consume it?(show answer)

The first thing I would pin down about reusable workflow pin and review is which merge or promote it is allowed to block.

Reusable workflows are consumed at a tag or SHA, and a platform change is a versioned PR in each product or a staged rollout of the tag, not @main.

Concretely, pin @v3.4.1 or a SHA, have platform publish a tag after its own CI, and let products upgrade by PR with a canary product moving first. Record the digest, the gate that ran, and the rollback command for reusable workflow pin and review on the same change, and store that record beside the promote job so an incident later can name the bits that shipped.

The reason for that specificity is a failure I have seen: A platform commit on Friday afternoon changed helm defaults; 40 services rolled with maxUnavailable 50% before anyone read the diff.

Consumption.

RefProducts on itBlast
@main4040
@v3.4.13939, chosen
@v3.5.0 canary11

I would not consider it settled without evidence: No product workflow references @main on a platform repo, proven by a search job, and a canary product is one tag ahead of the fleet.

If every product tracks platform main, you have one blast radius the size of the company. That record is part of the change, not a wiki page written afterwards.

Curated: · Written: · Reviewed:

QA-38A workflow uses pull_request_target to comment on PRs and checks out the PR head. Why is that a production incident waiting to happen?(show answer)

I would start untrusted pull_request_target from the artifact that actually ships, not from the job that happened to go green.

pull_request_target runs with the base's secrets; checking out and running untrusted PR code is remote execution in CI.

Concretely, use pull_request for untrusted code. pull_request_target may only use the base ref and must not execute PR scripts. Comments can use a different workflow with a read-only token. Record the digest, the gate that ran, and the rollback command for untrusted pull_request_target on the same change, and store that record beside the promote job so an incident later can name the bits that shipped.

The reason for that specificity is a failure I have seen: A fork PR replaced a helper script that the pull_request_target job ran; it printed org secrets to the log in 40 seconds.

Event choice.

EventSecretsCheckout PR head
pull_requestnone / fork-safeyes, isolated
pull_request_targetbase secretsnever run it
incidentbase secrets + PR script40 s leak

I would not consider it settled without evidence: A policy check: pull_request_target workflows contain no checkout of github.event.pull_request.head and no npm/pip install from the PR.

The adjective is target, not trusted. The PR is still the attacker. That record is part of the change, not a wiki page written afterwards.

Curated: · Written: · Reviewed:

QA-39The test job needs a staging API key. Fork PRs from the community cannot run without it, the team says. A maintainer typed /ok-to-test. What is still wrong?(show answer)

This is a place where a passing pipeline and a safe production change for fork PRs must not receive secrets are not the same event.

Fork PRs run without staging or production secrets. Replaying unreviewed fork code with those secrets, even from a throwaway ref, is still an exfiltration path.

Concretely, secrets stay environment-scoped to same-repo branches. Forks get unit tests and public fixtures. Privileged tests run only after the patch is reviewed and merged to an internal branch, using one-use credentials that cannot reach production data. Record the digest, the gate that ran, and the rollback command for fork PRs must not receive secrets on the same change, and store that record beside the promote job so an incident later can name the bits that shipped.

The reason for that specificity is a failure I have seen: A fork PR received STAGING_API_KEY after /ok-to-test and dumped it; staging shared the production database credential via a "temporary" tunnel.

Secret matrix.

PR sourceunitstaging keyprod key
same-repo branchyesyesno
forkyesnono
maintainer replay of unreviewed forkyesyesstill a leak

I would not consider it settled without evidence: A fork PR log shows masked/absent staging secrets, and a maintainer replay job has an audit trail.

If a stranger's PR can call staging, staging is on the internet with a key in CI. That record is part of the change, not a wiki page written afterwards.

Curated: · Written: · Reviewed:

QA-40A self-hosted runner has been registered since 2024 and reuses a dirty workspace. A job wrote /tmp/keys. What is the runner model?(show answer)

My answer to self-hosted runner as ephemeral VM begins at the promotion boundary, because that is where delivery either has a gate or has a story.

Self-hosted runners that see secrets are single-job ephemeral VMs (or equivalent), not pets with a workspace from last week.

Concretely, jIT registration, one job, destroy. No docker.sock from the host. Cache is remote and authenticated, not leftover /home/runner. Record the digest, the gate that ran, and the rollback command for self-hosted runner as ephemeral VM on the same change, and store that record beside the promote job so an incident later can name the bits that shipped.

The reason for that specificity is a failure I have seen: Job A wrote a GitHub PAT to /tmp; job B (a fork PR) read it 11 minutes later because the runner was persistent.

Runner lifetime.

ModelWorkspacedocker.sockFork jobs
pet since 2024dirtyyessame box
ephemeral VMemptynonew VM
leakPAT in /tmp—job B

I would not consider it settled without evidence: Each job's metadata shows a unique instance id, and a test file written in job A is absent in job B on the "same" runner label.

A runner that remembers is a shared workstation with every secret that ever passed through it. That record is part of the change, not a wiki page written afterwards.

Curated: · Written: · Reviewed:

QA-41The build job mounts /var/run/docker.sock because Kaniko "was slow". What did you grant the job?(show answer)

I would treat Docker-in-docker privilege as a control in the path to production rather than as a dashboard tile after the fact.

Mounting the host docker.sock is host root for the duration of the job; image builds use a rootless builder that cannot talk to the host daemon.

Concretely, kaniko, Buildah, or BuildKit in rootless/isolated mode. No sock mount. If a daemon is required, it is a sidecar with no host mount. Record the digest, the gate that ran, and the rollback command for Docker-in-docker privilege on the same change, and store that record beside the promote job so an incident later can name the bits that shipped.

The reason for that specificity is a failure I have seen: A compromised npm preinstall used docker.sock to start a privileged container and escaped to the runner host, which still had cloud keys in instance metadata.

Build isolation.

MethodHost docker.sockPrivilege
volume mount sockyeshost root
Kaniko in podnocontainer
BuildKit rootlessnouser ns

I would not consider it settled without evidence: Pod spec of the build job has no hostPath docker.sock, enforced by admission and by a unit test of the job YAML.

The socket is not a cache optimisation; it is the host. That record is part of the change, not a wiki page written afterwards.

Curated: · Written: · Reviewed:

QA-42Kaniko needs a registry write credential. The team put docker config.json in a cluster Secret used by all namespaces. What is wrong?(show answer)

The useful question for Kaniko or buildah without docker.sock is what a developer is still able to merge if the control is skipped.

The builder's push credential is scoped to the image repository it is allowed to write, in the namespace that builds, not a cluster-wide docker config.

Concretely, a per-repo robot account or OIDC to the registry, mounted only in the build namespace. Prod pull credentials cannot push. Record the digest, the gate that ran, and the rollback command for Kaniko or buildah without docker.sock on the same change, and store that record beside the promote job so an incident later can name the bits that shipped.

The reason for that specificity is a failure I have seen: The shared config.json could push to prod/checkout; a PR in a sandbox namespace retagged prod/checkout:stable.

Registry identities.

IdentityPull prodPush prodPush sandbox
cluster-wide docker jsonyesyesyes
sandbox robotnonoyes
prod-build OIDCyesyesno

I would not consider it settled without evidence: Registry audit: push to prod/checkout only from the prod-build identity, and a sandbox job that tries to push is 403.

A builder that can tag production is a production deploy. That record is part of the change, not a wiki page written afterwards.

Curated: · Written: · Reviewed:

QA-43gitleaks runs on a nightly cron. A key was pushed at 10:01 and used at 10:07. What is the scan's place in the path?(show answer)

I would settle secret scanning on every push against a reversible promote, so a bad change has a measured way back.

Secret scanning is a pre-receive or required PR/push check, not a nightly report, because the leak is already live by morning.

Concretely, server-side push protection plus a required gitleaks/trufflehog job on every PR and every push to main. Confirmed leaks revoke automatically where the API allows. Record the digest, the gate that ran, and the rollback command for secret scanning on every push on the same change, and store that record beside the promote job so an incident later can name the bits that shipped.

The reason for that specificity is a failure I have seen: A Google API key sat in a commit from 10:01; nightly scan opened a ticket at 02:14; the key had billed $18,400 of GPU time.

Time to detect.

ControlDetectRevoke
nightly cron16 hticket
required PR job90 smerge blocked
push protection0 (rejected)never landed

I would not consider it settled without evidence: A PR that adds a well-known test AWS key is red in <2 minutes, and push protection rejects it on main.

A nightly secret scan is an invoice, not a control. That record is part of the change, not a wiki page written afterwards.

Curated: · Written: · Reviewed:

QA-44The job runs npm install because npm ci "fails too often". What did you stop verifying?(show answer)

The judgement in lockfile integrity with npm ci is which check is required to ship, not which check looks impressive on a pull request.

The lockfile is the bill of materials; npm ci (or equivalent --frozen-lockfile) fails when it does not match, which is the point.

Concretely, cI uses the frozen install. Lockfile changes are reviewed as dependency changes. Enable lockfile-integrity / checksums where the ecosystem supports them. Record the digest, the gate that ran, and the rollback command for lockfile integrity with npm ci on the same change, and store that record beside the promote job so an incident later can name the bits that shipped.

The reason for that specificity is a failure I have seen: npm install resolved a different transitive on Tuesday; a typosquat in a nested dep ran in CI and in the image that shipped.

Install mode.

CommandLockfileTuesday tree
npm installupdated silentlynew transitive
npm cimust matchidentical
ci + checksummust matchidentical

I would not consider it settled without evidence: A PR that changes package.json without the lockfile is red, and two CI runs on the same lockfile produce the same node_modules tree hash.

If install can pick, the lockfile was a suggestion. That record is part of the change, not a wiki page written afterwards.

Curated: · Written: · Reviewed:

QA-45Internal package @acme/checkout-lib is also a name someone published to the public npm. CI resolved the public one. How do you bind the name?(show answer)

Where teams lose a week on dependency confusion in the private registry is usually a bypass that was left in for one hotfix and never removed.

Internal names are scoped and the installer is configured to never fall back to the public registry for that scope; missing internal packages fail, they do not substitute.

Concretely, use a company scope, registry config that maps @acme to the private registry only, and disable unknown-scope fallback. Prefer reserved namespaces on the public registry. Record the digest, the gate that ran, and the rollback command for dependency confusion in the private registry on the same change, and store that record beside the promote job so an incident later can name the bits that shipped.

The reason for that specificity is a failure I have seen: A developer published @acme/checkout-lib@9.9.9 to public npm; CI on a misconfigured runner installed it and exfiltrated the npm token.

Resolve rules.

ScopeRegistryMissing package
@acmenpmjs (bad)public substitute
@acmeartifactory onlyfail
reserved public @acmeemptyfail closed

I would not consider it settled without evidence: A canary public package of the internal name is not installed on a clean CI run, and a missing internal version fails the job.

If the public registry can satisfy an internal name, it will, on the worst day. That record is part of the change, not a wiki page written afterwards.

Curated: · Written: · Reviewed:

QA-46The registry holds an unsigned digest and a signed one. Production pulled the unsigned because the tag moved. What do you verify?(show answer)

I would answer Cosign keyless signing of the promoted digest by separating what the pipeline proved from what production has not yet seen.

Admission and the deploy job verify a signature on the digest being applied, bound to the expected OIDC issuer and identity, not "this repo has some signature".

Concretely, cosign sign the digest in CI with keyless or a stored key. Verify digest plus Fulcio/OIDC issuer and certificate identity (repo and workflow). The cluster's Binary Authorization / Kyverno repeats that verify before the pod runs. Record the digest, the gate that ran, and the rollback command for Cosign keyless signing of the promoted digest on the same change, and store that record beside the promote job so an incident later can name the bits that shipped.

The reason for that specificity is a failure I have seen: A tag was overwritten with an image keyless-signed by a different workflow; nodes that pulled after the overwrite ran it because the policy only checked that a signature existed.

Verify target.

PolicyPasses
repo has some signatureunsigned :stable after overwrite
cosign verify digest in manifestonly aa11
unsigned aa11denied
issuer/identity mismatch on aa11denied

I would not consider it settled without evidence: A staging unsigned digest is rejected, a signed digest that is not the manifest digest is rejected, and the correct digest signed by a different workflow identity is rejected.

A signature without the expected identity is not your builder. That record is part of the change, not a wiki page written afterwards.

Curated: · Written: · Reviewed:

QA-47A developer docker-pushed checkout:4.12 from a laptop and the cluster ran it. CI also produces 4.12. How do you tell them apart?(show answer)

The engineering content of SLSA provenance for the artifact is the lead time and the change-fail cost, not the number of stages.

Only artifacts with provenance from the trusted builder identity are promotable; a laptop push that reuses a tag is not that builder.

Concretely, generate SLSA provenance (or equivalent in-toto) in CI. Promote verifies builder id, source repo, and digest. Registry permissions deny laptop push to prod repositories. Record the digest, the gate that ran, and the rollback command for SLSA provenance for the artifact on the same change, and store that record beside the promote job so an incident later can name the bits that shipped.

The reason for that specificity is a failure I have seen: A laptop :4.12 overwrote CI's image; production ran an unsigned-from-CI binary that still "was 4.12" in the ticket.

Builder identity.

PushProvenance builderProd pull
CI workflowgithub.com/acme/checkout/.githuballowed
laptop docker pushnone403
tag 4.12 both—digest decides

I would not consider it settled without evidence: Provenance on the production digest names the CI workflow and SHA, and a laptop push to prod/checkout is 403.

If a laptop can write the production name, provenance is a PDF. That record is part of the change, not a wiki page written afterwards.

Curated: · Written: · Reviewed:

QA-48CI verifies cosign, then kubectl apply from a human bypassed GitOps. The unsigned image ran. Where must verify live?(show answer)

Before adding another job I would write what a good delivery of Binary Authorization at admit time looks like in DORA terms.

Signature verification that exists only in CI is skipped by every path that is not CI; admission on the cluster is the control that matches what runs.

Concretely, binary Authorization or Kyverno verify image signatures at admit. GitOps is the usual path; admission still catches the laptop. Record the digest, the gate that ran, and the rollback command for Binary Authorization at admit time on the same change, and store that record beside the promote job so an incident later can name the bits that shipped.

The reason for that specificity is a failure I have seen: An engineer kubectl set image to an unsigned debug tag during an incident; it stayed for 3 days because CI never saw it.

Path versus control.

PathCI cosignAdmit verify
GitOpsyesyes
kubectl set imageskippeddeny unsigned
incident debug tagskippeddeny

I would not consider it settled without evidence: kubectl set image with an unsigned digest is denied in prod, demonstrated in a drill, and the audit log shows the deny.

A gate on the pipeline is not a gate on the API server. That record is part of the change, not a wiki page written afterwards.

Curated: · Written: · Reviewed:

QA-49Git says replicas: 3. The cluster has 12 because someone scaled during an incident. What should the reconciler do, and what should humans do?(show answer)

The first thing I would pin down about GitOps reconcile versus kubectl apply from a laptop is which merge or promote it is allowed to block.

The cluster converges to Git; incident scaling is a Git change or a recorded override with a TTL, not a standing kubectl that Git silently loses or silently overwrites without a process.

Concretely, auto-sync with prune according to policy. Break-glass annotations expire. An incident scale is a PR (or a documented override) so the next reconcile does not surprise anyone. Record the digest, the gate that ran, and the rollback command for GitOps reconcile versus kubectl apply from a laptop on the same change, and store that record beside the promote job so an incident later can name the bits that shipped.

The reason for that specificity is a failure I have seen: Auto-sync reset replicas to 3 in the middle of a load spike because the 12 was only on the cluster; checkout 500ed for 9 minutes.

Source of truth.

WriterReplicasAfter 5 min
kubectl only123 (surprise) or 12 (drift)
Git PR1212
override TTL 2h123 after TTL

I would not consider it settled without evidence: A drill: scale in Git, reconciler matches; a kubectl scale without Git is either reverted on a known cadence or blocked by SSA/policy.

If Git is not the last writer, Git is a wiki. That record is part of the change, not a wiki page written afterwards.

Curated: · Written: · Reviewed:

QA-50dev built sha256:aa, staging rebuilt sha256:bb from the same SHA, production rebuilt sha256:cc. What promotion do you actually have?(show answer)

I would start same digest promoted through environments from the artifact that actually ships, not from the job that happened to go green.

Promotion is moving an already-tested digest along environments; three rebuilds are three releases with a shared git hash.

Concretely, a single build job; env overlays change replicas, URLs, and secrets, not Dockerfile inputs. The promote job is kustomize/helm pointing at the same digest. Record the digest, the gate that ran, and the rollback command for same digest promoted through environments on the same change, and store that record beside the promote job so an incident later can name the bits that shipped.

The reason for that specificity is a failure I have seen: staging's bb had a different npm mirror result than production's cc; tests in staging were not tests of production.

Train of one digest.

EnvGit SHADigest
wrong9f3caa, then bb, then cc
right9f3caa11 in all three
overlay—replicas, secrets, hostnames

I would not consider it settled without evidence: dev, staging, and prod running image IDs are identical for a given release train, shown in a table of three cluster queries.

If the digest changes at the environment border, you did not promote. That record is part of the change, not a wiki page written afterwards.

Curated: · Written: · Reviewed:

QA-51Someone deleted a Service from Git by accident. Auto-sync prune removed it from production in 3 minutes. What safeguard belongs in the app spec?(show answer)

This is a place where a passing pipeline and a safe production change for Argo auto-sync prune without a dry-run are not the same event.

Prune is a deletion controller; high-impact resources need a window, a dry-run/diff notification, or a manual sync, not silent delete.

Concretely, use prune with sync windows or require a human for deletions of Services, Ingress, and PVCs. Notify on diff that includes a delete. Recover from Git history, not from memory. Record the digest, the gate that ran, and the rollback command for Argo auto-sync prune without a dry-run on the same change, and store that record beside the promote job so an incident later can name the bits that shipped.

The reason for that specificity is a failure I have seen: A bad merge removed checkout's Service; prune deleted it; DNS still pointed at the old ClusterIP for 11 minutes while CI was green because the App was Synced.

Prune policy.

ResourceAuto pruneGuard
ConfigMapyesnone
Service / Ingressnomanual or 30 min window
PVCneverpreserve

I would not consider it settled without evidence: A staging App where deleting a Service in Git does not apply for 30 minutes or without approval, and a Slack/email diff lists the delete.

Synced and empty is not healthy; it is a successful deletion. That record is part of the change, not a wiki page written afterwards.

Curated: · Written: · Reviewed:

QA-52Image automation writes checkout:stable into Git every time the tag moves. How do you keep GitOps from becoming a tag follower?(show answer)

My answer to Flux image policy that still writes tags begins at the promotion boundary, because that is where delivery either has a gate or has a story.

Image policy should write a digest (or an immutable tag that includes the digest) into Git, so a moving :stable cannot retag production without a commit that names bits.

Concretely, flux ImagePolicy filter to digest. The automation PR is reviewable. Disable policies that only copy floating tags. Record the digest, the gate that ran, and the rollback command for Flux image policy that still writes tags on the same change, and store that record beside the promote job so an incident later can name the bits that shipped.

The reason for that specificity is a failure I have seen: :stable moved to a broken build; Flux committed it; production rolled before anyone saw the image because the policy was "latest tag".

What automation writes.

PolicyGit fieldCluster after retag
latest tagcheckout:stablenew bits, no review
digestcheckout@sha256:aaunchanged
immutable tagcheckout:4.12.7-aa11unchanged until PR

I would not consider it settled without evidence: The prod overlay contains sha256, and a retag of :stable without a new digest commit does not change the cluster.

If Git stores a tag that moves outside Git, GitOps is just a cron. That record is part of the change, not a wiki page written afterwards.

Curated: · Written: · Reviewed:

QA-53base sets memory: 256Mi. prod overlay used to set 1Gi and the line was dropped in a merge. How do you catch it before promote?(show answer)

I would treat kustomize overlay drift as a control in the path to production rather than as a dashboard tile after the fact.

CI renders every overlay and diffs against a committed golden or against policy ranges, because a dropped overlay line is a silent prod change.

Concretely, kustomize build prod in CI, kubeconform + policy (memory ≥ 1Gi for checkout). Fail if the rendered prod spec regresses below the floor. Record the digest, the gate that ran, and the rollback command for kustomize overlay drift on the same change, and store that record beside the promote job so an incident later can name the bits that shipped.

The reason for that specificity is a failure I have seen: The 1Gi line vanished; prod OOMKilled 18 times in 40 minutes; the PR that dropped it only touched a comment in base.

Rendered prod.

kustomize build overlays/prod > dist/prod.yaml
kubeconform dist/prod.yaml
conftest test dist/prod.yaml  # memory >= 1Gi

I would not consider it settled without evidence: A fixture overlay missing the memory floor is red, and the rendered prod YAML is an artifact of the pipeline.

An overlay you do not render is documentation. That record is part of the change, not a wiki page written afterwards.

Curated: · Written: · Reviewed:

QA-54A pre-upgrade Job runs migrations, and the chart author also put hook weight 0 on the Deployment "so they start together". What actually races?(show answer)

The useful question for Helm hook ordering for jobs is what a developer is still able to merge if the control is skipped.

Helm runs pre-install/pre-upgrade hooks and waits for hook Jobs to complete before it loads ordinary resources; hook weights only order hooks in the same phase, they do not put a Deployment in that race.

Concretely, keep migrate as a pre-upgrade hook with a deletion policy and retries. Test hook failure (release stops before the Deployment changes), not a fictional weight contest. Post-upgrade Jobs that must finish before traffic need --wait-for-jobs or an explicit wait. Record the digest, the gate that ran, and the rollback command for Helm hook ordering for jobs on the same change, and store that record beside the promote job so an incident later can name the bits that shipped.

The reason for that specificity is a failure I have seen: The migrate Job was a normal Job with no hook; 4.12 pods started against an unmigrated schema and crash-looped while the Job still ran.

Hook phase, not weight vs Deployment.

ObjecthookWhen Deployment starts
Job migratepre-upgradeafter Job succeeds
Deploymentnoneafter pre-hooks
Job without hooknonerace with pods

I would not consider it settled without evidence: A kind install where a failing pre-upgrade Job leaves the previous Deployment intact, proven in CI.

If migrate is not a pre-hook, production will pick the race Helm would have prevented. That record is part of the change, not a wiki page written afterwards.

Curated: · Written: · Reviewed:

QA-55Apply runs on every PR to "shift left". A preview applied a prod IAM change. What is the split, and whose plan do you apply?(show answer)

I would settle terraform plan on PR, apply on merge against a reversible promote, so a bad change has a measured way back.

Plan on the PR is speculative; apply happens after merge from a trusted workflow that generates a new plan for the merged revision, protects that artifact, and applies only that file.

Concretely, pR: terraform plan for review, no apply, no prod credentials. Merge to the env branch: trusted job plans the merged SHA, stores the plan with integrity protection, requires approval, then apply. Never apply a plan file produced by untrusted PR code. Record the digest, the gate that ran, and the rollback command for terraform plan on PR, apply on merge on the same change, and store that record beside the promote job so an incident later can name the bits that shipped.

The reason for that specificity is a failure I have seen: PR apply used a mis-selected workspace and attached a public IP to a prod RDS; the PR was closed unmerged and the IP stayed.

When apply runs.

Eventplanapplycreds
pull_requestyesnoread / plan
merge to envnew plan of merged SHAyesenv role
PR apply (bad)—yesprod

I would not consider it settled without evidence: A PR cannot assume the prod apply role (OIDC sub is not the production environment), and the apply job's plan checksum is produced after merge, not uploaded by the PR.

An apply of a PR-built plan is still the PR's code, even after merge. That record is part of the change, not a wiki page written afterwards.

Curated: · Written: · Reviewed:

QA-56On-call ran terraform apply -target=aws_security_group.checkout to "go faster". The next full apply destroyed 14 unrelated objects. Why?(show answer)

The judgement in targeted terraform apply as an anti-pattern is which check is required to ship, not which check looks impressive on a pull request.

Targeted apply updates state for a subset and leaves the rest of the graph unapplied; the next full apply then reconciles surprise.

Concretely, forbid -target in the pipeline except a named break-glass that still plans the whole workspace afterwards. Fix the graph instead of targeting around it. Record the digest, the gate that ran, and the rollback command for targeted terraform apply as an anti-pattern on the same change, and store that record beside the promote job so an incident later can name the bits that shipped.

The reason for that specificity is a failure I have seen: A targeted SG change left a dangling NIC in state; the next apply deleted 14 production interfaces that "should not have been there" in the config.

State after target.

CommandImmediateNext full apply
apply -target SGSG okdestroy 14 NICs
full applywhole graphmatches config
glass + full planSG + planno surprise

I would not consider it settled without evidence: CI grep fails on -target in job YAML, and a post-glass full plan is required green before the incident is closed.

A target is a lie you tell the graph; the next apply tells the truth. That record is part of the change, not a wiki page written afterwards.

Curated: · Written: · Reviewed:

QA-57Apply is stuck on a lock from a cancelled runner. An engineer force-unlocked from their laptop. What is the procedure?(show answer)

Where teams lose a week on state lock and who can break it is usually a bypass that was left in for one hotfix and never removed.

Force-unlock is a break-glass with the same two-person rule as prod apply, because unlocking during a live apply splits state.

Concretely, locks live in DynamoDB/GCS with restricted iam. Unlock is a pipeline job with a reason, not a local CLI. Never unlock while another apply's runner is still alive. Record the digest, the gate that ran, and the rollback command for state lock and who can break it on the same change, and store that record beside the promote job so an incident later can name the bits that shipped.

The reason for that specificity is a failure I have seen: Force-unlock during a still-running apply produced a state file that listed a half-created ALB; the next apply duplicated it and split traffic 50/50 with a dead target.

Unlock paths.

ActorWhile apply liveOutcome
laptop force-unlockyessplit state
pipeline unlock, runner deadnook
wait for timeoutyesapply finishes or fails

I would not consider it settled without evidence: Audit: unlock events are pipeline jobs with two persons, and CloudTrail shows no Unlock from user laptops.

The lock is not a nuisance; it is the mutex for the real world. That record is part of the change, not a wiki page written afterwards.

Curated: · Written: · Reviewed:

QA-58dev, staging, and prod are terraform workspaces in one state bucket. A workspace select mistake applied prod from a staging PR. How do you separate them?(show answer)

I would answer separate state per environment not workspaces by separating what the pipeline proved from what production has not yet seen.

Environments are separate states and separate credentials; workspaces in one backend are a naming convention that a wrong select can cross.

Concretely, one backend key per env, OIDC role per env, and directory or repo boundaries. CI never interpolates workspace from a PR title. Record the digest, the gate that ran, and the rollback command for separate state per environment not workspaces on the same change, and store that record beside the promote job so an incident later can name the bits that shipped.

The reason for that specificity is a failure I have seen: terraform workspace select prod on a staging job, because TF_WORKSPACE was copied from a secret named similarly; the plan was 400 lines of prod destroy.

Backend keys.

Envstate keyrole
stagingacme/staging/terraform.tfstatestaging-apply
prodacme/prod/terraform.tfstateprod-apply
workspace prod in same bucketdefault?too easy

I would not consider it settled without evidence: Staging role cannot s3:Put the prod state key, demonstrated by a denied apply, and there is no workspace select in the jobs.

If one credential can write every environment's state, the workspace name is a comment. That record is part of the change, not a wiki page written afterwards.

Curated: · Written: · Reviewed:

QA-59module "vpc" { source = "git::https://github.com/acme/terraform-vpc.git" } with no ref. Monday and Thursday applied different VPCs. What do you pin?(show answer)

The engineering content of module source pin to a tag or digest is the lead time and the change-fail cost, not the number of stages.

Module sources are versioned (tag, commit, or registry version); a floating git default branch is an unreviewed module upgrade.

Concretely, source = ...?ref=v1.8.2 or a commit SHA. Renovate PRs bump versions. CI fails on unpinned git sources. Record the digest, the gate that ran, and the rollback command for module source pin to a tag or digest on the same change, and store that record beside the promote job so an incident later can name the bits that shipped.

The reason for that specificity is a failure I have seen: The module's main added a destructive peering change; Thursday's apply deleted a peering prod still needed; CI of the app repo never saw the module diff.

Module ref.

module "vpc" {
  source = "git::https://github.com/acme/terraform-vpc.git?ref=v1.8.2"
}
# forbidden: no ref, tracks main

I would not consider it settled without evidence: A grep over *.tf fails unpinned git sources, and a bump of the module is a PR with the module's changelog.

Unpinned modules are someone else's apply inside yours. That record is part of the change, not a wiki page written afterwards.

Curated: · Written: · Reviewed:

QA-60A typo in a for_each dropped a production database from config. Apply wanted to destroy it. prevent_destroy was on the resource. Why is that not enough?(show answer)

Before adding another job I would write what a good delivery of terraform destroy protection looks like in DORA terms.

prevent_destroy only applies while the instance remains in configuration; removing it via a for_each typo takes the lifecycle rule with it, so the cloud API and a plan policy must still refuse the destroy.

Concretely, rDS deletion_protection (or equivalent) stays on in the provider. A plan analyser fails on destroy of aws_db_instance without an override file. Treat prevent_destroy as belt, not the only suspenders. Record the digest, the gate that ran, and the rollback command for terraform destroy protection on the same change, and store that record beside the promote job so an incident later can name the bits that shipped.

The reason for that specificity is a failure I have seen: The for_each typo planned destroy of prod-checkout-db; prevent_destroy was gone with the instance address; apply ran because the only remaining gate was "plan succeeded"; restore took 6 hours from backup.

Destroy gates.

LayerTypo for_each
no gateapply destroys DB
prevent_destroy onlygone with the instance, apply destroys
plan policyCI red before apply
cloud deletion_protectionAPI deny

I would not consider it settled without evidence: A fixture plan that destroys the DB is red in CI, and the cloud API still has deletion_protection true.

prevent_destroy cannot protect an instance that is no longer in the graph. That record is part of the change, not a wiki page written afterwards.

Curated: · Written: · Reviewed:

QA-61ClickOps added a security group rule. Git still describes the old SG. When do you find out?(show answer)

The first thing I would pin down about drift detection job that opens a PR is which merge or promote it is allowed to block.

A scheduled terraform plan (or cloud drift detector) opens a PR or an incident when live ≠ Git; waiting for the next human apply is how drift becomes the source of truth.

Concretely, daily plan against prod state. Non-empty plan without a Git change files a PR or a Sev ticket. Do not auto-apply the drift away without review; it might be an incident mitigation. Record the digest, the gate that ran, and the rollback command for drift detection job that opens a PR on the same change, and store that record beside the promote job so an incident later can name the bits that shipped.

The reason for that specificity is a failure I have seen: A ClickOps 0.0.0.0/0 rule lasted 47 days until a scheduled apply "fixed" it during a launch, which was also when anyone noticed.

Drift loop.

DetectorTime to PRAuto-apply
next human apply47 dsurprise
daily plan≤24 hno, review
auto-apply drift≤24 hcan undo incident

I would not consider it settled without evidence: A synthetic ClickOps change in staging produces a plan PR within the job's interval (for example 24 hours).

If only applies detect drift, ClickOps owns the days between applies. That record is part of the change, not a wiki page written afterwards.

Curated: · Written: · Reviewed:

QA-62kubectl edit on a Deployment is still possible. GitOps will overwrite in 3 minutes. What do you do for a real incident?(show answer)

I would start ClickOps after GitOps from the artifact that actually ships, not from the job that happened to go green.

Emergency cluster edits go through a recorded break-glass that either writes Git immediately or sets a TTL annotation the reconciler honours; raw kubectl edit is denied or reverted on purpose.

Concretely, rBAC: no patch on Deployments for humans except the break-glass role. Break-glass opens a Git PR from the live spec or applies an annotated override. Record the digest, the gate that ran, and the rollback command for ClickOps after GitOps on the same change, and store that record beside the promote job so an incident later can name the bits that shipped.

The reason for that specificity is a failure I have seen: kubectl edit raised replicas; 3 minutes later GitOps set them back in the middle of a sale; this happened three times in one incident.

Edit paths.

ActionGitAfter 3 min
kubectl editnoreverted
break-glass + PRyeskept
annotated TTLoverrideexpired

I would not consider it settled without evidence: A normal engineer kubectl edit is forbidden; a break-glass session is logged and produces a Git commit or an expiry.

An edit that Git will fight is not an incident tool; it is a race. That record is part of the change, not a wiki page written afterwards.

Curated: · Written: · Reviewed:

QA-63prod/.env is gitignored and copied by hand onto the runner. How do secrets enter GitOps?(show answer)

This is a place where a passing pipeline and a safe production change for SOPS or sealed secrets in git are not the same event.

Encrypted secrets live in Git (SOPS, Sealed Secrets, or equivalent); decryption happens in the cluster or the apply identity, not as a file on a laptop.

Concretely, commit *.enc.yaml. CI can plan without plaintext. The controller or SOPS with KMS decrypts in prod. Rotation is a PR. Record the digest, the gate that ran, and the rollback command for SOPS or sealed secrets in git on the same change, and store that record beside the promote job so an incident later can name the bits that shipped.

The reason for that specificity is a failure I have seen: A hand-copied .env on a persistent runner was 11 months old and contained a database password that had been "rotated" in the password manager only.

Where plaintext exists.

MethodGitRunner diskCluster
copied .envnoyesenv
SOPS + KMSciphertextnodecrypt
SealedSecretciphertextnocontroller

I would not consider it settled without evidence: A prod overlay contains only ciphertext, and a PR that adds a plaintext AWS key is red in scanning.

A gitignored secret is an unversioned, unreviewed, unrotated production dependency. That record is part of the change, not a wiki page written afterwards.

Curated: · Written: · Reviewed:

QA-64The image has ENV DATABASE_URL baked at build time per environment. You need to rotate the password. What was the design error?(show answer)

My answer to External Secrets versus baked env files begins at the promotion boundary, because that is where delivery either has a gate or has a story.

Runtime secrets are pulled at start or via a controller; baking them into the image makes rotation a rebuild and leaks them in layers.

Concretely, external Secrets / Vault agent / cloud identity. The image has the name of the secret, not the value. Rotation updates the store and then rolls or reloads the consumers; a Secret patch alone does not restart pods that already have the env var. Record the digest, the gate that ran, and the rollback command for External Secrets versus baked env files on the same change, and store that record beside the promote job so an incident later can name the bits that shipped.

The reason for that specificity is a failure I have seen: Password rotation required 3 image rebuilds (dev/stage/prod) and 40 minutes of pipeline; two environments still had the old password in an old digest that got rolled back.

Rotation work.

DesignRotate passwordNew digest?
baked ENVrebuild × envsyes
ExternalSecretupdate store + rolloutno
rollback old imageold password returns—

I would not consider it settled without evidence: The production image filesystem has no DATABASE_URL value, and a secret update is followed by a rollout whose new pods show the new value.

If rotation rebuilds the app, the secret was a compile-time constant. That record is part of the change, not a wiki page written afterwards.

Curated: · Written: · Reviewed:

QA-65CI uses a Vault token with 30-day TTL because "jobs retry overnight". What should the TTL be?(show answer)

I would treat Vault short-lived credentials in CI as a control in the path to production rather than as a dashboard tile after the fact.

CI tokens last for the job, not for the calendar; overnight retries get a fresh login, not a standing token in GitHub secrets.

Concretely, oIDC or AppRole per job, TTL minutes (for example 15–60), no token stored in repository secrets. Revoke on job end where possible. Record the digest, the gate that ran, and the rollback command for Vault short-lived credentials in CI on the same change, and store that record beside the promote job so an incident later can name the bits that shipped.

The reason for that specificity is a failure I have seen: A 30-day token in GitHub was copied into a fork PR log and had create-database on prod for 26 more days.

TTL.

AuthTTLStored in GitHub
static token30 dyes
OIDC per job20 minno
leaked 30 d26 d leftinvoice

I would not consider it settled without evidence: Vault audit: CI tokens live ≤60 minutes, and GitHub secrets list contains no VAULT_TOKEN.

A month-long CI token is a user account that never logs out. That record is part of the change, not a wiki page written afterwards.

Curated: · Written: · Reviewed:

QA-66set -x printed the AWS_SECRET_ACCESS_KEY because it was passed as an env var to a script. How do you stop that class of leak?(show answer)

The useful question for no secrets in workflow logs is what a developer is still able to merge if the control is skipped.

Secrets are masked, never echoed, and passed via the runner's secret API or a file descriptor, not interpolated into command lines that shells trace.

Concretely, register secrets with the CI masker. Disable debug on jobs that have secrets. Prefer OIDC so there is no long secret to print. Scan logs on every job. Record the digest, the gate that ran, and the rollback command for no secrets in workflow logs on the same change, and store that record beside the promote job so an incident later can name the bits that shipped.

The reason for that specificity is a failure I have seen: A debug rerun with ACTIONS_STEP_DEBUG printed the registry password; it was in the public log for 14 minutes before deletion, which does not delete forks' copies.

Leak paths.

PatternIn log
set -x + env secretyes
masker + no debug***
OIDC, no static secretnothing to print

I would not consider it settled without evidence: A job that echo $SECRET is masked in the log, and debug is forbidden on production environments.

Log deletion is not containment; the leak already left the building. That record is part of the change, not a wiki page written afterwards.

Curated: · Written: · Reviewed:

QA-67A Helm values change to a timeout shipped with three feature commits. The timeout was the outage. How do you separate them?(show answer)

I would settle config change as its own promote against a reversible promote, so a bad change has a measured way back.

Production config is a first-class promote with its own digest/revision and rollback; bundling it with features makes the timeout un-rollbackable without the features.

Concretely, config-only PRs, small values diffs, and a pipeline that can roll back the chart values without rolling back the app digest when they are versioned separately — or accept they share a revision and keep config PRs tiny. Record the digest, the gate that ran, and the rollback command for config change as its own promote on the same change, and store that record beside the promote job so an incident later can name the bits that shipped.

The reason for that specificity is a failure I have seen: One revision changed timeout 30s→2s and added a search feature; rollback removed search, which product refused, so the timeout stayed wrong for 2 hours.

Revision contents.

RevisionApptimeoutRollback
mixed+search30→2sblocked by product
config-onlysame digest30→2syes
app-only+search30syes

I would not consider it settled without evidence: A week of prod revisions: config-only changes are labelled and majority-small, and a drill rolls back values without an app revert when the design allows.

If config always rides with features, you will choose the feature and keep the outage. That record is part of the change, not a wiki page written afterwards.

Curated: · Written: · Reviewed:

QA-68NEW_CHECKOUT_V2 has been 100% for 4 months and is still in the code and the flag service. What fails CI?(show answer)

The judgement in feature-flag cleanup in the same pipeline is which check is required to ship, not which check looks impressive on a pull request.

A flag at 100% past its expiry is debt the pipeline can see: a catalogue of flags with owners and dates, and a failing check when expiry passes.

Concretely, each flag is a YAML entry with expires_on. CI fails on expired flags still referenced in code. Removing the flag is a normal PR, not a festival. Record the digest, the gate that ran, and the rollback command for feature-flag cleanup in the same pipeline on the same change, and store that record beside the promote job so an incident later can name the bits that shipped.

The reason for that specificity is a failure I have seen: The flag service outage defaulted NEW_CHECKOUT_V2 to off after 4 months at 100%; the old path had bit-rotted and checkout 500ed.

Flag catalogue.

Flag100% sinceexpires_onCI
NEW_CHECKOUT_V22026-05-012026-06-01red in Sep
TAX_EXPERIMENT10%2026-10-01green

I would not consider it settled without evidence: An expired flag in the catalogue fails main, and a code reference to a deleted flag fails unit tests.

A permanent flag is just another config file you forgot is a branch. That record is part of the change, not a wiki page written afterwards.

Curated: · Written: · Reviewed:

QA-69Policy bans Friday production. A Sev-1 on Saturday had no pipeline because "we don't deploy". What should exist?(show answer)

Where teams lose a week on Friday deploys with a documented hotfix path is usually a bypass that was left in for one hotfix and never removed.

A freeze is a default, not an absence of a path; hotfix uses the same artifact rules with a recorded freeze exception.

Concretely, document who can approve a freeze exception, that the digest was already built, and that smoke and rollback still run. Monday's first job audits weekend exceptions. Record the digest, the gate that ran, and the rollback command for Friday deploys with a documented hotfix path on the same change, and store that record beside the promote job so an incident later can name the bits that shipped.

The reason for that specificity is a failure I have seen: Saturday's fix was a kubectl set image of a laptop build because the pipeline was "closed"; it lacked the SBOM and could not be rolled back to a known digest.

Weekend path.

PathArtifactAudit
no pipelinelaptop imagenone
freeze exceptionexisting digestMonday
normal Fridayexisting digestsame gates

I would not consider it settled without evidence: A tabletop: freeze exception ticket, promote of an existing digest, smoke, and Monday audit; a laptop path is still denied.

If the only Saturday path is SSH, the freeze taught the wrong lesson. That record is part of the change, not a wiki page written afterwards.

Curated: · Written: · Reviewed:

QA-70CAB meets Thursdays. GitOps promotes 12 times a day. How do they coexist without a fake screenshot after the fact?(show answer)

I would answer change calendar versus GitOps promote by separating what the pipeline proved from what production has not yet seen.

If a human approval is required, it is an environment rule on the promote, not a committee that rubber-stamps last week's Git history.

Concretely, either drop CAB for standard-risk promotes that already have two-person env rules, or make CAB the reviewer of the production environment. Do not generate compliance PDFs from already-shipped revisions. Record the digest, the gate that ran, and the rollback command for change calendar versus GitOps promote on the same change, and store that record beside the promote job so an incident later can name the bits that shipped.

The reason for that specificity is a failure I have seen: 12 daily promotes and a Thursday CAB that approved a spreadsheet; an auditor asked for the 17:40 Friday digest and nobody could show a pre-approve.

Approval time versus live time.

DigestLiveCABEnv approve
17:40 Fri17:41next Thumissing
17:40 Fri17:41—17:39
high-risk12:0011:50 same day11:55

I would not consider it settled without evidence: Every production digest has a timestamped approve before the sync that made it live, in the same system that deploys.

A meeting cannot approve a digest it has not seen; the environment rule can. That record is part of the change, not a wiki page written afterwards.

Curated: · Written: · Reviewed:

QA-71Lead time is 5 days, 4 of them waiting for CAB. DORA wants faster. What do you move out of CAB?(show answer)

The engineering content of CAB as a bottleneck versus environment rules is the lead time and the change-fail cost, not the number of stages.

Standard-risk, already-gated digest promotes leave CAB; CAB keeps irreversible data, legal, and new high-risk systems.

Concretely, classify changes. Standard: env rules + automated gates. High: CAB before promote. Measure queue time of CAB as waste, not as quality. Record the digest, the gate that ran, and the rollback command for CAB as a bottleneck versus environment rules on the same change, and store that record beside the promote job so an incident later can name the bits that shipped.

The reason for that specificity is a failure I have seen: Every CSS padding waited for CAB; people batched 80 changes into the Thursday slot and the one bad SQL went with them.

Lead time split, p50.

StageBeforeAfter class
CI20 min20 min
CAB queue4.2 d0 for standard
promote15 min15 min
CFR9%8%

I would not consider it settled without evidence: 28-day lead time split: coding, CI, CAB queue, promote. After reclassifying, CAB queue drops and CFR does not rise.

A committee on every padding change is how you buy big-bang Thursdays. That record is part of the change, not a wiki page written afterwards.

Curated: · Written: · Reviewed:

QA-72Slack /deploy prod checkout runs a workflow_dispatch that skips required reviewers "because a human typed it". What did you skip?(show answer)

Before adding another job I would write what a good delivery of ChatOps deploy still goes through the same gates looks like in DORA terms.

ChatOps is another client of the same protected environment; a slash command is not a second, weaker pipeline.

Concretely, the slash command triggers the same job with the same environment rules. The Slack user must map to a reviewer who is not the author, or the command only queues. Record the digest, the gate that ran, and the rollback command for ChatOps deploy still goes through the same gates on the same change, and store that record beside the promote job so an incident later can name the bits that shipped.

The reason for that specificity is a failure I have seen: /deploy from a stolen Slack session shipped a debug digest; there was no second reviewer because the workflow_dispatch path had been marked skip.

Clients of prod env.

ClientReviewersStaging green
GitHub UIrequiredrequired
/deploy skipnoneskipped
/deploy same jobrequiredrequired

I would not consider it settled without evidence: An attempt to /deploy a digest that is not green on staging is rejected with the same message as the GitHub UI.

If Slack can do what GitHub cannot, attackers only need Slack. That record is part of the change, not a wiki page written afterwards.

Curated: · Written: · Reviewed:

QA-73Every PR creates a full stack. 220 namespaces are still running from merged PRs. What is the lifecycle?(show answer)

The first thing I would pin down about preview environments with a TTL is which merge or promote it is allowed to block.

Preview environments die when the PR dies, on a TTL, and on a cost budget; they are not a second production that nobody owns.

Concretely, namespace labelled with PR number and expiry (for example 3 days). A janitor deletes merged/closed PRs hourly. Sleep scale-to-zero after 2 hours idle if the platform allows. Record the digest, the gate that ran, and the rollback command for preview environments with a TTL on the same change, and store that record beside the promote job so an incident later can name the bits that shipped.

The reason for that specificity is a failure I have seen: 220 leftover stacks cost $18k/month and one of them still had a public ingress with last week's admin password.

Preview census.

Open PRsNamespacesCost
18220$18k
1818$1.4k
merged PR + 1 h0 for that PR—

I would not consider it settled without evidence: Count of preview namespaces equals open PRs ± janitor lag, and a merged PR's namespace is gone within an hour.

A preview without a death date is just unmanaged production. That record is part of the change, not a wiki page written afterwards.

Curated: · Written: · Reviewed:

QA-74Previews run at production replica counts. Finance asks why CI is a top-three cloud bill. What do you size?(show answer)

I would start cost of idle preview namespaces from the artifact that actually ships, not from the job that happened to go green.

Preview size is enough to exercise the path, not a copy of production capacity; idle previews scale to zero.

Concretely, replicas: 1, small instance classes, shared data fixtures, scale to zero on idle. Budget alerts per namespace. Record the digest, the gate that ran, and the rollback command for cost of idle preview namespaces on the same change, and store that record beside the promote job so an incident later can name the bits that shipped.

The reason for that specificity is a failure I have seen: Each preview ran 12 checkout pods like prod; 40 PRs × 12 = 480 pods idle overnight, $11k in compute before anyone opened a browser.

Replicas.

Envcheckout replicasOvernight
prod12needed
preview copy12 × 40 PRs$11k
preview1 + scale to 0$400

I would not consider it settled without evidence: A preview's 24-hour cost is a published number under a cap, and an over-cap namespace is scaled or deleted.

If the preview bill rivals production, you are load-testing the budget, not the PR. That record is part of the change, not a wiki page written afterwards.

Curated: · Written: · Reviewed:

QA-75A new service copies a 900-line workflow from checkout and changes the name. Six months later checkout got OIDC and the new service still has keys. How should a service start?(show answer)

This is a place where a passing pipeline and a safe production change for golden path templates versus copy-paste YAML are not the same event.

New services instantiate a versioned golden-path template (or repo generator) that receives platform updates as version bumps, not as folklore.

Concretely, a cookiecutter/Backstage template pins platform workflow @v3. Service CI fails if it is more than N versions behind the current golden path. Record the digest, the gate that ran, and the rollback command for golden path templates versus copy-paste YAML on the same change, and store that record beside the promote job so an incident later can name the bits that shipped.

The reason for that specificity is a failure I have seen: 14 copy-paste repos still had AWS keys 6 months after checkout moved to OIDC; two keys leaked.

Skew.

RepotemplateOIDC
checkoutv3.4yes
new-service copyv1.0 from Junkeys
generatedv3.4yes

I would not consider it settled without evidence: A scorecard: template version on each repo, and a failing check below the allowed skew.

Copy-paste is a one-time golden path and a permanent fork. That record is part of the change, not a wiki page written afterwards.

Curated: · Written: · Reviewed:

QA-76Teams open tickets for a namespace, quota, and ingress. Median wait is 6 days. What does the platform expose instead?(show answer)

My answer to platform API for namespace as a product begins at the promotion boundary, because that is where delivery either has a gate or has a story.

A namespace is a product with an API: identity, quota, network policy, and TTL, created in minutes without a ticket for the default path.

Concretely, a self-service API or Backstage action that creates a namespace from a golden spec. Tickets remain for exceptions. Measure request-to-Ready time. Record the digest, the gate that ran, and the rollback command for platform API for namespace as a product on the same change, and store that record beside the promote job so an incident later can name the bits that shipped.

The reason for that specificity is a failure I have seen: A 6-day ticket for a namespace made teams share a leftover preview namespace; a test dropped the other's tables.

Time to namespace.

Pathp50Failure mode
ticket6 dshared leftover NS
API golden spec8 minnone
exception (public ingress)1 dstill ticket

I would not consider it settled without evidence: p50 time from request to Ready namespace < 15 minutes on the golden path, published on the platform scorecard.

If the paved road is a ticket, the dirt road is production. That record is part of the change, not a wiki page written afterwards.

Curated: · Written: · Reviewed:

QA-77The golden path cannot run a privileged DaemonSet that security approved for one team. They copied a blank cluster. How do you allow the exception?(show answer)

I would treat escape hatch that is reviewed and expires as a control in the path to production rather than as a dashboard tile after the fact.

Escape hatches are named, owned, time-boxed, and reviewed; they are not a second undocumented platform.

Concretely, an exception object in Git: who, why, until, extra RBAC. CI fails expired exceptions. The default path stays default. Record the digest, the gate that ran, and the rollback command for escape hatch that is reviewed and expires on the same change, and store that record beside the promote job so an incident later can name the bits that shipped.

The reason for that specificity is a failure I have seen: A "temporary" privileged ServiceAccount from 2024 was still in the namespace and was the identity that deployed an unsigned image.

Exception record.

FieldValue
teampayments
extraprivileged DS
until2026-10-01
expired still presentCI red

I would not consider it settled without evidence: Every extra permission in prod namespaces appears in the exception catalogue with a future until date, or CI is red.

An exception without an expiry is the new golden path, just for one team. That record is part of the change, not a wiki page written afterwards.

Curated: · Written: · Reviewed:

QA-78Deploy frequency is 14/week, all on Thursday 16:00. Is that high frequency?(show answer)

The useful question for DORA deploy frequency without weekend batching is what a developer is still able to merge if the control is skipped.

DORA deployment frequency counts successful production deployments; fourteen independent Thursday deploys still count as fourteen. Concentration and batch size are separate risk diagnostics, not a reason to under-count.

Concretely, report successful prod deploys per day, plus a weekday histogram and median commits per promote. A Thursday train of 14 is high frequency and high concentration; the interview answer names both. Record the digest, the gate that ran, and the rollback command for DORA deploy frequency without weekend batching on the same change, and store that record beside the promote job so an incident later can name the bits that shipped.

The reason for that specificity is a failure I have seen: The Thursday train contained a bad migrate; 13 other changes rolled back with it; leadership heard "elite frequency" and missed the batch-size risk.

28 days of promotes.

PatternCountThu share
Thursday train56100%
spread weekdays5622%
batch size p5018 vs 3—

I would not consider it settled without evidence: A 28-day histogram of prod promotes by weekday/hour, and median batch size (commits or services) per promote, next to the DORA frequency number.

Fourteen deploys in one hour are still fourteen; they are also one blast radius if they share a migrate. That record is part of the change, not a wiki page written afterwards.

Curated: · Written: · Reviewed:

QA-79MTTR in the incident tool is 4 hours. DORA time to restore is asked in an interview. What clock do you start and stop?(show answer)

I would settle time to restore as a delivery metric against a reversible promote, so a bad change has a measured way back.

Time to restore is production-impact start to the digest that restored service being healthy, not ticket close and not "we rolled forward on Monday".

Concretely, stamp impact_start from the SLI, restore_time from the first healthy smoke of the fixing digest. Report p50/p90. Link the digest. Record the digest, the gate that ran, and the rollback command for time to restore as a delivery metric on the same change, and store that record beside the promote job so an incident later can name the bits that shipped.

The reason for that specificity is a failure I have seen: The incident closed when the doc was written on Monday; checkout was restored Saturday 19:12; leadership thought restore was 38 hours.

One failed promote.

ClockTime
SLI impactSat 18:40
digest healthySat 19:12
ticket closeMon 09:00
DORA restore32 min

I would not consider it settled without evidence: For the last 10 failed changes, restore clock matches SLI recovery, and the fixing digest is in the record.

If restore is when the ticket closes, you will write faster than you ship. That record is part of the change, not a wiki page written afterwards.

Curated: · Written: · Reviewed:

QA-80CI is "slow" at 35 minutes. Execution is 7 minutes. Where is the 28?(show answer)

The judgement in pipeline queue time versus execution time is which check is required to ship, not which check looks impressive on a pull request.

Lead time inside CI is queue + run; you scale runners for queue, and you split tests for run, and you do not confuse them.

Concretely, export queue_duration and run_duration. Alert when p50 queue > a budget (for example 5 minutes). Autoscale runners on queue, not on CPU of a busy job. Record the digest, the gate that ran, and the rollback command for pipeline queue time versus execution time on the same change, and store that record beside the promote job so an incident later can name the bits that shipped.

The reason for that specificity is a failure I have seen: The team spent a quarter sharding tests (run 12→7 min) while queue stayed 28 minutes on a 4-runner pool during the afternoon peak.

35-minute wall clock.

PartMinutesFix
queue28more runners
run7shards
shard-only projectstill 28+7—

I would not consider it settled without evidence: A dashboard of p50 queue vs p50 run for 14 days, and a change in runner count that moves queue without moving run.

Faster tests do not help a job that has not started. That record is part of the change, not a wiki page written afterwards.

Curated: · Written: · Reviewed:

QA-81concurrency: cancel-in-progress on main cancelled the required check of the merge commit when a later push arrived. What is the rule?(show answer)

Where teams lose a week on cancel-in-progress versus the merge queue is usually a bypass that was left in for one hotfix and never removed.

Cancel-in-progress is for disposable PR runs; the merge queue's job for a candidate SHA must not be cancelled by the next SHA.

Concretely, separate concurrency groups: pr-N may cancel; queue-SHA never cancels. Required checks reference the queue group. Record the digest, the gate that ran, and the rollback command for cancel-in-progress versus the merge queue on the same change, and store that record beside the promote job so an incident later can name the bits that shipped.

The reason for that specificity is a failure I have seen: A second merge into the queue cancelled the first SHA's integration job; GitHub reported missing checks and a human force-merged.

Concurrency groups.

Groupcancel-in-progressUse
pr-Nyesdraft pushes
queue-shanomerge queue
group: main (bad)yescancelled required

I would not consider it settled without evidence: Two queued SHAs both complete their required jobs, demonstrated in a fixture, and force-merge is still denied.

If the required job can be cancelled, the merge is racing the next push. That record is part of the change, not a wiki page written afterwards.

Curated: · Written: · Reviewed:

QA-82Two production deploys ran at once and helm interleaved. What serialises them?(show answer)

I would answer concurrency group per environment by separating what the pipeline proved from what production has not yet seen.

One in-flight promote per environment (or per service+env); concurrency is a lock, not a suggestion in a comment.

Concretely, cI concurrency group prod-checkout with cancel-in-progress: false (queue). GitOps has a similar sync lock. Do not start helm if a revision is progressing. Record the digest, the gate that ran, and the rollback command for concurrency group per environment on the same change, and store that record beside the promote job so an incident later can name the bits that shipped.

The reason for that specificity is a failure I have seen: Two GitHub jobs applied revision 88 and 89; the cluster ended on 88's values with 89's image for 15 minutes.

Lock.

concurrency:
  group: deploy-prod-checkout
  cancel-in-progress: false
# second job waits; it does not cancel the first

I would not consider it settled without evidence: A test that two overlapping prod jobs: the second waits, and helm history is monotonic.

Parallel production applies are a split brain you scheduled. That record is part of the change, not a wiki page written afterwards.

Curated: · Written: · Reviewed:

QA-83Load tests run in a nightly Jenkins and page Slack. Production shipped a 3× slower endpoint at noon. Where should the test sit?(show answer)

The engineering content of load test as a promote gate is the lead time and the change-fail cost, not the number of stages.

A performance budget that can block a digest belongs on the promote path (or a required pre-prod env), not on a nightly that cannot unship.

Concretely, staging soak against the candidate digest with a p99 budget. Fail the promote on budget breach. Nightly tests the fleet, not the candidate. Record the digest, the gate that ran, and the rollback command for load test as a promote gate on the same change, and store that record beside the promote job so an incident later can name the bits that shipped.

The reason for that specificity is a failure I have seen: Noon digest raised p99 180→640 ms; nightly reported it at 02:10; rollback was then mixed with 6 other changes.

When p99 is seen.

TestSees digestCan block prod
nightly Jenkinsyesterdayno
staging soak on candidatethis SHAyes
prod onlyuserstoo late

I would not consider it settled without evidence: A candidate that sleeps 500 ms extra is red on the promote job in staging, before production pointer moves.

A nightly that cannot stop yesterday's digest is a newspaper. That record is part of the change, not a wiki page written afterwards.

Curated: · Written: · Reviewed:

QA-84A protobuf field was reused with a new meaning. Old mobile clients crashed after the API deployed. What gate runs on the proto change?(show answer)

Before adding another job I would write what a good delivery of schema compatibility in CI looks like in DORA terms.

Compatibility tests (buf breaking, schema registry, or equivalent) run on every PR that touches the contract, against the last released version.

Concretely, buf breaking --against the production tag. Field reuse, type changes, and removals fail. Additive optional fields pass. Record the digest, the gate that ran, and the rollback command for schema compatibility in CI on the same change, and store that record beside the promote job so an incident later can name the bits that shipped.

The reason for that specificity is a failure I have seen: Field 7 went from int64 user_id to string country; old clients read garbage and crashed for 9% of sessions.

buf breaking.

ChangeCI
add optional field 18pass
reuse field 7fail
delete field 3fail unless reserved

I would not consider it settled without evidence: A PR that reuses a field number is red, and a PR that adds field 18 is green.

If the first client to see the new meaning is production, you did not version the contract. That record is part of the change, not a wiki page written afterwards.

Curated: · Written: · Reviewed:

QA-85The app digest rolled. index.html at the CDN edge was 45 minutes old. Users mixed 4.12 JS with 4.11 HTML. Where is purge?(show answer)

The first thing I would pin down about CDN purge as a release step is which merge or promote it is allowed to block.

A frontend promote includes cache invalidation (or hashed assets with a new HTML that is purged); the pipeline is red until the edge serves the new index.

Concretely, hashed JS/CSS, then purge or versioned HTML. The job fetches the public URL and asserts the new asset hash before succeeding. Record the digest, the gate that ran, and the rollback command for CDN purge as a release step on the same change, and store that record beside the promote job so an incident later can name the bits that shipped.

The reason for that specificity is a failure I have seen: App pods were 4.12; CloudFront still had 4.11 index.html; 18% of users hit 404 on hashed files that 4.11 HTML did not name.

Promote checklist.

StepDone
pods 4.12yes
purge index.htmljob
GET index has 4.12 hashrequired
skip purge18% 404

I would not consider it settled without evidence: curl of the public index after promote contains the new hash, or the job retries purge until it does, within a timeout.

The cluster is not the edge. Users meet the edge. That record is part of the change, not a wiki page written afterwards.

Curated: · Written: · Reviewed:

QA-86You deployed us-east-1 and eu-west-1 in parallel. The schema expand had only landed in us-east-1. What is the order?(show answer)

I would start multi-region promote order from the artifact that actually ships, not from the job that happened to go green.

Data-plane and schema changes follow a documented region order (expand everywhere, then dual-running app, then contract); parallel app deploys cannot outrun schema.

Concretely, pipeline stages: migrate expand all regions, then app canary region A, then remaining regions. Contract last, after all apps are new. Record the digest, the gate that ran, and the rollback command for multi-region promote order on the same change, and store that record beside the promote job so an incident later can name the bits that shipped.

The reason for that specificity is a failure I have seen: eu-west-1 pods 4.12 wrote a column that eu's database did not have yet; 4% of EU orders 500ed for 20 minutes.

Order.

Stageus-east-1eu-west-1
expandyesyes
app 4.12thenthen
parallel app (bad)yes500s
contractafter both appsafter both

I would not consider it settled without evidence: A DAG in the pipeline that refuses app 4.12 in a region whose expand job has not succeeded.

Simultaneous is not safer; it is a race between regions. That record is part of the change, not a wiki page written afterwards.

Curated: · Written: · Reviewed:

QA-87The migrate Job ran at 03:00. At 03:04 the Job was wrong. PITR was 24 hours. Restoring the 03:00 snapshot would drop four minutes of orders. What should the Job have done?(show answer)

This is a place where a passing pipeline and a safe production change for database backup before migrate job are not the same event.

Prefer expand/contract so rollback does not need a restore. If you snapshot before a destructive migrate, recovery is PITR or table-level restore plus a plan for writes that landed after the snapshot, not a blind whole-database rewind.

Concretely, pre-migrate: record a PITR mark or snapshot id in the Job log. Apply additive SQL first. On fail, restore only the affected objects and replay or reconcile post-mark writes; do not clobber the four minutes of good orders. Record the digest, the gate that ran, and the rollback command for database backup before migrate job on the same change, and store that record beside the promote job so an incident later can name the bits that shipped.

The reason for that specificity is a failure I have seen: A DROP went out at 03:00; on-call restored yesterday's snapshot and lost 25 hours of orders that binlogs then had to rebuild by hand.

Job steps.

1 record PITR mark / snapshot id checkout-pre-mig-19
2 apply additive SQL only
3 on fail: restore affected objects, reconcile writes after the mark
# do not restore yesterday's whole DB over four minutes of good orders

I would not consider it settled without evidence: Staging Job logs contain a snapshot id, and a drill restores that snapshot in under the RTO.

A migrate whose restore would discard later good writes is still a one-way command. That record is part of the change, not a wiki page written afterwards.

Curated: · Written: · Reviewed:

QA-884.12 is bad. 4.11 is compatible. A forward-fix will take 90 minutes to write. What does the pipeline prefer?(show answer)

My answer to forward-fix versus revert as a deploy begins at the promotion boundary, because that is where delivery either has a gate or has a story.

If the previous digest is compatible, revert/rollback is the first promote; a forward-fix is a second digest with its own gates, not an argument to keep the broken one live.

Concretely, rollback job is one click/command on the last known good digest. Forward-fix follows the normal pipeline. Do not hotfix by editing live. Record the digest, the gate that ran, and the rollback command for forward-fix versus revert as a deploy on the same change, and store that record beside the promote job so an incident later can name the bits that shipped.

The reason for that specificity is a failure I have seen: The team spent 90 minutes writing 4.12.1 while 4.12 stayed live; rollback to 4.11 was 4 minutes and was unused.

Clocks.

ActionTime to healthy
rollback 4.114 min
author 4.12.190 min
kubectl editunrecorded

I would not consider it settled without evidence: A drill: rollback wall time < 10 minutes for a compatible pair, documented next to the forward-fix path.

Keeping a bad digest live while you author is a choice, not a constraint. That record is part of the change, not a wiki page written afterwards.

Curated: · Written: · Reviewed:

QA-89Hotfix/tax-bug is built on the runner with --no-verify and pushed as :hotfix. Why is that not a hotfix?(show answer)

I would treat hotfix still uses the same artifact rules as a control in the path to production rather than as a dashboard tile after the fact.

A hotfix is a small change through the same build, sign, and promote path, usually cherry-picked onto the release digest's branch, not a second factory.

Concretely, cherry-pick onto the commit of the production digest, run the required jobs (possibly a reduced set that is still documented), sign, promote. Tag :hotfix is not a policy. Record the digest, the gate that ran, and the rollback command for hotfix still uses the same artifact rules on the same change, and store that record beside the promote job so an incident later can name the bits that shipped.

The reason for that specificity is a failure I have seen: :hotfix skipped tests and signature; it fixed tax and reintroduced the secret-in-layer bug CI had been catching for a month.

Hotfix path.

PathTestsSignProd repo
--no-verify :hotfixnonoyes (bad)
cherry-pick + CIrequired setyesyes
laptopnono403

I would not consider it settled without evidence: The hotfix digest has provenance from CI and a signature, and a --no-verify job cannot push to the prod repository.

If hotfix bypasses the factory, every outage is a chance to smuggle risk. That record is part of the change, not a wiki page written afterwards.

Curated: · Written: · Reviewed:

QA-90A junior changed .github/workflows/prod.yml to skip integration. CODEOWNERS listed /src. Who reviewed the skip?(show answer)

The useful question for CODEOWNERS covering the pipeline itself is what a developer is still able to merge if the control is skipped.

Workflow, Helm, Terraform, and policy paths have owners who are not the default "anyone with write"; the pipeline is production code.

Concretely, cODEOWNERS for .github/, deploy/, terraform/, policies/. Required reviews cannot be skipped on those paths. Platform team is the owner. Record the digest, the gate that ran, and the rollback command for CODEOWNERS covering the pipeline itself on the same change, and store that record beside the promote job so an incident later can name the bits that shipped.

The reason for that specificity is a failure I have seen: The skip merged with one application-developer approval; 9 days without integration; the next incident was a missed contract break.

CODEOWNERS.

/src/ @checkout-devs
/.github/ @platform
/deploy/ @platform
/terraform/ @platform

I would not consider it settled without evidence: A PR that only touches prod.yml cannot merge without a platform owner, shown by a branch-protection test.

If YAML that deploys production has no owner, production has no owner. That record is part of the change, not a wiki page written afterwards.

Curated: · Written: · Reviewed:

QA-91Required check is named integration. The job was renamed integration-v2. Merges now sit pending forever. What went wrong, and what is the actual bypass?(show answer)

I would settle required status contexts that match the jobs against a reversible promote, so a bad change has a measured way back.

A renamed job posts a new context; the old required context stays pending and blocks merge until protection is updated in the same PR. The hole is removing the old name from the required list before the new job is proven, or granting bypass.

Concretely, a repo test lists required contexts from the API and asserts each is produced by a job in the workflow files. Renames update branch protection in the same PR. Never delete a required context without adding its replacement. Record the digest, the gate that ran, and the rollback command for required status contexts that match the jobs on the same change, and store that record beside the promote job so an incident later can name the bits that shipped.

The reason for that specificity is a failure I have seen: Someone removed integration from the required list so merges would flow, then integration-v2 flaked and skipped; 20 merges went through with no integration status at all.

Context map.

RequiredJob postsMerge
integrationrenamed awayblocked pending
integration-v2yesmust update protection
both in same PRyesok

I would not consider it settled without evidence: CI fails if a required context is not in the workflow's job outputs, run on every workflow PR, and a protection diff is part of the rename PR.

A required check that nothing posts blocks the merge; deleting it from the list is the bypass. That record is part of the change, not a wiki page written afterwards.

Curated: · Written: · Reviewed:

QA-92The org has 6 people who can bypass branch protection. Two left the company. How do you keep the list true?(show answer)

The judgement in bypass actors inventory is which check is required to ship, not which check looks impressive on a pull request.

Bypass actors are a standing production permission; they are inventoried, owned, and removed on offboarding the same day as GitHub write.

Concretely, quarterly export of bypass actors vs HR/offboarding. Alert on add. Prefer merge queue over bypass. Apps that bypass are named. Record the digest, the gate that ran, and the rollback command for bypass actors inventory on the same change, and store that record beside the promote job so an incident later can name the bits that shipped.

The reason for that specificity is a failure I have seen: A contractor bypass token merged unsigned workflow changes 4 months after their last day.

Bypass census.

ActorEmployedBypass
4 platformyesyes
2 contractorsnostill yes (bad)
after jobnono

I would not consider it settled without evidence: A monthly artifact: bypass list equals the approved named set, and an offboarding ticket that removes GitHub also removes bypass.

A bypass list you have not read this month is an unknown admin team. That record is part of the change, not a wiki page written afterwards.

Curated: · Written: · Reviewed:

QA-93The GitHub prod role is AdministratorAccess "so helm and terraform both work". What do you split?(show answer)

Where teams lose a week on IAM for the deploy role least privilege is usually a bypass that was left in for one hotfix and never removed.

Deploy identities are scoped to the actions they perform: the helm role cannot iam:CreateUser, the terraform role is per-workspace, and neither is admin.

Concretely, separate OIDC roles: ecr:Get + eks:describe + helm's need; terraform's exact services. Access Analyzer / unused-permission reports feed removals. Record the digest, the gate that ran, and the rollback command for IAM for the deploy role least privilege on the same change, and store that record beside the promote job so an incident later can name the bits that shipped.

The reason for that specificity is a failure I have seen: A compromised workflow with AdministratorAccess created a new IAM user and a key in 90 seconds; helm only needed to patch a Deployment.

Roles.

RoleCan helmCan iam:CreateUser
AdministratorAccessyesyes
gha-helm-prodyesno
gha-tf-prodnono

I would not consider it settled without evidence: Policy simulator: the helm role cannot iam:CreateUser, proven in a test, and Access Analyzer finds no admin on CI roles.

If the deploy role can create users, the pipeline is an IdP. That record is part of the change, not a wiki page written afterwards.

Curated: · Written: · Reviewed:

QA-94The AWS trust policy matches token.actions.githubusercontent.com:sub like repo:acme/*:*. Any repo in the org can assume prod. What claim did you forget?(show answer)

I would answer OIDC audience and subject claims by separating what the pipeline proved from what production has not yet seen.

Trust conditions bind repo, ref or environment, and audience; a wildcard org is one compromised repo away from prod.

Concretely, sub is repo:acme/checkout:environment:production (or ref:refs/heads/main). aud is the expected audience. Review wildcards as incidents. Record the digest, the gate that ran, and the rollback command for OIDC audience and subject claims on the same change, and store that record beside the promote job so an incident later can name the bits that shipped.

The reason for that specificity is a failure I have seen: A public-by-mistake repo in acme ran Actions and assumed the prod role because sub was repo:acme/*.

Trust sub.

sub conditionOther acme repo
repo:acme/:assume prod
repo:acme/checkout:environment:productiondeny
repo:acme/checkout:ref:refs/heads/maindeny from PR

I would not consider it settled without evidence: A job from another repo in the org is denied AssumeRole, and the trust policy is committed and reviewed.

An org-wide sub is a single shared prod role with extra steps. That record is part of the change, not a wiki page written afterwards.

Curated: · Written: · Reviewed:

QA-95Engineers have kubectl edit in prod "for incidents". 80% of prod writes last month were kubectl, not GitOps. What is the default path?(show answer)

The engineering content of production access only through the pipeline is the lead time and the change-fail cost, not the number of stages.

The default production write is the pipeline; human kubectl is break-glass, metered, and rare, or GitOps is optional theatre.

Concretely, rBAC: view in prod, write via CI role. Break-glass role with 60-minute binding. Report the ratio of GitOps writes to human writes monthly. Record the digest, the gate that ran, and the rollback command for production access only through the pipeline on the same change, and store that record beside the promote job so an incident later can name the bits that shipped.

The reason for that specificity is a failure I have seen: 80% kubectl meant the rendered Git was fiction; a "rollback" via Git rolled to a state nobody had been running.

Write source, 30 days.

SourceMutations
GitOps / CI role20%
human kubectl80%
targetCI ≥ 95%, glass ≤ 5%

I would not consider it settled without evidence: Audit API: human writes < a small budget (for example 5% of mutations), and each is a break-glass record.

If humans write more than Git, Git is the backup copy. That record is part of the change, not a wiki page written afterwards.

Curated: · Written: · Reviewed:

QA-96Someone needs kubectl in prod now. What does the grant look like so it is not standing admin?(show answer)

Before adding another job I would write what a good delivery of break-glass kubectl with session recording and expiry looks like in DORA terms.

Break-glass is a short-lived, recorded, ticketed binding to a constrained role, not a permanent group membership.

Concretely, jIT: 60-minute ClusterRole binding, session recording (or API audit to SIEM), ticket id in the request, automatic expiry, post-review next business day. Record the digest, the gate that ran, and the rollback command for break-glass kubectl with session recording and expiry on the same change, and store that record beside the promote job so an incident later can name the bits that shipped.

The reason for that specificity is a failure I have seen: A "temporary" cluster-admin group membership from an incident in March was still there in September and was used to disable NetworkPolicies.

JIT grant.

FieldValue
roleprod-emergency-edit
ttl60 min
recordAPI audit + ticket
March group still theredaily query fails

I would not consider it settled without evidence: Bindings with cluster-admin in prod are zero except inside an active JIT window, proven by a daily query.

If the glass never expires, it is a window. That record is part of the change, not a wiki page written afterwards.

Curated: · Written: · Reviewed:

QA-97Conftest runs in a weekly audit. A privileged pod shipped on Tuesday. When does policy run?(show answer)

The first thing I would pin down about policy as code on every PR is which merge or promote it is allowed to block.

Admission-equivalent policy runs on the rendered manifest in CI on every PR, so Tuesday's pod never merges.

Concretely, kustomize build | conftest/OPA/Checkov in the required job. The same policy bundle is what admission uses, versioned together. Record the digest, the gate that ran, and the rollback command for policy as code on every PR on the same change, and store that record beside the promote job so an incident later can name the bits that shipped.

The reason for that specificity is a failure I have seen: Weekly audit opened 40 tickets on Friday; Tuesday's privileged pod had been in production for 3 days.

When privileged is caught.

ControlTuesday pod
weekly auditFriday ticket
CI conftestmerge blocked
admit onlyblocked if CI skipped

I would not consider it settled without evidence: A PR adding privileged: true is red in <5 minutes, and the policy file is the same bundle hash as Gatekeeper in the cluster.

A weekly policy run is an archaeology department. That record is part of the change, not a wiki page written afterwards.

Curated: · Written: · Reviewed:

QA-98CI has policies. A human applied a bad manifest with --validate=false. What still stops it?(show answer)

I would start Kyverno or Gatekeeper as a platform default from the artifact that actually ships, not from the job that happened to go green.

Cluster admission enforces the same baseline as CI; CI is the fast feedback, admission is the last gate for every client.

Concretely, install Kyverno/Gatekeeper with the baseline (non-root, no privileged, signed images). Audit then enforce. CI uses the same rules to fail early. Record the digest, the gate that ran, and the rollback command for Kyverno or Gatekeeper as a platform default on the same change, and store that record beside the promote job so an incident later can name the bits that shipped.

The reason for that specificity is a failure I have seen: kubectl --validate=false applied a hostNetwork pod; CI had never seen the YAML because it never entered Git.

Two gates, one bundle.

ClientCIAdmit
Git PRfail—
kubectlskipdeny
--validate=falseskipstill deny

I would not consider it settled without evidence: A denied apply in the audit log for hostNetwork, and a CI test of the same YAML also red.

Policy that lives only in CI is optional for anyone who does not use CI. That record is part of the change, not a wiki page written afterwards.

Curated: · Written: · Reviewed:

QA-99A preview Job spawned 400 pods and starved checkout in the shared cluster. What does the namespace come with?(show answer)

This is a place where a passing pipeline and a safe production change for resource quotas on the paved path are not the same event.

Every namespace the platform creates includes quota and limitRange; unbounded is not the default even for "just a Job".

Concretely, quota: cpu, memory, pods, jobs. LimitRange: default requests. Golden path applies them. CI of the platform tests a Job that would exceed quota is denied. Record the digest, the gate that ran, and the rollback command for resource quotas on the paved path on the same change, and store that record beside the promote job so an incident later can name the bits that shipped.

The reason for that specificity is a failure I have seen: 400 pods on a mis-set parallelism ate node CPU; checkout p99 went to 2.4s until the Job finished 50 minutes later.

Namespace defaults.

ObjectExample
ResourceQuota pods50
LimitRange cpu request100m
Job parallelism 400denied

I would not consider it settled without evidence: A new namespace has a ResourceQuota object, and a Job with parallelism 400 is forbidden.

A cluster without quota is a single tenant: whoever schedules first. That record is part of the change, not a wiki page written afterwards.

Curated: · Written: · Reviewed:

QA-100PR CI is 41 minutes. People push to main to "see if it works". What do you cut so the inner loop is under 10 minutes?(show answer)

My answer to inner-loop feedback under ten minutes begins at the promotion boundary, because that is where delivery either has a gate or has a story.

The required PR suite is the cheapest set that still protects main; slower jobs move to merge-queue or post-merge, not into every keystroke.

Concretely, target p50 PR feedback ≤10 minutes: affected tests, remote cache, parallel shards. Full integration on the merge queue. Measure and publish p50. Record the digest, the gate that ran, and the rollback command for inner-loop feedback under ten minutes on the same change, and store that record beside the promote job so an incident later can name the bits that shipped.

The reason for that specificity is a failure I have seen: 41-minute PR CI taught the team to push broken commits to main, where a blocked main took 2 hours to unwind.

Where minutes go.

JobPRMerge queue
unit affected4 min—
integration—12 min
before41 min allunused
p50 target≤10 minextra

I would not consider it settled without evidence: 14-day p50 of PR required jobs ≤10 minutes, and the rate of direct-to-main commits for "trying it" drops.

A 41-minute inner loop is a process that trains people to bypass it. That record is part of the change, not a wiki page written afterwards.

Curated: · Written: · Reviewed: