Skip to content
Tech Interview Prep home
Technical interview guide

CI/CD Pipeline Design

Continuous integration and continuous delivery — automating the path from commit to a shippable build.

Read
30 min
Practice MCQs
25
Interview QA
25
Edition
v4
Editorial status
Reviewed

Scope: GitHub Actions, GitLab CI/CD, NIST SSDF, SLSA v1.2, Sigstore, OCI, Kubernetes, Argo CD, and DORA guidance current 2026-08-31.

Overview

Curated: · Written: · Reviewed:

CI/CD pipelines turn reviewed source into traceable, reversible production change

The three commitments, and where interviews start

Interviewers usually open with "what's the difference between continuous delivery and continuous deployment?" and the weak answer is a synonym shuffle. The distinction is a policy decision, not a tool choice:

  • Continuous integration: every change is built and tested together with mainline frequently, so integration risk is paid in small installments. The guarantee: mainline is never knowingly broken for long, and you find out within minutes.
  • Continuous delivery: every change that passes the pipeline produces a deployable artifact. The guarantee: you could release the artifact at any time; releasing is a business decision (one click, one approval).
  • Continuous deployment: every qualifying change is released automatically. The guarantee: the pipeline's evidence is trusted enough to be the release authority — no human in the loop.

The follow-up that separates senior candidates: "what breaks when you move from delivery to deployment?" Answer: the pipeline's false-negative rate becomes your release failure rate. A flaky test that fails 1% of runs stops being an annoyance and becomes 1% of releases blocked or rolled back, so continuous deployment forces you to invest in test reliability, progressive delivery, and automated rollback in a way delivery does not.

Each pipeline stage holds a specific guarantee, and a good answer names them:

StageGuarantee it must holdWhat breaks if it lies
BuildSame source revision → same artifact (reproducible, content-addressed)You tested one artifact and shipped another
TestThe evidence the release decision rests onEscaped defects with no trace of what was verified
PackageArtifact is immutable, signed, and identified by digest"Latest" tag drifts between environments
ReleaseThe promoted artifact is the tested digestEnvironment-specific rebuilds diverge
DeployThe running system matches the declared state, or rollback is safeSuccess status with a broken release

The core design rule: build once, promote the immutable artifact. If each environment rebuilds from source, you have three artifacts and no chain of custody. Deploy and verify content digests, not tags — tags are convenient names, digests are identity.

Test strategy: what gates a merge versus what gates a release

A common weak answer treats "the pipeline" as one undifferentiated test run. Interviewers probe the split:

Merge gate (minutes, every PR): format, lint, type check, unit tests, focused security scans (dependency audit, secret detection). These are fast, deterministic, and cheap to run on every change. If this stage takes 40 minutes, developers batch changes to avoid the wait — which defeats CI's whole premise.

Release gate (longer, on merge or on schedule): integration tests against real dependencies, database migrations, performance/load tests, end-to-end journeys, accessibility, environment parity tests. These run where their dependencies are meaningful — an integration test against a stubbed database proves little.

The test pyramid shape matters for pipeline duration: many fast unit tests, fewer integration tests, few end-to-end tests. An inverted pyramid (hundreds of browser tests, thin unit coverage) gives you a 90-minute pipeline and flakiness proportional to surface area. If asked "how do you keep the pipeline tolerable?", the senior answer includes:

  • Fan-out/fan-in: shard the test suite across parallel runners, then join for the deploy decision. A 40-minute suite on 10 runners is ~4 minutes plus overhead.
  • Selective testing: run only tests affected by the change's dependency graph, not the whole suite every time.
  • Path rules with care: skipping jobs on docs-only changes is fine, but a skipped job must not silently remove required evidence — make the skip explicit and auditable.
  • Flaky-test policy: quarantine with an owner and a deadline, not deletion and not tolerance. "We retry flaky tests automatically" is a weak answer — retries hide real defects and teach the team that red can mean green on the second try.

The follow-up to expect: "a flaky test blocks a Friday release — what do you do?" Weak answer: retry until green. Better answer: quarantine with a tracked ticket, ship, fix the flake next week, and look at why the test is flaky — it often encodes a real race condition.

Pipeline-as-code and workflow mechanics

The pipeline is reviewed code with provenance like any other code. Interviewers probe whether you treat it that way:

  • Triggers: push, PR, schedule (cron), webhook, manual with approval. Each trigger carries different trust and blast radius — a scheduled trigger runs with no human in the loop at all, so it needs the same least-privilege and isolation discipline as any privileged run, and a manual trigger with production authority needs an auditable actor.
  • Explicit dependency graph: model the workflow as a DAG, not a linear script. Fast deterministic checks fail early; integration and deploy stages wait only on what they causally need. Parallelize independent work; retain strict ordering for build → attest → deploy → verify.
  • Matrix builds: one job definition, fan-out across axes (OS, language version, dependency versions). Fan-in with a final job that requires all matrix legs.
  • Reusable workflows: shared steps across repos, pinned to reviewed versions — which makes them a supply-chain dependency you inventory like any other.
  • Branch strategy: trunk-based development (short-lived branches, merge to main frequently) is what CI was designed for; GitFlow's long-lived develop/release branches reintroduces the integration risk CI exists to remove. If your team runs GitFlow and calls it CI, expect the follow-up "when does your code actually integrate?"

A concrete shape, in GitHub Actions syntax (checked against the documented workflow schema):

# .github/workflows/ci.yml — GitHub Actions
on:
  pull_request:
  push:
    branches: [main]

jobs:
  fast-checks:          # merge gate, ~4 min
    runs-on: ubuntu-latest
    steps:
      - uses: actions/checkout@<pinned full SHA>
      - run: make lint typecheck unit-tests
      - run: make secret-scan

  integration:          # release gate, ~12 min, real database
    needs: fast-checks
    runs-on: ubuntu-latest
    services:
      postgres:
        image: postgres:16
        env:
          POSTGRES_PASSWORD: test
        options: >-
          --health-cmd pg_isready --health-interval 5s
    steps:
      - uses: actions/checkout@<pinned full SHA>
      - run: make integration-tests

  package:
    needs: integration
    runs-on: ubuntu-latest
    outputs:
      digest: ${{ steps.push.outputs.digest }}
    steps:
      - uses: actions/checkout@<pinned full SHA>
      - run: make image-push   # pushes the image, prints the digest
        id: push

  deploy-staging:
    needs: package
    runs-on: ubuntu-latest
    steps:
      - run: make deploy DIGEST=${{ needs.package.outputs.digest }}
      - run: make smoke-tests

The needs: chain is the causal ordering; the digest output is the chain of custody. Traced through: a PR opens → pull_request fires and fast-checks runs against the PR's merge ref (the candidate merge of the PR into the base branch) → on merge to main, the push trigger runs the same chain on the merge commit → integration gets a real Postgres service container → package builds and pushes an image and exports its digest → deploy-staging deploys that digest and runs smoke tests. Nothing rebuilds between environments; the digest that passed integration is the one that lands.

Caching, dependencies, and reproducible builds

"How do you make builds fast without making them unsafe?" is the trade-off question here. Caching buys speed and sells trust if done naively.

  • What's safe to cache: dependency downloads keyed by lockfile hash, compiler/build outputs keyed by source tree hash. What isn't: anything a lower-trust change could write that a privileged run later reads — a poisoned cache is a supply-chain attack that bypasses review entirely.
  • Cache keys: derive from the lockfile or content hash so a dependency change invalidates automatically; scope caches by branch trust level so a fork PR never shares cache entries with a mainline run.
  • Hermetic inputs: pin toolchains and third-party actions to immutable versions or digests (actions/checkout@<full-sha>, not @v4), lock dependencies, and record tool versions in provenance. A build that pulls latest from the internet is not reproducible and cannot be audited.

The interview version: "your build works on your machine and in CI but not on the deploy box — walk me through debugging it." The answer is provenance: compare dependency versions, toolchain versions, and artifact digests between the three environments. If the deploy box rebuilt rather than promoted, that's the bug.

Supply chain and runner trust

Assume pull requests, branch names, commit messages, issue text, generated files, and dependency scripts are untrusted input. The classic vulnerability is interpolating them into shell: run: echo "${{ github.event.pull_request.title }}" executes whatever the PR author wrote. Pass untrusted context as structured, validated inputs to isolated jobs instead.

  • Grant the default token minimal permissions; elevate per job. Prefer short-lived workload identity federation over long-lived cloud keys.
  • Protect secrets from forks and untrusted runners. Masking output so secrets don't appear in logs is worth doing, but don't mistake it for containment: a fork PR that can run code can exfiltrate a secret through channels masking doesn't cover — writing it into an artifact, encoding it into a cache key, or sending it out via network egress. The control is not giving untrusted code access to secrets at all; masking is a log-hygiene measure, not an anti-exfiltration one.
  • Separate runner pools: public/untrusted, internal build, production deploy. A self-hosted runner that executes an untrusted PR can poison later jobs on the same machine even if that PR received no secret — workspace persistence is the attack vector. Ephemeral clean runners, restricted egress, and post-job teardown close it.
  • Signing and provenance attest identity and build claims — they do not prove the software is correct. A signed artifact can still have logic defects; the signature is what a policy verifies, so protect the signing identity and the verification roots.

Deployment strategies and progressive delivery

This is where design questions concentrate. Know each strategy's rollback story, not just its definition:

StrategyHow it worksRollbackCost / risk
RollingReplace instances batch by batchRoll back batches (slow — you pass through mixed versions)Cheap; mixed-version window requires backward compatibility
Blue/greenTwo full environments, switch trafficInstant traffic switch back2× infrastructure; state drift between colors
CanarySmall % of traffic to new version firstShift traffic back in secondsNeeds per-version metrics and traffic splitting
Feature flagsDeploy dark, expose to users graduallyDisable in seconds, no redeployFlag debt; flags become a second config system needing hygiene

The senior-level point: deployment and release are different events. With flags, you deploy code continuously and release features independently — which decouples rollback of a feature (flag off, seconds) from rollback of a deployment (redeploy previous digest, minutes). Interviewers probe this: "a release is causing errors — what do you do first?" Weak answer: roll back. Better answer: check whether a flag disable or traffic shift is faster and safer, because rollback re-enters the deployment path while a flag flip doesn't.

Guardrails must be predeclared: error rate, p99 latency, saturation, data-integrity checks — with automatic halt or rollback only when that action is known safe. Sometimes roll-forward, traffic isolation, or write pause is safer than rollback. And a deployment job marked successful proves orchestration completed, not that users received a correct release — verify from the user boundary: synthetic journeys, SLO/error budget, expected digest running.

State changes and the limits of rollback

Application rollback cannot undo a destructive database change. This is the failure mode interviewers love: "you rolled back but lost data — what went wrong?" The answer is expand/migrate/contract:

  1. Expand: additive, backward-compatible schema change, deployed with old code still running.
  2. Migrate: backfill or transform data incrementally, idempotent, resumable.
  3. Contract: remove old columns/paths only after no version reads them — a later, separate release.

Keep dual-read/write compatibility windows explicit and remove them deliberately, or they become permanent. Gate irreversible steps (column drops, data deletion) more strongly than reversible ones, and rehearse recovery before you need it.

GitOps and measuring the pipeline itself

GitOps moves desired-state authorization into reviewed version control; a reconciler applies and corrects drift. Two probes worth preparing:

  • "What's the failure mode of self-heal?" Automated prune and self-heal amplify an incorrect desired state — a bad commit deletes production resources automatically. Test deletion, partial sync, and controller outage.
  • "An engineer hot-fixed production directly — what happens?" Detect the drift, don't silently normalize it: reconcile the change into source through review, or revert it through an accountable path. Silent normalization teaches the team that bypassing review works.

Finally, measure outcomes, not pipeline theater: feedback time, deployment frequency, lead time, change failure rate, recovery time, flaky-test count with owners, rollback success. Don't reward raw deployment count or a universally faster pipeline at the expense of safety. And test the pipeline itself: provider outage, runner exhaustion, credential expiry, compromised dependency, partial deployment. The pipeline is a system; the mature version makes the safest path the easiest path while keeping explicit, auditable exceptions.

Quick self-check

  • Delivery vs deployment: who makes the release decision, and what does that imply about your flaky-test rate?
  • Why deploy by digest rather than tag? What does a signature prove and not prove?
  • What gates a merge versus a release, and how does the pyramid shape keep the merge gate under ten minutes?
  • What can a fork PR poison, and how do cache scoping and runner isolation stop it?
  • For each deployment strategy: how fast is rollback, and what does it cost when nothing is wrong?
  • Why can't application rollback undo a destructive migration, and what does expand/migrate/contract change about your release sequencing?