Overview
Curated: · Written: · Reviewed:
Container security
Container security interviews are usually about trade-offs, not product names. The strong candidates have one mental model — containers are namespaced processes on a shared kernel — and derive every control from it. The weak ones list tools: "we scan with Trivy and use distroless." Tool names without the control they buy you score poorly. Expect the interviewer to push on each lifecycle stage you name: why does scanning not make you safe, why is a digest not a vulnerability guarantee, what does restricted actually block.
The mental model: a container is a process, not a VM
A container is a process the kernel runs with namespaces (PID, network, mount, UTS, IPC, user) for separation and cgroups for resource limits. It shares the host's kernel, syscalls, devices and, unless configured otherwise, parts of its filesystem. That single fact drives the whole subject:
-
A kernel vulnerability or a misconfigured capability is shared attack surface across every pod on a node. There is no hypervisor boundary to save you.
-
Namespaces and cgroups are isolation primitives, not a security boundary equivalent to a VM. They limit what a normal compromised process can see, not what a privileged one can do.
-
Therefore security is layered across the workload lifecycle: source → dependencies → build → image → registry → admission → runtime → network/identity → detection → response. Each stage answers a different attacker action; no single stage covers all of them.
The threat model interviewers want you to articulate: a compromise starts with workload code or a dependency, and escalates laterally — pod → service account → Kubernetes API → node → cluster → cloud credentials (via the node's IAM role or the metadata endpoint). Each control in the sections below cuts one rung of that ladder. Where container security ends: kernel hardening, node OS patching, IAM policy, and network segmentation are host/network/IAM security. Container security interfaces with them (pod → node trust, workload identity → IAM role); it does not subsume them.
What a weak answer sounds like: "containers are isolated, so a compromised pod can't affect anything else." A follow-up probe: "what happens if that pod's service account can create pods?" If the candidate can't answer, the isolation claim collapses.
Image supply chain: from source to admission
The image is the unit of deployment; its integrity is a supply-chain problem.
- Minimal bases and multi-stage builds. Build in a stage with compilers and package managers, copy artifacts into a distroless or minimal final stage. Fewer packages means a smaller CVE surface and less tooling for an attacker who lands a shell. Trade-off: minimal images make debugging painful — no shell, no package tool. Plan for that with copied-in debug sidecars or ephemeral debug containers rather than a shell in the production image.
- No secrets in layers. A file deleted in a later layer still exists in earlier layers and in the image history. Anyone with pull access can recover it. Inject secrets at runtime (mounted volumes, external stores), never
RUN echo $TOKEN > file. - Scanning in CI (Trivy/Grype-class) on the final image, ideally on dependencies too. Scanning tells you what matches known vulnerability intelligence — it says nothing about zero-days, misconfiguration, or whether the CVE is reachable in your workload.
- SBOM (SPDX/CycloneDX) so you can answer "are we exposed to CVE-X" from an inventory instead of re-scanning.
- Signing and provenance (cosign/Sigstore-class): a signature proves a named identity signed this digest; a provenance attestation (SLSA) proves it was built by your pipeline from that source revision. Admission verifies identity and policy against the digest. Neither proves the software is vulnerability-free — a signature attests who built it, not what it does.
The chain interviewers probe: "deploy by digest or by tag?" Mutable tags (:latest, :v1.2) can be re-pointed; the digest (@sha256:...) pins exact content. Release identity should ride on digests. Weak answer: "we pin the tag so it's immutable" — tags are metadata, not content.
Follow-ups to expect: what happens when a base image CVE lands (see next section), how admission verifies signatures (policy controller checks the registry + key), and what a registry compromise gives an attacker (image overwrite → code execution at next deploy, which is why registries need immutable tags and separate publish/read auth).
Prioritizing vulnerabilities: reachability beats CVE counts
Scanner findings grow faster than teams can fix them. The engineering decision is triage.
- Severity (CVSS) measures the bug, not your exposure. A critical CVSS in a library your container never loads at runtime is lower priority than a medium in a network-facing parser.
- Weight by runtime reachability (is the package loaded/executed?), exposure (is the service internet-facing?), and exploit signals — EPSS scores probability of exploitation in the wild; CISA's KEV list flags vulnerabilities already being exploited. A CVE on both lists jumps the queue regardless of CVSS.
- Base-image refresh beats per-package patching: rebuilding on an updated base picks up hundreds of patched dependencies in one go, and a regular rebuild cadence (say weekly) keeps patch latency low. Per-package
apt upgradein Dockerfiles is slow, drifts, and re-opens the image-size argument. - Documented exceptions with owners and expiry dates, not ignored findings. "We have 400 criticals open" and "we closed every scanner finding" are both weak answers; the strong one is "N open, M internet-facing, all KEV-listed ones fixed within an SLA of X days, exceptions owned and reviewed."
Hardening the workload: securityContext and Pod Security Standards
This is the most concrete section, and the one most often probed line by line. A restricted-grade pod:
# Kubernetes 1.29+, Pod Security Admission "restricted" level
apiVersion: v1
kind: Pod
metadata:
name: hardened
spec:
securityContext:
runAsNonRoot: true # kubelet refuses if UID is 0
runAsUser: 10001
seccompProfile:
type: RuntimeDefault # container-runtime seccomp profile
automountServiceAccountToken: false # no token unless the pod calls the API
containers:
- name: app
image: registry.internal/app@sha256:47c3...
securityContext:
allowPrivilegeEscalation: false
capabilities:
drop: ["ALL"] # then add back only what's proven needed
readOnlyRootFilesystem: true # writable data goes to emptyDir/ PVC mounts
resources:
requests: {cpu: 100m, memory: 256Mi}
limits: {memory: 512Mi}
Pod Security Admission enforces three levels — privileged (near-node access), baseline (blocks obvious escalations), restricted (hardened defaults) — with enforce, audit and warn modes per namespace via labels. Roll out with warn/audit first, inventory existing violations, then enforce. An interviewer will ask how to upgrade a running fleet: label namespaces with warn, fix the flagged pods, flip to enforce, keep exemptions (pod-security.kubernetes.io/enforce-version) pinned and owned. Escalation vectors to be able to list on demand: privileged: true, hostPath mounts (node filesystem), hostPID/hostIPC/hostNetwork, mounting /var/run/docker.sock or the container-runtime socket (effective node control), and the automounted default service account token — the classic escalation from "I can create a pod" to "I can act as any service account, mount Secrets, or request privileged settings" unless admission and namespace design constrain it.
Trade-off to volunteer: hardening costs operational flexibility. readOnlyRootFilesystem: true breaks apps that write to /tmp or /var/log until you mount writable volumes. Restricted level blocks some legitimate images that run as root. The mature answer names the friction and how it was handled, not just the flag.
What a container cannot defend: node and isolation boundaries
Because the kernel is shared, place sensitive or mutually untrusted workloads on separate nodes or clusters, or on sandboxed RuntimeClasses (e.g. a microVM-isolating runtime) when the threat justifies it. Protect the node itself: OS and kernel patching, kubelet and runtime socket permissions, the cloud metadata endpoint (block pod access to 169.254.169.254 unless a workload needs it), and node instance identity. A container escape or a privileged pod can yield node credentials, which escalates to cluster and cloud. This is also the honest "when not to use containers" answer: workloads with strong regulatory isolation requirements, multi-tenant distrust, or unusual kernel-adjacent needs may need VM-grade isolation instead.
Identity, network, and secrets at runtime
- Workload identity: one dedicated service account per workload, minimal RBAC, cloud workload identity bound to that account rather than node-wide roles. Disable token mounting when the pod doesn't call the API; otherwise use projected, short-lived, audience-bound tokens.
- NetworkPolicy: default-deny ingress and egress, then allow the specific peers, ports and DNS. Two traps interviewers love: (a) NetworkPolicy is only enforced if the CNI implements it — an accepted object a plugin ignores gives zero isolation, so test connectivity; (b) it's layer 3/4 — it doesn't authenticate users or encrypt anything. Application identity and encryption come from mTLS/service-mesh or app-level authn, not NetworkPolicy.
- Secrets: Kubernetes Secret objects are base64, not encrypted, at the API level unless etcd encryption at rest is on. RBAC on
get/list/watch secretsand on pod creation (which can mount them) is the real gate. Avoid env vars for sensitive values when process listings or crash dumps could leak them; mount or inject instead. Rotate, revoke, and make sure logs, manifests and image layers never carry values.
Runtime detection and response
Preventive controls settle what may run; they cannot settle what a compromised process does. Runtime detection observes behavior: unexpected process execution, shell and package-tool use, writes to protected paths, new listeners, unusual egress, privileged syscalls, service-account token use — typically via syscall/eBPF sensors (Falco-class rules) or the runtime's own events.
Two signals worth naming unprompted:
- Drift: the process binary differs from the image the admission controller approved — a compromised pod commonly downloads a second stage. Detect and alert on it.
- Cryptomining: high CPU plus new external connections plus an unfamiliar process is a well-known signature; it's the canonical "what would you detect first" example.
Detection is not prevention. The trade-off: you can't prevent every exploit, so you bound the blast radius (hardening, isolation, identity) and invest in fast detection and containment instead of chasing zero exploits.
On response, the strong arc is: preserve evidence (audit logs, runtime events), isolate network and revoke credentials, identify the image digest, node, service account, secrets and downstream access, then replace the workload from a known-good artifact — do not try to clean a running compromised container. Rotate what was exposed, inspect peers and build history, fix the entry path, and validate policies before restoring traffic. If node escape is plausible, replace the node through a trusted provisioning path.
What changes at scale, and the metrics that matter
At tens of teams and thousands of pods, per-image manual review dies. What survives scaling: policy as code at admission (same rules for everyone), digest-pinned deploys with verified provenance, automated rebuild cadence on refreshed bases, and exceptions that are owned, time-boxed and visible. Sandbox or dedicated-node placement becomes a scheduling decision driven by workload class, not an argument per team.
The measure is not zero scanner findings. It's a small, explainable attack surface: verified build identity on deployed digests, vulnerability exposure weighted by reachability with KEV/EPSS-driven SLAs, few and owned policy exemptions, a short list of privileged/host-access workloads, default-deny network isolation with tested exceptions, runtime alerts wired to a rehearsed containment runbook, and patch latency and time-to-contain you can quote.
