Skip to content
Tech Interview Prep home
Technical interview guide

Secrets Management in the Cloud

Using cloud-native secrets services so credentials are never hardcoded, with automatic rotation.

Read
45 min
Practice MCQs
25
Interview QA
25
Edition
v4
Editorial status
Reviewed

Scope: AWS Secrets Manager, Azure Key Vault, Google Cloud Secret Manager, and OWASP guidance current 2026-09-01.

Overview

Curated: · Written: · Reviewed:

Secrets management

Review status: rewritten from reviewer feedback — framing failed, content depth retained; second revision fixing a policy-evaluation error. Not yet re-reviewed.

A secret is any value whose disclosure grants authority: passwords, API keys, signing keys, TLS private keys, connection strings, recovery tokens, shared webhook secrets. Systems and design interviews probe this topic at two levels. At the design level, the question is usually "where do credentials come from and who can read them" — and the strong answer starts by trying to eliminate the secret rather than store it better. At the incident level, the question is "a key landed in a public repo, what happens in the next hour" — and the weak answer stops at "delete it and rotate it," which fixes nothing because the credential is still valid wherever it was copied.

The arc to hold onto: eliminate → minimize → control delivery → control storage → rotate → prove. Every section below is a decision point on that arc, and interviewers probe each one.

Threat model: where secrets live and what exposure means

Before discussing controls, map where the value actually exists. A credential passes through source control and its history, container image layers, CI configuration, environment variables, runtime memory, log lines, crash dumps, developer laptops, tickets and chat threads. Each location is a separate exposure surface with a separate owner — the answer "it's encrypted at rest in the vault" tells you nothing about the copy pasted into a Slack thread or baked into a base image.

Two categories drive every design decision:

  • Static secrets: long-lived shared values, changed only by rotation (database passwords, API keys). The failure mode is slow: leaked, copied, forgotten, and still valid eighteen months later.
  • Dynamic secrets: short-lived credentials generated per consumer, with expiry designed in (database credentials minted by a broker, STS tokens, cert-based identity). The failure mode is different: expiry too short for the workload, or renewal silently failing so the workload dies mid-flight.

Blast radius is the design driver. Ask of every credential: if this leaks right now, what can an attacker do, for how long, and what does revoking it break? A leaked root-account static key with no rotation is a total compromise of indefinite duration. A leaked fifteen-minute scoped token is an incident report entry. When interviewers ask you to design credential handling for a system, working out blast radius per credential — and then shrinking it — is the analysis they're listening for.

First move: kill the credential instead of storing it

The strongest answer to "how do you secure this secret" is often "this secret shouldn't exist." Concretely:

  • Workload identity: the platform vouches for the workload. In Kubernetes, a service account can be projected and exchanged (via OIDC issuer) for cloud IAM roles; on AWS, an EC2 instance profile or ECS task role gives the workload a role without any stored credential; GKE and Azure have direct equivalents. No bootstrap secret, no rotation, no expiry to miss.
  • Federation: CI systems (GitHub Actions OIDC, GitLab) trust a cloud provider's STS with a scoped role assumption, constrained by condition keys on repository, branch, and environment. No long-lived key sitting in the CI settings page.
  • Certificate-based identity: mTLS client certs with automated issuance and short lifetimes replace shared API keys for service-to-service auth.
  • Brokered dynamic credentials: a tool like Vault mints per-consumer database credentials with TTLs measured in hours or days.

The residual problem is always the last static secret: the thing that bootstraps trust. Workload identity itself has to be constrained — an overly permissive OIDC trust policy on the cloud side turns every pod in the cluster into a deputy for the role. Interviewers who are good at this topic push exactly here: "you've federated CI to AWS — what stops a modified pull request from assuming the production role?" The answer is condition keys (repo, environment, ref), and requiring approval gates for production environments, not just having federation.

What you give up by eliminating static secrets: a dependency on the identity plane. If token issuance fails, the workload cannot authenticate. A static credential cached on disk works when the broker is down. That trade-off — the availability cost of short-lived identity — is a legitimate senior-level point to raise.

Delivery: how the secret reaches the process, and why env vars leak

If a secret must exist, delivery choice is a leakage trade-off, not a convenience choice.

Environment variables are the default everyone reaches for and the worst option in practice. The value is readable by anything that can inspect the process (/proc/<pid>/environ on Linux), inherited by every child process, captured in crash dumps and core files, and printed by the first stack trace or debug logger that dumps the environment. The classic production leak is a Python Exception handler or a debug middleware printing request.headers or the whole process environment into a log aggregator with a 400-day retention policy.

Better options, in rough order:

  • Direct SDK retrieval in-process: the app authenticates with workload identity and fetches the secret at the moment of use. No copy in the process environment, precise audit trail of the fetch, but a runtime dependency on the manager's availability.
  • Mounted files or tmpfs: the platform writes the secret to a file owned by the process user, with an OS-enforced permission boundary, refreshed on rotation (Kubernetes secret volume updates, ECS config files). Leaks require file read, not process inspection; cleanup on decommission is explicit.
  • Init container or sidecar injection: a component with its own identity fetches and writes the credential before the app starts, or rewrites it on rotation. Adds a moving part, but keeps the main image credential-free and lets the sidecar own renewal.

Whatever the mechanism: fetch only the version needed, don't grant or use list operations, keep values out of command-line arguments (visible in ps output), and redact by default in logs, traces, metrics labels, and support bundles.

Caching is the follow-up interviewers raise. Cache in the consuming process only, bound the TTL to the rotation and emergency-revocation objective — an indefinite cache silently defeats revocation; a zero cache turns every request into a dependency on the manager's availability and can overload its rate limits. State what the process does when the manager is unreachable: serve stale with an alert, or fail, decided per credential.

Storage: the key hierarchy and where plaintext exists

Three shapes of store, with different trust models:

  • Managed platform stores: AWS Secrets Manager / Parameter Store, GCP Secret Manager, Azure Key Vault secrets. Zero operational burden, IAM-gated, per-secret pricing and retrieval costs, region-scoped.
  • Self-hosted brokers (HashiCorp Vault): rich dynamic-secret engines, lease lifecycle, revocation — at the cost of running, unsealing, backing up, and DR-testing the broker itself. A down Vault is now an outage class of its own.
  • Parameter stores / config stores (SSM Parameter Store and similar): fine for low-rate config; usually lacking per-version audit granularity and rotation primitives.

Underneath sits envelope encryption, and interviewers frequently probe whether you can state it precisely: a data encryption key (DEK) encrypts the secret; a key encryption key (KEK) — often a cloud KMS key, potentially a customer-managed one — encrypts the DEK. The store holds ciphertext plus the wrapped DEK. Decrypting requires a KMS API call, which means every decrypt is a loggable, alertable, deniable-or-not event — that's the property that makes the hierarchy worth drawing on a whiteboard.

The questions that matter for the design: who holds decrypt rights on the KEK (separate principals from secret administrators?), where does plaintext exist (in the store's memory, in the consumer's memory, in any cache), and what breaks if you rotate the KEK or lose access to it. Customer-managed keys give you revocation-by-key-denial and satisfy regulatory separation, but add a permission dependency: an AWS KMS key policy misconfiguration can make every secret in the account unreadable. Saying that trade out loud, rather than defaulting to "CMKs are more secure," is what separates a staff-level answer.

Access control and separation of duties

The failure most real environments have: one principal class that can read, decrypt, administer, and rotate. Split it where the platform allows:

  • Read/decrypt (the workload's runtime identity) vs manage (ops, pipelines) vs rotate (the automation or a rotation service) vs administer keys (the KMS key policy holder).
  • Policy types and how they combine. On AWS, this is a place where stating the evaluation logic precisely is the differentiator. For a principal in the same account as the KMS key, use of the key needs both the key policy and the IAM policy to permit it (an explicit deny in either wins). For cross-account access, the other account's root must be enabled in the key policy, and within that grant either the key policy statement or the principal's own IAM policy permitting is enough — the two do not have to both allow. Secret Manager resource policies and IAM identity policies similarly scope who can reach a given secret. A candidate who can say which evaluation applies and why it differs is answering a real follow-up; one who states a single blanket rule is reciting rather than reasoning.
  • Condition keys: decrypt only from a given VPC endpoint, only with a given aws:PrincipalTag, only for a given encryption-context value. Encryption context in KMS deserves a mention — binding ciphertext to purpose (e.g. {"app": "billing"}) means ciphertexts can't be swapped between uses, and it lands in audit logs.

Also separate environments and blast-radius boundaries outright: production should not share vaults, projects, credentials, or administrative identities with dev, and cross-account access should be explicit on both sides and continuously reviewed. And keep audit of both management operations (who changed a policy) and data operations (who read which secret) — AWS KMS, for example, logs these differently, and a manager whose data-plane reads are invisible cannot answer "who saw this value during the exposure window."

Alert on the reads that matter: bulk listing, reads from new principals or networks, disabled audit logging, deletion or purge, repeated denies. Join manager events with workload and deployment evidence, because a successful read proves retrieval, not misuse.

Rotation: a distributed change, not a new vault version

This is where strong candidates visibly separate from weak ones, because rotation is a coordinated multi-system change and interviewers know most teams have never tested theirs.

The sequence: generate strong new material → install it in the target system (the database, the SaaS, the API provider) → publish or stage the new version in the manager → have consumers adopt it → verify real authentication with the new credential → revoke the old one → keep audit evidence of the whole path. The ordering matters: revoking first is an outage; verifying by checking that the new version exists in the vault is not verification — the provider-side credential is what actually changed or didn't.

Mechanisms for no-downtime overlap: dual credentials (two active passwords, common with database users), alternating users (credential A/B with a flip), or version aliases with staged adoption. Version labels are coordination, not enforcement — consumers that pin an old version, cache indefinitely, or never reload will silently keep authenticating with material you believe you revoked.

Test rotation before production and continuously after: idempotency when the job re-runs, concurrent runs, partial failure mid-sequence, target unavailability, rollback, replica lag on the credential store, consumers in every region. And monitor the things that prove rotation worked: last successful rotation time, credential age, disabled versions, and the interval until old credentials actually stop working.

Incidents, deletion, and metrics

Leak response. A secret in Git history, a fork, a container layer, or a chat log is a credential incident, not a cleanup task. Preserve when and where it appeared; identify every copy and access path; revoke the credential at the authority that honors it (the target system, not the vault — deleting the vault entry or rewriting Git history changes nothing about a copied password's validity); mint and distribute replacement material; verify the old value fails; audit what was read during the exposure window; narrow permissions found to be excessive. Note the asymmetry: a rotated vault secret with a still-valid upstream credential is the single most common false "resolved" in real incidents.

Deletion and recovery. Soft delete and purge protection defend against accidental or malicious destruction — but recovery permissions are themselves privileged, and restoring a secret can restore access nobody should have had. Test restoration before you need it, and ensure decommissioning revokes the target credential before deleting the record of it. Replication adds resilience and residency/key-policy/logging obligations at the same time.

Metrics that mean something. Secrets eliminated through identity, inventory coverage, unknown-owner secrets, overdue rotations, time from leak detection to revocation, readership breadth (secrets with dozens of readers), unused secrets, and audit coverage. Secret count alone is ambiguous — growing count might mean growing attack surface or just honest inventory. The end state you argue for: every remaining credential necessary, narrowly authorized, short-lived where possible, observable, safely replaceable, promptly revocable.

How this gets graded in interviews

When a system-design or security question touches credentials, the probes are predictable:

  • "Where does the app get its database password?" — Weak: an env var from CI. Strong: workload identity → broker → mounted file, with the trade-offs named.
  • "A key was pushed to a public repo. Walk me through your next hour." — Weak: delete it, rotate it, move on. Strong: history and forks still hold it, the credential is valid until revoked at the authority, audit reads during the exposure window, replacement + verification + old-credential rejection. The interviewer is testing whether you treat it as a distributed incident.
  • "How does rotation work without downtime?" — Weak: "we update it in the vault." Strong: dual credentials or aliases, consumer adoption, target-side change, verification, staged revocation, consumers that cache forever breaking it.
  • "Who can decrypt?" — Weak: "the admins." Strong: the KEK/DEK hierarchy, decrypt rights separated from administration, encryption context in the audit trail, what a key-policy misconfiguration takes down.

Follow-ups to expect: how the workload authenticates to the manager without a bootstrap secret (workload identity, or admit the bootstrap problem honestly), what happens when the manager is down (stale-cache policy), how untrusted PRs are kept from CI credentials (federation constraints + approvals), and what "audit" means for data-plane reads specifically. If you can answer those four without hedging, you have this topic at the level it's asked.