Skip to content
Tech Interview Prep home

Top 100 Cloud Security Engineer Interview Questions and Answers

The questions most likely to actually come up in your Cloud Security Engineer interview, ranked by likelihood — with detailed, senior-level answers covering what an interviewer is really listening for.

Curated: · Written: · Reviewed:

Reviewed 76Review pending 24
QA-1A team says the provider handles security because they use managed services. Where does that reasoning break?(show answer)

The first thing I would establish about the shared responsibility boundary is what an attacker gains if the control is absent.

The provider secures the infrastructure and the service implementation; the customer configures access, network exposure, data classification, and encryption choices. The boundary moves with the service model, so a single sentence about "the provider handles it" is wrong for some part of any real estate.

Concretely, write the boundary per service rather than per provider, name who owns patching, access control, encryption, logging, and backup for each, and check that every item on the customer side has an owner rather than an assumption.

The reason for that specificity is a failure I have seen: A managed database was assumed to be secured by the provider; it was reachable from the public internet with a default administrative user because network exposure and credentials were both on the customer side of the boundary.

Who owns what, by service model.

ItemIaaSManaged DBSaaSCustomer-owned of 3
host patchingcustomerproviderprovider1
network exposurecustomercustomerprovider2
access controlcustomercustomercustomer3
data classificationcustomercustomercustomer3

I would not consider it settled without evidence: Produce the per-service responsibility table with named owners, and check the customer-side items against the running configuration rather than the design.

The boundary is per service, and the customer side is never empty.

Curated: · Written: · Reviewed:

QA-2A policy grants read-only access. How do you know that is what the identity can actually do?(show answer)

I would start effective permissions versus written policy from the permissions that are actually effective, not from the policy document.

What an identity can do is the result of every policy that applies to it — identity policies, resource policies, group and role inheritance, permission boundaries, and organisation-level controls — evaluated together. Reading one document tells you about one input.

Concretely, evaluate effective permissions with the provider's own policy simulator or an equivalent analysis, run it for the identity and the specific action rather than for the policy, and re-run after any change to any of the inputs.

The reason for that specificity is a failure I have seen: An identity with a read-only identity policy could write to a bucket because the bucket's own resource policy granted the whole account write access; the review had read the identity policy and stopped there.

Where the effective grant came from.

InputGrants writeReviewed
identity policynoyes
resource policyyesno
permission boundaryno limitno
inputs evaluated31 of 3

I would not consider it settled without evidence: Simulate the specific action against the specific resource for the specific principal, and record the result rather than the policy text.

Permissions are the sum of every input, not the one you read.

Curated: · Written: · Reviewed:

QA-3An identity has no dangerous permissions of its own. Why might that not be enough?(show answer)

This is an area where a passing compliance scan and a correct handling of privilege escalation through role assumption are different events.

The ability to assume another role, pass a role to a service, or update a policy is transitively the permissions of everything reachable that way. An identity whose only permission is to create a function and pass an administrative role to it is an administrator.

Concretely, analyse the identity graph transitively rather than the permission list, treat policy-modification, role-passing, and role-assumption permissions as escalation primitives, and check what each identity can reach in one or two hops.

The reason for that specificity is a failure I have seen: A build role held only permission to create compute instances and pass roles; it could pass the administrative role to an instance it created and read its credentials, which made a nominally limited role fully privileged.

Reachable privilege by hop count.

HopsIdentities reaching admin
0, direct3
119
244

I would not consider it settled without evidence: Run a transitive privilege-escalation analysis across the identity graph and report reachable administrative access per identity.

Passing a role is holding it, one step later.

Curated: · Written: · Reviewed:

QA-4A policy grants an action with a wildcard resource. When is that acceptable?(show answer)

My answer to wildcards in policy statements begins with the blast radius: if I cannot state what one compromise reaches, I have no design.

A wildcard is acceptable where the action is genuinely global and harmless — some listing and describe operations — and dangerous everywhere else, because it grants the permission over resources that do not exist yet, including ones created by an attacker.

Concretely, scope resources explicitly wherever the provider supports it, use conditions on tags or paths where naming cannot be enumerated in advance, and treat a wildcard on any write, delete, or policy-modifying action as a finding rather than a style preference.

The reason for that specificity is a failure I have seen: A wildcard resource on a key-decryption action allowed a compromised service to decrypt every tenant's data rather than only its own, turning a single-tenant compromise into a whole-estate one.

What one wildcard matched.

ScopeKeys matchedTenants reachable
explicit ARN11
path prefix41
wildcard340all

I would not consider it settled without evidence: Enumerate what each wildcard actually matches in the current account and confirm the identity should reach every one.

A wildcard grants over resources that do not exist yet.

Curated: · Written: · Reviewed:

QA-5How do you let a team manage their own roles without letting them grant themselves administrator?(show answer)

I would treat permission boundaries and organisation policies as a claim about an adversary's options that has to survive being tested.

Delegating role creation without a ceiling means delegating administrator, because the delegate can create a role with any permission. A permission boundary or organisation-level control policy sets a maximum that a delegated grant cannot exceed.

Concretely, require every role a team creates to carry a boundary they cannot remove, enforce that requirement in the permission that lets them create roles, and set organisation-level guardrails for the controls no account should be able to disable.

The reason for that specificity is a failure I have seen: A team was granted role-creation to move faster; within a month one of their roles carried administrator, and the guardrail that would have prevented it existed only as a written standard.

What the delegate can create.

ControlCan create admin roleCan remove controlAdmin roles after 1 month
written standardyesn/a1
boundary, removableyesyes1
boundary, enforcednono0

I would not consider it settled without evidence: Attempt to create a role exceeding the boundary using the delegated identity and confirm the attempt is denied.

Delegated role creation without a ceiling is delegated administration.

Curated: · Written: · Reviewed:

QA-6Why is a static cloud access key a problem even when it is stored in a vault?(show answer)

The useful question for long-lived access keys is what still holds in the accounts nobody has looked at this year.

A static key is valid until someone revokes it, so its exposure window is unbounded and its compromise is silent. Storing it well reduces the chance of exposure without changing what happens when it does leak.

Concretely, replace static keys with federated short-lived credentials from workload or workforce identity, disable key creation at the organisation level so the path is unavailable rather than discouraged, and where a key is unavoidable, scope it tightly and alert on use from unexpected locations.

The reason for that specificity is a failure I have seen: A static key committed to a repository three years earlier was used from an unfamiliar network; it was still valid, held write access to production storage, and its use looked identical to legitimate automation in the logs.

Exposure window by credential type.

CredentialLifetimeSilent compromise
static keyuntil revokedyes
12-hour session12 hbounded
workload identity1 hbounded

I would not consider it settled without evidence: Report the count and age of static keys across the estate and drive it to zero rather than rotating on a schedule.

A static credential's exposure window ends when someone notices.

Curated: · Written: · Reviewed:

QA-7How do you stop a storage bucket from becoming publicly readable?(show answer)

I would settle public exposure of storage by attempting the action the control is supposed to stop.

A per-bucket setting relies on every future creator getting it right. An account-level block that cannot be overridden by a bucket policy removes the possibility rather than reducing its likelihood, which is the difference between a control and a convention.

Concretely, enable the account-level public-access block, enforce it with an organisation policy so an account owner cannot turn it off, and where genuine public content is required, serve it through a content delivery layer with its own origin access rather than by opening the bucket.

The reason for that specificity is a failure I have seen: A per-bucket setting was applied to 200 buckets; a new bucket created by an automation six months later was public, and the exposure was found by an external researcher rather than by the scanner.

Control strength against a new bucket.

ControlNew bucket safeOwner can disable
per-bucket settingnoyes
account blockyesyes
organisation policyyesno

I would not consider it settled without evidence: Attempt to make a bucket public with an account-admin identity and confirm the attempt fails.

Prevent the state rather than detecting it.

Curated: · Written: · Reviewed:

QA-8Why is the instance metadata service a high-value target, and what mitigates it?(show answer)

The judgement in the metadata service and server-side request forgery is which identity can assume what, not which box the checklist ticks.

The metadata endpoint issues credentials for the instance's role to anything that can make a local HTTP request. An application vulnerable to server-side request forgery can therefore be induced to fetch those credentials without any code execution.

Concretely, require the session-oriented metadata protocol that a simple forged GET cannot satisfy, set the response hop limit so a container cannot reach the host's endpoint, keep instance roles minimal, and alert on credential use from outside the instance.

The reason for that specificity is a failure I have seen: A URL-fetching feature was induced to request the metadata endpoint and returned the instance role's credentials in an error message; the role had broad read access and the data was exfiltrated with valid credentials that looked legitimate.

What each mitigation stops.

MitigationSimple GETFull SSRFContainer reach
noneworksworksworks
session-orientedblockedharderworks
+ hop limit 1blockedharderblocked

I would not consider it settled without evidence: Attempt the forgery against the running application and confirm the metadata service refuses it.

Any request the server makes on your behalf can be aimed inward.

Curated: · Written: · Reviewed:

QA-9What do you check before allowing another account to assume a role in yours?(show answer)

Where candidates lose the interview on cross-account trust policies is reasoning from the architecture diagram.

A trust policy naming an external account trusts everything in that account, including identities created later by people you have never met. Without an external identifier or a condition on the specific principal, the grant is to the account rather than to the party you negotiated with.

Concretely, condition the trust on the specific principal where possible, require an external identifier for third-party access so a confused-deputy attack fails, scope the role's own permissions to the minimum the integration needs, and review external trusts on a schedule.

The reason for that specificity is a failure I have seen: A vendor integration trusted the vendor's whole account with no external identifier; another of that vendor's customers could have induced the vendor to assume the role, and nobody had reviewed the trust in two years.

What the trust actually grants.

Trust statementPrincipals trustedConfused deputy
whole accountall, present and futurepossible
specific role1possible
+ external ID1prevented

I would not consider it settled without evidence: Enumerate every external principal trusted by any role in the account and confirm each is still intended and still conditioned.

Trusting an account is trusting everyone in it, now and later.

Curated: · Written: · Reviewed:

QA-10Data is encrypted with a managed key. What does that actually protect against?(show answer)

I would answer key management and encryption context by separating what is configured from what is enforced at request time.

Provider-managed encryption at rest protects against physical media loss and against access to the storage layer beneath the service. It does not protect against an identity that is authorised to call the service, because the service decrypts transparently for exactly those callers.

Concretely, decide the threat first: use a customer-managed key with its own policy where you need decryption to be separately authorised and auditable, bind encryption context so a ciphertext is valid only in its intended scope, and treat key policy as an access control rather than a configuration.

The reason for that specificity is a failure I have seen: A team cited encryption at rest as the control against data theft; the actual compromise was a role with read access to the service, for which the encryption was entirely transparent.

What each layer stops.

AttackerProvider-managed keyCustomer key + policyField encryption
stolen diskstopsstopsstops
storage-layer accessstopsstopsstops
authorised service rolenowith key policystops
of 3 attackers stopped233

I would not consider it settled without evidence: Name the specific attacker the encryption stops and confirm the compromise you fear is on that list.

Encryption at rest stops the attacker below the service, not the one calling it.

Curated: · Written: · Reviewed:

QA-11Your platform injects secrets as environment variables. What is the risk and what would you do instead?(show answer)

The engineering content of secrets in environment variables is the detection and the revocation path, not the control's name.

Environment variables are visible to the whole process tree, are commonly captured in crash reports and debug endpoints, and are frequently logged by frameworks at startup. The secret is protected in transit and then placed somewhere many things read.

Concretely, mount secrets as files with restrictive permissions or fetch them at runtime through the SDK, keep them out of process arguments entirely, and check the error-reporting integration's redaction rather than assuming it.

The reason for that specificity is a failure I have seen: A database password in an environment variable was captured by an error tracker's automatic environment snapshot and was visible to every engineer with access to that tool, including contractors.

Where the secret appeared.

SurfaceEnv varMounted file
crash reportyesno
/proc for other usersyesno
framework startup logsometimesno
surfaces exposed of 330

I would not consider it settled without evidence: Trigger a crash in a lower environment and inspect what the error tracker captured.

An environment variable is read by more things than you listed.

Curated: · Written: · Reviewed:

QA-12A credential is found in a repository's history. What is the correct response order?(show answer)

Before calling detecting secret exposure in source control done I would write down the attack it does not stop.

Rotation comes first because history rewriting does not un-publish anything that was cloned, forked, or cached. A secret that has been pushed to a shared remote must be treated as compromised regardless of what happens to the history afterwards.

Concretely, rotate immediately, then investigate use of the old credential across the log retention window, then remove from history and add prevention at the push path. Treat the removal as hygiene rather than as remediation.

The reason for that specificity is a failure I have seen: A team rewrote history and considered the matter closed; the credential had been valid for 11 days, appeared in three forks, and had been used from an unrecognised address on day 6, which nobody looked for.

Response order and what it addresses.

StepStops further useFinds prior use
rewrite historynono
rotateyesno
audit log searchnoyes

I would not consider it settled without evidence: Search the audit log for use of the exposed credential across its full valid lifetime, not just since discovery.

Once pushed, it is published; rotate before you tidy.

Curated: · Written: · Reviewed:

QA-13How do you know the image running in production is the one your pipeline built?(show answer)

The first thing I would establish about container image provenance is what an attacker gains if the control is absent.

A tag is a mutable pointer, so an image reference by tag says nothing about which bytes are running. Without signed provenance the chain from source to running container rests on the assumption that nothing in between was modified.

Concretely, reference by digest, generate and sign provenance attesting to the source commit and builder, verify the signature at admission so an unattested image cannot run, and keep the signing key out of the build's own reach.

The reason for that specificity is a failure I have seen: An attacker with registry write access replaced the image behind a stable tag; the deployment pulled it, the digest changed, and nothing in the pipeline compared digests or verified a signature.

What each reference proves.

ReferenceImmutableTies to sourceDetects tamperProperties of 3
tagnonono0
digestyesnoyes2
digest + signed provenanceyesyesyes3

I would not consider it settled without evidence: Attempt to deploy an unsigned or tampered image and confirm admission rejects it.

A tag names a pointer, not the bytes you built.

Curated: · Written: · Reviewed:

QA-14What does a container boundary actually protect against?(show answer)

I would start container escape and workload isolation from the permissions that are actually effective, not from the policy document.

A container is a process isolation mechanism sharing the host kernel, so a kernel vulnerability or an over-privileged configuration puts the host and every other container on it within reach. It is a strong boundary against accident and a weaker one against a determined attacker.

Concretely, drop privileges and capabilities by default, refuse privileged containers and host namespace sharing through admission control, use a read-only root filesystem, and where the workload is genuinely untrusted, use a stronger boundary such as a micro-VM rather than hardening the container.

The reason for that specificity is a failure I have seen: A privileged container running a build was compromised through a dependency; it mounted the host filesystem and read the node's credentials, reaching every workload on that node.

Reach after compromising one container.

ConfigurationHost accessOther workloads reachable on a 24-pod node
privilegedfull23
defaultlimitedup to 23
dropped caps, read-onlyminimal0
micro-VMnone0

I would not consider it settled without evidence: Attempt to run a privileged container and confirm admission rejects it, then verify the same for host namespace and hostPath mounts.

A shared kernel is a shared boundary.

Curated: · Written: · Reviewed:

QA-15Your cloud security posture tool reports 98 percent compliant. What does that leave open?(show answer)

This is an area where a passing compliance scan and a correct handling of posture management and what it cannot see are different events.

A posture tool evaluates configuration against rules it has. It cannot see application logic, data sensitivity, or whether a permitted permission is appropriate, and it evaluates only the accounts it has been connected to.

Concretely, track scanner coverage as a first-class number — accounts connected, regions enabled, resource types supported — and pair posture with reachability analysis and with periodic adversarial testing that starts from an attacker's position rather than from a rule list.

The reason for that specificity is a failure I have seen: A 98 percent score covered 31 of 44 accounts; the breach began in one of the 13 unconnected accounts, which held a legacy environment nobody had onboarded.

Coverage behind the score.

MeasureValue
compliance score98%
accounts connected31 of 44
regions enabled4 of 17
account-region pairs scanned124 of 748, ~17%

I would not consider it settled without evidence: Report accounts and regions covered against accounts and regions that exist, from the organisation's own inventory rather than the tool's.

A score describes what was scanned.

Curated: · Written: · Reviewed:

QA-16Would you rather prevent a misconfiguration or detect it?(show answer)

My answer to guardrails versus detection begins with the blast radius: if I cannot state what one compromise reaches, I have no design.

Prevention removes the state and detection shortens its life. Prevention is stronger where the control is unambiguous and the false-positive cost is low; detection is the right answer where legitimate exceptions exist and blocking would stop real work.

Concretely, prevent the small set of things that are never acceptable through organisation-level policy, detect and remediate the larger set that is usually wrong, and measure the time between creation and remediation for the detected class since that is the exposure you are accepting.

The reason for that specificity is a failure I have seen: Public storage was handled by detection alone with a 24-hour scan interval; a bucket was public for 19 hours and was indexed by a search engine in that window.

Exposure by control type.

ControlExposure window
preventionnone
detection, 5 minup to 5 min
detection, daily scanup to 24 h

I would not consider it settled without evidence: Measure the actual time from misconfiguration to remediation, and compare it against how long the exposure needs to be useful to an attacker.

Detection is prevention with a delay you have to be able to state.

Curated: · Written: · Reviewed:

QA-17Which cloud activity does your audit log not contain?(show answer)

I would treat cloud audit log coverage as a claim about an adversary's options that has to survive being tested.

Control-plane logging records API calls that change or describe resources. Data-plane access — reading an object, querying a database, decrypting with a key — is usually a separate log that is off by default and charged separately, which is why it is missing exactly when an exfiltration investigation needs it.

Concretely, enable data-plane logging for the resources that hold sensitive data, accept the cost as the price of being able to investigate, and confirm coverage by performing a read and finding it in the log rather than by reading the configuration.

The reason for that specificity is a failure I have seen: An investigation could show that a role had permission to read a sensitive bucket and could not show whether it had, because object-level logging had never been enabled for that bucket.

What each log answers.

QuestionControl planeData plane
who changed the policyyesno
who read the objectnoyes
enabled by defaultyesno
buckets with it enabled44 of 443 of 44

I would not consider it settled without evidence: Read an object and confirm the read appears in the log, per bucket class rather than per account.

Permission logs are not access logs.

Curated: · Written: · Reviewed:

QA-18How quickly would you know an identity was being abused?(show answer)

The useful question for detection latency in cloud telemetry is what still holds in the accounts nobody has looked at this year.

Detection latency is the sum of log delivery, aggregation, rule evaluation, and human response, and each stage is usually measured optimistically in isolation. The number that matters is end to end, from the action to somebody acting on it.

Concretely, measure it by performing a benign version of the action you claim to detect and timing the alert, do it in production rather than in a lab, and set the target from how long the attacker needs rather than from what is convenient.

The reason for that specificity is a failure I have seen: A team believed detection was near real time; a controlled test showed 14 minutes of log delivery, 5 minutes of batch evaluation, and a 40-minute median acknowledgement, so the real figure was about an hour.

Where the hour goes.

StageLatency
log delivery14 min
rule evaluation5 min
acknowledgement40 min
total59 min

I would not consider it settled without evidence: Run the controlled test on a schedule and publish the end-to-end latency distribution.

Detection latency is measured to the person acting, not to the rule firing.

Curated: · Written: · Reviewed:

QA-19You believe a session token is compromised. How do you stop it being used?(show answer)

I would settle revoking a compromised session by attempting the action the control is supposed to stop.

Short-lived credentials cannot be revoked individually in most cloud identity systems; disabling the identity or attaching a deny policy is what stops them, and the token remains valid until one of those takes effect. Rotating the key does nothing to a session already issued.

Concretely, know the revocation mechanism per credential type in advance, attach a time-conditioned deny policy that invalidates sessions issued before now, and rehearse it so the first attempt is not during an incident.

The reason for that specificity is a failure I have seen: During an incident a team rotated the access key and believed the session was stopped; the attacker's session token remained valid for another 4 hours and was used throughout.

What each action stops.

ActionStops new sessionsStops existing session
rotate keyyesno
delete identityyesyes, with lag
time-conditioned denyyesyes

I would not consider it settled without evidence: Rehearse revocation for each credential type and measure the time from decision to the credential actually failing.

Rotating the key does not end the session it issued.

Curated: · Written: · Reviewed:

QA-20Your build system can deploy to production. What makes that safe?(show answer)

The judgement in least privilege for CI systems is which identity can assume what, not which box the checklist ticks.

A build system executes code from every contributor, so its permissions are effectively available to anyone who can influence a build. It is usually the most powerful and least scrutinised identity in the estate.

Concretely, separate build from deploy so the identity that runs untrusted code cannot deploy, scope deploy credentials per environment and per branch through the identity provider's own claims, require review for changes to pipeline definitions, and log every deploy with the initiating human identity.

The reason for that specificity is a failure I have seen: A pull request from a fork modified the pipeline definition and ran with the production deploy role, because the workflow was triggered on pull request and the role was not conditioned on branch.

What the CI credential is conditioned on.

ConditionFork PR can deployBranches able to deploy
noneyesall
repository onlyyesall
repository + protected branchno1

I would not consider it settled without evidence: Attempt a deploy from a non-protected branch with the CI identity and confirm the credential is not issued.

The build system runs everyone's code with your permissions.

Curated: · Written: · Reviewed:

QA-21A vendor asks for read-only access to your cloud accounts. What do you grant?(show answer)

Where candidates lose the interview on third-party integrations and their access is reasoning from the architecture diagram.

Read-only across an estate is a substantial grant: it includes configuration, and often data, from every account. The vendor's own security becomes part of yours, and the access usually outlives the commercial relationship.

Concretely, scope to the specific services the integration needs rather than a broad read-only policy, require an external identifier, log the vendor's activity separately, set a review date, and confirm the removal path before granting.

The reason for that specificity is a failure I have seen: A vendor retained broad read-only access 14 months after the contract ended; nobody owned the removal, and their own breach exposed the configuration of every account they could read.

Grant scope and what it exposes.

GrantServices readableData readable
broad read-onlyalloften
scoped to service3no
scoped + external ID3no

I would not consider it settled without evidence: Review external access on a schedule with a named owner and a stated expiry per grant.

A vendor's security posture becomes part of yours at the moment you grant.

Curated: · Written: · Reviewed:

QA-22How do you find the resources that are reachable from the internet?(show answer)

I would answer network exposure and security group drift by separating what is configured from what is enforced at request time.

Reachability is a property of the whole path — security group, network access list, route table, load balancer, and any peering — so reading security groups alone answers a different question. A resource can be exposed by a change several layers away from it.

Concretely, use a reachability analyser that evaluates the full path rather than a rule linter, run it continuously rather than at review time, and reconcile against an external scan since the provider's view and the internet's view can differ.

The reason for that specificity is a failure I have seen: A security group review found nothing; a database was nevertheless reachable through a load balancer in a peered account, which the group-level review had no visibility of.

What each check found.

MethodExposed resources found
security group review0
path reachability analysis3
external scan4

I would not consider it settled without evidence: Compare the provider's reachability analysis against an external scan of your own address space, and investigate every difference.

Exposure is a path property, not a rule property.

Curated: · Written: · Reviewed:

QA-23Traffic to a managed service stays inside the provider network. Is that private?(show answer)

The engineering content of private connectivity to managed services is the detection and the revocation path, not the control's name.

Traffic that does not traverse the public internet is still traffic to a service endpoint that other customers also reach, and network-level privacy says nothing about who is authorised. Private connectivity reduces exposure and does not replace authorisation.

Concretely, use private endpoints to remove internet exposure, keep identity and resource policy as the authorisation control, and add conditions restricting access to the specific endpoint so a credential leaked outside the network cannot be used from elsewhere.

The reason for that specificity is a failure I have seen: A team used a private endpoint and left the resource policy permitting any principal in the account; a compromised workload in an unrelated account used its own path to reach the same service.

What each control stops.

ControlInternet accessWrong principalRight principal, wrong pathStops of 3
private endpointstopsnono1
resource policynostopsno1
both + conditionstopsstopsstops3

I would not consider it settled without evidence: Attempt access with valid credentials from outside the private path and confirm the resource policy refuses it.

Private networking narrows the path, not the permission.

Curated: · Written: · Reviewed:

QA-24How would you stop data leaving through a compromised workload?(show answer)

Before calling egress control and exfiltration done I would write down the attack it does not stop.

Most estates restrict what can reach a workload and allow the workload to reach anything, which is the path exfiltration and command-and-control both use. Egress control is the harder half and the one usually missing.

Concretely, route outbound traffic through a proxy or firewall that logs destination and volume, allowlist the destinations workloads genuinely need, and alert on volume anomalies since a novel destination is easy to hide and a novel volume is not.

The reason for that specificity is a failure I have seen: A compromised workload transferred 400 GB to an external storage provider over 6 days; ingress rules were tight, egress was unrestricted, and there was no record of the destination.

Detection by egress posture.

PostureDestination recordedVolume alertDays to detect
unrestrictednononever
logged proxyyesno6
allowlist + volume alertyesyesunder 1

I would not consider it settled without evidence: Attempt an outbound transfer to an unapproved destination from a workload and confirm it is blocked or at least logged.

The data leaves the way nobody was watching.

Curated: · Written: · Reviewed:

QA-25Why separate environments into different accounts rather than different resource groups?(show answer)

The first thing I would establish about multi-account structure as a blast-radius control is what an attacker gains if the control is absent.

An account is the strongest natural isolation boundary a cloud provider offers: identity, quotas, billing, and most blanket permissions stop at it. Separation within an account depends on getting every policy right, which is a much larger surface to be correct on.

Concretely, put each environment and each significant trust boundary in its own account, use organisation-level controls for the rules that must hold everywhere, and keep cross-account access explicit and reviewable rather than implicit.

The reason for that specificity is a failure I have seen: Production and development shared an account with separation by tag; a development role with a tag-condition error listed and read production storage, and the condition had been wrong since it was written.

What a policy error costs.

StructureEffect of one policy errorControls that must all be right
single account, tagscross-environment access1
separate accountsdenied at boundary2
separate accounts + SCPdenied twice3

I would not consider it settled without evidence: Attempt cross-environment access with a development identity and confirm it is denied at the account boundary rather than by a policy condition.

Separation by policy is separation you have to keep getting right.

Curated: · Written: · Reviewed:

QA-26What can an organisation-level control policy not protect you from?(show answer)

I would start organisation control policies and their limits from the permissions that are actually effective, not from the policy document.

An organisation policy sets a maximum on what identities in member accounts may do, but it applies to those identities rather than to the management account or to service-linked roles in every case. It also cannot express intent — it can deny an action and cannot know whether a permitted action was appropriate.

Concretely, treat organisation policy as a guardrail for the small set of things no account should ever do, protect the management account separately since it is usually outside the policy's reach, and never rely on it as the only control for an action the business legitimately performs.

The reason for that specificity is a failure I have seen: A control policy denied disabling audit logging in every member account; the management account was exempt, and logging there was disabled for 5 months without any alert.

Where the policy applies.

IdentityCovered by SCP
member account roleyes
member account rootyes
management account roleno

I would not consider it settled without evidence: Test the denied action from the management account as well as a member account and record both results.

The account that owns the guardrails is usually outside them.

Curated: · Written: · Reviewed:

QA-27Single sign-on is in place. What has that concentrated?(show answer)

This is an area where a passing compliance scan and a correct handling of identity federation and the identity provider as a target are different events.

Federation replaces many credentials with one trust relationship, which improves lifecycle management and makes the identity provider the highest-value target in the estate. Anyone who can mint assertions there can become anyone.

Concretely, protect the provider itself with phishing-resistant authentication for administrators, restrict who can modify trust and claim mappings, alert on changes to those, and validate assertion conditions on the cloud side rather than accepting whatever arrives.

The reason for that specificity is a failure I have seen: An attacker with administrative access to the identity provider added a claim mapping granting themselves an administrative cloud role; the cloud audit log showed a normal, correctly federated login.

What the cloud log shows.

EventAnomalous in cloud logAnomalous at the IdP
forged assertion loginnono
claim mapping changenoyes
role assumptionnono
detectable of 301

I would not consider it settled without evidence: Alert on every change to federation trust and claim mapping, and route it to a person rather than to a log.

Federation makes one system able to impersonate everyone.

Curated: · Written: · Reviewed:

QA-28Is multi-factor authentication enough for administrative cloud access?(show answer)

My answer to conditional access and device posture begins with the blast radius: if I cannot state what one compromise reaches, I have no design.

Multi-factor authentication stops password reuse and most bulk phishing, and push-based or one-time-code factors remain phishable in real time through a relay. What resists that is a factor bound to the origin, such as a hardware-backed passkey.

Concretely, require phishing-resistant factors for administrative access, add device posture and network conditions where the workforce allows it, and treat a fallback path that permits a weaker factor as the actual strength of the control.

The reason for that specificity is a failure I have seen: An administrator approved a push notification during a real-time relay attack; the attacker held a valid session and the account's controls recorded a successful, compliant login.

Resistance by factor.

FactorBulk phishingReal-time relayAttacks resisted of 2
password onlynono0
push approvalresistsno1
one-time coderesistsno1
origin-bound passkeyresistsresists2

I would not consider it settled without evidence: Attempt a relay attack against your own login flow in a controlled test and confirm the factor refuses it.

A control is as strong as its weakest permitted fallback.

Curated: · Written: · Reviewed:

QA-29How should emergency administrative access work?(show answer)

I would treat break-glass access as a claim about an adversary's options that has to survive being tested.

A break-glass account exists for the case where normal access fails, so it must not depend on the systems that might be failing — the identity provider in particular. That independence is exactly what makes it dangerous if it is not tightly controlled.

Concretely, keep a small number of break-glass identities outside federation with hardware factors held under split control, alert on any use immediately and to more than one person, review every use, and test on a schedule so the credentials are known to work.

The reason for that specificity is a failure I have seen: A break-glass account existed and its factor had expired; during an identity-provider outage nobody could obtain administrative access for 3 hours, and the outage could not be remediated from the cloud side.

Break-glass readiness.

PropertyUntestedExercised quarterly
credential validunknownverified
alert firesunknownverified
time to access3 h, failed8 min

I would not consider it settled without evidence: Exercise the break-glass path quarterly, confirm the alert fires, and record how long the whole procedure took.

Emergency access that depends on the emergency system is not emergency access.

Curated: · Written: · Reviewed:

QA-30Where does access accumulate in a cloud estate?(show answer)

The useful question for joiners, movers, and leavers in cloud access is what still holds in the accounts nobody has looked at this year.

Access is granted on joining and on moving and is rarely removed on moving, so long-tenured people accumulate the union of every role they have held. Leaver processes usually cover the identity provider and miss standing grants that do not flow from it.

Concretely, drive access from group membership so a move changes it automatically, review entitlements on a cadence with the manager rather than centrally, and inventory the grants that do not flow from the identity provider since those are the ones a leaver process misses.

The reason for that specificity is a failure I have seen: An engineer who had moved teams twice retained production database access from a role held 3 years earlier; the access was granted directly rather than through a group and was invisible to the joiner-mover-leaver process.

Access against current need.

TenureGrants heldGrants justified
1 year66
3 years147
5 years228

I would not consider it settled without evidence: Reconcile effective cloud access against current role for a sample of long-tenured people, and count grants that no group explains.

Moving adds access; almost nothing removes it.

Curated: · Written: · Reviewed:

QA-31Should engineers hold standing production access?(show answer)

I would settle just-in-time privileged access by attempting the action the control is supposed to stop.

Standing access is available to an attacker at all times, whereas elevation on request with an expiry is available only during the window and leaves a record of why. The cost is friction at exactly the moment when friction is least welcome.

Concretely, make elevation fast and self-service with an approval only where the risk justifies one, set a short expiry with automatic revocation, log the stated reason, and ensure the elevation path itself works during an incident since that is when it is needed.

The reason for that specificity is a failure I have seen: Twelve engineers held standing administrative access; a phished laptop had it immediately, and the investigation could not narrow the window because there was no elevation record to correlate against.

Exposure by access model.

ModelHours privileged per engineer/monthAttacker window
standing730always
JIT, 4-hour grants1212 h

I would not consider it settled without evidence: Report standing privileged grants against elevations per month, and drive the standing count toward zero.

Standing privilege is privilege the attacker inherits on arrival.

Curated: · Written: · Reviewed:

QA-32A production change is recorded as performed by a deployment role. What is missing?(show answer)

The judgement in logging the human behind the automation is which identity can assume what, not which box the checklist ticks.

An audit trail that stops at the automation's identity cannot answer who caused the change, which is the question every investigation and every access review actually asks. The initiating identity has to be carried through, because it cannot be recovered afterwards.

Concretely, propagate the initiating human identity through the pipeline into the cloud request context via session tags or an equivalent, and require it so a request without it is refused rather than performed anonymously.

The reason for that specificity is a failure I have seen: Every production change appeared under one deployment role; attributing a specific destructive change to a person required correlating pipeline logs with cloud logs by timestamp and took 2 days.

Attribution by design.

Log contentNames personTime to attribute
role onlyno2 days
role + session tagyes2 min

I would not consider it settled without evidence: Take one cloud audit entry and confirm it names the person, without consulting another system.

An automation identity names the tool, not the actor.

Curated: · Written: · Reviewed:

QA-33An attacker gains administrative access. What happens to your logs?(show answer)

Where candidates lose the interview on immutable audit storage is reasoning from the architecture diagram.

An administrator can usually delete or stop the logs in their own account, so the record of the intrusion is within reach of the intruder. Logs are evidence only when they are outside the control of the environment they describe.

Concretely, deliver logs to a separate account with write-only access from the source, apply an object-lock retention that even the destination account's administrator cannot shorten, and alert on any interruption in delivery since silence is the first sign.

The reason for that specificity is a failure I have seen: An attacker with administrative access stopped log delivery and deleted 40 days of history in the same account; the investigation had nothing from before the intrusion to compare against.

What survives an account compromise.

DestinationAttacker can deleteAttacker can stop delivery
same accountyesyes
separate accountnoyes, detected
+ object locknoyes, detected

I would not consider it settled without evidence: Attempt to delete a delivered log object with the highest privilege available and confirm the retention lock refuses it.

Logs an administrator can delete protect you from everyone but an administrator.

Curated: · Written: · Reviewed:

QA-34Which identity signals are actually worth alerting on?(show answer)

I would answer anomaly detection on cloud identities by separating what is configured from what is enforced at request time.

Volume-based anomaly detection on cloud APIs produces mostly noise because legitimate automation is bursty. The signals worth alerting on are the ones with low legitimate frequency and high attacker value: first use of a credential from a new region, enumeration of permissions, and changes to security controls.

Concretely, alert on the small set of high-signal events with full context, keep the broad behavioural analytics for investigation rather than paging, and tune against your own baseline rather than a vendor default.

The reason for that specificity is a failure I have seen: A default rule set produced 340 alerts a week; the team disabled the noisiest category, which was the one that later fired on genuine credential enumeration.

Precision by signal.

SignalAlerts/weekActioned
API volume anomaly2901%
new region for credential1240%
security control disabled2100%

I would not consider it settled without evidence: Measure alert precision per rule and disable or demote the ones nobody actions, rather than muting a category wholesale.

An alert nobody actions is a rule that will be muted.

Curated: · Written: · Reviewed:

QA-35You confirm an attacker has valid credentials in a production account. What is the order of operations?(show answer)

The engineering content of responding to a compromised cloud account is the detection and the revocation path, not the control's name.

Containment and evidence preservation compete, and the order has to be decided in advance because the incident is a bad time to debate it. Snapshotting before terminating preserves what happened; terminating first stops the bleeding and loses it.

Concretely, revoke sessions with a time-conditioned deny, snapshot volatile state where the systems justify it, isolate rather than destroy, preserve logs and their retention, and only then remediate. Have the decision authority named beforehand.

The reason for that specificity is a failure I have seen: A team terminated the compromised instances immediately; the memory that would have shown the second-stage payload was gone, and the same intrusion recurred 3 weeks later through the persistence they had not found.

Order and what it costs.

OrderAttack stoppedPersistence found
terminate firstfastno
snapshot, isolate, then remediatefastyes

I would not consider it settled without evidence: Rehearse the sequence in a tabletop exercise with the real revocation and snapshot commands, and time it.

Destroying the host destroys the answer to how it recurs.

Curated: · Written: · Reviewed:

QA-36Where does a cloud workload's supply chain actually begin?(show answer)

Before calling supply chain risk in cloud dependencies done I would write down the attack it does not stop.

It begins with the base image and every package in it, extends through application dependencies, and includes the build system, the registry, and any infrastructure module pulled from a public source. A module fetched from a public registry runs with the permissions of whoever applies it.

Concretely, pin dependencies including infrastructure modules by digest or version, mirror what you depend on so an upstream deletion is not an outage, review modules that carry permissions, and keep the build system's own credentials scoped so a compromised dependency cannot reach production.

The reason for that specificity is a failure I have seen: A public infrastructure module was updated to add an outbound provisioner; teams that pinned only a major version picked it up and it exfiltrated credentials from every apply for 6 days.

Exposure by pinning discipline.

PinningPicks up malicious update
major version rangeyes
exact versionno, until upgrade
digestno

I would not consider it settled without evidence: Report the share of dependencies, including infrastructure modules, pinned by digest rather than by range.

An unpinned module runs whatever it becomes tomorrow.

Curated: · Written: · Reviewed:

QA-37Where in the pipeline should infrastructure security checks run?(show answer)

The first thing I would establish about infrastructure-as-code security scanning is what an attacker gains if the control is absent.

Checking the plan before apply catches issues while they are cheap and reversible; checking the running estate catches what was created outside the pipeline. Neither alone is sufficient, because the pipeline is not the only way resources appear.

Concretely, run policy checks against the plan as a merge gate, run continuous checks against the running estate, and reconcile the two so that a resource passing the gate and failing in production indicates a drift or an out-of-band change rather than a scanner disagreement.

The reason for that specificity is a failure I have seen: A team ran only plan-time checks; 61 resources created directly in the console over 2 years were never evaluated, and 4 of them had open network exposure.

Coverage by check location.

CheckPipeline resourcesConsole resources
plan-time4120
runtime posture41261

I would not consider it settled without evidence: Report the count of running resources never evaluated by the plan-time gate, and treat it as a coverage gap.

Plan-time checks see what the pipeline made.

Curated: · Written: · Reviewed:

QA-38Your application serves many customers from one account. How do you prevent one reaching another's data?(show answer)

I would start tenant isolation in a multi-tenant application from the permissions that are actually effective, not from the policy document.

Isolation enforced only in application code fails to a single missing filter, and that failure is silent. The stronger designs make cross-tenant access impossible at a layer beneath the application, so a code defect cannot express it.

Concretely, scope credentials per tenant where the architecture allows, enforce tenant predicates in the data layer rather than per query, use encryption context or per-tenant keys so a leaked object is unreadable, and test cross-tenant access as an explicit case rather than assuming.

The reason for that specificity is a failure I have seen: A single query missing its tenant predicate returned another customer's records; every other query in the codebase had it, which is exactly why nobody found the one that did not.

Where isolation is enforced.

LayerOne defect leaks dataDetectablePlaces to get right
per-query filteryesno340
data-layer predicatenoyes1
per-tenant keynoyes1

I would not consider it settled without evidence: Write a test that authenticates as tenant A and attempts to read tenant B's object by identifier, for each access path.

One missing filter is the entire control.

Curated: · Written: · Reviewed:

QA-39Why classify data before choosing controls?(show answer)

This is an area where a passing compliance scan and a correct handling of data classification driving controls are different events.

Without classification every control is applied uniformly, which means either over-spending everywhere or under-protecting the data that matters. Classification is what lets the strong controls go where the loss would be severe.

Concretely, classify at the data store rather than per record where possible, tag resources with their classification so policy can act on it, and drive retention, encryption, access review cadence, and logging depth from the tag.

The reason for that specificity is a failure I have seen: Uniform controls left a store of identity documents with the same access review cadence as a cache of public content; the review found the access list had been stale for 2 years.

Controls by classification.

ClassAccess reviewData-plane loggingKey
publicannualoffprovider
internal6-monthlyonprovider
restrictedquarterlyoncustomer-managed

I would not consider it settled without evidence: Report the share of data stores carrying a classification tag and confirm the controls actually differ by tag.

Uniform controls are wrong somewhere by construction.

Curated: · Written: · Reviewed:

QA-40What makes a backup useful against an attacker rather than against a mistake?(show answer)

My answer to backups as a ransomware control begins with the blast radius: if I cannot state what one compromise reaches, I have no design.

An attacker with administrative access deletes the backups first, so a backup in the same account under the same credentials protects against accident and not against intrusion. Immutability and separation are what change that.

Concretely, keep backups in a separate account with write-only access from the source, apply an immutability period no administrator can shorten, and test restoration on a schedule since an untested backup is an assumption.

The reason for that specificity is a failure I have seen: Snapshots lived in the same account with the same credentials; the attacker deleted them before encrypting, and the only surviving copy was a 6-week-old export that predated a schema change and could not be restored cleanly.

What survives an intrusion.

Backup designAttacker can deleteRestore testedData loss
same accountyesnototal
separate accountnonounknown
separate + immutable + testednomonthlyunder 1 h

I would not consider it settled without evidence: Restore from the immutable copy on a schedule and record the elapsed time and the data loss.

Backups an administrator can delete do not survive an administrator.

Curated: · Written: · Reviewed:

QA-41A key is suspected compromised. What does rotating it achieve?(show answer)

I would treat cloud key compromise and re-encryption as a claim about an adversary's options that has to survive being tested.

Rotating a key changes what future data is encrypted with; it does not re-encrypt existing data, and anything already exfiltrated remains readable under the old key material. Rotation limits forward exposure only.

Concretely, know in advance which data would need re-encryption and how long it would take, keep key version recorded per object so re-encryption is possible, and treat destroying the old key version as a decision with an availability consequence rather than a cleanup step.

The reason for that specificity is a failure I have seen: A team rotated a key and considered the incident closed; 40 TB encrypted under the previous version remained readable to anyone holding the exfiltrated material, and re-encryption had never been scoped.

What rotation changes.

DataReadable with old key after rotation
new writesno
existing objectsyes
exfiltrated copiesyes

I would not consider it settled without evidence: Estimate re-encryption time from a measured sample before the incident, and record which key version encrypted each object.

Rotation protects the future, not the copies already taken.

Curated: · Written: · Reviewed:

QA-42What is different about securing a serverless function?(show answer)

The useful question for serverless function permissions is what still holds in the accounts nobody has looked at this year.

The host and runtime are the provider's, so the customer's surface is almost entirely identity, event source, and dependencies. The function's role is invoked by whatever can trigger it, which makes the trigger's own access control part of the function's.

Concretely, scope the function's role to exactly what it needs, control who and what can invoke it, validate the event rather than trusting its shape, and keep the dependency set small since it is the main remaining attack surface.

The reason for that specificity is a failure I have seen: A function with broad write access could be invoked by any principal in the account because its resource policy was permissive; a compromised low-privilege identity used the function as a privilege escalation path.

Escalation via invoke permission.

Invoker scopeFunction roleEscalationPrincipals able to invoke
any account principalwriteyes214
one servicewriteno1
any account principalread-onlylimited214

I would not consider it settled without evidence: Enumerate who can invoke each function and confirm the invoker set is as narrow as the role is powerful.

A powerful function is as exposed as its trigger.

Curated: · Written: · Reviewed:

QA-43A message arrives on a queue from another service. What do you validate?(show answer)

I would settle event-driven trust boundaries by attempting the action the control is supposed to stop.

A queue is a trust boundary in the same way an HTTP endpoint is: the consumer must not assume the producer's identity, the message's integrity, or its freshness from the fact that it arrived. Anything that can write to the queue can write anything.

Concretely, restrict who can publish, validate the message against a schema, carry and verify the originating identity rather than inferring it, and check freshness so a replayed message cannot re-trigger an action.

The reason for that specificity is a failure I have seen: A consumer trusted a tenant identifier in the message body; a service with publish permission but no business relationship set the field to another tenant and the consumer acted on it.

What the consumer verified.

CheckPresentStops forged tenantPublishers permitted
schema validationyesno12
publisher restrictionpartialpartly4
signed origin identitynoyes1

I would not consider it settled without evidence: Publish a hostile message from an identity with publish permission and confirm the consumer refuses it.

Arriving on your queue is not authentication.

Curated: · Written: · Reviewed:

QA-44How do you threat model a system that is mostly managed services?(show answer)

The judgement in threat modelling a cloud workload is which identity can assume what, not which box the checklist ticks.

The interesting boundaries move from hosts and networks to identities, event sources, and data flows. The useful question stops being which port is open and becomes which principal can invoke what, and what each one reaches if compromised.

Concretely, enumerate principals and trust boundaries rather than network zones, walk each data flow asking what an attacker at that point could do, and record the assumptions so a later change that invalidates one is visible.

The reason for that specificity is a failure I have seen: A model built around network zones missed the entire identity plane; the eventual intrusion used a compromised CI credential and never crossed a network boundary the model had drawn.

Where the model looked.

Model basisBoundaries drawnCovered actual intrusion
network zones6no
identity and data flow19yes

I would not consider it settled without evidence: Check that the model names every principal that can reach the data, not every network segment.

In cloud, the boundary is usually an identity.

Curated: · Written: · Reviewed:

QA-45What makes a security review useful rather than a gate people route around?(show answer)

Where candidates lose the interview on security in the design review is reasoning from the architecture diagram.

A review that arrives after the design is fixed can only object, which makes it expensive and adversarial. A review that arrives while the design is fluid can change it cheaply, which is what makes teams invite it rather than avoid it.

Concretely, engage at the point where the trust boundaries are being decided, bring specific alternatives rather than objections, write down the decision including accepted risks with an owner and a review date, and keep the standing rules available so most designs need no review at all.

The reason for that specificity is a failure I have seen: A review at implementation-complete required a change to the authorisation model; it cost 6 weeks, and the next three teams did not book a review until after launch.

Cost of a finding by stage.

StageFindingsMedian rework
design42 days
implementation33 weeks
pre-launch26 weeks

I would not consider it settled without evidence: Measure how early reviews happen and the share of findings that require rework, and treat late findings as a process defect.

A late review can only be expensive.

Curated: · Written: · Reviewed:

QA-46A team wants to accept a risk you raised. What makes that acceptable?(show answer)

I would answer risk acceptance that means something by separating what is configured from what is enforced at request time.

Risk acceptance is legitimate and needs an owner with the authority to carry the consequence, a stated expiry, and a description specific enough that a later reader can tell whether it still applies. Without those it is a way of closing a finding rather than a decision.

Concretely, record the specific technical condition, who accepted it, what would change the answer, and when it will be revisited, and report accepted risks that are past their review date as overdue rather than as accepted.

The reason for that specificity is a failure I have seen: A register held 140 accepted risks with no expiry; 31 referred to systems that had been decommissioned and 4 to controls that had since become mandatory, and nobody could tell which mattered.

State of the register.

CategoryCount
current and valid105
system decommissioned31
control now mandatory4

I would not consider it settled without evidence: Report accepted risks by age and by whether their stated conditions still hold.

An acceptance without an expiry is a decision nobody revisits.

Curated: · Written: · Reviewed:

QA-47Which numbers would tell you the cloud security programme is working?(show answer)

The engineering content of measuring a security programme is the detection and the revocation path, not the control's name.

Counting findings measures the scanner, and counting controls measures the checklist. What indicates the programme is working is exposure over time — how long a misconfiguration lives, how much privilege is standing, how much of the estate is covered — because those are what an attacker experiences.

Concretely, track mean time to remediate by severity, standing privileged access, coverage of accounts and data-plane logging, and time to detect a controlled test action. Report trends rather than absolute counts, since the denominator moves.

The reason for that specificity is a failure I have seen: A programme reported findings closed per quarter, which rose steadily while median time to remediate a critical finding also rose from 6 to 19 days; the reported number improved while exposure worsened.

Two views of the same programme.

MeasureQ1Q4Reads as
findings closed220410improving
median time to remediate6 d19 dworsening
accounts covered3134slow

I would not consider it settled without evidence: Report exposure duration and coverage as the headline, with finding counts as context.

Count exposure time, not findings closed.

Curated: · Written: · Reviewed:

QA-48How do you know your cloud detections work?(show answer)

Before calling purple teaming cloud controls done I would write down the attack it does not stop.

A detection rule that has never fired on a real technique is a hypothesis. Generating the technique deliberately is the only way to learn whether the telemetry exists, the rule matches, and the alert carries enough context to act on.

Concretely, run the techniques you claim to detect in a controlled way against production telemetry, record which stage failed when the alert does not appear — missing log, missing rule, missing context — and fix that stage rather than adding another rule.

The reason for that specificity is a failure I have seen: Of 40 mapped techniques, a controlled exercise fired alerts for 11; of the rest, 18 had no telemetry enabled at all, which no amount of rule writing would have fixed.

Where the coverage claim failed.

StageTechniques
alerted correctly11
telemetry missing18
rule did not match8
alert lacked context3

I would not consider it settled without evidence: Report detection coverage as techniques exercised and confirmed, not as rules written.

A rule with no telemetry beneath it detects nothing.

Curated: · Written: · Reviewed:

QA-49A requirement says data must stay in one jurisdiction. What actually enforces that?(show answer)

The first thing I would establish about regional and sovereignty constraints is what an attacker gains if the control is absent.

Choosing a region for the primary store is the easy part. Backups, replicas, logs, telemetry, support access, and managed-service control planes all move data or metadata, and several of them default to somewhere else.

Concretely, enumerate every place the data or its derivatives can land, restrict regions at the organisation level so a new resource cannot be created elsewhere, and check the managed services' own documentation for control-plane and support-access behaviour rather than assuming.

The reason for that specificity is a failure I have seen: A workload was correctly deployed in-region and its logs were delivered to a bucket in another jurisdiction by a default that nobody had changed, which the compliance attestation had not covered.

Where the data went.

DestinationIn jurisdictionChecked
primary storeyesyes
backupyesyes
audit logsnono
telemetrynono
of 4 destinations2 compliant2 checked

I would not consider it settled without evidence: Inventory every destination the data and its logs reach, and confirm each against the requirement.

The data moves in more places than the primary store.

Curated: · Written: · Reviewed:

QA-50You find a serious misconfiguration in another team's system. How do you raise it?(show answer)

I would start communicating a security finding from the permissions that are actually effective, not from the policy document.

A finding is only useful once someone acts on it, so how it is framed determines the outcome as much as the technical content. A finding delivered as an accusation, or as a scanner output with no context, produces defensiveness or is ignored.

Concretely, state the specific exposure and what an attacker gains, give the smallest change that closes it, offer to help implement it, and agree a date. Escalate on the date rather than at the discovery.

The reason for that specificity is a failure I have seen: A finding was filed as a ticket with a scanner rule identifier and no explanation; it sat for 4 months because the receiving team could not tell what it meant or how serious it was.

Remediation time by framing.

FramingMedian time to fix
scanner rule ID only4 months
exposure + impact3 weeks
+ specific fix offered4 days

I would not consider it settled without evidence: Measure time to remediation by how the finding was communicated, and treat a long queue as a communication problem before a compliance one.

A finding nobody understands is a finding nobody fixes.

Curated: · Written: · Reviewed:

QA-51A system verifies who the caller is on every request. What can still go wrong?(show answer)

This is an area where a passing compliance scan and a correct handling of authentication versus authorisation are different events.

Authentication establishes identity and says nothing about entitlement. The recurring failure is an endpoint that authenticates correctly and then acts on an object identifier supplied by the caller without checking that the caller owns it.

Concretely, make the authorisation check part of the data access rather than a separate step an endpoint can forget, fail closed when ownership cannot be determined, and test the negative case for every object-accessing path rather than the positive one.

The reason for that specificity is a failure I have seen: An endpoint authenticated the user and fetched the record by identifier; changing the identifier in the request returned another customer's record, and 340 endpoints shared the same pattern.

What each check catches.

CheckWrong passwordWrong object
authenticationstopsno
role checknopartly
ownership checknostops

I would not consider it settled without evidence: Test that an authenticated user cannot read another user's object by identifier, once per access path.

Knowing who they are does not say what they may reach.

Curated: · Written: · Reviewed:

QA-52How long should an access token live?(show answer)

My answer to token lifetime and refresh design begins with the blast radius: if I cannot state what one compromise reaches, I have no design.

Lifetime trades exposure against revocation cost. A short-lived token limits the window in which a stolen one is useful, at the cost of more refreshes; a long-lived one avoids the traffic and remains valid long after the session should have ended.

Concretely, keep access tokens short and refresh tokens revocable and bound to the client, check revocation at refresh rather than on every request, and bind tokens to a client or device where the sensitivity justifies it so a stolen token is not usable elsewhere.

The reason for that specificity is a failure I have seen: An access token with a 24-hour lifetime was stolen from a log; it remained valid for 19 hours after the account was disabled, because nothing in the request path consulted revocation.

Useful window after account disable.

Access token lifeRevocation checkedWindow
24 hat refresh onlyup to 24 h
15 minat refresh onlyup to 15 min
15 minper requestseconds

I would not consider it settled without evidence: Disable an account and measure how long an already-issued token continues to work.

The lifetime is how long a theft stays useful.

Curated: · Written: · Reviewed:

QA-53An internal service requests broad scopes because narrowing them is awkward. What is the argument against?(show answer)

I would treat OAuth scopes and consent in internal systems as a claim about an adversary's options that has to survive being tested.

A token is only as safe as its widest scope, and a broad token stolen from any of its holders grants everything it names. Scope narrowing is what limits what a leaked token can do, and it is the only control that operates after the leak.

Concretely, issue per-purpose tokens with the narrowest scope, use audience restriction so a token for one service is rejected by another, and prefer exchanging a broad token for a narrow one at the boundary rather than passing it onward.

The reason for that specificity is a failure I have seen: A broad token was passed from a gateway to a downstream service that logged request headers; the token was valid for every API the gateway could reach, not just the one being called.

What a leaked token reaches.

DesignAPIs reachable
broad, passed through14
audience-restricted1
exchanged at boundary1

I would not consider it settled without evidence: Present a token issued for one service to another and confirm the audience check rejects it.

A passed-through token grants everything it was ever for.

Curated: · Written: · Reviewed:

QA-54What causes most TLS certificate outages, and how do you prevent them?(show answer)

The useful question for certificate lifecycle and expiry is what still holds in the accounts nobody has looked at this year.

Certificates expire on a known date, so an expiry outage is a monitoring and automation failure rather than a surprise. The ones that cause outages are typically the certificates nobody knew existed — internal, on appliances, or issued outside the standard process.

Concretely, automate issuance and renewal, discover certificates from the network rather than from an inventory people maintain, alert well before expiry with an owner attached, and treat a certificate that cannot be automatically renewed as a risk to be removed.

The reason for that specificity is a failure I have seen: An internal service certificate issued by hand 2 years earlier expired on a Saturday; nothing monitored it because it was not in the inventory, and the outage lasted 5 hours.

Certificates by discovery method.

SourceCertificatesMonitored
inventory210210
network scan247210
unmonitored370

I would not consider it settled without evidence: Discover certificates by scanning your own estate and reconcile against the inventory; investigate every certificate the scan finds and the inventory does not.

The certificate that takes you down is the one not in the list.

Curated: · Written: · Reviewed:

QA-55When is mutual TLS worth its operational cost?(show answer)

I would settle mutual TLS between services by attempting the action the control is supposed to stop.

Mutual TLS authenticates both ends at the transport layer, which is valuable where the network is not trusted and where service identity must be established without an application-level token. It costs certificate lifecycle management for every workload, which is where it usually fails.

Concretely, adopt it where a service mesh or platform issues and rotates the certificates automatically, keep the identity short-lived, and do not adopt it manually across a large estate, since expiry management by hand becomes the dominant failure mode.

The reason for that specificity is a failure I have seen: A hand-managed mutual TLS rollout across 40 services produced 3 expiry outages in a year, more downtime than the threat it was introduced to address had ever caused.

Cost by issuance model.

IssuanceServicesExpiry outages/year
manual403
automated, 24 h certs400

I would not consider it settled without evidence: Confirm automated issuance and rotation exists before adopting mutual TLS, and measure certificate-related incidents afterwards.

Mutual TLS is a certificate lifecycle problem wearing a security hat.

Curated: · Written: · Reviewed:

QA-56How do you tell a genuine second layer from a duplicate first one?(show answer)

The judgement in defence in depth without theatre is which identity can assume what, not which box the checklist ticks.

A second layer is only independent if it fails for different reasons. Two controls that share an input, a configuration source, or a code path fail together, so the second one adds cost and no assurance.

Concretely, for each pair of controls, name what would have to fail for both to fail together; where the answer is a single shared element, the layers are one. Prefer controls at different levels — identity, network, data — over two at the same level.

The reason for that specificity is a failure I have seen: A network rule and an application check both derived their allowlist from one configuration file; a bad deploy of that file removed both at once, and the design had been described as defence in depth.

Independence check.

PairShared elementIndependentFailures to lose both
network rule + app checkconfig fileno1
identity policy + networknoneyes2
WAF + input validationnoneyes2

I would not consider it settled without evidence: Fault-inject the shared element and confirm at least one layer still holds.

Two controls sharing an input are one control.

Curated: · Written: · Reviewed:

QA-57A vendor says their product delivers zero trust. What would you ask?(show answer)

Where candidates lose the interview on zero trust in practice is reasoning from the architecture diagram.

Zero trust is a design principle — do not grant access on network location, verify explicitly per request, and assume breach — rather than a product. A single product can implement part of it and cannot deliver the principle, since the principle is about every access path.

Concretely, ask which access paths the product covers and which remain location-trusted, check that the legacy paths are actually removed rather than merely deprioritised, and measure the share of access decisions made with full identity and device context.

The reason for that specificity is a failure I have seen: A zero-trust rollout covered the web applications; the VPN remained, granting flat network access to anyone who connected, and 40 percent of access still flowed that way after the project was declared complete.

Access paths after the rollout.

PathLocation-trustedShare of access
web via proxyno60%
VPNyes38%
direct legacyyes2%

I would not consider it settled without evidence: Report the share of access paths that still grant on network location, and drive it down.

The old path is the one that gets used.

Curated: · Written: · Reviewed:

QA-58You have 4,000 open vulnerabilities. Which do you fix first?(show answer)

I would answer vulnerability management prioritisation by separating what is configured from what is enforced at request time.

Severity score alone is a property of the vulnerability rather than of your exposure. What matters is whether the vulnerable code is reachable, whether the asset is exposed, whether exploitation is observed in the wild, and what the asset holds.

Concretely, combine severity with reachability, internet exposure, known exploitation, and data classification, and work the resulting small set. Report the unprioritised backlog separately so it does not consume attention it does not deserve.

The reason for that specificity is a failure I have seen: A team worked strictly by severity score; a medium-severity flaw in an internet-facing service with observed exploitation sat unpatched for 6 weeks behind hundreds of high-severity findings in internal-only systems.

Same 4,000 findings, two rankings.

RankingTop-20 internet-facingTop-20 exploited in wild
severity only31
exposure-weighted2014

I would not consider it settled without evidence: Rank by exposure and known exploitation as well as severity, and check the top of the list against what an attacker would try first.

Severity describes the flaw and not your exposure to it.

Curated: · Written: · Reviewed:

QA-59Which patching is still yours when the provider manages the service?(show answer)

The engineering content of patching managed services is the detection and the revocation path, not the control's name.

The provider patches the service implementation on its own schedule, and the customer usually still chooses when to take a major version and is responsible for anything running inside — container images, function runtimes, and application dependencies.

Concretely, track end-of-support dates for every managed engine version and runtime you use, plan upgrades before forced migration rather than during it, and keep an inventory of runtimes so a deprecation announcement can be assessed in hours rather than weeks.

The reason for that specificity is a failure I have seen: A function runtime reached end of support; the provider stopped allowing updates to those functions, and 14 of them could not be changed until they were migrated, including one that needed an urgent security fix.

Estate against support dates.

ServiceResourcesWithin 6 months of EOL
function runtime6114
database engine123
container base2100

I would not consider it settled without evidence: Report the count of resources on versions within 6 months of end of support, per service.

Managed means patched by them, on their schedule, not yours.

Curated: · Written: · Reviewed:

QA-60Your DR environment is a copy of production. What does that mean for security?(show answer)

Before calling security of the disaster recovery environment done I would write down the attack it does not stop.

A DR environment holds the same data with the same sensitivity and usually receives a fraction of the attention, which makes it the softer of two doors into the same room. Controls applied to production and not to DR leave the data protected by whichever is weaker.

Concretely, apply the same controls, the same review cadence, and the same monitoring to DR, and include it in scanner coverage and in access reviews rather than treating it as non-production.

The reason for that specificity is a failure I have seen: A DR account held a full copy of customer data with an access list that had not been reviewed in 3 years; it was excluded from the posture tool because it was classified as non-production.

Coverage by environment.

EnvironmentSame dataIn scanner scopeAccess reviewed
productionyesyesquarterly
DRyesnonever

I would not consider it settled without evidence: Confirm DR accounts are in scanner scope and in access review scope, and check the access list against production's.

A copy of the data deserves a copy of the controls.

Curated: · Written: · Reviewed:

QA-61Should production data be used in a test environment?(show answer)

The first thing I would establish about non-production data is what an attacker gains if the control is absent.

Copying production data into a test environment copies the sensitivity and not the controls, and test environments are by design more accessible. The reason teams want it — realistic shape and volume — can usually be met without the real values.

Concretely, generate synthetic data that preserves distribution and referential integrity, or mask irreversibly at the point of copy rather than after, and where genuinely real data is required, apply production controls to that environment and say so.

The reason for that specificity is a failure I have seen: A production copy in a test account was accessible to 140 engineers and contained live payment identifiers; the masking script had failed silently 8 months earlier and nothing checked its output.

Exposure by approach.

ApproachReal identifiersPeople with access
production copyyes140
masked, uncheckedsometimes140
syntheticno140

I would not consider it settled without evidence: Sample the non-production dataset and confirm no real identifiers survive, on a schedule rather than at setup.

Copying the data copies the obligation.

Curated: · Written: · Reviewed:

QA-62Why classify incident severity before investigating?(show answer)

I would start incident severity classification from the permissions that are actually effective, not from the policy document.

Severity determines who is woken, what authority they have, and which clocks start, and those decisions cannot wait for the investigation to conclude. Classifying on current known impact with an explicit willingness to re-classify is what lets the response begin.

Concretely, define severities by observable impact rather than by cause, allow anyone to raise the severity and require a named role to lower it, and start the regulatory and contractual clocks at declaration since they usually run from discovery.

The reason for that specificity is a failure I have seen: An incident was held at low severity while the team established whether data had been accessed; by the time it was raised, 30 of the 72 available notification hours had elapsed.

Notification budget consumed before declaration.

Declared atHours elapsedBudget remaining
discovery072
after triage666
after root cause3042

I would not consider it settled without evidence: Rehearse the classification decision with an ambiguous scenario and confirm the clocks are understood to start at discovery.

Classify on impact known now, and revise.

Curated: · Written: · Reviewed:

QA-63What has to be true before an incident for a cloud investigation to be possible?(show answer)

This is an area where a passing compliance scan and a correct handling of forensic readiness in cloud are different events.

Cloud forensics depends on evidence that must already exist — logs with adequate retention, snapshot capability, and an account to work in — because the resources involved may be terminated by autoscaling before anyone looks.

Concretely, retain logs beyond the expected dwell time, keep an isolated forensics account with pre-agreed access, automate snapshot capture on a security alert so the evidence survives autoscaling, and know which provider logs are unavailable retroactively.

The reason for that specificity is a failure I have seen: An investigation began 5 days after the alert; the instance had been replaced by autoscaling on day 2, no snapshot had been taken, and the 7-day log retention covered only part of the intrusion.

Evidence available at day 5.

SourceRetentionAvailable
instance diskreplaced day 2no
memorynot capturedno
control-plane log7 dayspartly
data-plane lognot enabledno

I would not consider it settled without evidence: Run an investigation drill on a synthetic alert and record which evidence was actually obtainable.

Evidence has to exist before you know you need it.

Curated: · Written: · Reviewed:

QA-64What makes a security exercise more than a meeting?(show answer)

My answer to tabletop exercises that find something begins with the blast radius: if I cannot state what one compromise reaches, I have no design.

An exercise is useful when it forces real decisions against real systems under time pressure and produces findings that change something. A walkthrough of a document confirms the document exists.

Concretely, use a scenario nobody has prepared for, require participants to actually retrieve the evidence and run the commands rather than describe them, inject a complication partway, and record every point at which the answer was "I would look that up".

The reason for that specificity is a failure I have seen: An annual tabletop confirmed the plan was understood; the real incident revealed that nobody had the access to isolate an account out of hours, which the exercise had not required anyone to demonstrate.

Findings by exercise style.

StyleSteps performedGaps found
walkthrough01
hands-on, with injection149

I would not consider it settled without evidence: Require each step to be performed rather than described, and count the steps that could not be completed.

An exercise where nothing is attempted proves nothing works.

Curated: · Written: · Reviewed:

QA-65How do you keep security debt from becoming permanent?(show answer)

I would treat security debt and its owner as a claim about an adversary's options that has to survive being tested.

Security debt persists because its cost falls on a future incident rather than a current sprint, so it never competes successfully against work with a visible deadline. Making it visible in the same terms as other work is what changes that.

Concretely, record each item with the exposure it creates, an owner with the authority to fix it, and a review date; report items past their date as overdue; and express the risk in operational terms rather than as a severity label.

The reason for that specificity is a failure I have seen: A register of 90 items had a named owner for only 22 and a review date for only 14; 6 of them appeared in the post-incident analysis as contributing conditions, all raised more than a year earlier.

Register health.

PropertyCount
items open90
with named owner22
with review date14
appeared in an incident6

I would not consider it settled without evidence: Report the age distribution of open security debt and how many items appeared in incidents after being raised.

Debt without an owner and a date is a note.

Curated: · Written: · Reviewed:

QA-66A penetration test returns no critical findings. What does that tell you?(show answer)

The useful question for reading a penetration test report is what still holds in the accounts nobody has looked at this year.

It tells you what the testers found in the time and scope they had, against the systems they were pointed at. Scope, duration, and the level of access they were given determine the result at least as much as the security of the estate does.

Concretely, read the scope and the hours before the findings, note what was explicitly excluded, check whether the testers had credentials and at what privilege, and treat an unscoped system as untested rather than as clean.

The reason for that specificity is a failure I have seen: A clean report covered the customer-facing application; the identity plane and the CI system were out of scope, and the eventual intrusion came through CI.

What the engagement covered.

SystemIn scopeAssuredTester days spent
web applicationyesyes10
identity providernono0
CI systemnono0

I would not consider it settled without evidence: List what was excluded from scope and confirm those systems have their own assurance rather than inheriting the report's conclusion.

A clean report is a statement about the scope.

Curated: · Written: · Reviewed:

QA-67You are certified against a recognised framework. What does that establish?(show answer)

I would settle compliance frameworks and actual security by attempting the action the control is supposed to stop.

Certification establishes that a set of controls was designed and, for some frameworks, operating over a period, as assessed by a third party against a defined scope. It is evidence about process rather than about whether a determined attacker can get in.

Concretely, use the framework as a floor and a communication tool, keep the scope statement visible so nobody reads it as covering more than it does, and maintain a separate technical assurance programme aimed at the attacker rather than the auditor.

The reason for that specificity is a failure I have seen: A certification covering the production environment was cited in response to a customer's security question; the breach originated in a corporate system entirely outside the assessed scope.

Where the data lives against the scope.

SystemHolds customer dataIn certification scope
productionyesyes
analytics warehouseyesno
corporate file storeyesno
of 3 holding data31 in scope

I would not consider it settled without evidence: Read the scope statement alongside the certificate, and map it against where the sensitive data actually lives.

Certification describes a scope and a period.

Curated: · Written: · Reviewed:

QA-68A team wants to adopt a SaaS product that will hold customer data. What do you assess?(show answer)

The judgement in security review of a third-party SaaS is which identity can assume what, not which box the checklist ticks.

The assessment that matters is what the product can reach, how access to it is controlled, and what happens to the data if the vendor is breached or the relationship ends. A questionnaire answered by the vendor is evidence of their answers.

Concretely, check the authentication options — whether federation and phishing-resistant factors are available on the tier you are buying — the audit log availability, the data export and deletion path, the sub-processor list, and the breach notification terms.

The reason for that specificity is a failure I have seen: A product was adopted on a tier that did not support federation; 60 individual passwords were created, and when an employee left, their account remained active because it was outside the identity provider.

Capability by tier.

CapabilityStandard tierEnterprise tier
SSO federationnoyes
audit log exportnoyes
deletion on requestyesyes

I would not consider it settled without evidence: Confirm federation and audit log export are available on the specific tier being purchased, before signing.

Check the tier you are buying, not the feature list.

Curated: · Written: · Reviewed:

QA-69How do you address insider risk proportionately?(show answer)

Where candidates lose the interview on insider risk without surveillance is reasoning from the architecture diagram.

Most insider incidents are mistakes and departures rather than malice, so controls that assume malice are both disproportionate and poorly targeted. Access minimisation and good leaver processes address more of the actual risk than monitoring does.

Concretely, reduce standing access, require elevation with a reason, separate duties for the highest-consequence actions, and monitor high-signal events rather than general behaviour, keeping monitoring proportionate and transparent.

The reason for that specificity is a failure I have seen: A team deployed broad behavioural monitoring while 12 engineers held standing production administrative access; the eventual incident was an accidental deletion by one of the 12, which access minimisation would have prevented and monitoring only recorded.

Where the risk actually was.

ControlAddresses accidentAddresses maliceCost
access minimisationyesyeslow
separation of dutiesyesyesmedium
behavioural monitoringnopartlyhigh

I would not consider it settled without evidence: Compare standing privileged access against monitoring investment, and address the access first.

Minimise what an insider holds before watching what they do.

Curated: · Written: · Reviewed:

QA-70Why is the build environment often the weakest part of a well-secured estate?(show answer)

I would answer security of the build environment by separating what is configured from what is enforced at request time.

It runs untrusted code by design, holds credentials to production, and is usually exempted from the controls applied elsewhere because those controls slow builds down. That combination makes it the highest-value and least-defended target.

Concretely, isolate build execution from deploy credentials, run untrusted builds in ephemeral isolated environments, restrict egress from builders, require review on pipeline definition changes, and apply the same logging and access review as production.

The reason for that specificity is a failure I have seen: A dependency's post-install script ran in a shared build environment with production credentials in the environment; it exfiltrated them, and the build system had no egress restriction to record the destination.

Build environment posture.

PropertyTypicalHardened
runs untrusted codeyesyes, isolated
holds prod credentialsyesno
egress restrictednoyes
access reviewednoquarterly
controls present of 404

I would not consider it settled without evidence: Attempt an outbound connection from a build to an unapproved destination and confirm it is blocked or recorded.

The build system runs strangers' code next to your credentials.

Curated: · Written: · Reviewed:

QA-71What is in an infrastructure state file and why does it matter?(show answer)

The engineering content of protecting the terraform state is the detection and the revocation path, not the control's name.

State contains the full configuration of the managed estate and, in many cases, the plaintext of resource attributes marked sensitive — generated passwords and keys among them. It is a credential store that most teams treat as a build artefact.

Concretely, store state in an encrypted backend with access restricted to the pipeline identity, never in a repository, avoid generating secrets in the configuration where possible, and audit reads of the state bucket as you would a secret store.

The reason for that specificity is a failure I have seen: A state file in a repository contained a generated database password in plaintext; it had been readable by every engineer for 14 months and was never rotated because nobody thought of state as holding secrets.

What state contained.

AttributePlaintext in stateRotated
generated DB passwordyesno
API key resourceyesno
KMS key materialnon/a

I would not consider it settled without evidence: Search state for sensitive attributes and confirm the storage location's access list matches a secret store's.

State is a secret store nobody labelled as one.

Curated: · Written: · Reviewed:

QA-72Can a billing anomaly indicate a compromise?(show answer)

Before calling cloud cost as a security signal done I would write down the attack it does not stop.

Cryptomining and data exfiltration both cost money, and billing is sometimes the first control-plane-independent signal available. It is a slow detector and a useful corroborator, particularly in accounts nobody watches closely.

Concretely, alert on unusual spend per account and per service with a short evaluation window, treat compute spend in an account with no compute workloads as an incident signal rather than a finance one, and route the alert to security as well as to finance.

The reason for that specificity is a failure I have seen: An unused account accrued 40,000 dollars of compute over 5 weeks; the finance team queried it at month end, and the mining had been running since the first week.

Detection latency by route.

RouteEvaluationDetected after
monthly finance reviewmonthly5 weeks
daily anomaly alertdaily2 days
hourly, per accounthourly4 h

I would not consider it settled without evidence: Route per-account spend anomalies to security with a daily evaluation, and test with a controlled spike.

Someone is paying for the attacker's compute.

Curated: · Written: · Reviewed:

QA-73What is the security risk in a resource nobody uses?(show answer)

The first thing I would establish about decommissioning cloud resources is what an attacker gains if the control is absent.

An unused resource keeps its permissions, its network exposure, and its data while losing the attention that would notice a change. Abandoned accounts and forgotten subscriptions are consistently where intrusions begin.

Concretely, require an owner on every account and resource, expire ownership so an unclaimed resource is flagged, and run a decommissioning process that removes credentials and data rather than merely stopping compute.

The reason for that specificity is a failure I have seen: A project account from a cancelled initiative retained a role trusted by production and a snapshot of the customer database; nobody had owned it for 2 years and it was not in the posture tool's scope.

Estate by ownership.

AccountsOwner respondsIn scanner scope
31yesyes
9noyes
4nono

I would not consider it settled without evidence: Reconcile the account inventory against active ownership and investigate every account with no owner responding.

Unowned resources keep their permissions and lose their supervision.

Curated: · Written: · Reviewed:

QA-74What security requirements would you attach to a new internet-facing service?(show answer)

I would start security requirements for a new service from the permissions that are actually effective, not from the policy document.

The requirements worth stating are the ones that are cheap now and expensive later: identity, logging, data handling, and the boundary. Anything that can be added afterwards without redesign does not need to be a launch requirement.

Concretely, require federated identity with a phishing-resistant factor for administrative access, control-plane and data-plane logging to immutable storage, a stated data classification with matching key policy, egress restriction, and a named owner for the account.

The reason for that specificity is a failure I have seen: A service launched without data-plane logging because it could be enabled later; the later incident needed the history from before it was enabled, which did not exist.

Cost of adding after launch.

RequirementCost at designCost after launch
data-plane logging10 minhistory unrecoverable
federated identity4 hmigration of 1,200 users
egress restriction6 h3 weeks of discovery

I would not consider it settled without evidence: Check each requirement against the running service rather than the design document, before launch.

The cheap-now, expensive-later items are the launch requirements.

Curated: · Written: · Reviewed:

QA-75A team needs to do something that breaks a security standard to meet a deadline. What do you do?(show answer)

This is an area where a passing compliance scan and a correct handling of saying no to a business request are different events.

A refusal with no alternative moves the work outside your visibility rather than stopping it. The useful response finds the smallest deviation that meets the need, bounds it, and makes the risk owner explicit.

Concretely, understand the actual requirement rather than the proposed solution, offer the narrowest exception with a compensating control and an expiry, name the accepting owner, and follow up on the expiry rather than letting it lapse.

The reason for that specificity is a failure I have seen: A blanket refusal led the team to provision in a personal cloud account; the workload held customer data outside every control and was found 7 months later through a billing query.

Outcome by response.

ResponseIn scopeControls appliedExpiry
refusenononen/a
bounded exceptionyescompensating90 days

I would not consider it settled without evidence: Track exceptions granted and expired, and check that refusals did not simply relocate the work.

A refusal without an alternative relocates the risk.

Curated: · Written: · Reviewed:

QA-76A web tier calls an app tier that calls a managed database, and every call is authorised. Where do the network boundaries go so one compromised tier cannot reach the rest of the estate?(show answer)

My answer to network boundaries and lateral movement begins with the blast radius: if I cannot state what one compromise reaches, I have no design.

Authorisation governs which caller may invoke an API; it never governs which packets reach a socket, and some services have none at all. Network isolation bounds blast radius for an attacker already inside a tier, so design for them, not just for legitimate traffic.

Concretely, give each tier its own security group whose inbound rules name the source group as peer, not a VPC CIDR; expose the database only via a private endpoint with public access off, and keep the implicit default deny. Test reachability from inside each tier — a diagram shows intent, not paths.

The reason for that specificity is a failure I have seen: An app-tier security group still carried a temporary 'allow all from 10.0.0.0/8' rule left over from a migration three years earlier. When the app host was compromised through a deserialization flaw, the attacker swept the VPC in under a minute, reached 1,847 hosts, and found a Redis instance with no authentication on 6379 holding 2.1 GB of cached session tokens. The workload's IAM policy correctly limited it to two services and was never consulted, because Redis had no authoriser to consult.

Reachable targets from a compromised app-tier host.

Rule setHosts reachableDatabase exposure
allow all from 10.0.0.0/81,847public and private endpoints
per-tier groups, private endpoint, public access off3private endpoint only

I would not consider it settled without evidence: Name the reachability test you would run from a host inside each tier and what passing it looks like, and name the setting that proves the database has no public path — the endpoint resolving to private addresses with public access disabled on the service itself.

Identity answers who may call; the network answers what can be reached at all.

Curated: · Written: · Reviewed:

QA-77A service needs a database password: where does it live, who can read it, and how does it change?(show answer)

I would treat secret storage and rotation as a claim about an adversary's options that has to survive being tested.

A managed store earns its place through identity-scoped reads, an audit trail, and rotation with no redeploy — not encryption at rest, which every option already has.

Concretely, keep the secret in a managed store — AWS Secrets Manager, GCP Secret Manager or Azure Key Vault — fetch it at runtime under the workload's own identity, and grant read to that identity alone plus break-glass. Rotate so the new version publishes while the old still works, and measure trigger-to-cutover time.

The reason for that specificity is a failure I have seen: A database password was handed to 22 services as a CI-injected environment variable. When it surfaced in a contractor's exported VM image, rotation meant editing 22 pipeline definitions and redeploying every service: 7 hours of coordinated work across three accounts. The password had last been changed 22 months earlier and stayed valid for that entire window.

What each placement costs when the secret leaks.

PlacementWho can read itRotation costReads attributable
in repo or image layeranyone with clone or registry pullrebuild and redeploy every consumerno
CI variable injected as env varevery process in the workload, and anything that dumps the environmentedit every pipeline, redeploy every consumerrarely
managed store, workload identitythe workload's own role plus named break-glass principalspublish a new version; consumers pick it up on next fetchyes, per read in the store's audit log

I would not consider it settled without evidence: Name one production secret, show the date it last changed and the measured time one full rotation takes end to end, then name the single identity authorised to read it.

The value of a secret store is that changing the secret is cheap.

Curated: · Written: · Reviewed:

QA-78A developer's long-lived AWS access key for a production account is found in a public GitHub repository. What happens in the first hour, what happens by the end of the first day, and what tells you the incident is over?(show answer)

The useful question for leaked long-lived cloud access key is what still holds in the accounts nobody has looked at this year.

Disabling the key is a ten-minute action; the incident is whatever that identity could reach and did reach while it was valid — the key is a symptom, the account-level compromise assessment is the real question. Revoking it stops future use, not the sessions it already minted or the access it established.

Concretely, start by disabling the key rather than deleting it, revoke the temporary sessions it already minted, and rotate every other secret in the same commit; delete the key once nothing depends on it. Walk CloudTrail from the commit timestamp on that key ID to bound what it read and which identities it created or roles it assumed, treat each of those as compromised until evidence clears it, and retire the standing key in favour of SSO or role assumption before you close.

The reason for that specificity is a failure I have seen: A key committed at 09:14 UTC was reported by an external researcher five days later; the team deleted it within six minutes and closed the ticket. Eleven days after that, a bucket access review surfaced 2.4 TB pulled from an analytics bucket between 02:10 and 05:40 on the night after the commit, and CloudTrail held iam:CreateUser and iam:CreateAccessKey calls 41 minutes after the commit — a second access key belonging to the attacker's new user was still valid and appeared nowhere in the incident report.

What each containment step stops, and what it leaves open.

StepWhat it stopsWhat it leaves open
disable the key (Status=Inactive)new requests signed with the secretsessions it already minted: AssumeRole 15 min–12 h (default 1 h), GetSessionToken up to 12 h, 36 h with MFA
deny credentials issued before now (aws:TokenIssueTime)live temporary credentials from that keycalls already made, and anything minted outside that key
delete the key and scrub the committhe secret at its originclones and forks already taken; public repos are scraped by bots within minutes
walk CloudTrail on the key ID from commit timenothing — this is measurementanything before trail retention, outside its regions, or in an event type it does not record
hunt persistence: new users, keys, roles, trust policies, Lambda, CloudTrail stopswhat the attacker left runningthe incident stays open until this list is empty and evidenced

I would not consider it settled without evidence: Show the CloudTrail query that filters events on that access key ID, name the timestamps you correlate — commit time, key creation, first and last use, and the creation time of every identity created inside the window — then name the check that proves no session it minted is still live.

Deleting the key closes the credential; proving what it touched closes the incident.

Curated: · Written: · Reviewed:

QA-79An access key attached to an identity with broad read and write across production has been in a public repository for six weeks. How do you size the blast radius?(show answer)

I would settle blast radius of a leaked access key by attempting the action the control is supposed to stop.

Permissions bound what could have happened; the audit trail bounds what you can prove happened, and the gap between them is unmeasured exposure. Scope starts from every account the identity can reach through role assumption, not from the repository or the account that issued the key.

Concretely, start by deactivating the key and exporting the audit trail before anything else, so the window stops growing and the evidence survives your own response. Then measure worst case from the identity's reach in every account and truth from events on that key ID since the commit timestamp, report the two numbers separately, and replace the standing key with short-lived credentials.

The reason for that specificity is a failure I have seen: An access key on a CI identity had been public for 19 days. IAM showed no last-used date and the credential report could not be regenerated for another three hours, so the team reported "no evidence of use" and spent two days on a least-privilege rewrite instead of a compromise assessment. The identity could assume a deploy role in the data account through a trust policy with no external ID and no source-account condition, and that account's S3 data events had been switched off to save roughly $400 a month in CloudTrail charges. When egress billing later showed 1.1 TB leaving an analytics VPC inside the window, there was no object-level record to attribute it or exclude it. Silence had been filed as a finding.

Two bounds on the same exposure: what each measurement source actually establishes.

SourceEstablishesBlind spot
IAM evaluation: identity policy, boundary, SCP, session policy, then every cross-account trustthe upper bound — actions and resources the key could reach, per accountwhether any of it happened
CloudTrail management events (Event history keeps 90 days; trails keep what you retain)calls signed with that key ID: AssumeRole, CreateUser, PutBucketPolicydata-plane calls, regions or services outside the trail, anything past retention
CloudTrail data events (S3 object, Lambda invoke) — off by default, list price ≈ $0.10 per 100,000 recordedwhich objects were read and which functions were invokednothing at all if never enabled; silence is not evidence
GetAccessKeyLastUsed and the credential reportkey age, last-used date, service and region, whether it was ever rotateda report you can regenerate only once every four hours
VPC flow logs and NAT/egress billingdata leaving the estate — often the first signal when the trail is thinattribution to a key once the session assumed other roles

I would not consider it settled without evidence: Show the two bounds side by side: the CloudTrail query that filters events on the access key ID from the commit timestamp, and the permission evaluation that lists every account and role the identity can reach. Then name the services where no data-plane event exists at all, so the report states an upper bound and a proven set rather than one confident number.

Permissions bound the worst case, evidence bounds the truth — report both or you have mis-sized the incident.

Curated: · Written: · Reviewed:

QA-80How do you keep service and pipeline secrets out of code and configuration, and when should a static secret be replaced by a short-lived credential?(show answer)

The judgement in secret sprawl and credential lifetime is which identity can assume what, not which box the checklist ticks.

A value copied into image layers, state files and logs is a static secret wherever it lives, and it stays valid until something revokes it. Replace it with a short-lived credential once the consumer can prove its identity to the issuer — the exposure window is bounded by the session, not by rotation.

Concretely, start by inventorying every surface a secret landed on — git history, image layers, retained IaC state, CI logs — then move each value to a managed store with runtime injection and a named owner. Replace pipeline static keys with OIDC federation pinned to one repository and ref.

The reason for that specificity is a failure I have seen: A team 'fixed' secret sprawl by moving a static deploy access key from GitHub Actions variables into Secrets Manager and rotating it on a 30-day schedule. The same key had been passed into a Terraform provider block in 2021 and so sat in 74 retained versions of the state file in a versioned S3 bucket that roughly 60 engineers could read; sensitive = true had redacted the plan output and nobody had looked inside the state itself. A scanner then found the key in a public fork. The responder rotated it, but rotation only rewrote the Secrets Manager value, so three of nine consumers kept retrying with the dead credential and two undocumented ones — a nightly batch and a contractor's laptop — kept working on the old key for 11 days until it was deleted outright. Sessions the key had minted before deactivation stayed valid for up to an hour afterwards. The lesson the team drew was that rotation works; the accurate reading is that one credential had nine consumers and nobody held the list.

The same leaked value under five credential designs: how far it spreads and what ends the exposure.

DesignWhere the value copies toStill valid after a leakWhat ends the exposure
Static key in the repositorygit history, forks, clones, CI cachesuntil the key is deleted — weeks if nobody owns itmanual rotation, and only if every consumer is known
Static key in a vault, injected at runtimejob memory, debug logs, image layers if baked in, IaC state if passed to a providerunchanged — the value is identical wherever it landsthe same rotation; the vault adds an audit trail, not a shorter life
Static key with 7-day automated rotationas above, but each copy dies at the next rotationup to 7 days, plus any session minted before deactivation (valid to its own expiry)rotation — provided no consumer caches and none are missed
OIDC federation: pipeline assumes a cloud rolenothing durable; the run holds an STS sessionrun length plus the session TTL (AssumeRoleWithWebIdentity accepts 900–43,200 s, capped by the role's max session duration, default 3,600 s)editing the trust policy or deleting the role revokes every future session
Dynamic database credential issued by the vaultprocess memory onlythe lease TTL (e.g. 1 hour)revoking the lease kills the live credential immediately

I would not consider it settled without evidence: Name the sweep that proves the value is gone from every copy surface — a history scan of each repository, a layer-by-layer scan of every pushed image, and a read of every retained state version — and show the IAM query that returns zero static keys on any CI identity alongside the OIDC trust policy's aud and sub conditions.

Judge a credential by how fast it stops being valid and how far it reaches, not by which vault it was sitting in.

Curated: · Written: · Reviewed:

QA-81When do you let the provider hold the encryption keys, and when do you bring or hold your own — and what does each option actually buy you?(show answer)

Two questions hide in this one: who can authorise decryption, and who ends up holding plaintext. Assume the data sits in a managed store and the workloads reading it run in the provider's compute — that assumption decides most of it.

Provider-managed key (SSE-S3, Google default, Azure PMK)Customer-managed / BYOK (KMS CMK, Cloud KMS, Key Vault)HYOK / external key manager (KMS custom key store + CloudHSM, AWS XKS, GCP EKM)
Protects againstlost or scrapped media, snapshots crossing a boundarysame, plus misuse of keys by your own adminssame, plus provider-side access to key material
Separation of dutiesnonekey admin ≠ data reader, enforced in the key policyyou hold the authorisation point
Key-use auditnothing attributable to your identities (GCP states Google-managed key use is not in Cloud Audit Logs)every GenerateDataKey/Decrypt with caller identityyour HSM's log plus the cloud-side call log
Revocationnonedisable or schedule deletion → crypto-shredstop serving key unwrap
Costzero opspolicy discipline, quota managementHSM fleet in the request path

Default keys buy encryption at rest and nothing about identity: whoever the service lets read the object gets plaintext, because the service decrypts for them. Against stolen credentials or a confused-deputy read they are worth nothing.

A customer-managed key earns its keep in the key policy — the one place that says who may decrypt, independent of who may read the blob:

{ "Sid": "WorkloadMayDecrypt", "Effect": "Allow",
  "Principal": {"AWS": "arn:aws:iam::111122223333:role/orders-api"},
  "Action": ["kms:Decrypt", "kms:GenerateDataKey"], "Resource": "*",
  "Condition": {"StringEquals": {"aws:SourceVpce": "vpce-08a1..."}} }

Key administration — PutKeyPolicy, ScheduleKeyDeletion — goes to a separate break-glass role, so the team running the database cannot rewrite who decrypts it. That split is the separation of duties; an account boundary alone does not give it to you.

Rotation and revocation do different jobs. AWS KMS rotates a symmetric customer-managed key on a schedule (annual by default) and keeps the old backing key, so existing ciphertext still decrypts: rotation bounds how long a stolen data key is useful for new data and re-encrypts nothing. Only revocation makes existing ciphertext unreadable — disable the key (immediate, reversible) or schedule deletion (AWS KMS 7–30 days, default 30; Azure Key Vault soft-delete retention 7–90 days, default 30). Plan for that lag. The same key covers replicas and backups, so crypto-shredding a departing customer's key takes those too: the point in an offboarding, a catastrophe in a mistake.

With your own keys you get a key-use trail — CloudTrail, Key Vault diagnostics or Cloud KMS audit logs record each call with the caller's identity. The alert that matters is a change of pattern: a role that never called kms:Decrypt doing 40,000 of them in an hour. With provider-managed keys there is nothing to alert on.

Cost is real. Every envelope operation is a call to the key service: 2,000 object reads per second, each decrypting a per-object data key, is 2,000 Decrypt/s against a per-account, per-region quota (AWS documents 5,500 symmetric operations/s; raisable via support). Data key caching — one data key reused for 60 seconds, capped at 100 messages per entry — takes that to 20 Decrypt/s, at the price of holding a plaintext data key in memory for up to a minute. HYOK adds a network hop to your HSM on every unwrap and makes that HSM a tier-0 dependency: when it is down, reads fail and the data is unreadable. Multi-region HSMs and a tested break-glass, or you have swapped a security question for an outage.

The misconception is "we hold the keys, so the provider cannot read our data." The workload decrypts inside the provider's compute; a compromised workload, stolen role credentials or plaintext in logs read exactly the same whichever option you chose. BYOK buys authorisation control, separation of duties, key-use audit and crypto-shredding. HYOK additionally buys the defensible claim that the provider cannot decrypt without your key manager — that is a key-custody regulation or a contractual commitment, not general hardening.

Failure modes I have seen: a key policy allowing kms:Decrypt on Resource * so any role that can read the blob can decrypt it; rotation scheduled with no re-encrypt job, leaving old ciphertext under old material; a deletion scheduled during an incident that also destroyed the backups; and a Lambda fan-out throttling on KMS at peak because nobody counted the Decrypt calls.

I would not call it decided without the key policy diff, a test proving a data-reader role cannot decrypt outside the workload path, and the measured Decrypt rate at peak against the quota.

Curated: · Written: · Reviewed:

QA-82Role-based or attribute-based access control?(show answer)

I would answer designing an authorisation model by separating what is configured from what is enforced at request time.

Roles are simple to reason about and multiply as exceptions accumulate; attributes express fine-grained rules and are harder to audit because the effective grant depends on data. Most real systems end up with roles for coarse structure and attributes for the specific conditions.

Concretely, keep roles few and meaningful, express genuinely data-dependent rules as attributes with the policy in one place rather than scattered through the code, and make the effective decision queryable so a review can answer who can reach what.

The reason for that specificity is a failure I have seen: A role-only model grew to 240 roles, most held by one person each, and the quarterly review became a formality because nobody could reason about the set.

Model growth over 3 years.

ModelRolesRoles held by 1 personReview feasible
role-only240180no
roles + attributes140yes

I would not consider it settled without evidence: Count roles against people and check that a reviewer can state what each role grants without reading policy documents.

Roles multiply at the rate you add exceptions.

Curated: · Written: · Reviewed:

QA-83Explain the confused deputy problem in a cloud context.(show answer)

The engineering content of the confused deputy is the detection and the revocation path, not the control's name.

A privileged service acting on a request from a less-privileged party can be induced to use its own authority on that party's behalf. The service is not compromised; it is doing exactly what it was asked, without checking whether the requester was entitled to the effect.

Concretely, have the deputy check the requester's entitlement to the specific resource rather than only its own, use an external identifier so a third party's request cannot be replayed by another of their customers, and pass the requester's identity through rather than acting purely as the deputy.

The reason for that specificity is a failure I have seen: A logging service with write access to every customer bucket accepted a bucket name from the caller; one customer supplied another customer's bucket name and the service wrote to it with its own permissions.

What the deputy verified.

CheckPresentStops cross-customer write
deputy has permissionyesno
requester authenticatedyesno
requester owns resourcenoyes
of 3 checks2 present1 that mattered

I would not consider it settled without evidence: Test that the deputy refuses a request naming a resource the requester does not own.

The deputy's permissions become the requester's unless it checks.

Curated: · Written: · Reviewed:

QA-84A control checks a condition and then acts on it. What can go wrong?(show answer)

Before calling time-of-check to time-of-use done I would write down the attack it does not stop.

Any gap between checking a condition and relying on it is a window in which the condition can change. In cloud systems the window is often large — a policy evaluated at request time and a resource modified moments later — and the check's result becomes stale rather than wrong.

Concretely, make the check and the action atomic where the platform supports it, use conditional writes with a precondition, and where a gap is unavoidable, re-verify at the point of use rather than relying on the earlier result.

The reason for that specificity is a failure I have seen: An authorisation check read the resource's owner tag, and the tag was changed between the check and the write; the write proceeded under an authorisation that was no longer valid.

Window between check and use.

DesignWindowRace exploitable
check then write40-300 msyes
conditional write0no
re-verify at useunder 1 msrarely

I would not consider it settled without evidence: Test the race deliberately by changing the condition between check and use, and confirm the action fails.

A check is true at the moment it ran.

Curated: · Written: · Reviewed:

QA-85Your authorisation service is unreachable. What should the application do?(show answer)

The first thing I would establish about failing closed is what an attacker gains if the control is absent.

Failing open keeps the system available and removes the control at exactly the moment when something is already wrong. Failing closed preserves the control at the cost of availability, and the choice must be deliberate rather than an accident of the error handling.

Concretely, default to closed for authorisation, use a short-lived local cache of prior decisions to ride out brief outages rather than falling through, alarm loudly when the cache is in use, and make the deliberate exceptions explicit and few.

The reason for that specificity is a failure I have seen: An exception handler around the authorisation call logged and continued; during a 20-minute outage every request was treated as authorised, and the behaviour had been in place for 2 years untested.

Behaviour during a 20-minute outage.

DesignRequests authorisedAvailability
fail openallfull
fail closednonenone
cached decisions + closedprior userspartial

I would not consider it settled without evidence: Make the authorisation service unreachable in a lower environment and observe what the application actually does.

An unhandled error is a decision nobody made.

Curated: · Written: · Reviewed:

QA-86How do secrets end up in logs, and how do you stop it?(show answer)

I would start logging what should not be logged from the permissions that are actually effective, not from the policy document.

Secrets reach logs through generic mechanisms rather than deliberate logging: whole-request dumps, exception serialisation that includes arguments, and debug modes left on. Redaction lists fail because they enumerate known names and the leak comes from an unanticipated path.

Concretely, wrap secret values in a type whose string representation is redacted so the leak is impossible rather than filtered, keep whole-object serialisation out of the logging path, and scan logs for high-entropy strings as a detective backstop.

The reason for that specificity is a failure I have seen: An exception handler serialised the request object including the authorisation header; the redaction list covered a field named password and not one named authorization.

Leak paths and what catches them.

PathRedaction listRedacting type
explicit log callcatchescatches
exception serialisationmissescatches
debug request dumpmissescatches
paths covered of 313

I would not consider it settled without evidence: Trigger errors on each path in a lower environment and scan the output for known test secrets.

A redaction list covers the names you thought of.

Curated: · Written: · Reviewed:

QA-87Why can an internal package name be dangerous?(show answer)

This is an area where a passing compliance scan and a correct handling of dependency confusion are different events.

When a resolver is configured with both a public and a private source, a package published publicly under an internal name can be preferred if its version is higher. The attack requires only knowing the name, which appears in every artefact that ships the dependency list.

Concretely, reserve internal names publicly or use a scope you control, configure the resolver to source internal names only from the internal registry rather than falling through, and pin by digest so a substituted package fails verification.

The reason for that specificity is a failure I have seen: An internal utility name was published publicly at a high version number; builds resolved to the public package for 3 days before anyone noticed, and it ran with the build system's credentials.

Resolution behaviour.

ConfigurationPublic higher version wins
both sources, version priorityyes
scoped to internal registryno
pinned by digestno

I would not consider it settled without evidence: Confirm the resolver refuses to fetch an internally-scoped name from a public source, by attempting it.

Your dependency list is a list of names an attacker can register.

Curated: · Written: · Reviewed:

QA-88Artifacts are signed. Where does that stop being useful?(show answer)

My answer to signing and verifying at the right point begins with the blast radius: if I cannot state what one compromise reaches, I have no design.

A signature is only worth what the verification is. Signing without verifying at admission, or verifying with a key the build system itself controls, produces a chain that an attacker inside the build can complete.

Concretely, verify at the point of use rather than only in the pipeline, keep signing keys outside the reach of the code being built, and check that verification fails closed rather than logging a warning and proceeding.

The reason for that specificity is a failure I have seen: Images were signed and verification was configured in warn mode during a rollout; it stayed in warn mode for 8 months, and an unsigned image deployed without objection.

Enforcement state.

StageConfiguredEnforcing
signingyesyes
pipeline checkyesyes
admissionyeswarn only

I would not consider it settled without evidence: Deploy a deliberately unsigned artefact and confirm it is refused rather than warned about.

A signature with no enforcing verifier is metadata.

Curated: · Written: · Reviewed:

QA-89A dependency has a critical vulnerability. Do you have to fix it now?(show answer)

I would treat reachability in dependency findings as a claim about an adversary's options that has to survive being tested.

The finding is about the package; your exposure depends on whether the vulnerable function is reachable from your code and whether the path can be driven by untrusted input. Treating every finding as equally urgent means the genuinely urgent one waits its turn.

Concretely, use call-graph reachability analysis to separate reachable from present, prioritise reachable findings on internet-facing services, and record the reason a finding was deprioritised so the decision survives the next scan.

The reason for that specificity is a failure I have seen: A team upgraded 340 packages for findings of which 12 were reachable; the work took 3 weeks and one of the 12, on a public endpoint, was done in week 3.

Findings by reachability.

CategoryCountMedian fix priority
present, unreachable328routine
reachable, internal8days
reachable, internet-facing4hours

I would not consider it settled without evidence: Report reachable findings separately and work them first, with the unreachable set tracked but not urgent.

Present is not reachable, and reachable is what you have.

Curated: · Written: · Reviewed:

QA-90How would you detect that a stored object was modified?(show answer)

The useful question for hashing and integrity for stored objects is what still holds in the accounts nobody has looked at this year.

Storage-level checksums detect corruption and not modification by an authorised writer, because an attacker with write access updates the checksum too. Detecting deliberate modification needs a signature or a hash recorded somewhere the writer cannot reach.

Concretely, record content hashes in a separate store with different write permissions, or sign objects at write with a key the storage identity does not hold, and verify at read for the classes of data where silent modification matters.

The reason for that specificity is a failure I have seen: An audit archive relied on the storage service's own checksums; an identity with write access altered records and the checksums updated with them, leaving no evidence of the change.

What detects a modification.

MechanismCorruptionAuthorised writerThreats detected of 2
storage checksumyesno1
hash in separate storeyesyes2
signature, external keyyesyes2

I would not consider it settled without evidence: Modify an object with the storage identity and confirm the independent record detects it.

A checksum the writer can update detects accidents only.

Curated: · Written: · Reviewed:

QA-91How would you decide which of a hundred findings actually matters?(show answer)

I would settle graph reasoning about attack paths by attempting the action the control is supposed to stop.

Findings in isolation are conditions; what matters is whether they compose into a path from something an attacker can reach to something they want. A path of three medium findings can be worse than an isolated critical one.

Concretely, model the estate as a graph of identities, resources, and network reachability, run shortest-path analysis from external entry points to sensitive assets, and prioritise the findings that appear on the shortest paths rather than by individual severity.

The reason for that specificity is a failure I have seen: A hundred findings were worked by severity; the actual intrusion used three findings rated medium that composed into a two-hop path from a public endpoint to the customer database.

Findings on the shortest path.

Path lengthFindings involvedIndividual severity
2 hops3all medium
4 hops1critical
unreachable61mixed

I would not consider it settled without evidence: Report the shortest attack paths to your most sensitive assets and the findings on them.

Attackers compose findings; scanners rank them separately.

Curated: · Written: · Reviewed:

QA-92Your automated remediation ran twice on the same finding. What should have happened?(show answer)

The judgement in queues and idempotency in security automation is which identity can assume what, not which box the checklist ticks.

Security automation acts on live infrastructure, so a duplicate execution is a duplicate change. Automation that quarantines, revokes, or deletes must be idempotent, because the alerting path that triggers it is at-least-once by nature.

Concretely, key remediation on the finding identity, record the outcome so a repeat is recognised, make each action safe to repeat, and require a human decision for anything genuinely destructive rather than automating it.

The reason for that specificity is a failure I have seen: An automation revoked a role's policy on a finding and ran again on a duplicate event; the second run removed a policy that had been re-added correctly in between, and it took an hour to work out what had happened.

Second execution outcome.

DesignSecond runCollateralMinutes to diagnose
no keyrepeats actionreverts valid change60
keyed, outcome storedno-opnone0

I would not consider it settled without evidence: Replay the same finding event and confirm the second execution is a no-op.

At-least-once delivery means your remediation runs twice.

Curated: · Written: · Reviewed:

QA-93What makes a security runbook usable at three in the morning?(show answer)

Where candidates lose the interview on writing a security runbook people can follow is reasoning from the architecture diagram.

A runbook is read by someone tired, under pressure, and possibly unfamiliar with the system. Prose describing an approach fails in that condition; exact commands with expected output succeed.

Concretely, write numbered steps with the exact command, the expected result, and what to do when it differs; state the decision points and who has authority; and test it by having someone who did not write it follow it during a drill.

The reason for that specificity is a failure I have seen: A runbook said to isolate the affected account; during the incident nobody knew which command did that or who could approve it, and 40 minutes went to finding out.

Drill outcome by runbook style.

StyleSteps completed unaidedTime
prose description4 of 1240 min
exact commands12 of 126 min

I would not consider it settled without evidence: Have someone unfamiliar with the system execute the runbook during a drill and record every step they could not complete.

Write it for the person who has never done it.

Curated: · Written: · Reviewed:

QA-94A customer sends a 300-question security questionnaire. How do you handle it well?(show answer)

I would answer customer security questionnaires by separating what is configured from what is enforced at request time.

The questionnaire is a proxy for the questions the customer actually cares about, and answering it accurately matters more than answering it favourably, because an inaccurate answer becomes a contractual representation.

Concretely, maintain a reviewed answer library tied to evidence, answer accurately including where a control is partial with the compensating measure stated, and route the genuinely novel questions to the people who know rather than approximating.

The reason for that specificity is a failure I have seen: An answer stating that all data was encrypted with customer-managed keys was true for the primary store and not for backups; the discrepancy surfaced during a customer audit and cost more credibility than an accurate qualified answer would have.

Answer accuracy against scope.

StoreCustomer-managed key
primaryyes
replicayes
backupno
of 3 stores2 covered

I would not consider it settled without evidence: Tie each library answer to the evidence that supports it and re-verify on a cadence.

An inaccurate answer becomes a commitment.

Curated: · Written: · Reviewed:

QA-95Your controls are correct and developers work around them. What do you change?(show answer)

The engineering content of security and developer experience is the detection and the revocation path, not the control's name.

A control routed around provides no protection and consumes the credibility needed for the next one. Where the safe path is slower than the unsafe one, the friction is the finding.

Concretely, measure how long the safe path takes against the alternative, invest in making it faster rather than in enforcement, and where enforcement is genuinely required, ship the faster path first.

The reason for that specificity is a failure I have seen: Secret access required a ticket with a 2-day turnaround; engineers copied secrets into a shared document instead, which was worse than the access the process was protecting.

Path chosen by cost.

PathTimeShare of use
ticket for access2 days10%
shared document2 min90%
self-service, audited30 s100%

I would not consider it settled without evidence: Measure the time cost of the safe path and compare against the workaround people actually use.

The unsafe path wins whenever it is faster.

Curated: · Written: · Reviewed:

QA-96A small security team supports forty engineering teams. How do you scale?(show answer)

Before calling building a security champions network done I would write down the attack it does not stop.

A central team cannot review everything, so it must scale by making the safe path default and by distributing judgement rather than by reviewing more. Champions work when they are given authority and time, and fail when the role is an unpaid addition.

Concretely, give champions real decision authority for defined categories, allocate their time explicitly, provide the escalation path for the cases beyond that, and measure the share of decisions made without central involvement.

The reason for that specificity is a failure I have seen: A champions programme named 40 people with no allocated time; participation collapsed within 2 quarters and the review queue returned to the central team, longer than before.

Programme outcome.

DesignActive after 2 quartersCentral queue
named, no time4 of 40longer
10% time, authority31 of 4060% shorter

I would not consider it settled without evidence: Report the share of security decisions made locally and the time actually allocated to champions.

A role with no time is a title.

Curated: · Written: · Reviewed:

QA-97Name a security control you would argue against.(show answer)

The first thing I would establish about when a control is not worth it is what an attacker gains if the control is absent.

Controls have costs in engineering time, developer friction, and operational risk, and a control whose cost exceeds the loss it prevents makes the organisation worse. Being able to argue against one is what makes the arguments for the others credible.

Concretely, estimate the loss prevented and the cost imposed, including the friction cost across every affected team and the operational risk the control itself introduces, and be explicit when the answer is that it is not worth it.

The reason for that specificity is a failure I have seen: A mandatory 90-day password rotation cost measurable support load and drove predictable password patterns, and the guidance it was based on had been withdrawn years earlier.

Cost against benefit.

ControlAnnual costLoss prevented
90-day rotation340 support hoursnegligible
phishing-resistant MFA40 hours setupsubstantial

I would not consider it settled without evidence: Cost the control across every affected team and compare against the loss it prevents before mandating it.

A control that costs more than it prevents is a net loss.

Curated: · Written: · Reviewed:

QA-98How do you decide whether a newly announced vulnerability matters to you?(show answer)

I would start keeping current without chasing headlines from the permissions that are actually effective, not from the policy document.

A named vulnerability with a logo is not automatically relevant, and a quiet advisory in a dependency you actually run may be. Relevance is determined by whether you run the affected version in a reachable configuration, not by coverage.

Concretely, maintain an inventory good enough to answer "do we run this" in minutes, check the affected configuration rather than just the version, and publish the assessment internally so forty teams do not each investigate separately.

The reason for that specificity is a failure I have seen: A widely covered vulnerability consumed 3 days across 12 teams; the affected configuration was not in use anywhere, and the inventory that would have answered it in an hour did not exist.

Time to assess an advisory.

Inventory qualityTime to answerTeams involved
none3 days12
version inventory2 h1
version + configuration20 min1

I would not consider it settled without evidence: Measure how long it takes to answer whether you run an affected version, and treat a slow answer as an inventory defect.

The inventory is what turns an advisory into an answer.

Curated: · Written: · Reviewed:

QA-99How do you present cloud security risk to an executive audience?(show answer)

This is an area where a passing compliance scan and a correct handling of explaining risk to non-technical leadership are different events.

Severity labels and finding counts do not translate into decisions. What does is the specific scenario, its plausibility given what exists today, the loss it would cause, and the concrete option being asked for.

Concretely, present the two or three scenarios that actually matter with the evidence behind each, state what would have to happen and what already has, give the cost of the mitigation, and make a recommendation rather than presenting a menu.

The reason for that specificity is a failure I have seen: A quarterly report of 4,000 findings by severity produced no decision for 3 quarters; a single slide describing one plausible path to the customer database with its cost was funded the same week.

Format against outcome.

FormatQuarters presentedDecisions
findings by severity30
one scenario with cost11

I would not consider it settled without evidence: Measure whether the presentation produced a decision, and change the format when it did not.

Executives fund scenarios, not severities.

Curated: · Written: · Reviewed:

QA-100What mistake do experienced cloud security engineers make?(show answer)

My answer to the security engineer's own failure mode begins with the blast radius: if I cannot state what one compromise reaches, I have no design.

The recurring one is optimising for the controls that are measurable and defensible rather than for the paths an attacker would take, because the first produces reports and the second produces arguments. The result is a strong posture score and an unaddressed path.

Concretely, spend a fixed share of effort starting from the attacker's position rather than from the control list, and treat any exercise that finds nothing as evidence the exercise was too narrow rather than that the estate is secure.

The reason for that specificity is a failure I have seen: A team spent a year raising a posture score from 74 to 98 while the CI system, out of scope for the scanner, kept production credentials in environment variables throughout.

Where the year went.

ActivityEffortPaths closed
posture score work80%3
attacker-perspective work20%11

I would not consider it settled without evidence: Reserve time for attacker-perspective work and report what it finds separately from control coverage.

The attacker does not read your control list.

Curated: · Written: · Reviewed: