Top 100 Security Architect Interview Questions and Answers
The questions most likely to actually come up in your Security Architect interview, ranked by likelihood — with detailed, senior-level answers covering what an interviewer is really listening for.
Curated: · Written: · Reviewed:
QA-1The estate uses Confidential, Internal, and Public. Why is that not yet a data-protection architecture?(show answer)
The first thing I would establish about classification labels versus handling rules is which production path still violates the invariant after the diagram looks finished.
A label is a name. Architecture is the handling that follows the name at every copy: encryption, access, retention, export, support, and deletion. Three adjectives with no binding rules leave backups, analytics extracts, and vendor tickets outside the intended control.
Concretely, publish four to six labels, bind each to encryption, audience, retention days, allowed processors, and deletion method, then refuse any flow that carries a label without those bindings in the schema and the pipeline.
The reason for that specificity is a failure I have seen: Payroll sat under Confidential while nightly CSV extracts to a BI share had no audience rule. Support copied 18,400 salary rows into a ticket system over 11 days, and the copies outlived the primary table by 6 years.
Where Confidential actually travelled.
| Hop | Handling bound | Rows copied | Retention remaining |
|---|---|---|---|
| primary payroll table | field encryption, 25 named roles | 18,400 | 7 years |
| BI extract share | none | 18,400 | unbounded |
| support ticket attachments | none | 2,160 | 6 years after close |
I would not consider it settled without evidence: sample 20 labelled flows from create through backup, export, restore, and delete, and require each hop to name the handling rule that fired.
A label without a handling rule is interior decoration.
Curated: · Written: · Reviewed:
QA-2A workshop listed 140 critical systems. How do you decide which few the architecture actually has to hold?(show answer)
I would start crown-jewel inventory from residual risk, named authority, and expiry, not from the workshop minutes.
Crown jewels are the small set whose compromise ends the business: signing keys, recovery credentials, regulated master records, and the identity plane that can mint all of the above. A 140-row critical list is a CMDB dump, and it dilutes prevention onto assets that cannot produce that harm.
Concretely, rank services by customer harm, legal duty, and irreplaceable recovery, keep a working set under 15, fund isolation and monitoring for that set first, and treat everything else as important without pretending it is existential.
The reason for that specificity is a failure I have seen: The critical list had 140 entries including a cafeteria kiosk. Signing HSMs shared a VLAN with 38 of those entries. An unpatched kiosk became the foothold, and 4 code-signing keys were used for 9 days before anyone looked at the HSM logs.
Workshop list versus funded isolation.
| Set | Count | Isolation tested in 90 days | Keys reached in the drill |
|---|---|---|---|
| workshop critical | 140 | 6 | 4 signing keys |
| funded crown jewels | 12 | 12 | 0 |
| cafeteria kiosk | 1 | 0 | foothold at hour 3 |
I would not consider it settled without evidence: name the 15 jewels, the blast radius of each, and the last date an isolation test was run against a neighbouring low-value host.
If everything is a crown jewel, nothing is funded like one.
Curated: · Written: · Reviewed:
QA-3The context diagram shows six boxes and TLS everywhere. What is still missing?(show answer)
This is an area where a control inventory and a held production path are different artefacts.
A trust boundary is a change of identity, authority, tenant, protocol, or administrative domain, not a box outline. TLS on every arrow can hide the credential exchange that makes one service an implicit broker for the rest.
Concretely, enumerate every crossing as who authenticates, who authorises, which data class moves, which logs fire, and what happens when the far side is hostile, then add a sequence view for the five highest-consequence abuse cases.
The reason for that specificity is a failure I have seen: A payments sidecar minted service tokens for 22 microservices after a single mTLS handshake. When the sidecar's client certificate leaked from a debug dump, 9 downstream APIs accepted the broker as gospel for 31 hours and moved £2.4 million.
The missing broker on the tidy diagram.
| Crossing | Identity checked | Authorisation checked | Hours until contained | Value moved |
|---|---|---|---|---|
| client to sidecar | mTLS | none beyond presence | 31 | £2.4 million |
| sidecar to ledger | sidecar cert only | none | 31 | £2.4 million |
| after broker removal | per-workload SPIFFE | per-action policy | — | £0 in replay |
I would not consider it settled without evidence: walk five abuse cases end to end and account for identity, data, policy, telemetry, and failure at each crossing, including the broker nobody drew.
Boxes and padlocks are not a trust model.
Curated: · Written: · Reviewed:
QA-4The team produced a 90-row STRIDE spreadsheet. Why might you still reject the design?(show answer)
My answer to STRIDE theatre versus scenario-led threat modelling begins with prevention, detection, response, and recovery as one system, not a purchase.
STRIDE is a prompt, not a proof. Architecture needs a small number of feasible scenarios that name the attacker, the invariant they break, and the design change that stops them. A checklist that labels every box Spoofing while the support workflow remains unmodelled is theatre.
Concretely, pick the three to seven scenarios that can actually reach money, identity, or regulated data, write the sequence, attach a named design change to each, and refuse to close the model while any high-consequence path still has only a label.
The reason for that specificity is a failure I have seen: The spreadsheet had 90 STRIDE tags and a green workshop. A contractor support tool could impersonate any tenant by swapping an unauthenticated customer-id header. Testers found it on day 4 of a later pentest, after 12,700 accounts were already reachable.
Tags versus the path that mattered.
| Artefact | Count | Support impersonation covered | Accounts reachable |
|---|---|---|---|
| STRIDE rows | 90 | no | 12,700 |
| funded scenarios | 5 | yes, after redesign | 0 in retest |
| days from workshop to pentest finding | 46 | — | — |
I would not consider it settled without evidence: replay the highest-consequence scenarios against the implemented design and close each mitigation with a failed-then-passed test, not a tag count.
A label is not a mitigation.
Curated: · Written: · Reviewed:
QA-5Risk scoring paints this as red 9. How do you decide whether to spend £400,000 on a control?(show answer)
I would treat FAIR-style ranges versus a single red score as a claim about every sibling service and template, not about the one application that was reviewed.
A single ordinal score compresses unlike losses and hides the assumptions. An investment decision needs a loss scenario, a range on frequency and magnitude, and a threshold the risk owner already agreed. Two red-9 items can differ by two orders of magnitude once you write pounds and days.
Concretely, write the scenario, bound annual frequency and loss in 90-percent ranges, show how a stronger control moves the distribution, and require the named owner to pick treat, transfer, or accept against a published threshold.
The reason for that specificity is a failure I have seen: A red-9 ransomware row and a red-9 badge-printer row received the same budget split. The printer got a £180,000 camera system. Ransomware later halted order entry for 14 days at £1.1 million per day because the identity-plane control had been deferred.
Same colour, unlike loss.
| Scenario | Score | 90% annual loss range | Control funded | Outage days |
|---|---|---|---|---|
| identity-plane ransomware | 9 | £2.0m–£18m | deferred | 14 |
| badge printer theft | 9 | £4k–£40k | £180,000 cameras | 0 |
| order-entry halt | — | £1.1m per day | — | 14 |
I would not consider it settled without evidence: show how changing likelihood, control strength, or impact assumptions alters the selected treatment, with pounds attached.
Colour is not a capital allocation.
Curated: · Written: · Reviewed:
QA-6Engineers say they accept the leftover risk so the release can ship. What is wrong with that?(show answer)
The useful question for residual risk with named authority and expiry is which trust assumption remains true after the next deployment.
Implementers may describe residual risk; they may not accept business consequence. Acceptance needs a named authority whose remit covers the harm, an expiry, compensating monitoring, and a closure condition. Otherwise delivery pressure silently becomes the risk policy.
Concretely, record the unmet control, the exposure in numbers, the compensations, the accepting officer, a calendar expiry not exceeding 90 days unless the board extends it, and the metric that will force shutdown or the real control.
The reason for that specificity is a failure I have seen: A payments team accepted missing object-level checks in a spreadsheet. The note had no officer and no date. Eight renewals later the gap was 19 months old, and a guessed UUID wrote 6,200 refunds totalling £890,000.
Ungoverned acceptance versus a dated decision.
| Record | Officer | Expiry | Monitor probes | Refunds issued |
|---|---|---|---|---|
| sprint spreadsheet | none | none | 0 | 6,200 |
| after architecture standard | CFO delegate | 60 days | 14 of 14 blocked | 0 |
| age of the gap | 19 months | — | — | £890,000 |
I would not consider it settled without evidence: produce the acceptance record with officer, expiry, monitor, and the last time the monitor fired on a probe.
Shipping is not a risk appetite.
Curated: · Written: · Reviewed:
QA-7Network monitoring will compensate for missing segmentation until next year. How do you govern that?(show answer)
I would settle compensating controls that expire by exercising the abuse case against the running path, not against the slide.
A compensating control is a time-boxed substitute with a detection duty, not a permanent exception wearing nicer clothes. Without expiry and a measured substitute, monitoring becomes the architecture and segmentation never arrives.
Concretely, name the missing preventive control, the substitute's coverage and false-negative rate, a 90-day maximum unless re-justified with fresh numbers, and an automatic ticket that reopens the preventive work on expiry.
The reason for that specificity is a failure I have seen: East-west monitoring compensated for a flat payments VLAN for 3 years. The sensor sampled 1 in 256 flows. A worm used the unsamped majority and reached 41 hosts in 22 minutes, including the card-data subnet.
Sampled monitoring as a fake wall.
| Control | Duration | Sample rate | Hosts reached | Time to 41 hosts |
|---|---|---|---|---|
| flow sensor compensation | 3 years | 1/256 | 41 | 22 minutes |
| default-deny segments | not yet | n/a | 0 in later drill | — |
| planted technique catch | — | — | 2 of 50 trials | — |
I would not consider it settled without evidence: show the substitute's measured catch rate against a planted east-west technique, and the calendar date the preventive control replaces it.
A compensation without an end date is the new design.
Curated: · Written: · Reviewed:
QA-8SSO is green and HR disabled the account. Why can the leaver still refund customers?(show answer)
The judgement in workforce identity versus application roles is who may accept leftover business consequence, and until when.
Central login is not application authorization. Roles, refresh tokens, and local sessions live at relying parties. A disabled directory row that leaves a refund role intact for 36 hours is an architecture that stopped at the IdP.
Concretely, drive joiner-mover-leaver from HR into both the IdP and every high-consequence application's role store, revoke refresh grants and sessions on the disable event, and measure time-to-denied-action rather than time-to-directory-row.
The reason for that specificity is a failure I have seen: HR disabled an agent at 17:02. The IdP showed inactive. A cached refund role on the payments app survived until 09:14 the next day, and 84 refunds for £126,000 posted from a still-valid refresh token.
Directory row versus the refund that still posted.
| Event | Clock | IdP state | App role | Refunds after disable |
|---|---|---|---|---|
| HR disable | 17:02 | inactive at 17:03 | still agent | 84 |
| refresh reuse | 08:41 next day | inactive | still agent | 12 of the 84 |
| after RP revocation hook | — | inactive | removed in 4 minutes | 0 |
I would not consider it settled without evidence: time joiner, mover, leaver, recovery, and session-revocation on three critical apps, and require the leaver's next action to fail within the published minutes.
A green SSO tile is not a revoked grant.
Curated: · Written: · Reviewed:
QA-9The board was told MFA is deployed. The second factor is SMS. What do you tell them?(show answer)
Where candidates lose the interview on SMS MFA is not phishing-resistant is calling a passed architecture workshop the control.
SMS and voice OTPs are interceptable and promptable. Phishing-resistant authenticators bind the origin, typically FIDO2 or certificate-based, so a lookalike login page cannot relay the proof. Coverage that counts SMS as MFA is a metric, not a control for the high-consequence plane.
Concretely, require phishing-resistant authenticators for administrators, finance, and remote production access, keep SMS only as a degraded recovery with extra monitoring, and stop reporting SMS users in the MFA-complete numerator.
The reason for that specificity is a failure I have seen: An adversary-in-the-middle kit collected SMS codes from 310 staff over 6 days. 19 of those sessions reached the finance VPN, and 3 initiated supplier-bank changes totalling £3.8 million before step-up was redesigned.
Relay results by authenticator.
| Authenticator | Board metric | Relay success in drill | Supplier changes |
|---|---|---|---|
| SMS OTP | 94% MFA | 19 of 25 | 3, £3.8 million |
| FIDO2 | excluded from the 94% | 0 of 25 | 0 |
| days of collection | 6 | 310 staff prompted | — |
I would not consider it settled without evidence: attempt a relay against SMS and against a FIDO2 policy on the same role, and show one succeeding and one failing.
A text message is a second secret, not a bound authenticator.
Curated: · Written: · Reviewed:
QA-10Helpdesk can reset MFA by reading a date of birth. How do you architect recovery?(show answer)
I would answer MFA recovery as the back door by separating the target invariant from the product that is supposed to implement it.
Recovery is an authentication ceremony with the same consequence as the primary factor. A low-assurance helpdesk reset converts phishing-resistant login into a social-engineering API.
Concretely, require a second registered authenticator or a time-delayed, dual-control recovery with identity proofing matched to the role, log every recovery as a privileged event, and freeze high-consequence roles until a live owner confirms.
The reason for that specificity is a failure I have seen: A contractor helpdesk reset MFA for a treasury admin after matching a date of birth from a breach dump. The new authenticator enrolled in 11 minutes, and £1.6 million left in two Faster Payments before the dual-control rule existed.
Helpdesk reset versus dual control.
| Recovery path | Proofing | Time to new authenticator | Value moved |
|---|---|---|---|
| DOB via helpdesk | 1 knowledge item | 11 minutes | £1.6 million |
| dual-control + delay | 2 officers, 4 hours | 4 hours 12 minutes | £0 in retest |
| privileged freeze | until owner confirm | — | sessions killed: 3 |
I would not consider it settled without evidence: attempt a helpdesk-only reset on a privileged role and require refusal plus a ticket that names two officers.
Recovery that is easier than login is the real login.
Curated: · Written: · Reviewed:
QA-11Domain admins log in daily with the same accounts they have had for years. What is the architectural defect?(show answer)
The engineering content of standing privileged accounts is the reusable pattern and its evidence, not the local ticket that closed.
Standing privilege turns one stolen workstation session into lasting control-plane authority. Privileged access architecture makes high-consequence rights eligible, approved, time-bound, isolated, and recorded, then absent when unused.
Concretely, move production, directory, KMS, and recovery administration onto just-in-time elevation from a PAM or PIM plane, put those sessions on isolated workstations, and delete standing admin groups except a monitored break-glass set of two.
The reason for that specificity is a failure I have seen: A helpdesk laptop with a cached Domain Admin token was stolen from a car. The token was valid for 10 hours. The thief reset 64 privileged passwords and planted a GPO that survived 13 days.
Standing token versus JIT.
| Design | Daily standing admins | Stolen-session window | Passwords reset |
|---|---|---|---|
| Domain Admin on the helpdesk laptop | 22 | 10 hours | 64 |
| JIT eligible, 60-minute grant | 0 standing | grant refused | 0 |
| GPO persistence | 13 days | — | 1 malicious GPO |
I would not consider it settled without evidence: attempt unapproved elevation, emergency access, session kill, and post-use revocation in a rehearsal with timestamps.
A permanent admin is a permanent incident waiting for a laptop.
Curated: · Written: · Reviewed:
QA-12PIM is licensed. Users still keep Owner on production subscriptions. What did we actually buy?(show answer)
Before calling just-in-time elevation done I would write down the template, region, or future service that still ships the old design.
A PIM licence is not JIT. JIT is the absence of standing assignment plus an approval and an expiry that the control plane enforces. Eligible-but-always-active is standing privilege with extra screens.
Concretely, remove standing Owner and Contributor from humans, require eligible plus approval for production, cap activation at 8 hours, and alert on any assignment that is permanent or that skips approval.
The reason for that specificity is a failure I have seen: The PIM dashboard showed 400 eligible users. 37 still had standing Owner because a migration script granted it. Those 37 created 14 public storage accounts over 5 weeks, one of which leaked 8.2 million customer emails.
Licence versus standing Owner.
| Signal | Count | Public accounts created | Emails exposed |
|---|---|---|---|
| PIM eligible users | 400 | — | — |
| standing Owner remaining | 37 | 14 | 8.2 million |
| after standing removal | 2 break-glass | 0 in 30-day watch | 0 |
I would not consider it settled without evidence: list every standing privileged assignment, require the count to be the break-glass set only, and activate a grant that must expire without human memory.
Eligible forever is still standing.
Curated: · Written: · Reviewed:
QA-13Two emergency admin accounts live in a password manager everyone in SRE can open. Is that break-glass?(show answer)
The first thing I would establish about break-glass dual control is which production path still violates the invariant after the diagram looks finished.
Break-glass is rare, dual-controlled, monitored, and painful enough that it is not the daily path. A shared vault that 40 people can open is a standing admin cluster with a nicer name.
Concretely, split the emergency secret across two officers, alert on any use within minutes, rotate after every use, and test the path quarterly without leaving the credentials cached on jump hosts.
The reason for that specificity is a failure I have seen: An intern used the shared break-glass to debug a ticket at 02:11. The use produced no page. The intern's laptop later mined 2,400 additional API calls against the billing export over 17 days.
Shared vault versus split control.
| Design | People who can open | Page on use | Extra API calls |
|---|---|---|---|
| shared SRE vault | 40 | none | 2,400 |
| split secret, two officers | 2 + 2 | 3 minutes | 0 in later drill |
| intern debug at 02:11 | 1 | missed | 17 days dwell |
I would not consider it settled without evidence: open break-glass in a drill and require two people, a page in under 5 minutes, and a rotation before the drill closes.
If forty people can open it, it is not glass.
Curated: · Written: · Reviewed:
QA-14Microservices share one API key in Kubernetes secrets, rotated yearly. What should replace that?(show answer)
I would start workload identity versus copied secrets from residual risk, named authority, and expiry, not from the workshop minutes.
A copied static secret authenticates from any host that obtained the bytes, and it cannot name the workload instance. Workload identity issues short-lived, audience-bound credentials from an attestation of the runtime, which is what SPIFFE and cloud workload identity exist to do.
Concretely, issue SPIFFE IDs or cloud workload tokens bound to namespace, service account, and audience, set lifetimes in minutes, refuse the old shared key, and log issuance so a token maps to a replica.
The reason for that specificity is a failure I have seen: The yearly key was copied into 11 forks, 3 laptops, and a Slack thread. It authenticated from a developer's home IP for 11 months, pulling 44 million events from a telemetry API that thought it was talking to production.
Yearly key versus attested identity.
| Credential | Lifetime | Distinct holders found | Events pulled from home IP |
|---|---|---|---|
| shared API key | 365 days | 15 | 44 million |
| SPIFFE SVID | 60 minutes | 1 replica | 0 |
| audience mismatch rejects | — | — | 100 of 100 test calls |
I would not consider it settled without evidence: demonstrate issuance, audience rejection, rotation, revocation, and workload replacement without a static secret remaining in the cluster.
A secret that can be copied is a secret that will be copied.
Curated: · Written: · Reviewed:
QA-15The Dockerfile copies a .env file so local and prod behave the same. Why is that an architecture defect?(show answer)
This is an area where a control inventory and a held production path are different artefacts.
An image is a redistributable artefact. A secret in the image is a secret in every registry, cache, developer laptop, and forgotten tag. Runtime injection from a manager or from workload identity keeps the artefact free of credentials.
Concretely, fail the build if scanners find high-entropy secrets, inject at runtime from a manager with per-replica identity, and treat historical tags that contain secrets as compromised until rebuilt and the secrets rotated.
The reason for that specificity is a failure I have seen: A public fork of an internal base image still contained a Stripe live key from layer 7. The key processed 1,120 unauthorised charges totalling $84,000 over 9 days before rotation.
Layer 7 versus runtime injection.
| Location | Tags affected | Unauthorised charges | Hours to rotate |
|---|---|---|---|
| Dockerfile COPY .env | 12 | 1,120 ($84,000) | 38 |
| runtime manager | 0 | 0 | n/a |
| public fork clones | 4 | included in the 1,120 | — |
I would not consider it settled without evidence: pull the latest five tags, scan them, and require zero live credentials plus a rotation record for any historical hit.
If it is in the layer, it is in the world.
Curated: · Written: · Reviewed:
QA-16Services accept any JWT signed by the internal issuer. What goes wrong?(show answer)
My answer to audience-bound workload tokens begins with prevention, detection, response, and recovery as one system, not a purchase.
A signature proves an issuer minted a token. Audience proves it was minted for this API. Tokens without audience checks are bearer tickets for the whole estate, which is how a reporting job impersonates payments.
Concretely, put a distinct audience on every API, reject tokens whose aud does not match, keep lifetimes short, and never accept an ID token or a sibling service's access token as a substitute.
The reason for that specificity is a failure I have seen: A batch reporter's token, signed correctly, called the payout API 2,900 times because the API checked signature and expiry only. £410,000 left to attacker-controlled accounts over a weekend.
Signature-only versus audience.
| Check | Batch reporter token | Payouts issued | Value |
|---|---|---|---|
| signature + expiry only | accepted | 2,900 | £410,000 |
| audience match | rejected | 0 | £0 |
| weekend dwell | 47 hours | — | — |
I would not consider it settled without evidence: present a correctly signed token with the wrong audience and require a distinct rejection reason in the logs.
A valid signature is the start of authorisation, not the end.
Curated: · Written: · Reviewed:
QA-17We put an identity-aware proxy in front of the flat office network and called it zero trust. Did we?(show answer)
I would treat zero trust is per-request authorisation as a claim about every sibling service and template, not about the one application that was reviewed.
Zero trust authorises each request with verified identity, device, and context against that resource. A proxy login in front of a flat network is a VPN with extra branding: once past the login, east-west remains implicit trust.
Concretely, move policy to the resource, evaluate identity, device posture, and least privilege on every call, shorten session binding so posture drift re-evaluates, and stop treating network location as a permit.
The reason for that specificity is a failure I have seen: The proxy required SSO then dumped users onto 10.20.0.0/16. A stolen session cookie from an unmanaged laptop reached the HR database at 03:40 because no per-resource check remained. 62,000 records left over 4 hours.
Proxy-then-flat versus per-resource.
| Design | After login | Stolen cookie to HR DB | Records exported |
|---|---|---|---|
| IAP + flat /16 | implicit east-west | allowed | 62,000 |
| per-resource policy | evaluated each call | refused | 0 |
| cookie age | 12 hours | 03:40 use | 4 hours |
I would not consider it settled without evidence: test stolen-session, unmanaged-device, risky-location, device-drift, and resource-revocation, and require the resource to refuse independently of the proxy login.
A login page is not a request policy.
Curated: · Written: · Reviewed:
QA-18Conditional access checks the device at login. Why can a later-compromised laptop still keep the session?(show answer)
The useful question for device posture drift after admission is which trust assumption remains true after the next deployment.
Admission is a point in time. Posture drifts: disk encryption off, EDR uninstalled, OS rollback. Architecture re-evaluates posture on a short cadence or on each sensitive action, and it kills the session when the device no longer qualifies.
Concretely, bind sessions to a device signal that is refreshed at least every 60 minutes for high-consequence apps, refuse sensitive actions when posture is stale, and inventory unmanaged endpoints that still hold refresh tokens.
The reason for that specificity is a failure I have seen: A laptop passed at 09:00. EDR was uninstalled at 12:12. The 12-hour session still opened payroll at 18:05. An attacker who later used that laptop exported 9,400 payslips.
Login check versus mid-session drift.
| Event | Time | Posture | Payslips exported |
|---|---|---|---|
| login | 09:00 | healthy | — |
| EDR removed | 12:12 | unhealthy, session live | — |
| payroll open | 18:05 | still admitted | 9,400 |
| 60-minute re-eval | — | action refused | 0 in retest |
I would not consider it settled without evidence: uninstall the agent mid-session and require the next sensitive action to fail within the published refresh window.
Yesterday's healthy device is today's untrusted endpoint.
Curated: · Written: · Reviewed:
QA-19We have PCI, corporate, and vendor VLANs. Why can a vendor jump still reach card data?(show answer)
I would settle segmentation versus transitive shared services by exercising the abuse case against the running path, not against the slide.
Subnet labels are not isolation if shared services — DNS, jump hosts, monitoring, CI, file shares — recreate transitive reachability. The effective graph is the union of explicit allows and every helper that both zones can touch.
Concretely, derive the effective graph from policy and flow logs, default-deny between consequence zones, give shared services identity-aware, protocol-narrow fronts, and attempt prohibited paths from every major zone including the helpers.
The reason for that specificity is a failure I have seen: PCI and vendor VLANs looked separate. Both mounted the same logging NFS share with write. A vendor engineer wrote a cron that read card files copied there for debugging. 1.1 million PANs sat on the share for 8 days.
VLAN list versus the NFS bridge.
| Path | Intended | Effective | PANs exposed |
|---|---|---|---|
| vendor to PCI direct | deny | deny | 0 |
| vendor to logging NFS | allow | allow | 1.1 million |
| PCI debug copy to NFS | undocumented | allow | 8 days |
| after identity-aware log sink | — | vendor cannot read PCI objects | 0 |
I would not consider it settled without evidence: attempt prohibited paths from each zone and from each shared helper, and publish the effective graph rather than the intended VLAN list.
A shared helper is a bridge you forgot to draw.
Curated: · Written: · Reviewed:
QA-20Security groups allow 10.0.0.0/8 on 443 everywhere because services need to talk. What is the alternative?(show answer)
The judgement in default-deny east-west is who may accept leftover business consequence, and until when.
A /8 permit is a flat network with extra YAML. East-west architecture enumerates service-to-service flows by identity and port, denies the rest, and treats a new flow as a change with a ticket, not as a reason to widen the CIDR.
Concretely, start from deny, allow named service identities on named ports, review unused allows every 30 days, and alert on any rule whose source is a supernet larger than a /24 without a named exception.
The reason for that specificity is a failure I have seen: The /8 rule let a compromised CI runner scan 14,000 hosts on 443. It found an unauthenticated admin UI on an internal package registry and pushed a backdoored library that 26 production services pulled within 6 hours.
/8 permit versus named flows.
| Rule | Hosts reachable on 443 | Backdoored pulls | Hours to 26 services |
|---|---|---|---|
| 10.0.0.0/8 | 14,000 | 26 | 6 |
| named identities only | 4 intended peers | 0 | — |
| unused allows removed | 118 rules deleted in 30 days | — | — |
I would not consider it settled without evidence: from a host in an unrelated subnet, attempt 443 to a sensitive service and require a policy drop with a logged identity.
Need to talk is not a CIDR.
Curated: · Written: · Reviewed:
QA-21Every endpoint sits behind authentication middleware. Why did tenant A download tenant B's invoices?(show answer)
Where candidates lose the interview on object-level authorisation is calling a passed architecture workshop the control.
Authentication answers who is present. Object-level authorisation answers whether this principal may touch this record. Middleware that only checks a session leaves IDOR and broken object-level authorisation as the write path.
Concretely, enforce ownership or tenant membership on every read and write of a protected object, centralise the check so new endpoints inherit it, and test cross-tenant identifiers on every resource type before release.
The reason for that specificity is a failure I have seen: Invoice PDFs keyed by sequential IDs were fetchable with any logged-in session. A script walked 80,000 IDs in 3 hours and retrieved 12,400 foreign invoices including 2,100 with bank details.
Middleware versus object check.
| Control | Foreign invoices fetched | Bank details in set | Walk duration |
|---|---|---|---|
| session middleware only | 12,400 | 2,100 | 3 hours |
| per-object tenant check | 0 of 5,000 probes | 0 | — |
| sequential IDs remaining | yes, until UUID migration | — | — |
I would not consider it settled without evidence: exercise cross-tenant identifiers, replay, concurrency, and alternate endpoints against the same ownership invariant, and require 0 foreign objects returned.
A session is not a title deed.
Curated: · Written: · Reviewed:
QA-22The parent project is tenant-checked. Nested comments are not. Is the API isolated?(show answer)
I would answer BOLA on nested resources by separating the target invariant from the product that is supposed to implement it.
Tenant isolation that stops at the parent leaks through children, attachments, search, and export jobs. Architecture applies the same membership check to every object that can name a parent, including asynchronous workers.
Concretely, inherit tenant from a server-side lookup of the parent, never from a client-supplied tenant-id, and cover list, get, patch, delete, search, and export for every nested type.
The reason for that specificity is a failure I have seen: Comments accepted a projectId from the body without membership. An attacker iterated projectIds and scraped 310,000 comments, 4,800 of which contained access tokens pasted by developers.
Parent check versus nested scrape.
| Endpoint | Tenant check | Comments scraped | Tokens in comments |
|---|---|---|---|
| GET /projects/{id} | yes | — | — |
| GET /comments?projectId= | no | 310,000 | 4,800 |
| after inherited membership | yes | 0 of 20,000 probes | 0 |
I would not consider it settled without evidence: request nested objects under another tenant's parent from list, get, and export, and require empty results plus a uniform error that does not confirm existence.
The child is the parent for an attacker.
Curated: · Written: · Reviewed:
QA-23Partner webhooks are HMAC-signed. Why did we pay a payout twice?(show answer)
The engineering content of webhook replay and idempotency is the reusable pattern and its evidence, not the local ticket that closed.
A signature proves the sender and the payload bytes, not that this delivery may be applied again. Financial and state-changing webhooks need a replay window, a nonce or event id, and an idempotent receiver so a retry cannot mint a second transfer.
Concretely, verify HMAC, timestamp, and audience, store processed event ids with a 72-hour replay cache, and make the payout unique on a business key so a second delivery returns the first result.
The reason for that specificity is a failure I have seen: A partner retried a signed payout webhook five times after a 504. The original plus those five retries each created a transfer. £72,000 left six times, £432,000 total, before the idempotency key existed.
HMAC without idempotency.
| Delivery | HMAC | Transfer created | Amount |
|---|---|---|---|
| original | valid | 1 | £72,000 |
| retries 2–6 | valid | 5 more | £360,000 |
| after event-id cache | valid | 0 extra | £0 |
| 80-hour delay | valid, expired ts | rejected | £0 |
I would not consider it settled without evidence: replay a valid signed webhook and a delayed copy outside the window, and require one business effect and one rejection.
Signed twice is not authorised twice.
Curated: · Written: · Reviewed:
QA-24A programme wants to treat a signed OAuth access token as proof the user is at the keyboard. Why is that the wrong artefact?(show answer)
Before calling OAuth access token is not authentication done I would write down the template, region, or future service that still ships the old design.
An access token answers whether a client may call an API under a grant. An ID token communicates an authentication event to its intended client; it does not, by itself, prove the user is at the keyboard. Recent presence needs an explicit reauthentication policy, not a silent refresh.
Concretely, reject resource-server access tokens at the session edge. Validate issuer, client audience, signature, nonce, and expiry on the ID token. For recent presence, require max_age or prompt=login and check auth_time, amr, and acr. Keep access tokens on the API path only.
The reason for that specificity is a failure I have seen: A mobile BFF accepted resource-server access tokens as session credentials. Tokens minted for a public weather API, valid 60 minutes, opened 410 intranet sessions over 8 days because the BFF never checked audience.
BFF session from the wrong artefact.
| Artefact at the BFF | Audience | Sessions opened | Days |
|---|---|---|---|
| weather API access token | weather.api | 410 | 8 |
| ID token for the BFF | bff.intranet | 0 foreign | — |
| userinfo used as login | no auth_time check | 140 of the 410 | — |
I would not consider it settled without evidence: present a foreign-audience access token at the session endpoint and require a refusal that names audience mismatch. Also prove silent SSO cannot satisfy a forced recent-authentication requirement.
An API grant is not an authentication ceremony.
Curated: · Written: · Reviewed:
QA-25The CRM questionnaire passed. The integration requested Mail.ReadWrite for every mailbox. What did diligence miss?(show answer)
The first thing I would establish about SaaS OAuth overscope is which production path still violates the invariant after the diagram looks finished.
A questionnaire describes intent. The grant describes power. SaaS architecture constrains scopes to the operations the business needs, reviews admin-consent, and tests support and exit paths. Overscope is a standing impersonation of the tenant.
Concretely, inventory every OAuth grant, refuse tenant-wide mail and files scopes unless a named owner accepts a dated exception, prefer resource-specific consent, and rehearse token revocation on vendor exit.
The reason for that specificity is a failure I have seen: A calendar plugin with Mail.ReadWrite exfiltrated 2.3 million messages after its vendor was breached. Diligence had a green questionnaire dated 14 months earlier.
Questionnaire versus mailbox grant.
| Artefact | Result | Messages reachable | Time since diligence |
|---|---|---|---|
| vendor questionnaire | pass | not measured | 14 months |
| Mail.ReadWrite tenant-wide | granted | 2.3 million | — |
| resource-specific calendars only | after redesign | 0 mailboxes | revoke in 8 minutes |
I would not consider it settled without evidence: list scopes, compare them to documented operations, and revoke a test grant while proving the plugin loses mailbox access within the SLO.
A passed PDF is not the grant.
Curated: · Written: · Reviewed:
QA-26We generate an SBOM at build and store it. Why can we still not prove what ran in production?(show answer)
I would start SBOM is not provenance from residual risk, named authority, and expiry, not from the workshop minutes.
An SBOM lists ingredients someone claimed. Provenance attests who built what from which source on which isolated builder, typically SLSA-style attestations verified at deploy. A generated JSON file next to an artefact is a packing list, not a chain of custody.
Concretely, build in isolated, non-mutable runners, pin inputs, sign provenance, verify at admission, and refuse images whose attestation does not match the source commit and builder identity.
The reason for that specificity is a failure I have seen: A protected repo still deployed an image built on a mutable shared runner. The SBOM looked complete. The runner had been executing a malicious pre-step for 21 days, and 7 production services ran the tainted digest.
SBOM on disk versus attested build.
| Control | Isolated builder | Tainted services | Days the pre-step ran |
|---|---|---|---|
| SBOM generated | no | 7 | 21 |
| SLSA provenance verified at admit | yes | 0 after block | — |
| mutable runner still in pool | 3 | — | — |
I would not consider it settled without evidence: rebuild from pinned inputs and verify source-to-artefact provenance, signature enforcement, and a compromised-runner containment test.
A packing list is not a signature over the cook.
Curated: · Written: · Reviewed:
QA-27Self-hosted runners are persistent VMs with Docker cache to save minutes. What is the supply-chain cost?(show answer)
This is an area where a control inventory and a held production path are different artefacts.
A persistent runner is a privileged workstation that can poison every later job. Isolated, ephemeral builders throw away state so a compromised job cannot teach the next one its secrets or its binaries.
Concretely, use ephemeral VMs or pods per job, forbid Docker-in-Docker on shared hosts, scope tokens to the job, and treat any long-lived runner as a production admin endpoint with JIT and monitoring.
The reason for that specificity is a failure I have seen: A persistent runner's cache contained a backdoored gcc. 48 subsequent builds compiled it in. The resulting binaries called home from 12 production hosts for 16 days.
Persistent VM versus ephemeral job.
| Builder | Jobs after compromise | Binaries calling home | Days |
|---|---|---|---|
| persistent VM + Docker cache | 48 | 12 hosts | 16 |
| ephemeral per job | 0 inherited | 0 | — |
| minutes saved by cache | 6 per job | — | — |
I would not consider it settled without evidence: compromise a job in a rehearsal and show the next job on a fresh builder does not inherit the payload, while the persistent pool does.
A warm cache is a warm persistence.
Curated: · Written: · Reviewed:
QA-28Encryption uses AES-256-GCM. The data key is stored next to the ciphertext in S3. What failed?(show answer)
My answer to KMS non-exportable keys begins with prevention, detection, response, and recovery as one system, not a purchase.
Confidentiality is secrecy of the keying material relative to the attacker. A data key beside the object is labelling, not encryption. Non-exportable keys in KMS or an HSM unwrap data keys only to authorised identities, with encryption context binding the object.
Concretely, generate data keys via KMS, persist only the wrapped blob, bind encryption context to object id and tenant, deny kms:Decrypt except to the workload identity, and never write plaintext keys to disk or logs.
The reason for that specificity is a failure I have seen: An open bucket listed objects and their adjacent .key files. 640 GB of "encrypted" exports were readable in 40 minutes because decrypt was a local AES call with the neighbouring file.
Adjacent key files versus KMS wrap.
| Design | Plaintext keys on disk | Data readable without IAM | GB taken |
|---|---|---|---|
| AES + .key sibling | yes | yes | 640 |
| KMS wrap + context | no | no | 0 of 50 probes |
| minutes to enumerate bucket | 40 | — | — |
I would not consider it settled without evidence: attempt decrypt without KMS and with the wrong encryption context, and require both to fail while the workload identity succeeds.
The key beside the ciphertext is a caption, not a lock.
Curated: · Written: · Reviewed:
QA-29One HMAC key signs cookies, webhooks, and password-reset tokens. Why split them?(show answer)
I would treat cryptographic purpose separation as a claim about every sibling service and template, not about the one application that was reviewed.
A key's purpose is part of its threat model. Reusing one HMAC across cookies and resets lets a webhook oracle or a cookie become a reset token. Architecture gives each purpose its own key, algorithm, and rotation.
Concretely, allocate distinct keys in the manager, name them by purpose, refuse cross-use in code by type, and rotate on independent schedules so a webhook leak does not mint sessions.
The reason for that specificity is a failure I have seen: A leaked webhook HMAC forged a password-reset token for 1,900 accounts in 2 hours because verification used the same key and a compatible encoding.
Shared HMAC versus purpose keys.
| Use | Shared key forges reset | Accounts | Hours |
|---|---|---|---|
| webhook HMAC reused | yes | 1,900 | 2 |
| distinct purpose keys | 0 of 500 forgeries | 0 | — |
| rotation independent | — | webhook rotated, cookies kept | 24 hours |
I would not consider it settled without evidence: use a webhook signature as a cookie and as a reset token in tests, and require both to fail on purpose mismatch.
One key, many ceremonies, one breach.
Curated: · Written: · Reviewed:
QA-30Legal asked for erasure. Engineering deleted the users row. Why is that not done?(show answer)
The useful question for GDPR erasure is not a primary-row DELETE is which trust assumption remains true after the next deployment.
Erasure is removal of identifiability from every store that can still reconstitute the person: replicas, search, features, caches, logs in policy, backups by process, and vendors. Deleting the primary key while those copies remain is a UI lie.
Concretely, keep an inventory of personal-data stores, fan erasure with a completion token per store, block restores that rehydrate erased subjects without a legal basis, and sample a subject until derived stores are empty.
The reason for that specificity is a failure I have seen: The users row vanished. The recommendation feature store still held 14 months of events, a vendor marketing copy held 9 months, and a support export reconstituted 8,200 erased profiles in 20 minutes from those features, including 1,400 that overlapped the vendor copy.
Primary row versus derived copies.
| Store | After DELETE | Profiles reconstituted | Months of events |
|---|---|---|---|
| users table | gone | — | — |
| feature store | present | 8,200 | 14 |
| vendor marketing copy | present | 1,400 overlap | 9 |
| after fan-out erasure | gone in 6 of 6 stores | 0 of 50 samples | — |
I would not consider it settled without evidence: follow twenty selected records through the inventory and verify access, purpose, retention, export, and deletion on the primary, the feature store, the cache, the vendor, and a restore test that must not revive them.
A primary-key delete is only the first store.
Curated: · Written: · Reviewed:
QA-31Backups keep 35 days. Erasure requests arrive daily. How do you reconcile them?(show answer)
I would settle backup retention versus erasure duty by exercising the abuse case against the running path, not against the slide.
Backup is a recovery control with a stated retention. Erasure is a legal control that must either wait for expiry with access blocked, or use targeted purge where the medium allows. Pretending backups are out of scope while they remain restorable into production is a dual-use store.
Concretely, document which backup generations can still rehydrate a subject, block ad-hoc restores of erased identities, purge or expire on a published clock, and test that a 35-day restore cannot put an erased subject back into the serving path without a logged legal exception.
The reason for that specificity is a failure I have seen: Ops restored a cluster to debug an outage and brought back 3,100 erased customers into the live app for 11 hours, including 400 who had cited threat-to-life grounds.
35-day backup meeting an erasure.
| Action | Erased subjects live again | Hours exposed | Legal-exception log |
|---|---|---|---|
| cluster restore for debug | 3,100 | 11 | none |
| restore with erasure filter | 0 | 0 | n/a |
| threat-to-life subset | 400 | 11 | none |
I would not consider it settled without evidence: restore a backup that contains an erased test subject and require the serving path to keep that subject unavailable.
A restore is a processing activity.
Curated: · Written: · Reviewed:
QA-32The provider's page says they manage the cloud of the cloud. Why do we still own IAM and data?(show answer)
The judgement in shared responsibility on IaaS is who may accept leftover business consequence, and until when.
On IaaS the provider secures the facilities, hypervisor, and hardware. The customer still owns identities, guest configuration, network exposure, encryption choices, and data. Reading managed infrastructure as managed security is how public security groups survive.
Concretely, write a service-by-service matrix naming who configures, monitors, patches, restores, and accepts each material risk, then wire preventive and detective controls to the customer-owned rows.
The reason for that specificity is a failure I have seen: A team treated EC2 as a managed server. A 0.0.0.0/0 SSH rule sat for 71 days. Credential stuffing hit 14 instances; 3 had reused passwords and became miners.
IaaS split on one account.
| Control | Owner | Days misconfigured | Instances compromised |
|---|---|---|---|
| hypervisor patching | provider | — | — |
| SSH 0.0.0.0/0 | customer | 71 | 3 of 14 hit |
| instance IAM role | customer | standing admin | miners for 9 days |
I would not consider it settled without evidence: sample each service model and demonstrate who configures, monitors, restores, patches, and accepts each material risk, including IAM.
The provider's fence is not your IAM.
Curated: · Written: · Reviewed:
QA-33It is a managed database, so security is included. What still has to be designed?(show answer)
Where candidates lose the interview on PaaS public-access settings is calling a passed architecture workshop the control.
PaaS shifts patching of the engine, not the customer's network exposure, IAM, encryption flags, or backup access. Public-access toggles, firewalls, and admin identities remain customer architecture.
Concretely, default private endpoints, disable public network access in policy, require customer-managed keys where the data class demands it, and alert on any PaaS resource that becomes public.
The reason for that specificity is a failure I have seen: A managed Redis was created with the public default to save a peering ticket. It held session tokens. Shodan found it in 6 hours; 22,000 sessions were hijacked over a weekend.
Public Redis default.
| Setting | Default in console | Sessions stolen | Hours to discovery |
|---|---|---|---|
| public network access | on | 22,000 | 6 |
| paved private endpoint | off, denied if on | 0 | — |
| weekend window | 52 hours | — | — |
I would not consider it settled without evidence: create a PaaS instance through the paved path and through a raw console, and require the console path to fail policy if public access is on.
Managed engine, unmanaged front door.
Curated: · Written: · Reviewed:
QA-34We bought a certified SaaS. Who owns tenant isolation, SSO, and support access?(show answer)
I would answer SaaS tenant configuration as the architecture by separating the target invariant from the product that is supposed to implement it.
Certification speaks to the provider's common controls. The customer's architecture is tenant settings: SSO enforcement, SCIM, admin MFA, support-access approval, data residency, and exit. An unconfigured certified product is an ungoverned estate with a logo.
Concretely, baseline tenant settings as code or a checked inventory, enforce SSO and disable local passwords for staff, wrap support access in time-boxed approval, and test data return and deletion annually.
The reason for that specificity is a failure I have seen: Local passwords remained enabled beside SSO. 180 staff still used them. 12 were in a credential dump; 4 of those were tenant admins. The attacker created a forwarding rule that silently copied 90,000 messages over 28 days.
Certified product, local passwords.
| Control | Configured | Admin accounts in dump | Messages forwarded |
|---|---|---|---|
| ISO report | yes | — | — |
| local passwords | still on | 4 | 90,000 |
| SSO-only after lockdown | enforced | 0 new | forwarding removed in 2 hours |
I would not consider it settled without evidence: verify SSO-only, support paths, logs, incident duties, data return, deletion, and a replacement tabletop.
The certificate is the provider's; the tenant is yours.
Curated: · Written: · Reviewed:
QA-35The SIEM ingests 40 TB a month and the dashboard is green. How do you know Valid Accounts (T1078) is covered?(show answer)
The engineering content of ATT&CK technique coverage not ingestion volume is the reusable pattern and its evidence, not the local ticket that closed.
Ingestion is a cost. Coverage is whether a named technique produces an event, a detection, a triage context, and a containment path. Healthy GB totals can hide the one audit event a privilege-escalation technique never emits.
Concretely, pick the techniques that matter for the architecture, generate them in production-like conditions, and measure event arrival, rule fire, analyst context, and containment, recording blind spots as design work.
The reason for that specificity is a failure I have seen: Ingestion was 40 TB. Azure AD privileged-role activation was not logged to the SIEM. An attacker used T1078 for 19 days; the first ticket came from a customer who saw a strange mailbox rule.
Volume versus one technique.
| Signal | Value | T1078 visible | Dwell |
|---|---|---|---|
| monthly ingestion | 40 TB | no | 19 days |
| privileged-role log to SIEM | missing | — | — |
| after log + detection | 12 of 12 drills | yes | 18 minutes median |
I would not consider it settled without evidence: generate the technique and measure event arrival, rule firing, triage context, containment, and the named blind spots.
Terabytes are not T1078.
Curated: · Written: · Reviewed:
QA-36Backups restored the ERP in 4 hours, inside RTO. Why was that still a failed recovery?(show answer)
Before calling restore-of-compromise done I would write down the template, region, or future service that still ships the old design.
RTO and RPO are time and data-loss bounds. They are not proof the restored system is free of the persistence that caused the incident. Restore-of-compromise replays the attacker's foothold from backup and from identity.
Concretely, restore into an isolated clean room, rotate credentials and keys before exposure, validate integrity and authorisation paths, then cut over, and refuse a restore that brings back the compromised identity plane.
The reason for that specificity is a failure I have seen: ERP came back in 4 hours. The backup included the attacker's SSO app registration. Persistence returned at minute 40, and a second encryption event started at hour 6.
RTO met, persistence restored.
| Step | Clock | Attacker app registration | Second encryption |
|---|---|---|---|
| production restore | 4 hours | restored | hour 6 |
| clean-room + identity rebuild | 9 hours | omitted | none |
| RPO | 15 minutes data | — | — |
I would not consider it settled without evidence: exercise a destructive compromise through clean-room recovery, credential rotation, validation, backlog handling, and return to service.
A green restore clock can resurrect the attacker.
Curated: · Written: · Reviewed:
QA-37The backup job is green every night. How do you know RPO is 15 minutes?(show answer)
The first thing I would establish about backup job success is not RPO is which production path still violates the invariant after the diagram looks finished.
A successful job means the scheduler ran. RPO is the age of the newest recoverable consistent copy that you have actually restored. Green jobs can backup an empty volume, skip the database, or never have been restored.
Concretely, restore on a cadence, measure the timestamp of the newest transaction present, compare it to the stated RPO, and treat an unrestored backup as an unproven control.
The reason for that specificity is a failure I have seen: Nightly jobs succeeded for 11 months. The selected dataset omitted the ledger volume after a mount change. The restore during an incident was 17 days stale, against a 15-minute RPO, and 2.4 million postings were unrecoverable.
Green job, empty ledger.
| Check | Result | Newest ledger tx | Gap versus 15-minute RPO |
|---|---|---|---|
| job status | success 334 nights | missing volume | 17 days |
| actual restore | 17 days stale | 2.4 million lost | 17 days |
| after dataset monitor | restore monthly | within 9 minutes | pass |
I would not consider it settled without evidence: restore the newest copy and record the newest business transaction time versus the RPO clock.
The scheduler's smile is not a recoverability proof.
Curated: · Written: · Reviewed:
QA-38The programme closed 12,000 low CVSS findings this quarter. Why might risk have gone up?(show answer)
I would start exposure-led vulnerability priority from residual risk, named authority, and expiry, not from the workshop minutes.
CVSS without reachability, identity, and consequence is a queue, not a strategy. Architecture ranks by internet reach, available credentials, and business harm. Closing thousands of unscored internals while one reachable identity flaw remains is a vanity burn-down.
Concretely, join scanner output to asset consequence and attack-path reachability, SLA the reachable high-consequence set first, and report that set's age rather than ticket volume.
The reason for that specificity is a failure I have seen: Teams closed 12,000 lows. An internet-facing SSO bypass on a forgotten subdomain stayed open 61 days. It became the entry for a ransomware affiliate; 70% of endpoints encrypted.
12,000 closes versus one subdomain.
| Work | Count | Internet reachable | Endpoints encrypted |
|---|---|---|---|
| low CVSS closed | 12,000 | no | — |
| SSO bypass on old shop | 1 | yes, 61 days | 70% |
| after exposure SLA | reachable highs in 7 days | 0 open >14 days | — |
I would not consider it settled without evidence: validate inventory coverage and retest the highest attack paths rather than reporting ticket closure or raw counts.
A burn-down of trivia can hide a door.
Curated: · Written: · Reviewed:
QA-39Security's slot is a pentest two weeks before launch. What should have happened earlier?(show answer)
This is an area where a control inventory and a held production path are different artefacts.
A pentest finds implementation bugs cheaply and architectural authorisation defects expensively, after clients and schemas are frozen. SDLC architecture puts threat-informed authorisation requirements and tests on the paved road before code hardens, and uses pentest as confirmation, not discovery of the data model.
Concretely, require an authorisation model and abuse-case tests at design, automate object-level tests in CI, and refuse to schedule pentest as the first security review of a new product.
The reason for that specificity is a failure I have seen: Pentest week found missing tenant checks on 9 write APIs. Mobile clients were already in stores. A breaking fix took 6 weeks and a forced update; 3 tenants had already been crossed in the wild.
First security look at week −2.
| Gate | When | Tenant bugs found | Forced-update weeks |
|---|---|---|---|
| pentest only | launch −14 days | 9 write APIs | 6 |
| design + CI object tests | sprint 1 | 0 at pentest | 0 |
| tenants crossed before fix | 3 | — | — |
I would not consider it settled without evidence: sample releases and trace threats to requirements, code controls, tests, exceptions, and production evidence, with pentest as a later confirmation.
A pentest is a late, expensive design review.
Curated: · Written: · Reviewed:
QA-40The board approved a diagram 18 months ago. Why is that not still the architecture?(show answer)
My answer to architecture decision records with expiry begins with prevention, detection, response, and recovery as one system, not a purchase.
A decision is a time-bounded choice among alternatives with named assumptions, residual risk, and a review date. Diagrams without expiry become invisible production law after the assumptions die.
Concretely, write ADRs with context, options, consequences, evidence, authority, and an expiry not longer than 12 months for high-consequence choices, and verify implementation still matches before each extension.
The reason for that specificity is a failure I have seen: An approved diagram assumed a private partner link. The partner moved to the internet 11 months later. The ADR was never revisited. 4.6 million records crossed an unmonitored path for 7 weeks.
Expired assumption, live path.
| ADR | Age | Assumption still true | Records on the new path |
|---|---|---|---|
| partner private link | 18 months | no for 11 months | 4.6 million |
| after expiry reviews | 12-month max | checked quarterly | 0 unmonitored weeks |
| weeks unmonitored | 7 | — | — |
I would not consider it settled without evidence: select ten decisions and verify current implementation, assumptions, risk acceptance, expiry, and superseding records.
Approval without a calendar is folklore.
Curated: · Written: · Reviewed:
QA-41SOC 2 sampling passed. Why do you still not trust access revocation?(show answer)
I would treat continuous control evidence versus annual sample as a claim about every sibling service and template, not about the one application that was reviewed.
An annual sample proves a few moments. Access revocation fails in the months nobody sampled. Architecture collects continuous evidence of the control population — joiners, leavers, exceptions — with provenance, not a once-a-year screenshot.
Concretely, reperform representative controls on a short cadence, reconcile scope and exceptions automatically, and treat a passed sample that hides failed months as a detection miss.
The reason for that specificity is a failure I have seen: The sample of 25 leavers in April looked clean. Between May and March, 1,140 leavers waited a median 11 days. 22 still had VPN certificates when a former contractor connected in February.
April sample versus the other eleven months.
| Window | Leavers sampled | Median revoke | Live VPN certs |
|---|---|---|---|
| April SOC sample | 25 | 1 day | 0 |
| May–March | 1,140 | 11 days | 22 |
| contractor reconnect | 1 | still valid | February |
I would not consider it settled without evidence: reperform representative controls continuously and reconcile scope, population, exceptions, evidence provenance, and corrective action.
A clean April is not a clean year.
Curated: · Written: · Reviewed:
QA-42IT connected the acquired company's directory to ours so people can work on day one. What is the architectural risk?(show answer)
The useful question for M&A trust bridges is which trust assumption remains true after the next deployment.
A merger creates two trust models that were never designed to meet. A forest trust expands reachable resources and makes foreign identities security principals; it does not automatically confer parent-domain administration. Architecture uses selective authentication, SID filtering, constrained delegation, explicit ACLs, segmentation, and a funded retirement.
Concretely, inventory identities and flows, put a constrained identity bridge with MFA and reduced groups, segment networks, rehearse compromise of an acquired identity against parent resources, and give the trust an expiry aligned to migration.
The reason for that specificity is a failure I have seen: A two-way forest trust for "temporary" coexistence lasted 4 years with SID filtering off and unconstrained delegation. Compromised acquired accounts reached 9 parent admin-capable systems. The acquired estate still ran 12 unpatched domain controllers.
Temporary forest trust.
| Bridge | Intended duration | Actual | Parent admin-capable systems reached |
|---|---|---|---|
| two-way forest trust | 6 months | 4 years | 9 |
| constrained sync + MFA | 6 months | retired on time in later deal | 0 |
| unpatched acquired DCs | 12 | — | referral path |
I would not consider it settled without evidence: compromise an acquired identity and test parent resources, delegation, and privileged groups, then rehearse trust removal, staged migration, rollback, and final decommission.
A forest trust expands authorization; it does not merge domain authority.
Curated: · Written: · Reviewed:
QA-43The programme reports 14 tools deployed, 8,000 findings closed, and 97% training. How should a roadmap read instead?(show answer)
I would settle roadmap measures versus tool-count vanity by exercising the abuse case against the running path, not against the slide.
Tools, closures, and training are inputs. A security roadmap is funded work against named attack paths, with leading coverage of those paths and lagging harm — incidents, exception age, time-to-revoke. Vanity counts can rise while attack-path cost stays flat.
Concretely, tie each quarter's spend to a small set of path measures, publish exception age and residual-risk movement, and drop any KPI that can be gamed by buying a logo or closing trivia.
The reason for that specificity is a failure I have seen: Tool count went from 9 to 14. Training hit 97%. A phishing-to-admin path still needed 2 stolen SMS codes and 40 minutes, unchanged in 5 quarters, and it was used in the incident that followed.
Vanity versus path cost.
| KPI | Q1 | Q5 | Phishing-to-admin minutes |
|---|---|---|---|
| tools deployed | 9 | 14 | 40 unchanged |
| training complete | 81% | 97% | 40 unchanged |
| FIDO on admins | 12% | 12% | 40 unchanged |
| after FIDO mandate | — | 100% admins | relay failed 25 of 25 |
I would not consider it settled without evidence: compare quarterly control coverage and adversarial outcomes against residual-risk movement, exceptions, incidents, and spend.
A logo is not a shorter attack path.
Curated: · Written: · Reviewed:
QA-44Operators use the same workstation for email and for Kubernetes admin. What should the architecture separate?(show answer)
The judgement in admin plane versus data plane isolation is who may accept leftover business consequence, and until when.
The admin plane can change identity, keys, and clusters. The data plane serves customers. Mixing them on a general-purpose laptop makes every phishing mail a control-plane event. Privileged access workstations and separate identities exist to split those planes.
Concretely, issue separate admin identities, require hardened workstations or virtual admin desktops, block mail and web on those endpoints, and monitor admin-plane logins from non-admin devices.
The reason for that specificity is a failure I have seen: A cluster-admin clicked a payload in Outlook. kubectl still had a cached token. 180 workloads were wiped in 12 minutes, and the same token pulled secrets from 40 namespaces.
Shared laptop versus PAW.
| Endpoint | kubectl token | Workloads wiped | |
|---|---|---|---|
| general laptop | yes | cached 12 hours | 180 |
| admin desktop | blocked | JIT 60 minutes | 0 in phishing drill |
| namespaces with secrets pulled | 40 | — | 12 minutes |
I would not consider it settled without evidence: attempt admin API use from a mail-enabled workstation and require a policy drop, then succeed only from the admin desktop.
Inbox and cluster-admin are two different rooms.
Curated: · Written: · Reviewed:
QA-45The mesh requires mTLS everywhere. Why can service A still read service B's tenant data?(show answer)
Where candidates lose the interview on mTLS is not authorisation is calling a passed architecture workshop the control.
mTLS authenticates sockets. Authorisation decides actions on objects. A mesh that stops at encryption-in-transit leaves every authenticated workload as a peer with implied read of everything behind the port.
Concretely, pair mTLS with identity-aware authorisation on methods and tenants, default-deny in the policy store, and test that a correctly authenticated wrong-workload cannot read another tenant's rows.
The reason for that specificity is a failure I have seen: mTLS was universal. A debug service with a valid identity enumerated B's API and pulled 5.5 million tenant records because authorisation was allow-any-authenticated-peer.
mTLS plus allow-authenticated.
| Policy | Socket | Tenant rows from debug svc | Records |
|---|---|---|---|
| mTLS only | encrypted | allowed | 5.5 million |
| mTLS + method ACL | encrypted | denied | 0 of 1,000 |
| allow authenticated | — | the actual rule | — |
I would not consider it settled without evidence: call a sensitive method with a valid mesh identity that is not in the allow list and require a deny with an identity in the log.
A certificate is a nametag, not a permission.
Curated: · Written: · Reviewed:
QA-46The gateway checks JWT and rate limits, so services trust every request they receive. What breaks?(show answer)
I would answer API gateway as the only PEP by separating the target invariant from the product that is supposed to implement it.
A gateway is one enforcement point and one bypass away from being decorative. Services that trust internal headers or mesh identity without repeating authorisation fail when a peer, a job, or a mis-routed call skips the gateway.
Concretely, keep the gateway for edge concerns, repeat object-level checks in the service, reject unsanitised internal trust headers, and test calls that never touch the gateway.
The reason for that specificity is a failure I have seen: A cluster-local cron posted to an internal URL with a spoofed X-User-Id. The service trusted the header because the gateway normally set it. 16,000 account merges ran overnight.
Gateway-only PEP.
| Path | JWT checked | X-User-Id trusted | Account merges |
|---|---|---|---|
| through gateway | yes | set by gateway | — |
| cluster cron internal | no | yes | 16,000 |
| service-side membership | yes | ignored | 0 in retest |
I would not consider it settled without evidence: send a forged identity header on the internal path and require the service to ignore it and deny.
The edge is not the only door.
Curated: · Written: · Reviewed:
QA-47We encrypt PAN at rest. Reports still show full PAN to 200 analysts. What did encryption not do?(show answer)
The engineering content of tokenisation versus field encryption is the reusable pattern and its evidence, not the local ticket that closed.
Encryption at rest defends stolen disks and backups. It does not implement least privilege for live readers. Tokenisation or format-preserving vaults keep live applications and analysts on surrogates unless a rare, logged detokenisation is authorised.
Concretely, store tokens in the warehouse, detokenise only in a controlled service with dual control for bulk, and measure how many humans can still see a live PAN.
The reason for that specificity is a failure I have seen: Disk encryption was correct. 200 analysts had SELECT on the live table. A departing analyst exported 1.8 million PANs the week before HR disablement completed.
Encrypted disks, live SELECT.
| Control | Analysts with PAN | PANs exported | HR disable lag |
|---|---|---|---|
| AES at rest | 200 | 1.8 million | 6 days |
| warehouse tokens | 3 vault operators | 0 | — |
| bulk detokenise | dual control | 0 unaudited | — |
I would not consider it settled without evidence: query the warehouse as an analyst and require tokens, then detokenise one PAN only through the vault with a ticket.
At-rest encryption is not a read model.
Curated: · Written: · Reviewed:
QA-48Backups are on the same domain with the same admins. How do you make recovery survive the identity plane?(show answer)
Before calling immutable backups against ransomware done I would write down the template, region, or future service that still ships the old design.
Ransomware that owns directory also owns writable backups. Immutable, off-identity copies with separate custody and tested restore are the recovery architecture. A second folder on the same NAS is not a second world.
Concretely, write backups to an account the production identity cannot delete, lock qualifying generations for the retention window (for example 30–90 days), size copy frequency separately to meet RPO, split admin identities, and restore quarterly from that copy.
The reason for that specificity is a failure I have seen: The affiliate who took Domain Admin deleted 90 days of backups in 25 minutes. Negotiations started from zero copies. Downtime lasted 21 days.
Writable NAS versus immutable vault.
| Copy | Production DA can delete | Minutes to empty | Days down |
|---|---|---|---|
| domain NAS backups | yes | 25 | 21 |
| immutable object lock | no | denied 10 of 10 | 0 in drill restore |
| immutable retention | 30 days | RPO 15 minutes separately | — |
I would not consider it settled without evidence: from a simulated production admin, attempt delete of the immutable copy and require failure, then restore a file from it.
The same admins cannot guard the last copy.
Curated: · Written: · Reviewed:
QA-49Logs go to a hot SIEM the SOC admins also administer. Why is that a problem in a dispute?(show answer)
The first thing I would establish about forensic integrity of logs is which production path still violates the invariant after the diagram looks finished.
Evidence that the suspected operator can rewrite is advocacy, not evidence. Architecture writes an immutable, separately custodied stream with time sync, then lets the SIEM hold a working copy.
Concretely, dual-write to WORM storage with a different cloud account, alert on gaps, NTP-pin collectors, and hash batches.
The reason for that specificity is a failure I have seen: An insider truncated SIEM indices covering 14 hours of transfers. Legal had no second copy. £6.1 million in disputed payouts had no contemporaneous log.
Hot SIEM versus WORM.
| Store | Insider can truncate | Hours missing | Disputed value |
|---|---|---|---|
| hot SIEM | yes | 14 | £6.1 million |
| WORM second account | no | 0 in drill | hashes match |
| collector NTP offset | 8 seconds worst | — | — |
I would not consider it settled without evidence: delete a hot index in a drill and still produce the WORM batch hashes for that hour.
The investigator's database is not the archive.
Curated: · Written: · Reviewed:
QA-50The SaaS vendor can jump into our tenant for troubleshooting. How do you architect that?(show answer)
I would start vendor support access paths from residual risk, named authority, and expiry, not from the workshop minutes.
Support access is privileged identity with a foreign employer. Architecture makes it requestable, time-boxed, recorded, and scoped, and it treats always-on vendor admin as standing privilege you do not employ.
Concretely, disable standing vendor admin, require a ticket and expiry measured in hours, record sessions, and review every use the next business day.
The reason for that specificity is a failure I have seen: Always-on vendor admin was used from a contractor laptop at the vendor. 44,000 customer records were exported over 3 nights; the first notice was a DPA complaint.
Always-on vendor admin.
| Mode | Standing | Records exported | Nights |
|---|---|---|---|
| always-on vendor admin | yes | 44,000 | 3 |
| requestable 2-hour | no | 0 unaudited | recorded 12 of 12 drills |
| DPA complaint lag | 11 days | — | — |
I would not consider it settled without evidence: attempt vendor login without a live request and require denial, then approve a 2-hour window and confirm recording.
Their jumphost is still your tenant.
Curated: · Written: · Reviewed:
QA-51The pipeline still has an AWS access key in a secret variable, rotated twice a year. What should the architecture be?(show answer)
This is an area where a control inventory and a held production path are different artefacts.
A long-lived cloud key in CI is a production admin credential that lives in a ticket system, a logs sink, and every forked workflow. Workload federation issues a short-lived role from a signed OIDC identity of that job, so there is no static key to leak.
Concretely, delete standing access keys from pipelines, trust only the org and repo in the cloud role, constrain the role to the job's needed actions, and alert on any CreateAccessKey from humans who are not break-glass.
The reason for that specificity is a failure I have seen: A public fork of a workflow printed the secret in a debug step. The key was 7 months old and had AdministratorAccess. It created 40 GPU instances in 3 hours at $28,000 before the budget alarm.
Static pipeline key versus OIDC.
| Credential | Age | Privileges | GPU spend in 3 hours |
|---|---|---|---|
| AKIA in GitHub secret | 7 months | AdministratorAccess | $28,000 |
| OIDC job role | 1 hour token | 4 actions | $0 in replay |
| public fork debug | 1 print | — | key harvested |
I would not consider it settled without evidence: run a job that assumes the federated role, then present the old static key and require it to be disabled with AccessDenied.
A six-month key in CI is a six-month incident.
Curated: · Written: · Reviewed:
QA-52State lives in an S3 bucket the whole platform team can read. Why is that an architecture defect?(show answer)
My answer to Terraform state as a secret store begins with prevention, detection, response, and recovery as one system, not a purchase.
State files contain plaintext secrets, resource identifiers, and sometimes database passwords. Treating state as shared documentation is treating the highest-privilege map of the estate as a wiki.
Concretely, encrypt state with a restricted KMS key, lock ACLs to the pipeline identity, enable object versioning and access logs, and move secrets out of state into a manager referenced by handle.
The reason for that specificity is a failure I have seen: A contractor downloaded state to debug an apply. The file held 63 database passwords. 4 were still valid 11 months later and opened the payments replica.
World-readable state.
| Reader | Secrets in state | Still valid at month 11 | Replica opened |
|---|---|---|---|
| whole platform group | 63 | 4 | yes |
| pipeline role only | 12 handles, 0 passwords after move | 0 | no |
| contractor laptop copy | 1 local file | 4 | payments |
I would not consider it settled without evidence: attempt GetObject as a human platform role and require denial, then show the pipeline role succeeding and the secret count in state trending to zero.
State is a credential dump until you design it not to be.
Curated: · Written: · Reviewed:
QA-53The app fetches URLs the user supplies for previews. The metadata service is a hop away. What do you design?(show answer)
I would treat SSRF egress as an architecture control as a claim about every sibling service and template, not about the one application that was reviewed.
User-controlled fetch is an egress capability. Architecture denies link-local and cloud metadata, pins allowed schemes and hosts, and gives the fetcher its own identity with no role that can mint credentials.
Concretely, put the preview worker on a network with no IMDS or with hop limit 1 and IMDSv2 required, allow-list outbound domains, and strip credentials from any follow.
The reason for that specificity is a failure I have seen: A preview URL of 169.254.169.254 returned instance-role credentials. Those credentials listed 2,200 S3 buckets; 8 were world-readable from the role and held 900 GB of exports.
Preview to IMDS.
| Target | Blocked | Role creds returned | Buckets listed |
|---|---|---|---|
| 169.254.169.254 | no | yes | 2,200 |
| after IMDS hop + allow-list | yes | no | 0 |
| world-readable among them | 8 | 900 GB | — |
I would not consider it settled without evidence: request a preview of the metadata IP and of an internal admin host, and require connection refused plus an empty body.
A fetch feature is a proxy you did not mean to ship.
Curated: · Written: · Reviewed:
QA-54A file-conversion service receives a URL and an OAuth token to write the result. How can it be abused?(show answer)
The useful question for confused deputy in integrations is which trust assumption remains true after the next deployment.
A deputy with a powerful identity will act on source and destination the attacker chooses. Architecture binds both source and destination to the requesting tenant and operation, and it does not accept arbitrary bearer tokens as job parameters.
Concretely, do not accept caller-supplied bearer tokens. Exchange for a token restricted to specific source and destination objects, validate tenant ownership before any read, and write only with a service identity that cannot reach customer buckets.
The reason for that specificity is a failure I have seen: A caller selected a foreign-tenant finance object as source and an attacker-controlled destination. The converter's broad service identity read 14 GB from finance and wrote it to the attacker.
Converter as deputy.
| Binding | Token accepted | Invoices uploaded | GB |
|---|---|---|---|
| any caller-supplied token | yes | finance bucket | 14 |
| downscoped to one object | no for foreign | 0 | 0 |
| service identity write | none to customer buckets | 0 | 0 |
I would not consider it settled without evidence: attempt both a foreign source and a foreign destination and require denial before any read or write.
A deputy must validate the job, not merely possess authority.
Curated: · Written: · Reviewed:
QA-55When the authorisation service times out, the API allows the request so availability SLOs hold. What did we choose?(show answer)
I would settle fail closed on policy errors by exercising the abuse case against the running path, not against the slide.
Failing open on authorisation is a business decision to prefer availability over confidentiality, and it must be explicit, time-bounded, and monitored. Silent fail-open is an architecture that turns every dependency blip into a data incident.
Concretely, fail closed on authz timeout for high-consequence actions, cache last-known-deny more eagerly than last-known-allow, and if a degraded mode exists, name the officer, the max minutes, and the extra logging.
The reason for that specificity is a failure I have seen: A 7-minute policy-service outage failed open. Scrapers pulled 3.2 million profile records that would have been denied. The SLO dashboard stayed green.
Fail-open during a 7-minute outage.
| Mode | Authz timeout | Profiles scraped | Availability SLO |
|---|---|---|---|
| fail open | allow | 3.2 million | met |
| fail closed | 503 | 0 | 7 minutes missed |
| officer-approved degrade | 15 minutes max | extra logging | named |
I would not consider it settled without evidence: kill the policy service in a drill and require high-consequence endpoints to return 503, not 200.
A green availability tile can be an open door.
Curated: · Written: · Reviewed:
QA-56Engineering skipped prepared statements because the WAF has an SQLi rule pack. What do you say?(show answer)
The judgement in WAF as a substitute architecture is who may accept leftover business consequence, and until when.
A WAF is a filter on some HTTP shapes, not a data-access architecture. Query construction still has to be parameterised. Rule packs miss encodings, alternate endpoints, and non-HTTP workers.
Concretely, require parameterised queries in the paved data layer, treat the WAF as defence in depth, and test the API with encodings the WAF does not claim to cover.
The reason for that specificity is a failure I have seen: JSON-in-header injection bypassed the WAF SQLi pack. A single endpoint concatenated the header into SQL and dumped 440,000 customer rows in 18 minutes.
WAF pack versus concatenation.
| Layer | Blocked the JSON header | Rows dumped | Minutes |
|---|---|---|---|
| WAF SQLi pack | no | 440,000 | 18 |
| parameterised query | n/a, no injection | 0 | — |
| worker path without WAF | never inspected | same bug class | — |
I would not consider it settled without evidence: send a payload the WAF allows and the database layer must still refuse, with 0 rows returned.
A rule pack is not a query planner.
Curated: · Written: · Reviewed:
QA-57Login is limited to 10 requests per IP per minute. Why is that not an authentication architecture?(show answer)
Where candidates lose the interview on rate limiting as the only brute-force control is calling a passed architecture workshop the control.
Per-IP limits fold under botnets and shared NATs, and they punish offices. Authentication architecture uses hashing with appropriate cost, phishing-resistant MFA, per-account lockout or step-up, and credential-stuffing detection on the username, not only the address.
Concretely, rate-limit by account and by IP, require MFA on failures from new devices, block known-breached passwords, and measure stuffing success rather than IP 429 counts.
The reason for that specificity is a failure I have seen: A 4,000-IP botnet stayed under 10/min each. It tried 2.1 million passwords in 5 hours and hit 1,140 accounts that had SMS MFA off.
Per-IP cap versus botnet.
| Control | IPs | Passwords tried | Accounts hit |
|---|---|---|---|
| 10/min per IP | 4,000 | 2.1 million | 1,140 |
| per-account step-up | — | stuffing stopped at 5 failures | 0 in retest |
| SMS-off subset | 1,140 | — | 5 hours |
I would not consider it settled without evidence: run a distributed stuffing replica and require account-level controls to trip while a single office NAT still works for honest users.
Ten per address is not ten per identity.
Curated: · Written: · Reviewed:
QA-58Objects stay in eu-west-1. The KMS key is in us-east-1 because that was the first account. Is residency held?(show answer)
I would answer data residency versus key location by separating the target invariant from the product that is supposed to implement it.
Residency distinguishes data storage, plaintext processing, key control, and administrator jurisdiction. A US-region KMS key unwraps data keys; object plaintext is processed where the application decrypts. Key location alone does not locate that processing.
Concretely, keep workloads and plaintext processing in approved regions, use regional keys where policy requires, deny foreign data-plane roles, and ensure key administration cannot grant object read.
The reason for that specificity is a failure I have seen: A US workload received data-read and decrypt authority and fetched 210,000 EU-resident records into Virginia for 6 hours. The objects had remained stored in eu-west-1.
EU objects, US workload processing.
| Material | Region | US workload data-plane | EU records processed |
|---|---|---|---|
| objects stored | eu-west-1 | yes read+decrypt | 210,000 |
| after regional data-plane deny | eu-west-1 | denied 20 of 20 | 0 |
| hours of US processing | 6 | — | — |
I would not consider it settled without evidence: test a US workload's complete data-plane access, not merely a Decrypt call, and require denial; succeed only on the EU processing path.
Key location affects control and jurisdiction; it does not by itself locate plaintext processing.
Curated: · Written: · Reviewed:
QA-59Staging copies production nightly so bugs are realistic. Who is the audience of that copy?(show answer)
The engineering content of production data in non-production is the reusable pattern and its evidence, not the local ticket that closed.
Non-production has weaker identity, more engineers, and more integrations. A nightly prod clone is a second production with extra readers. Architecture uses subset, synthetic, or irreversibly masked data with a documented residual risk.
Concretely, stop full clones, generate masked subsets with tested irreversibility, and scan non-prod for live PAN, secrets, and customer email.
The reason for that specificity is a failure I have seen: Staging had 100% of production. An intern's personal AWS key was in a staging env var with read to the clone. 6.4 million emails left to a personal bucket over a weekend.
Nightly clone versus masked subset.
| Environment | Rows | Live emails | Personal-bucket export |
|---|---|---|---|
| staging clone | 6.4 million | 6.4 million | weekend |
| masked 2% subset | 128,000 | 0 | 0 |
| intern key age | 41 days | — | — |
I would not consider it settled without evidence: query staging for a known production email and PAN and require 0 hits, then show the masked subset still reproduces the bug class.
Realistic bugs are not a licence for a second warehouse.
Curated: · Written: · Reviewed:
QA-60Incident response wants a kill switch for a dangerous export. Where does that belong in the architecture?(show answer)
Before calling feature flags as containment done I would write down the template, region, or future service that still ships the old design.
A flag that only hides a button leaves the API live. Containment architecture puts authorisation and a server-side flag on the capability, with a named owner who can disable it without a deploy, and with an audit trail.
Concretely, gate high-consequence actions on a server-evaluated flag plus authorisation, test disablement in IR drills, and refuse flags that only affect CSS.
The reason for that specificity is a failure I have seen: The UI hid Export All. The JSON endpoint still ran. An attacker called it 40 times and pulled 1.2 million rows during the incident the flag was meant to stop.
CSS hide versus server gate.
| Control | API still live | Rows pulled | Calls |
|---|---|---|---|
| UI hide | yes | 1.2 million | 40 |
| server flag + authz | no | 0 | 40 of 40 403 |
| IR disable time | 2 minutes | — | no deploy |
I would not consider it settled without evidence: disable the flag and call the API directly, requiring 403, then re-enable with a logged change.
A hidden button is not a closed capability.
Curated: · Written: · Reviewed:
QA-61Hosts drift by up to 40 seconds. Why does that matter to architecture, not only to NTP?(show answer)
The first thing I would establish about time synchronisation for evidence is which production path still violates the invariant after the diagram looks finished.
Ordering of authentication, authorisation, and data-change events is how you reconstruct an incident. Unsynchronised clocks make tokens look valid, logs unmergeable, and legal timelines arguable.
Concretely, require authenticated NTP, alert beyond 50 milliseconds for identity and payment planes, and refuse to accept logs from hosts that cannot prove sync.
The reason for that specificity is a failure I have seen: A host 11 minutes fast issued JWTs that looked unexpired to a verifier 11 minutes slow. 860 requests after a revocation still passed. The merger of SIEM events put the attacker after the lockout.
Eleven minutes of lie.
| Clock | Post-revocation accepts | SIEM order | Offset |
|---|---|---|---|
| unsynced pair | 860 | attacker after lockout | 11 minutes |
| authenticated NTP | 0 | correct | 12 ms worst |
| payment plane alert | >50 ms | — | 3 pages in 30 days |
I would not consider it settled without evidence: offset a host by 5 minutes in a lab and show token and log correlation failing, then passing with pinned NTP.
A timeline is a security control.
Curated: · Written: · Reviewed:
QA-62NTLM remains for one print server. Why is that an estate-wide architecture issue?(show answer)
I would start legacy protocol islands from residual risk, named authority, and expiry, not from the workshop minutes.
A legacy protocol is a downgrade oracle and a credential-theft surface that attackers coerce other hosts into using. A captured NetNTLM challenge-response is relay or cracking material, not an NT hash that can be pass-the-hashed. Architecture isolates the island, requires SMB signing and LDAP protections, forbids outbound NTLM from user VLANs, and funds replacement.
Concretely, put the print server on an isolated VLAN with NTLM allowed only from a print spooler identity, disable or restrict NTLM elsewhere, require signing, prevent coercion, and date the replacement.
The reason for that specificity is a failure I have seen: A coerced NetNTLM response from a user VLAN was relayed to an unsigned SMB service. That path opened 3 file shares holding 70,000 contracts.
NTLM island without isolation.
| Path | NTLM allowed | Hashes captured | Contracts opened |
|---|---|---|---|
| user VLAN to DC | yes (legacy default) | 1 privileged | 70,000 |
| isolated spooler only | DC no, island yes | 0 in drill | 0 |
| replacement date | missing for 3 years | — | — |
I would not consider it settled without evidence: exercise coercion plus relay from a user subnet and require protocol protections to stop it, while the print island stays contained.
A challenge response is relay or cracking material, not directly an NT hash.
Curated: · Written: · Reviewed:
QA-63ZTNA covers HTTPS apps. Developers set DoH and talk to internal APIs by IP. What did the architecture miss?(show answer)
This is an area where a control inventory and a held production path are different artefacts.
An HTTP proxy is not a network policy. If DNS and IP paths still reach the service, the proxy is optional. Zero-trust network architecture has to close alternate names and addresses, not only browser PAC files.
Concretely, deny direct IP to internal services, resolve internal names only through the access plane, and detect DoH to public resolvers from managed devices.
The reason for that specificity is a failure I have seen: A contractor used 10.8.2.44 by IP and skipped the identity proxy. The API still trusted RFC1918 as corp. 55,000 records left without an SSO log.
HTTPS ZTNA, IP still open.
| Path | SSO log | Records | RFC1918 trust |
|---|---|---|---|
| browser via proxy | yes | — | — |
| curl by IP | none | 55,000 | yes |
| after deny-direct | dropped | 0 | removed |
I would not consider it settled without evidence: curl the service by IP from a managed laptop and require a drop, then succeed only through the identity path.
A PAC file is not a packet filter.
Curated: · Written: · Reviewed:
QA-64The architecture review passed. How do you know the controls fire together?(show answer)
My answer to purple-team evidence of architecture begins with prevention, detection, response, and recovery as one system, not a purchase.
A review scores intent. A purple exercise generates the technique and watches prevention, detection, and response on the live path. Architecture that cannot be exercised is still a diagram.
Concretely, schedule purple tests against the funded attack paths, file design bugs when prevention is absent even if an alert fired, and track time-to-detect and time-to-contain as architecture metrics.
The reason for that specificity is a failure I have seen: The review listed EDR, SSO, and segmentation. Purple testers walked the phishing-to-DA path in 55 minutes with no prevention and a detection at minute 49. The board had been told the path was closed.
Review versus 55-minute path.
| Control | Review | Purple result | Minutes |
|---|---|---|---|
| phishing-resistant admin MFA | claimed | not enforced | — |
| EDR | deployed | detection | 49 |
| segmentation | drawn | unused jump | 55 to DA |
| after redesign | FIDO + JIT | path stopped | 0 DA |
I would not consider it settled without evidence: run the named path and record which control stopped it, which only alerted, and which never saw it.
A passed review is a hypothesis.
Curated: · Written: · Reviewed:
QA-65We ran a ransomware tabletop with Legal and Comms. Why do you still want a restore rehearsal?(show answer)
I would treat tabletop versus technical rehearsal as a claim about every sibling service and template, not about the one application that was reviewed.
A tabletop tests conversation. A technical rehearsal tests backups, identity rebuild, out-of-band comms, and decision rights on real systems. Architecture needs both; only one of them proves RTO.
Concretely, run tabletops for decisions and a yearly destructive restore of a crown-jewel system into a clean account, with the people who would actually type.
The reason for that specificity is a failure I have seen: The tabletop scored well. The first real restore needed a DNS record nobody had, and identity rebuild took 29 hours against an 8-hour RTO.
Workshop score versus typed restore.
| Exercise | RTO claimed | RTO measured | Missing DNS |
|---|---|---|---|
| tabletop | 8 hours | not measured | not found |
| technical rehearsal | 8 hours | 29 hours | 1 record |
| after runbook fix | 8 hours | 6 hours 40 minutes | 0 |
I would not consider it settled without evidence: produce the last technical rehearsal's measured RTO, RPO, and the surprise that the tabletop never named.
Talking through a restore is not restoring.
Curated: · Written: · Reviewed:
QA-66The board pack is a heat map of 40 red boxes. What should a security architect put instead?(show answer)
The useful question for board reporting of residual risk is which trust assumption remains true after the next deployment.
Boards allocate capital and appetite. They need a short list of scenarios with ranges, funded treatments, exception age, and residual they are being asked to hold. Forty colours train them to ignore all of them.
Concretely, bring five scenarios, pounds and days, last purple result, and a decision requested, and retire the 40-box slide as an appendix at most.
The reason for that specificity is a failure I have seen: The heat map had been red in the same 12 cells for 6 quarters. Nobody noticed identity-plane ransomware had no funded owner until the 14-day outage.
Forty boxes versus five scenarios.
| Pack | Items | Funded owner for identity ransomware | Outage |
|---|---|---|---|
| heat map | 40 | none for 6 quarters | 14 days |
| five scenarios | 5 | CFO + CISO, dated | — |
| capital moved after numbers | £2.1m | — | next quarter |
I would not consider it settled without evidence: show the last board decision that changed spend or appetite, with the scenario numbers that drove it.
A wall of red is a request to look away.
Curated: · Written: · Reviewed:
QA-67Architecture review is a 90-minute slot in the last sprint. Champions exist on paper. How should design work actually flow?(show answer)
I would settle security champions versus a late gate by exercising the abuse case against the running path, not against the slide.
A late gate discovers frozen mistakes. Champions with a paved-road SDK, linters, and a fast consult move authorisation and crypto decisions to the week they are cheap. The architect's scarce hours go to exceptions and new threat classes.
Concretely, publish paved patterns, give champions merge rights on security tests, measure time-to-first-consult, and keep the late review for residual risk not for discovering IDOR.
The reason for that specificity is a failure I have seen: Every service waited for the 90-minute slot. 22 services shipped with copied insecure defaults. The architect found the same object-level bug 22 times in a year.
Late slot versus paved SDK.
| Path | Services | Identical IDOR | Weeks to fix each |
|---|---|---|---|
| last-sprint review | 22 | 22 | 3–6 |
| paved SDK + champion tests | 19 of 22 next year | 1 | 4 days |
| consult latency | 9 days median | — | — |
I would not consider it settled without evidence: count how many services inherited the paved object-check versus copied a raw handler in the last quarter.
Twenty-two identical findings are a platform gap.
Curated: · Written: · Reviewed:
QA-68The CMDB is 70% accurate. How do you design isolation if you cannot name the neighbours?(show answer)
The judgement in CMDB completeness as blast-radius truth is who may accept leftover business consequence, and until when.
Segmentation and crown-jewel isolation need an inventory that matches the running estate. A 70% CMDB means 30% of the blast radius is folklore. Architecture funds discovery that writes the system of record, then designs on that graph.
Concretely, reconcile cloud, k8s, and identity inventories daily, page on unknown internet-facing assets, and refuse isolation designs that assume the CMDB is complete.
The reason for that specificity is a failure I have seen: An isolation project skipped 40 unknown EC2 instances. 6 of them were internet-facing admin panels. One became the ransomware beachhead.
70% CMDB in the jewel neighbourhood.
| Source | Assets | Unknown internet-facing | Beachhead |
|---|---|---|---|
| CMDB | 70% of discovered | 6 | 1 |
| daily reconcile | 100% tagged in 45 days | 0 | — |
| skipped EC2 | 40 | 6 | ransomware |
I would not consider it settled without evidence: compare discovered internet-facing assets to the CMDB and require unknown count of zero for the crown-jewel neighbourhood.
You cannot isolate a ghost.
Curated: · Written: · Reviewed:
QA-69The old VPN was turned off. Why could partners still reach an internal API?(show answer)
Where candidates lose the interview on decommission leftover trust is calling a passed architecture workshop the control.
Decommission is the removal of identity, DNS, certificates, firewall holes, and vendor accounts, not the power-off of a box. Leftover trust is a second production.
Concretely, checklist identities, keys, DNS, network, vendors, and backups; prove each is gone; keep a watch for traffic to retired names.
The reason for that specificity is a failure I have seen: VPN concentrators were powered off. An old partner IP allow-list still hit a billing API with a 3-year-old certificate. 18,000 invoices were pulled over 9 days.
VPN off, allow-list on.
| Artefact | After power-off | Invoices pulled | Days |
|---|---|---|---|
| concentrators | off | — | — |
| partner IP allow-list | still present | 18,000 | 9 |
| 3-year cert | still valid | — | — |
| after full decommission | all gone | 0 of 100 probes | — |
I would not consider it settled without evidence: from a retired partner address space, attempt the old API and require drop, and show certificate revocation status.
Powered off is not untrusted.
Curated: · Written: · Reviewed:
QA-70Release signing uses a KMS SIGN_VERIFY key that CI can Sign without an approval gate. What is wrong?(show answer)
I would answer code-signing keys in software HSM versus general KMS by separating the target invariant from the product that is supposed to implement it.
Code-signing keys must be non-exportable, usage-separated, and dual-controlled. A SIGN_VERIFY key cannot Decrypt, and that is not the defect. Unrestricted Sign on the CI principal makes the workflow a worldwide publisher.
Concretely, use a dedicated asymmetric SIGN_VERIFY key in KMS or an HSM, deny direct Sign from CI, and place approval and release policy before signing. Private-key export remains impossible.
The reason for that specificity is a failure I have seen: A poisoned workflow abused unconditional kms:Sign and published a malicious artefact as the company. 31,000 agents updated in 8 hours.
Unconditional Sign on CI.
| Permission | CI role | Agents updated | Hours |
|---|---|---|---|
| Sign without approval | yes | 31,000 | 8 |
| Sign only through release gate | direct Sign denied | 0 in drill | — |
| poisoned workflow | 1 | — | — |
I would not consider it settled without evidence: direct CI Sign must fail; an approved release signs successfully; private key export remains impossible.
The publishing boundary is authorization to sign.
Curated: · Written: · Reviewed:
QA-71The mobile app pins a leaf certificate. Rotation is in 20 days. What will happen?(show answer)
The engineering content of pinning versus rotation is the reusable pattern and its evidence, not the local ticket that closed.
Pinning a leaf without a rotation programme bricks the fleet or forces a skipped pin. Architecture pins a backup SPKI or uses a short-lived pin set distributed in-app, and it rehearses rotation before expiry.
Concretely, pin two SPKIs, ship an update 30 days before primary expiry, and have a remotely killable pin only under dual control.
The reason for that specificity is a failure I have seen: Leaf expiry hit with 40% of users un-updated. 2.1 million sessions failed. Support told people to disable pinning via an unofficial APK, which also disabled TLS checks.
Leaf pin at expiry.
| Pin | Users on old build | Failed sessions | Unofficial APK |
|---|---|---|---|
| single leaf | 40% | 2.1 million | circulated |
| backup SPKI | 40% still old | 0 | not needed |
| days of warning unused | 20 | — | — |
I would not consider it settled without evidence: rotate the leaf in staging against the current app build and require success with the backup pin.
A pin without a rehearsal is a time bomb.
Curated: · Written: · Reviewed:
QA-72Procurement bought a SASE suite. Sales said it is zero trust. How do you tell?(show answer)
Before calling SASE branding versus ZTNA resource policy done I would write down the template, region, or future service that still ships the old design.
SASE is a packaging of SWG, CASB, and sometimes ZTNA. Zero trust is still per-request policy on the resource. A SASE pop that then hairpins onto a flat data-centre VLAN is a new office network in the cloud.
Concretely, require resource-level authorisation independent of the pop, measure whether RFC1918 remain reachable after login, and refuse to record ZT complete based on licence seats.
The reason for that specificity is a failure I have seen: Seats showed 100% SASE. After the pop, users had /16 reach. A stolen laptop browsed the HR file share without a per-resource check. 8,800 files synced in 70 minutes.
SASE seats, flat /16.
| Metric | Value | HR share without entitlement | Files synced |
|---|---|---|---|
| SASE seats | 100% | allowed | 8,800 |
| resource policy | not deployed | — | 70 minutes |
| after ZTNA on the share | denied | 0 | — |
I would not consider it settled without evidence: after SASE login, attempt a resource the user should not get and a flat-subnet scan, and require both to fail.
A seat count is not a request policy.
Curated: · Written: · Reviewed:
QA-73The threat model is only external attackers. Who else can mint a payout?(show answer)
The first thing I would establish about insider-capable paths in the design is which production path still violates the invariant after the diagram looks finished.
Anyone with standing production access is an insider-capable path: SRE, vendors, broken deputies. Architecture minimises those paths, dual-controls money movement, and monitors them as first-class threats, without turning the workplace into surveillance theatre.
Concretely, list humans and vendors who can cause the top losses, put dual control on those actions, and alert on break-glass and bulk export.
The reason for that specificity is a failure I have seen: A single SRE role could approve payouts in the admin UI to test. That role exported a 14,000-row payout file to personal email the week the person resigned.
SRE as payout oracle.
| Actor | Dual control | Rows to personal email | Days before last day |
|---|---|---|---|
| SRE admin UI | no | 14,000 | 7 |
| after dual control | yes | 0 | blocked |
| vendor copies of the role | 3 | — | — |
I would not consider it settled without evidence: show dual control on a test payout and a blocked bulk export from an SRE identity.
External-only models leave the operator unmodelled.
Curated: · Written: · Reviewed:
QA-74Privacy signed a DPIA. Security signed a threat model. Neither mentioned the analytics sidecar. What failed?(show answer)
I would start DPIA and security review as one crossing from residual risk, named authority, and expiry, not from the workshop minutes.
Personal data crossings are security crossings. A DPIA that ignores how the sidecar authenticates, and a threat model that ignores lawful basis, produce two incomplete architectures.
Concretely, use one data-flow inventory for privacy and security, require both sign-offs on new processors, and test the sidecar's identity and retention.
The reason for that specificity is a failure I have seen: The sidecar used a shared write key and stored raw emails for 13 months. A researcher scraped 1.1 million addresses from an open project bucket the sidecar used.
Unmentioned sidecar.
| Artefact | Sidecar present | Retention | Addresses scraped |
|---|---|---|---|
| DPIA | no | — | — |
| threat model | no | — | — |
| actual sidecar | yes | 13 months | 1.1 million |
| shared write key | 1 | open bucket | yes |
I would not consider it settled without evidence: list processors from the DPIA and from flow logs, and require the sets to match plus an identity model for each.
Two reviews of two diagrams is zero reviews of the sidecar.
Curated: · Written: · Reviewed:
QA-75Leadership says insurance covers ransomware so identity work can wait. How do you answer?(show answer)
This is an area where a control inventory and a held production path are different artefacts.
Insurance transfers some residual financial loss under conditions. It does not restore customers, keep regulators away, or replace MFA. Architecture still owns prevention and recovery; insurance is not a control owner.
Concretely, map policy exclusions to the actual architecture, fund the controls the policy silently requires, and never list the insurer as the residual-risk owner.
The reason for that specificity is a failure I have seen: The policy excluded unpatched internet-facing SSO. That was the entry. The insurer denied £18 million of the claim after the 14-day outage.
Exclusion meeting the real path.
| Item | Status | Claim | Outage |
|---|---|---|---|
| ransomware premium | paid | denied on SSO exclusion | 14 days |
| internet SSO unpatched | 61 days | £18 million refused | — |
| identity work deferred | 5 quarters | — | — |
I would not consider it settled without evidence: read the exclusions against the current attack path and show which funded control closes each exclusion.
A premium is not a privileged-access programme.
Curated: · Written: · Reviewed:
QA-76There are 310 open security exceptions. What does that say about the standard?(show answer)
My answer to exception register as living architecture begins with prevention, detection, response, and recovery as one system, not a purchase.
Exceptions are the real architecture if they never expire. A register without owners, dates, and monitors is how a standard dies. Architecture caps concurrent exceptions, ages them in public, and funds closure.
Concretely, require owner, expiry, compensation, and monitor on every exception, auto-close or escalate at 90 days, and report the oldest 10 to the risk committee.
The reason for that specificity is a failure I have seen: 310 exceptions, median age 14 months. One was public RDP for a vendor. It was the beachhead for the incident that followed, still marked temporary.
310 temporaries.
| Register | Count | Median age | Public RDP |
|---|---|---|---|
| open exceptions | 310 | 14 months | still open |
| after 90-day max | 40 | 41 days | closed |
| beachhead | that RDP | — | incident |
I would not consider it settled without evidence: produce the age histogram, the oldest exception's last monitor result, and the last exception that actually closed on time.
Temporary at year two is the design.
Curated: · Written: · Reviewed:
QA-77Endpoint DLP is deployed with a 10,000-regex pack. Why do analysts still email PAN?(show answer)
I would treat DLP without classification as a claim about every sibling service and template, not about the one application that was reviewed.
DLP without a classification and handling model is a noisy filter. Architecture labels data at creation, restricts channels by label, and uses DLP as a detective backstop, not as the only definition of sensitive.
Concretely, bind labels to channels, reduce regex to the few patterns that match the labels, and measure true-positive rate on planted documents.
The reason for that specificity is a failure I have seen: The regex pack generated 18,000 weekly alerts. Analysts ignored them. 400 spreadsheets with PAN left on USB in a month, none of which matched a brittle regex.
Regex pack versus labelled USB.
| Control | Weekly alerts | PAN on USB | True positives on plants |
|---|---|---|---|
| 10,000 regex | 18,000 | 400 sheets | 1 of 20 |
| label-bound channels | 40 | 0 | 20 of 20 |
| analyst ignore rate | high | — | — |
I would not consider it settled without evidence: plant labelled and unlabelled PAN files on a USB path and show the labelled flow blocks while the regex-only estate misses.
Ten thousand patterns are not a data model.
Curated: · Written: · Reviewed:
QA-78CASB found 900 SaaS apps. 80 have tenant tokens. How do you architect the response?(show answer)
The useful question for shadow SaaS and OAuth sprawl is which trust assumption remains true after the next deployment.
Shadow SaaS is unsanctioned processing and unsanctioned identity. Architecture has a sanctioned catalogue, blocks high-risk OAuth classes, and gives teams a paved path so they stop pasting tokens into random apps.
Concretely, revoke unknown high-scope grants, publish a 30-app sanctioned set with SSO, and detect new admin-consent weekly.
The reason for that specificity is a failure I have seen: A to-do app with Files.ReadWrite.All synced 2.8 million documents to a consumer account after an intern clicked Allow.
Allow on a to-do app.
| Apps | Tenant tokens | Documents synced | Intern click |
|---|---|---|---|
| discovered | 900 | 80 | — |
| to-do with Files.ReadWrite.All | 1 | 2.8 million | Allow |
| sanctioned catalogue | 30 | 0 unknown high-scope | blocked |
I would not consider it settled without evidence: list admin-consent grants, revoke a lab grant, and show the app loses Files access within the SLO.
Nine hundred logos are nine hundred processors.
Curated: · Written: · Reviewed:
QA-79The mesh has a wide-open allow, but SPIFFE IDs exist. Are we identity-aware?(show answer)
I would settle service-mesh policy versus workload identity by exercising the abuse case against the running path, not against the slide.
An identity that is never consulted is inventory. Mesh architecture evaluates SPIFFE IDs on each method. Wide-open policy with pretty identities is mTLS with extra YAML.
Concretely, default-deny mesh policy, allow named SPIFFE IDs on named methods, and test a valid identity on the wrong method.
The reason for that specificity is a failure I have seen: Every SPIFFE ID could call payments.Settle. A canary workload with a valid identity settled 9,200 test pennies against live merchants for 3 hours.
SPIFFE with open mesh.
| Policy | SVID valid | Live settlements | Hours |
|---|---|---|---|
| allow all | yes | 9,200 pennies | 3 |
| named ID on Settle | yes, wrong ID | 0 | — |
| canary identity | valid | the caller | — |
I would not consider it settled without evidence: call Settle from a workload not on the allow list and require deny despite a valid SVID.
Issued identities are not enforced identities.
Curated: · Written: · Reviewed:
QA-80The diagram says private buckets. A decade of ACLs says otherwise. Which is the architecture?(show answer)
The judgement in object ACLs that outlive the architecture diagram is who may accept leftover business consequence, and until when.
Effective access is the union of identity policies, bucket policies, and leftover ACLs. Diagrams that ignore ACLs are fiction. Architecture forbids object ACLs, uses bucket policies plus IAM, and inventories leftovers.
Concretely, block public ACLs at the org, migrate to bucket-owner-enforced, and alert on any ACL grant to AllUsers or a foreign account.
The reason for that specificity is a failure I have seen: A 2015 ACL granted AllUsers read on a prefix thought empty. It held 12 years of scanned passports, 41,000 files, discovered by a student researcher.
2015 AllUsers prefix.
| Source | Public objects | Passport scans | Years sitting |
|---|---|---|---|
| architecture diagram | 0 claimed | — | — |
| ACL union | 41,000 | 41,000 | 12 |
| after owner-enforced | 0 | 0 public | — |
I would not consider it settled without evidence: enumerate ACL grants estate-wide and require zero AllUsers plus a dated exception for any foreign-account grant.
The oldest ACL is the current design.
Curated: · Written: · Reviewed:
QA-81Workloads use IMDS for roles. Why must hop count and v2 be architecture standards?(show answer)
Where candidates lose the interview on instance metadata as a hop identity is calling a passed architecture workshop the control.
Instance metadata is an identity provider on loopback. IMDSv1 is reachable from an SSRF that can GET the link-local address; IMDSv2 raises the bar but is not a complete defence when the fetcher can PUT and set headers. Hop limit 1 can also break containers because TTL is consumed crossing the network hop.
Concretely, disable IMDS where the workload does not need it. Otherwise require IMDSv2, least-privilege instance roles, and application-level blocking of metadata egress. Choose hop limit for the real topology — containers often need more than 1 — and do not treat hop limit as the SSRF boundary.
The reason for that specificity is a failure I have seen: A CMS preview fetcher supported redirects and the IMDSv2 token PUT. It obtained the instance role and AssumeRole into billing. £210,000 of crypto miners appeared over 11 hours.
SSRF that could PUT for IMDSv2.
| IMDS | Fetcher | Billing AssumeRole | Miner spend |
|---|---|---|---|
| v2 + header-capable SSRF | yes | yes | £210,000 |
| IMDS disabled on CMS | n/a | no | £0 |
| hours | — | — | 11 |
I would not consider it settled without evidence: test through the vulnerable URL-fetch interface, including redirects, alternate IP encodings, and header-capable requests, and require the role not to be reachable.
IMDSv2 raises the bar; it does not sanitize an arbitrary fetcher.
Curated: · Written: · Reviewed:
QA-82A partner can AssumeRole into our logs account to debug. What must the architecture state?(show answer)
I would answer cross-account assume as a trust boundary by separating the target invariant from the product that is supposed to implement it.
Every AssumeRole trust is a foreign identity plane. Architecture sets external IDs, constrains actions, logs the partner principal, and expires the trust.
Concretely, require sts:ExternalId, deny without it, scope to GetObject on one prefix, and review trusts quarterly.
The reason for that specificity is a failure I have seen: The trust had no external ID. Another customer of the partner's AWS account assumed our role and read 8 million log lines including session tokens.
Partner role without ExternalId.
| Check | Other customer of partner | Log lines | Tokens in logs |
|---|---|---|---|
| no ExternalId | assumed | 8 million | 240 |
| ExternalId + prefix | denied | 0 | 0 |
| quarterly review | missed 5 quarters | — | — |
I would not consider it settled without evidence: assume without the external ID and from a foreign account, requiring both denied.
A trust policy is a door with a name on it.
Curated: · Written: · Reviewed:
QA-83Debug logs include request bodies for support. What is the architectural conflict?(show answer)
The engineering content of logging that creates a second copy of secrets is the reusable pattern and its evidence, not the local ticket that closed.
Bodies contain passwords, tokens, and health data. Architecture redacts at the edge, forbids body logging on auth and payment routes, and treats a log store as a regulated database if it still holds them.
Concretely, allow-list fields, drop Authorization and password, encrypt the log store, and scan for jwt and PAN.
The reason for that specificity is a failure I have seen: A debug flag stayed on for 26 days. 1.4 million Authorization headers landed in a SIEM that 60 contractors could query.
Debug bodies in SIEM.
| Flag | Days | Authorization headers | Contractor readers |
|---|---|---|---|
| debug bodies | 26 | 1.4 million | 60 |
| allow-listed fields | off | 0 canaries | — |
| payment routes | still logged | 44,000 PANs | 60 |
I would not consider it settled without evidence: send a login with a canary password and require 0 hits in logs after 5 minutes.
A helpful body dump is a credential warehouse.
Curated: · Written: · Reviewed:
QA-84SIEM retention is 30 days to save money. Legal hold needs 2 years for a case. Who wins?(show answer)
Before calling legal hold versus SIEM retention done I would write down the template, region, or future service that still ships the old design.
Retention is a per-class decision. Cheap default retention that destroys hold-relevant evidence is a legal architecture failure. Security architecture defines classes: ephemeral debug, operational 90 days, hold-capable WORM.
Concretely, keep an immutable hold store that can accept a case identifier, and stop using the 30-day hot SIEM as the only copy.
The reason for that specificity is a failure I have seen: A case landed on day 45. The only copy had aged out at day 30. 11,000 relevant events were gone.
Day-30 drop versus hold.
| Store | Retention | Events at day 45 | Case |
|---|---|---|---|
| hot SIEM | 30 days | 0 | failed |
| WORM hold | 2 years | 11,000 | readable |
| money saved | £40k/year | — | £ lost in discovery |
I would not consider it settled without evidence: place a test hold and show events still readable at day 400 in WORM while hot SIEM is empty.
A cost optimisation is not a preservation policy.
Curated: · Written: · Reviewed:
QA-85API CORS is Access-Control-Allow-Origin star with credentials. Why is that an architecture bug?(show answer)
The first thing I would establish about CORS wildcard as a browser trust boundary is which production path still violates the invariant after the diagram looks finished.
CORS is a browser-enforced trust boundary. A wildcard with credentials is not a legal configuration in browsers, and a reflected origin with credentials hands tokens to any origin that can induce a request.
Concretely, reflect only an allow-list of exact origins, never star with cookies, and test a foreign origin.
The reason for that specificity is a failure I have seen: The API reflected any Origin and sent Access-Control-Allow-Credentials true. A malicious site read 19,000 user records from sessions of people who visited it.
Reflected Origin with credentials.
| Header | Foreign origin | Records read | Users lured |
|---|---|---|---|
| reflect + credentials | allowed | 19,000 | 640 |
| exact allow-list | blocked | 0 | — |
| star + credentials | invalid / ignored | — | — |
I would not consider it settled without evidence: from a foreign origin with credentials, require the browser to hide the body, proven in an automated test.
Any origin is every origin.
Curated: · Written: · Reviewed:
QA-86Security groups and WAF are complete for IPv4. IPv6 is enabled for future-proofing. What is missing?(show answer)
I would start IPv6 control parity from residual risk, named authority, and expiry, not from the workshop minutes.
A second address family is a second attack surface. Architecture applies the same default-deny, logging, and identity-aware access to v6, or it disables v6 until that exists.
Concretely, inventory v6 routes, clone deny policies, and test the same abuse cases on v6.
The reason for that specificity is a failure I have seen: SSH was denied on v4 and open on ::/0. 22 instances were brute-forced; 2 fell. The IPv4 dashboard stayed clean.
v4 deny, v6 ::/0.
| Family | SSH | Instances fallen | Dashboard |
|---|---|---|---|
| IPv4 | deny | 0 | clean |
| IPv6 | ::/0 | 2 of 22 | no v6 tiles |
| after parity | deny | 0 | both |
I would not consider it settled without evidence: nmap the dual-stack VIP on v6 for the ports v4 denies, requiring the same deny.
Future-proofing is not an open protocol.
Curated: · Written: · Reviewed:
QA-87DR is a cold copy with last year's IAM and public snapshots to make restore easy. What did we create?(show answer)
This is an area where a control inventory and a held production path are different artefacts.
DR is production with a worse on-call. Public snapshots, standing keys, and skipped detections make it the preferred target. Architecture applies the same identity, encryption, and monitoring bars, plus a tested failover that does not lower them.
Concretely, private snapshots, same org policies, monitoring on, and a failover rehearsal that keeps identity standards.
The reason for that specificity is a failure I have seen: A public DR snapshot of the customer DB sat for 200 days. It contained 4.9 million rows. A scanner found it; a journalist called before SOC did.
Public DR snapshot.
| DR control | State | Rows | Days public |
|---|---|---|---|
| snapshot ACL | public | 4.9 million | 200 |
| IAM | last year's standing Owner | — | — |
| after org block | private | 0 public | 0 |
I would not consider it settled without evidence: list snapshots and DR IAM, requiring 0 public and 0 standing admin beyond break-glass.
Easy restore is not public restore.
Curated: · Written: · Reviewed:
QA-88The IdP rotated signing keys. Two APIs still cache JWKS for 24 hours. What happens to tokens?(show answer)
My answer to JWKS rotation and stale verifiers begins with prevention, detection, response, and recovery as one system, not a purchase.
Token verification is part of the identity architecture. Planned rollover overlaps keys so honest tokens keep verifying. Compromise rotation revokes the old key immediately and accepts disruption. A stale cache that keeps a compromised kid is a standing grant to the attacker.
Concretely, for planned rollover, publish overlapping kids sized to token lifetime and cache TTL, cache JWKS no more than 15 minutes, and refresh on unknown kid. On compromise, push or invalidate caches and stop honouring the old kid immediately.
The reason for that specificity is a failure I have seen: After a compromise rotation, 3 APIs kept a 24-hour JWKS cache and accepted the old kid. 7,200 tokens minted by the attacker still passed.
24-hour JWKS cache after compromise.
| Verifier cache | Old kid after compromise | Attacker tokens | Hours |
|---|---|---|---|
| 24 hours, no invalidate | accepted | 7,200 | 24 |
| push invalidate | rejected | 0 after push | — |
| planned overlap (honest rollover) | 12 hours both kids | honest tokens OK | not used in compromise |
I would not consider it settled without evidence: test planned rollover for continuity and emergency revocation separately, and show compromised kids rejected as soon as caches are invalidated.
Overlap is for planned rollover, never for a compromised key.
Curated: · Written: · Reviewed:
QA-89Fraud wants the full payment graph. Marketing wants the same table. Can one lake serve both?(show answer)
I would treat purpose limitation in analytics pipelines as a claim about every sibling service and template, not about the one application that was reviewed.
Purpose limitation is an access architecture. One lake with one IAM role for all analysts collapses purposes. Architecture separates datasets or attribute views, with contracts and monitoring for each purpose.
Concretely, split fraud and marketing views, drop PAN from marketing, and log purpose on query.
The reason for that specificity is a failure I have seen: Marketing queried the fraud table because the lake role allowed it. 880,000 full PANs sat in a campaign tool for 6 weeks.
Shared lake role.
| Purpose | Table used | PANs in campaign tool | Weeks |
|---|---|---|---|
| marketing | fraud graph | 880,000 | 6 |
| after view split | marketing view | 0 | — |
| fraud role | full graph | stays in fraud VPC | — |
I would not consider it settled without evidence: query as marketing and as fraud, requiring PAN only in fraud, with purpose tags on the audit log.
One role cannot hold two purposes.
Curated: · Written: · Reviewed:
QA-90The checkout page loads four marketing tags. Who can read PAN at the keyboard?(show answer)
The useful question for third-party JavaScript as a control plane is which trust assumption remains true after the next deployment.
Script on a payment origin is in the trusted computing base. Architecture minimises tags, uses subresource integrity or a tag manager with CSP, and keeps PAN on a separate origin or fields the browser isolates.
Concretely, cSP allow-list, SRI where possible, separate payment origin, and a change control on any new script.
The reason for that specificity is a failure I have seen: A tag manager published a malicious update. It skated 64,000 PANs to an external host for 31 hours before CSP would have blocked it, had it existed.
Four tags on checkout.
| Origin | Marketing tags | PANs skated | Hours |
|---|---|---|---|
| checkout | 4 | 64,000 | 31 |
| after CSP + pay origin | 0 on pay | 0 | — |
| tag manager update | unsigned | the vector | — |
I would not consider it settled without evidence: load checkout with a foreign script and require CSP to block, and show PAN origin has 0 marketing tags.
A tag is remote code on the till.
Curated: · Written: · Reviewed:
QA-91The CEO refuses a security key because travel is hard. SMS is allowed as an exception. What is the estate impact?(show answer)
I would settle executive hardware-key exception by exercising the abuse case against the running path, not against the slide.
Executives are high-value phishing targets. An SMS exception on that identity is the preferred path into the company. Architecture offers a second roaming authenticator, not a weaker factor, and it does not let status cancel phishing resistance.
Concretely, issue two FIDO devices plus a tightly controlled recovery, and refuse SMS on executive and finance roles.
The reason for that specificity is a failure I have seen: The CEO SMS exception was phished. The session approved a $4.2 million wire from the finance tool that trusted executive SSO.
CEO SMS exception.
| Factor | Relay | Wire approved | Amount |
|---|---|---|---|
| SMS exception | yes | yes | $4.2 million |
| two FIDO keys | no | no in drill | $0 |
| finance tool trust | executive SSO | — | — |
I would not consider it settled without evidence: show relay against the executive account fails with FIDO and succeeds with SMS in a lab, then keep SMS off.
Rank is not a second factor.
Curated: · Written: · Reviewed:
QA-92Plant historians now sit on the corporate IdP so dashboards are easy. What trust did we create?(show answer)
The judgement in OT and IT identity bridging is who may accept leftover business consequence, and until when.
OT outages harm people and physical process. Bridging corporate identity into control networks imports phishing and ransomware into safety systems. One-way historian export uses a data diode or unidirectional gateway. Interactive administration is bidirectional and needs a separately brokered, MFA/JIT, recorded jump — not a "unidirectional jump."
Concretely, deny corporate workstations on OT, replicate historian data through a diode, use a monitored jump with MFA/JIT and no mail where administration is unavoidable, and keep historians from being domain-joined to the corporate forest.
The reason for that specificity is a failure I have seen: Ransomware on the corporate IdP disabled historian logons and, through a wide trust, stopped a packing line for 16 hours. Safety was luck, not design.
Corporate IdP on the plant.
| Bridge | Packing line | Hours | Corporate ransomware |
|---|---|---|---|
| historians on corp IdP | stopped | 16 | yes |
| diode export + recorded jump | running in drill | 0 | cannot reach OT |
| domain join | yes | — | ticket path |
I would not consider it settled without evidence: from a corporate malware-like identity, attempt OT logon and historian admin, requiring deny.
Easy dashboards are a safety decision.
Curated: · Written: · Reviewed:
QA-93Admin APIs use cookies and have no CSRF token because they are JSON. Is that enough?(show answer)
Where candidates lose the interview on CSRF on state-changing admin APIs is calling a passed architecture workshop the control.
Cookie-authenticated state change is CSRF-vulnerable unless SameSite, custom headers, or tokens bind the request to the app's origin. Requiring application/json helps only when the server rejects simple form and text types and CORS does not authorize the attacker origin. Browsers send cookies cross-site; arbitrary JSON usually needs a preflight.
Concretely, reject text/plain and form content types for the admin API, require a CSRF token or a custom header that CORS would preflight, set SameSite=strict on admin cookies, and audit CORS.
The reason for that specificity is a failure I have seen: The invite API accepted text/plain bodies as JSON. A forum page issued POSTs that the browser treated as simple requests. 120 staff were invited to an attacker tenant over 2 days using the victim's admin cookie.
Cookie admin API.
| Defence | Cross-site invite | Staff added | Days |
|---|---|---|---|
| JSON only | succeeded | 120 | 2 |
| CSRF token + SameSite | 403 | 0 | — |
| CORS preflight header | required | 0 | — |
I would not consider it settled without evidence: test both a foreign form POST and a preflighted JSON fetch, and require 403 with cookies still present.
JSON helps only when content-type enforcement and CORS make it non-simple.
Curated: · Written: · Reviewed:
QA-94Workers pickle jobs from a Redis queue that several apps write. What is the trust model?(show answer)
I would answer deserialization of untrusted jobs by separating the target invariant from the product that is supposed to implement it.
Deserializing attacker-controlled objects is remote code execution. Architecture uses a constrained schema (JSON with explicit types), signs jobs, and never pickle.load untrusted bytes.
Concretely, replace pickle with JSON schema, HMAC the job, and isolate the worker identity.
The reason for that specificity is a failure I have seen: An app with queue write injected a pickle gadget. The worker ran as cluster-admin and created a cron that lasted 20 days.
Pickle jobs as RCE.
| Format | Gadget | Worker role | Persistence days |
|---|---|---|---|
| pickle | executed | cluster-admin | 20 |
| signed JSON schema | rejected | least privilege | 0 |
| writers of the queue | 6 apps | — | — |
I would not consider it settled without evidence: enqueue a gadget payload and require reject at parse, with 0 code execution.
A queue is an RPC, not a toy.
Curated: · Written: · Reviewed:
QA-95Users upload CVs as PDF to a bucket that CloudFront serves. What can go wrong?(show answer)
The engineering content of file-upload architecture is the reusable pattern and its evidence, not the local ticket that closed.
Uploads are untrusted content. Architecture checks type by content not name, stores outside the app origin, sets nosniff and content-disposition, and scans. Serving user files from the app origin is XSS and malware distribution.
Concretely, separate origin, rewritten names, malware scan, and never execute-capable MIME from that bucket.
The reason for that specificity is a failure I have seen: A CV named resume.pdf was HTML with a script. It was served from the app origin and stole 8,400 session cookies from recruiters.
PDF name, HTML body.
| Serve from | MIME sniffed | Cookies stolen | Recruiters |
|---|---|---|---|
| app origin | text/html | 8,400 | 210 |
| isolated origin + nosniff | application/octet-stream | 0 | — |
| scan | missed HTML | — | — |
I would not consider it settled without evidence: upload HTML as PDF and require stored type not executable and a different origin, with CSP blocking inline.
A bucket on your origin is your origin.
Curated: · Written: · Reviewed:
QA-96Sessions are 12-hour random cookies over TLS. Why can a stolen cookie from another country still work?(show answer)
Before calling session cookies without binding done I would write down the template, region, or future service that still ships the old design.
A bearer cookie is a password with a TTL. Architecture binds sessions to device or network signals for high-consequence apps, rotates on privilege change, and revokes on leaver in minutes.
Concretely, bind a device signal, shorten TTL for admin, and push revocation to a denylist checked on each request.
The reason for that specificity is a failure I have seen: A cookie dumped from a phishing page was used from another ASN. It approved 44 refunds in 90 minutes because nothing bound the session.
Stolen 12-hour cookie.
| Binding | Foreign ASN | Refunds | Minutes |
|---|---|---|---|
| none | accepted | 44 | 90 |
| device + step-up | challenged | 0 | — |
| leaver revoke | denylist | 0 after 4 minutes | — |
I would not consider it settled without evidence: replay a cookie from a second ASN and require step-up or deny on the finance app.
TLS in transit does not bind the cookie to a person.
Curated: · Written: · Reviewed:
QA-97The API trusts any client that presents a valid user token. A cloned APK talks to prod. What is missing?(show answer)
The first thing I would establish about mobile attestation versus a cloned app is which production path still violates the invariant after the diagram looks finished.
User auth is not app integrity. High-consequence mobile APIs need attestation of an unmodified app and a genuine device, plus pinning, or they are a documented API for malware.
Concretely, require Play Integrity or DeviceCheck on sensitive calls, refuse emulators for those calls, and rotate API keys out of the APK.
The reason for that specificity is a failure I have seen: A cloned APK automated 21,000 loyalty-point drains using stolen user tokens. The API could not tell it from the store build.
Cloned APK, real user tokens.
| Client | Attestation | Point drains | Tokens |
|---|---|---|---|
| cloned APK | none | 21,000 | stolen |
| store build + integrity | pass | 0 fraud in drill | — |
| API key in APK | extractable | — | 1 static |
I would not consider it settled without evidence: call the sensitive endpoint from an emulator without attestation and require deny, then pass from a genuine debug build in a lab profile.
A user token inside malware is still malware.
Curated: · Written: · Reviewed:
QA-98Product wants full DOM replay including payment fields for UX. Security wants nothing. What is the architecture?(show answer)
I would start privacy-preserving telemetry versus raw payloads from residual risk, named authority, and expiry, not from the workshop minutes.
Telemetry is a processing purpose. Architecture samples, masks payment and health fields at the SDK, and keeps a separate, dual-controlled store if raw replay is truly required for a tiny population.
Concretely, block payment selectors in the SDK, hash identifiers, and contract a 7-day max for any raw replay with named tickets.
The reason for that specificity is a failure I have seen: DOM replay captured CVV in 2,400 sessions because the selector list missed iframe fallbacks. PCI scope exploded overnight.
Replay catching CVV.
| SDK mask | Sessions with CVV | PCI scope | Days retained |
|---|---|---|---|
| missed iframe | 2,400 | exploded | 30 |
| payment selectors + iframe | 0 of 500 tests | unchanged | 7 if ticketed |
| dual-control raw | 12 sessions | ticketed | 7 |
I would not consider it settled without evidence: run a payment in a test session and require 0 CVV or PAN in the telemetry store.
UX replay is not a second cardholder-data environment — unless you let it be.
Curated: · Written: · Reviewed:
QA-99A segmentation change ships Friday with no documented rollback. Why is that a security-architecture issue?(show answer)
This is an area where a control inventory and a held production path are different artefacts.
A failed security change can lock out IR or fail open. Architecture requires a rehearsed rollback that restores a known-good policy, not a hope to unpick YAML at 02:00.
Concretely, pair every high-consequence policy change with a tested rollback artefact, a time-to-undo SLO, and a person on call who did the rehearsal.
The reason for that specificity is a failure I have seen: A deny rule shipped with a typo that blocked the EDR channel. Rollback took 6 hours. An active attacker noticed the blind window and dumped 3 file servers.
Typo deny, 6-hour undo.
| Change | EDR | Rollback | Servers dumped |
|---|---|---|---|
| Friday deny typo | blocked 6 hours | untested | 3 |
| rehearsed rollback | restored in 11 minutes | artefact | 0 in drill |
| SLO | 30 minutes | missed | — |
I would not consider it settled without evidence: apply the change in a clone, break it on purpose, and measure rollback under the SLO.
A control you cannot undo will be undone by the outage.
Curated: · Written: · Reviewed:
QA-100A broken object check was patched on the service that paged. Twenty-three siblings still copy the old handler. Are we done?(show answer)
My answer to reusable patterns after a class of failure begins with prevention, detection, response, and recovery as one system, not a purchase.
A local patch treats an incident as a ticket. Architecture treats it as a class: the template, the SDK, the paved module, and the next service that will be generated on Monday. Until those change, the estate still ships the weakness.
Concretely, fix the shared library or template first, block merges that instantiate the old handler, backfill siblings on a dated plan, and prove a newly generated service inherits the control.
The reason for that specificity is a failure I have seen: The paged service was patched in 6 hours. The cookie-cutter still emitted the old handler. 19 new services in the next quarter copied it, and a second tenant-cross hit 7 of them before the template changed.
One patch, nineteen new copies.
| Scope | Object check | New services next quarter | Tenants crossed |
|---|---|---|---|
| paged service | patched in 6 hours | — | first incident |
| cookie-cutter template | still old | 19 | 7 of 19 |
| after template + merge gate | inherited | 8 more, all pass | 0 |
I would not consider it settled without evidence: scaffold a service from the current template, run the cross-tenant test, and require a pass without a bespoke patch, then list remaining siblings with owners and dates.
The next clone is the incident you already had.
Curated: · Written: · Reviewed:
