Top 100 Cybersecurity Engineer Interview Questions and Answers
The questions most likely to actually come up in your Cybersecurity Engineer interview, ranked by likelihood — with detailed, senior-level answers covering what an interviewer is really listening for.
Curated: · Written: · Reviewed:
QA-1A dashboard says you have 85 percent detection coverage. What would you check before believing it?(show answer)
The first thing I would establish about what detection coverage actually means is what it costs an attacker to get past it.
Coverage claimed from a rule inventory says a rule exists, not that the telemetry beneath it is present, that the rule matches real behaviour, or that the alert carries enough context to act on. Each of those can fail independently while the cell stays green.
Concretely, generate the technique in a controlled way, then record which stage failed when nothing useful appeared — telemetry missing, rule did not match, or alert lacked context — and fix that stage rather than adding another rule.
The reason for that specificity is a failure I have seen: A team claiming 85 percent coverage across 40 techniques ran an exercise and alerted correctly on 11; 18 had no telemetry enabled at all, which no additional rule would have fixed.
Where the 85 percent went.
| Stage | Techniques |
|---|---|
| alerted with context | 11 |
| telemetry missing | 18 |
| rule did not match | 8 |
| alert lacked context | 3 |
I would not consider it settled without evidence: Report coverage as techniques exercised and confirmed end to end, with the failing stage named for each that did not.
A rule with no telemetry beneath it detects nothing.
Curated: · Written: · Reviewed:
QA-2A rule fires 200 times a week and 3 are real. Would you keep it?(show answer)
I would start alert precision and analyst trust from what the telemetry can actually show, not from the rule list.
A rule with low precision consumes the attention that a high-precision rule needs, and it trains analysts to dismiss its category. Past a threshold the rule is worse than nothing, because it degrades the response to everything else.
Concretely, measure precision per rule, tune with context that removes the common benign cause rather than raising a threshold blindly, and demote a rule that cannot be tuned to a hunting input rather than leaving it paging.
The reason for that specificity is a failure I have seen: A noisy rule at 1.5 percent precision was left in place; when it fired on a genuine intrusion the analyst closed it in 40 seconds as another false positive, and the intrusion ran for a further 9 days.
Rule performance over one quarter.
| Rule | Fires/week | True positives | Precision | Median triage |
|---|---|---|---|---|
| A | 200 | 3 | 1.5% | 40 s |
| B | 12 | 7 | 58% | 9 min |
| C | 2 | 2 | 100% | 22 min |
I would not consider it settled without evidence: Report precision and analyst dwell time per rule, and act on the rules at the bottom.
A rule nobody believes is worse than a rule nobody has.
Curated: · Written: · Reviewed:
QA-3How would you know a log source stopped reporting?(show answer)
This is an area where owning a tool for log source health and detecting it are different things.
A silent log source looks identical to a quiet environment, so its failure is invisible to every rule that depends on it. It is the most common reason a detection that used to work no longer does.
Concretely, alert on absence rather than on content: track expected event volume per source with a baseline, alarm when a source falls outside it, and treat a stopped source as an incident rather than as an operational ticket.
The reason for that specificity is a failure I have seen: Endpoint telemetry from 340 of 1,200 hosts stopped after an agent upgrade; nothing alerted for 6 weeks because every rule was written to fire on events rather than on their absence.
Hosts reporting after the upgrade.
| Week | Hosts reporting | Alerted |
|---|---|---|
| 0 | 1200 | n/a |
| 1 | 860 | no |
| 6 | 855 | no, found by hand |
I would not consider it settled without evidence: Alarm on per-source volume falling below its baseline, and test by deliberately stopping a source.
Silence and safety look the same in a log pipeline.
Curated: · Written: · Reviewed:
QA-4How long should you retain security logs?(show answer)
My answer to dwell time and retention begins with the condition rather than the alert it produced.
Retention has to exceed the time an intrusion is likely to go undetected, because the investigation needs the period before discovery. Retention chosen from storage cost regularly leaves an investigation with nothing from the window that matters.
Concretely, set hot retention from the investigation window you need and archive the rest cheaply but retrievably, and confirm the archive can actually be searched within the time an incident allows rather than in theory.
The reason for that specificity is a failure I have seen: A 30-day retention met the budget; the intrusion had begun 4 months earlier and the initial access, the persistence mechanism, and the scope of the data touched were all outside the window.
Retention against the intrusion timeline.
| Event | Days before discovery | In 30-day retention |
|---|---|---|
| initial access | 121 | no |
| persistence installed | 118 | no |
| lateral movement | 40 | no |
| exfiltration | 3 | yes |
I would not consider it settled without evidence: Compare retention against your own observed and industry dwell times, and test retrieval from archive under time pressure.
Retention shorter than dwell time investigates the wrong period.
Curated: · Written: · Reviewed:
QA-5A rule matches a specific command line. Why is that fragile?(show answer)
I would treat writing a detection that survives evasion as a claim about an adversary that has to survive being attempted.
Matching a literal string detects one spelling of a technique, and the spelling is the cheapest thing for an attacker to change. Detections anchored on behaviour or on the required effect survive variation that string matching does not.
Concretely, anchor on what the technique must do — a process relationship, a file write to a specific location, a network destination — rather than on how it was typed, and test with several variants including obfuscated ones before trusting the rule.
The reason for that specificity is a failure I have seen: A rule matching an exact encoded command missed the same technique with different casing and whitespace; the variants were the default output of a widely available tool.
Variants caught by rule type.
| Rule anchor | Exact string | Case-varied | Obfuscated | Variants caught of 3 |
|---|---|---|---|---|
| literal match | catches | misses | misses | 1 |
| normalised string | catches | catches | misses | 2 |
| parent-child behaviour | catches | catches | catches | 3 |
I would not consider it settled without evidence: Test each rule against at least three variants of the technique, including one an off-the-shelf tool produces.
Match the effect, not the spelling.
Curated: · Written: · Reviewed:
QA-6You want to alert on unusual administrative activity. Where do you start?(show answer)
The useful question for baselining before alerting is what still holds on the hosts nobody has looked at.
An anomaly is defined against a baseline, and most environments have never measured theirs. Deploying a vendor's default thresholds produces alerts against someone else's normal, which is why they arrive in volume and get muted.
Concretely, measure your own distribution for a period long enough to include weekly and monthly cycles, set thresholds from it, and re-baseline after any significant environmental change rather than treating the first measurement as permanent.
The reason for that specificity is a failure I have seen: A default threshold alerted on any administrative login outside business hours; the environment ran a nightly maintenance window, and the rule produced 60 alerts a night until it was disabled.
Alert volume by threshold source.
| Threshold | Alerts/night | True positives |
|---|---|---|
| vendor default | 60 | 0 |
| local baseline | 2 | 0.3 |
| baseline + context | 0.4 | 0.3 |
I would not consider it settled without evidence: Publish the measured baseline alongside the threshold, and show the expected alert volume before enabling.
An anomaly is relative to a baseline you have actually measured.
Curated: · Written: · Reviewed:
QA-7Walk through how you would run an incident from the first alert.(show answer)
I would settle the incident response lifecycle by generating the technique and seeing whether anything noticed.
Response is a sequence of decisions — declare, assign roles, contain, preserve, eradicate, recover, learn — and the value comes from having decided the order in advance. An incident is a bad time to work out who has the authority to disconnect a system.
Concretely, declare early and stand down cheaply, name commander, communicator, and operator as distinct people, contain in a way that preserves evidence, and keep a timeline as you go rather than reconstructing it from chat afterwards.
The reason for that specificity is a failure I have seen: An incident ran for 3 hours with no named commander; three people investigated in parallel, two containment actions conflicted, and the timeline had to be rebuilt from scrollback that had lost the reasoning entirely.
Time to key milestones by preparation.
| Milestone | Unrehearsed | Rehearsed |
|---|---|---|
| commander named | 45 min | 3 min |
| first containment | 3 h | 22 min |
| timeline usable | no | yes |
I would not consider it settled without evidence: Rehearse the sequence and record how long it takes to reach a named commander and a first containment decision.
Decide the order before you need it.
Curated: · Written: · Reviewed:
QA-8You find a compromised host. Do you power it off?(show answer)
The judgement in containment that preserves evidence is which control removes the class, not which one closes the ticket.
Powering off stops the activity and destroys volatile memory, which frequently holds the only record of what was running and how it persisted. Network isolation stops the spread while keeping that evidence available.
Concretely, isolate at the network rather than shutting down, capture memory before any change where the system justifies it, and decide in advance which classes of system get forensic capture and who can authorise the delay.
The reason for that specificity is a failure I have seen: A host was powered off immediately; the in-memory second stage was gone, the persistence mechanism was never identified, and the same intrusion returned 3 weeks later through the same route.
What each containment preserves.
| Action | Stops activity | Memory preserved | Persistence found |
|---|---|---|---|
| power off | yes | no | no |
| network isolate | yes | yes | yes |
| leave running, monitor | no | yes | yes |
I would not consider it settled without evidence: Rehearse isolation and memory capture on the actual platform and record how long each takes.
Powering off answers the wrong question permanently.
Curated: · Written: · Reviewed:
QA-9You have one compromised host. How do you decide whether there are more?(show answer)
Where candidates lose the interview on scoping an intrusion is reaching for user training first.
Remediating the host you found and stopping there is the most common way an intrusion returns, because the attacker's other footholds are unaffected. Scoping means searching the estate for the indicators and behaviours from this host before remediating any of it.
Concretely, extract indicators and, more importantly, the behaviours from the known host, sweep the estate for both, and remediate everything at once so the attacker cannot use the surviving access to re-establish the rest.
The reason for that specificity is a failure I have seen: A single host was rebuilt on day 1; the attacker retained access on 4 other systems and re-established on the rebuilt host within 48 hours using a credential taken before the rebuild.
Outcome by remediation order.
| Approach | Hosts found | Re-compromise |
|---|---|---|
| remediate as found | 1, then 5 | yes, 48 h |
| scope then remediate together | 5 | no |
I would not consider it settled without evidence: Sweep the whole estate for the behaviours before remediating, and remediate simultaneously rather than as each is found.
Remediating one host tells the attacker you are looking.
Curated: · Written: · Reviewed:
QA-10An attacker has valid credentials. What do you have to invalidate?(show answer)
I would answer credential invalidation during an incident by separating what was detected from what was merely logged.
Passwords, session tokens, API keys, certificates, and any credential derived from them are separate things with separate lifetimes, and resetting one leaves the others working. An incomplete invalidation returns access to the attacker within hours.
Concretely, enumerate every credential type the compromised identity holds or could have taken, invalidate sessions as well as secrets, rotate anything the host could read, and verify by attempting use of the old credential rather than assuming.
The reason for that specificity is a failure I have seen: A password reset left the attacker's refresh token valid; access resumed 20 minutes later and the team believed they were investigating a second intrusion.
What survived the reset.
| Credential | Reset password | Full invalidation |
|---|---|---|
| password | invalid | invalid |
| active session | valid | invalid |
| refresh token | valid | invalid |
| API key on host | valid | invalid |
I would not consider it settled without evidence: Attempt to use each invalidated credential type after the action and confirm each fails.
Resetting the password is one of several credentials.
Curated: · Written: · Reviewed:
QA-11Your phishing simulation click rate is 12 percent. What would you do?(show answer)
The engineering content of phishing defence beyond training is the containment path and its rehearsal, not the framework name.
Training moves click rates modestly and never to zero, because a well-built phishing email is designed to be indistinguishable. Controls that make a click harmless — phishing-resistant authentication, attachment detonation, link rewriting — address the outcome rather than the behaviour.
Concretely, deploy origin-bound authentication factors so a harvested credential is unusable, block or detonate attachments, and treat the simulation rate as a measure of how much you need those controls rather than as a target to reduce.
The reason for that specificity is a failure I have seen: Three years of quarterly training moved the click rate from 18 to 12 percent; the eventual compromise came through a credential harvest that phishing-resistant authentication would have made worthless at any click rate.
Outcome of a successful phish.
| Control | Click rate | Credential usable |
|---|---|---|
| training only | 12% | yes |
| + push MFA | 12% | with relay |
| + origin-bound passkey | 12% | no |
I would not consider it settled without evidence: Test whether a harvested credential is actually usable, rather than measuring how many people clicked.
Make the click harmless rather than rarer.
Curated: · Written: · Reviewed:
QA-12What password policy would you set today?(show answer)
Before calling password policy that reflects current guidance covered I would write down the technique it does not stop.
Forced periodic rotation and composition rules produce predictable transformations and support load without a measurable security gain, which is why current guidance dropped them. Length, blocklisting known-breached passwords, and rate limiting are what actually raise the cost.
Concretely, require a reasonable minimum length, check against a breached-password list at set time, rate-limit and monitor authentication failures per account, rotate only on evidence of compromise, and put the effort into removing passwords where a stronger factor exists.
The reason for that specificity is a failure I have seen: A 90-day rotation with composition rules produced predictable patterns and 340 support hours a year; the eventual compromise was credential stuffing that a breach-list check would have blocked at set time.
Cost against effect.
| Control | Annual cost | Blocks credential stuffing |
|---|---|---|
| 90-day rotation | 340 h | no |
| composition rules | 90 h | no |
| breach-list check | 8 h | yes |
I would not consider it settled without evidence: Measure support cost and check whether set passwords appear in breach corpora, rather than measuring policy compliance.
Rotation trains predictability; blocklisting removes the cheap attack.
Curated: · Written: · Reviewed:
QA-13Which second factor would you deploy for a workforce, and why not the others?(show answer)
The first thing I would establish about multi-factor authentication that resists relay is what it costs an attacker to get past it.
Push approvals and one-time codes stop bulk credential reuse and remain relayable in real time, because the user can be induced to approve or type a code for the attacker's session. A factor bound to the origin cannot be replayed against a different site.
Concretely, deploy origin-bound authenticators for anything privileged, remove weaker fallbacks rather than leaving them as recovery paths, and treat number matching and push fatigue controls as improvements to a phishable factor rather than as equivalents.
The reason for that specificity is a failure I have seen: An administrator approved a push during a real-time relay; the login was recorded as compliant with multi-factor policy and the attacker held a valid session for 6 hours.
Attacks resisted by factor.
| Factor | Reused password | Bulk phish | Real-time relay |
|---|---|---|---|
| password only | no | no | no |
| one-time code | yes | yes | no |
| push with number match | yes | yes | mostly no |
| origin-bound passkey | yes | yes | yes |
I would not consider it settled without evidence: Run a relay test against your own login flow and confirm the factor refuses it.
A control is as strong as the weakest factor you still permit.
Curated: · Written: · Reviewed:
QA-14An attacker has one workstation. What determines how far they get?(show answer)
I would start lateral movement and flat networks from what the telemetry can actually show, not from the rule list.
Reachability and credential reuse determine the blast radius, not the initial access. A flat network with shared local administrator credentials turns one workstation into every workstation, and neither condition is visible in a vulnerability scan.
Concretely, randomise local administrator passwords per host, block workstation-to-workstation traffic on the ports used for administration, tier administrative accounts so a workstation credential cannot log into a server, and test by attempting the movement.
The reason for that specificity is a failure I have seen: A shared local administrator password across 1,400 workstations meant one compromised host yielded credentials valid on all of them; the movement took under an hour and crossed no network boundary.
Reach from one workstation.
| Control state | Hosts reachable | Hosts accepting credential |
|---|---|---|
| flat, shared local admin | 1400 | 1400 |
| flat, randomised local admin | 1400 | 1 |
| segmented + randomised | 12 | 1 |
I would not consider it settled without evidence: Attempt lateral movement from a compromised test workstation and record how many hosts are reachable and how many accept the credential.
Initial access is cheap; what matters is what it reaches.
Curated: · Written: · Reviewed:
QA-15Why should a domain administrator never log into a workstation?(show answer)
This is an area where owning a tool for tiered administration and detecting it are different things.
Authenticating to a host exposes credential material to that host. A high-tier credential used on a low-tier system is recoverable by anyone who compromises it, which inverts the hierarchy the tiering was meant to create.
Concretely, define tiers by the systems a credential can reach, block high-tier accounts from authenticating to lower tiers at the identity layer rather than by policy, and provide dedicated administrative workstations so the legitimate need is met.
The reason for that specificity is a failure I have seen: A domain administrator logged into a helpdesk workstation to fix a problem; the workstation was already compromised, and the credential material was recovered and used to reach a domain controller within 20 minutes.
Exposure by where the account is used.
| Account | Used on workstations | Systems reachable if stolen |
|---|---|---|
| helpdesk | yes | 1400 |
| server admin | no | 240 |
| domain admin | blocked | all |
I would not consider it settled without evidence: Attempt to authenticate a tier-zero account to a workstation and confirm the identity layer refuses it.
Where the credential is used is where it can be stolen.
Curated: · Written: · Reviewed:
QA-16Where do service account credentials usually leak from?(show answer)
My answer to service accounts and their credentials begins with the condition rather than the alert it produced.
Service accounts tend to have high privilege, no interactive owner, passwords that never change, and credentials stored on every host that runs the service. That combination makes them the most reliable escalation path in most estates.
Concretely, use managed identities where the platform provides them, scope each account to one service and one host set, remove interactive logon rights, monitor for authentication from unexpected sources, and rotate on a schedule you have actually exercised.
The reason for that specificity is a failure I have seen: A service account with domain administrator rights ran a backup agent on 240 servers; a single compromised server yielded the credential, which was valid everywhere and had not changed in 5 years.
Exposure by account design.
| Design | Hosts holding credential | Privilege |
|---|---|---|
| shared, domain admin | 240 | full |
| per-service, scoped | 12 | limited |
| managed identity | 0 | limited |
I would not consider it settled without evidence: Enumerate service accounts with high privilege and the number of hosts holding their credentials, and reduce both.
A service account's credential lives on every host that runs it.
Curated: · Written: · Reviewed:
QA-17How would you detect an attacker extracting credentials from memory?(show answer)
I would treat detecting credential dumping as a claim about an adversary that has to survive being attempted.
The technique requires access to a specific process's memory, so the reliable signal is the access itself rather than the tool used. Detections keyed on tool names catch the default build and miss a recompile.
Concretely, alert on suspicious handle requests to the credential process, enable the platform's credential protection so the material is unavailable, and test with several tools including one built from source rather than a downloaded binary.
The reason for that specificity is a failure I have seen: A detection keyed on a well-known tool's file hash missed a recompiled version; the technique was identical and the alert never fired.
Detections by anchor.
| Anchor | Stock tool | Recompiled | Living-off-the-land | Caught of 3 |
|---|---|---|---|---|
| file hash | catches | misses | misses | 1 |
| tool name | catches | misses | misses | 1 |
| process handle access | catches | catches | catches | 3 |
I would not consider it settled without evidence: Generate the technique with a recompiled tool and confirm the behavioural detection still fires.
Detect the access, not the executable.
Curated: · Written: · Reviewed:
QA-18You have removed the malware. How do you know the attacker is out?(show answer)
The useful question for persistence mechanisms is what still holds on the hosts nobody has looked at.
Removing the payload does not remove the persistence, and there are many places to persist — scheduled tasks, services, registry run keys, startup items, account creation, and legitimate remote-access tooling. Missing one returns the attacker on the next reboot.
Concretely, enumerate persistence locations systematically rather than by memory, compare against a known-good baseline, and prefer rebuilding from a trusted image over cleaning where the system's role justifies it.
The reason for that specificity is a failure I have seen: A cleaned host reappeared in alerts 9 days later; a scheduled task created during the intrusion had been missed because the investigation focused on the service the malware had installed.
Persistence found by method.
| Method | Locations checked | Mechanisms found |
|---|---|---|
| remove observed payload | 1 | 1 |
| baseline comparison | 14 | 3 |
| rebuild from image | n/a | all removed |
I would not consider it settled without evidence: Enumerate every persistence location against a baseline and record the count checked, rather than removing what was found.
The payload is the part they expect you to find.
Curated: · Written: · Reviewed:
QA-19Where does an endpoint agent not see?(show answer)
I would settle endpoint detection and its blind spots by generating the technique and seeing whether anything noticed.
An agent sees the hosts it is installed on, in the states where it is running. Unmanaged devices, appliances, contractor laptops, systems booted from other media, and the interval before installation are all outside it, and those are where intrusions concentrate.
Concretely, reconcile agent coverage against an independent asset inventory rather than against the agent console, alarm on hosts that appear on the network without an agent, and treat coverage as a tracked number with a target rather than an assumption.
The reason for that specificity is a failure I have seen: An agent console showed 100 percent coverage of the hosts it knew about; a network inventory found 180 devices with no agent, and the intrusion began on one of them.
Coverage by inventory source.
| Source | Hosts | With agent |
|---|---|---|
| agent console | 1200 | 1200 |
| directory | 1310 | 1200 |
| network discovery | 1380 | 1200 |
I would not consider it settled without evidence: Compare agent inventory against network and directory inventories, and investigate every host present in one and not the other.
The console counts the hosts it already knows.
Curated: · Written: · Reviewed:
QA-20Is application allowlisting worth the operational cost?(show answer)
The judgement in application allowlisting is which control removes the class, not which one closes the ticket.
Allowlisting removes an entire class of attack — arbitrary executable delivery — and its cost is entirely in the exception process. It is strong on servers and fixed-function systems and expensive on developer workstations, so it is a per-population decision.
Concretely, deploy on servers and fixed-function endpoints first, use publisher or path rules that survive updates rather than hashes that break on every patch, run in audit mode long enough to size exceptions, and provide a fast exception path so it is not routed around.
The reason for that specificity is a failure I have seen: An estate-wide rollout including developer workstations generated 400 exception requests in the first week, and the control was disabled after 3 weeks with no population protected.
Exception volume by population.
| Population | Exceptions/week | Enforcement viable |
|---|---|---|
| fixed-function kiosks | 0 | yes |
| servers | 4 | yes |
| developer workstations | 400 | no |
I would not consider it settled without evidence: Run in audit mode per population and publish the expected exception volume before enforcing.
Scope it where the exception rate is survivable.
Curated: · Written: · Reviewed:
QA-21What stops a malicious document from executing?(show answer)
Where candidates lose the interview on macro and script execution controls is reaching for user training first.
Most document-borne intrusions need a scripting or macro capability that the majority of users never use. Disabling it centrally removes the path, whereas warning the user relies on a judgement the attacker has designed to defeat.
Concretely, block macros from files originating outside the organisation at the platform level, restrict scripting engines to signed scripts, and allow the small population with a genuine need through an exception rather than leaving it on for everyone.
The reason for that specificity is a failure I have seen: A warning banner was the only control; users clicked through it because they saw it on legitimate documents daily, and the banner had trained them to dismiss it.
Execution outcome by control.
| Control | Executes if user clicks | Users needing exception |
|---|---|---|
| warning banner | yes | 0 |
| block external macros | no | 12 of 1400 |
| signed scripts only | no | 30 of 1400 |
I would not consider it settled without evidence: Send a benign macro-bearing document from outside and confirm it cannot execute regardless of what the user clicks.
A prompt is a control that delegates the decision to the target.
Curated: · Written: · Reviewed:
QA-22What do SPF, DKIM, and DMARC each do, and what do they not do?(show answer)
I would answer email authentication by separating what was detected from what was merely logged.
SPF authorises sending addresses for a domain, DKIM signs the message so it can be attributed to a domain, and DMARC ties one of those to the visible sender and tells receivers what to do when it fails. None of them stop a lookalike domain, which is what most targeted phishing uses.
Concretely, publish all three and move DMARC to a rejecting policy after monitoring, then address lookalikes separately through registration monitoring and by treating external mail visibly, since the authentication stack has nothing to say about a domain you do not own.
The reason for that specificity is a failure I have seen: A rejecting DMARC policy was cited as anti-phishing coverage; the successful campaign came from a domain one character different, which was fully authenticated for its own domain.
What each stops.
| Attack | SPF/DKIM/DMARC | Lookalike monitoring | Share of targeted phishing |
|---|---|---|---|
| spoofed your domain | stops | no | 4% |
| lookalike domain | no | detects | 61% |
| compromised partner | no | no | 35% |
I would not consider it settled without evidence: Test with a lookalike domain and confirm what the stack does, which is to authenticate it correctly.
The authentication stack protects your domain, not your users.
Curated: · Written: · Reviewed:
QA-23An email from the finance director asks for an urgent payment. What control catches it?(show answer)
The engineering content of business email compromise is the containment path and its rehearsal, not the framework name.
The attack targets a process rather than a system, and the email may be entirely genuine from a compromised account. Technical controls help at the margin; what stops it is a payment process that does not depend on the content of a message.
Concretely, require out-of-band verification against a known number for any payment change or unusual payment, make that requirement unconditional so urgency cannot waive it, and design the process so the person under pressure is not the one who can approve alone.
The reason for that specificity is a failure I have seen: A payment was made from a genuine compromised director's account; every technical control passed because the message was authentic, and the verification step existed but was waived for urgency.
Where each control acts.
| Control | Spoofed sender | Compromised account | Attacks stopped of 2 |
|---|---|---|---|
| email authentication | stops | no | 1 |
| external sender banner | flags only | no | 0 |
| unconditional callback | stops | stops | 2 |
I would not consider it settled without evidence: Test the process with a realistic urgent request and confirm the verification step is performed rather than waived.
The message can be genuine and the request still fraudulent.
Curated: · Written: · Reviewed:
QA-24Your scanner reports 12,000 findings. Where do you start?(show answer)
Before calling vulnerability scanning versus exposure covered I would write down the technique it does not stop.
A finding count measures the scanner. What determines urgency is whether the affected service is reachable by an attacker, whether exploitation is being observed, and what the host holds — none of which the severity score captures.
Concretely, filter by internet exposure and known exploitation first, then by asset value, and work that set. Report the remainder as a backlog with a stated policy rather than as an emergency.
The reason for that specificity is a failure I have seen: A team worked strictly by severity; a medium-rated flaw on an internet-facing appliance with observed exploitation waited 6 weeks behind internal-only criticals and was the eventual entry point.
Top 20 by ranking method.
| Ranking | Internet-facing in top 20 | Known-exploited in top 20 |
|---|---|---|
| severity only | 2 | 1 |
| exposure-weighted | 20 | 15 |
I would not consider it settled without evidence: Rank by exposure and observed exploitation, and check the top of the list against what an attacker would reach first.
Severity describes the flaw; exposure describes your risk.
Curated: · Written: · Reviewed:
QA-25How quickly must you patch?(show answer)
The first thing I would establish about patching cadence and the exploitation window is what it costs an attacker to get past it.
The target is set by how quickly exploitation appears for the class of system, not by a policy number. Internet-facing appliances are commonly exploited within days of disclosure, while an internal library flaw may reasonably wait for the next cycle.
Concretely, set differentiated targets by exposure and asset class, measure achieved time to patch rather than compliance with the policy, and build the capability to patch an internet-facing system in days before promising it.
The reason for that specificity is a failure I have seen: A uniform 30-day policy was met; an internet-facing appliance was exploited on day 9 of a widely known vulnerability, comfortably inside the compliant window.
Window against achieved patching.
| Asset class | Typical exploitation | Achieved patch time |
|---|---|---|
| internet appliance | 3-10 days | 30 days |
| internet application | 1-3 weeks | 21 days |
| internal library | months | 30 days |
I would not consider it settled without evidence: Measure actual time to patch by asset class against observed exploitation timelines for that class.
The attacker's timeline sets the target, not the policy.
Curated: · Written: · Reviewed:
QA-26Why is an asset inventory a security control rather than an IT chore?(show answer)
I would start asset inventory as a security control from what the telemetry can actually show, not from the rule list.
Every other control has a denominator, and the inventory is it. Patch coverage, agent coverage, and scan coverage are all percentages of a number that is wrong if the inventory is, and the systems missing from it are exactly the ones nobody maintains.
Concretely, build the inventory from several independent sources — network discovery, directory, cloud APIs, agent consoles — and reconcile them, treating anything present in one and absent from another as a finding rather than a data-quality issue.
The reason for that specificity is a failure I have seen: A 98 percent patch compliance figure was computed against an inventory missing 180 hosts; the true figure was 86 percent, and the intrusion started on an unlisted host.
Coverage against inventory quality.
| Denominator | Hosts | Patch compliance |
|---|---|---|
| agent console | 1200 | 98% |
| reconciled inventory | 1380 | 86% |
I would not consider it settled without evidence: Reconcile at least three independent inventory sources and report the disagreement count as a tracked metric.
Every coverage percentage is a fraction of the inventory.
Curated: · Written: · Reviewed:
QA-27How would you segment an existing flat network without breaking it?(show answer)
This is an area where owning a tool for network segmentation for defence and detecting it are different things.
The value of segmentation is containment, and it is realised only where the boundaries follow what an attacker would want to reach. Applying segments from an architecture diagram breaks undocumented dependencies and gets reverted.
Concretely, capture the actual traffic graph over a period long enough to include periodic jobs, derive the initial policy from it, enforce one segment at a time starting with the highest-value assets, and keep the observation running so a new dependency is visible.
The reason for that specificity is a failure I have seen: A segmentation project enforced from a diagram broke 12 undocumented dependencies including a nightly reconciliation job, and was reverted within the day.
Documented against observed.
| Source | Edges | Undocumented |
|---|---|---|
| architecture diagram | 34 | — |
| 30-day observation | 46 | 12 |
| 90-day observation | 51 | 17 |
I would not consider it settled without evidence: Compare the observed traffic graph against the documented one and enumerate every edge the policy would cut, before enforcing.
Segment along the observed graph, not the drawn one.
Curated: · Written: · Reviewed:
QA-28Where do you put network monitoring in an encrypted world?(show answer)
My answer to intrusion detection placement begins with the condition rather than the alert it produced.
Most traffic is encrypted, so payload inspection at the network is largely unavailable without interception that has its own risks. What remains valuable is metadata — who talked to whom, when, how much — which is unencrypted and often sufficient.
Concretely, collect flow and DNS data broadly rather than pursuing full decryption, place inspection at boundaries where interception is justified and lawful, and use endpoint telemetry for what the network can no longer see.
The reason for that specificity is a failure I have seen: A team invested in inline decryption for internal traffic; it broke certificate pinning in 6 applications, created a new interception risk, and the detections that fired could all have come from flow data.
Detections by data requirement.
| Detection class | Needs payload | Available from metadata |
|---|---|---|
| beaconing | no | yes |
| exfiltration volume | no | yes |
| DNS tunnelling | no | yes |
| specific exploit signature | yes | no |
I would not consider it settled without evidence: Check which of your detections actually require payload rather than metadata before investing in decryption.
Metadata survives encryption and answers most questions.
Curated: · Written: · Reviewed:
QA-29Why is DNS logging disproportionately useful?(show answer)
I would treat DNS as a security signal as a claim about an adversary that has to survive being attempted.
Almost every action reaches a name before it reaches an address, so DNS records intent at a point where it is still unencrypted and cheap to collect. It sees the endpoints that network inspection cannot decrypt and the hosts endpoint agents do not cover.
Concretely, log resolver queries with the originating host, retain them for the investigation window, alert on newly observed and low-reputation domains and on anomalous query volume, and route internal resolution through resolvers you control so the log exists.
The reason for that specificity is a failure I have seen: An investigation could not determine which host had contacted a command-and-control domain because DNS logs recorded only the forwarder's address rather than the originating client.
What the DNS log answered.
| Log content | Identifies host | Median investigation time |
|---|---|---|
| forwarder address only | no | 2 days |
| client address | yes | 20 min |
| client + process | yes | 4 min |
I would not consider it settled without evidence: Confirm your DNS logs name the originating host, not the resolver, by checking a single query end to end.
Names are resolved before anything is encrypted.
Curated: · Written: · Reviewed:
QA-30How would you use threat intelligence without drowning in it?(show answer)
The useful question for threat intelligence that changes something is what still holds on the hosts nobody has looked at.
An indicator feed is the least durable form of intelligence: addresses and hashes change cheaply, so blocking them buys hours. Behavioural intelligence about how a group operates is durable and translates into detections and control decisions.
Concretely, use indicators for retrospective search across your retained logs rather than only for blocking, prioritise intelligence about groups plausibly targeting your sector, and convert each report into either a detection to test or a control to check rather than filing it.
The reason for that specificity is a failure I have seen: A feed of 200,000 indicators was blocked at the perimeter and produced no detections; the group in question had changed infrastructure before the feed was published.
Value by intelligence type.
| Type | Useful life | Actions produced |
|---|---|---|
| IP indicators | days | blocking only |
| file hashes | days | blocking only |
| behavioural TTPs | years | detections, controls |
I would not consider it settled without evidence: For each intelligence report consumed, record the detection tested or the control checked as a result.
Indicators expire; behaviour persists.
Curated: · Written: · Reviewed:
QA-31What separates threat hunting from browsing logs?(show answer)
I would settle threat hunting with a hypothesis by generating the technique and seeing whether anything noticed.
A hunt starts from a specific, falsifiable hypothesis about how an adversary would operate in your environment, which determines what data would confirm or refute it. Without one, the activity is unbounded and produces neither findings nor coverage.
Concretely, state the hypothesis and the data that would settle it, run it, and record the outcome including the negative result, since a refuted hypothesis with good data is a coverage statement worth keeping.
The reason for that specificity is a failure I have seen: A team spent 3 weeks browsing logs with no hypothesis, found nothing, and could not say what had been ruled out, so the next quarter repeated the same unbounded search.
Hunt records over a quarter.
| Approach | Hunts | Findings | Documented coverage |
|---|---|---|---|
| unstructured browsing | 1 | 0 | none |
| hypothesis-driven | 9 | 2 | 9 statements |
I would not consider it settled without evidence: Record each hunt's hypothesis, data used, and outcome so that a negative result is a documented coverage claim.
A hunt you cannot fail is a hunt that concludes nothing.
Curated: · Written: · Reviewed:
QA-32You can fund one exercise. Red team or purple team?(show answer)
The judgement in purple teaming over red teaming alone is which control removes the class, not which one closes the ticket.
A red team measures whether one path succeeded against a prepared defender and often produces a single narrative. A purple team exercises many techniques with the defenders present and produces coverage data across all of them, which usually improves detection faster.
Concretely, run purple team exercises regularly to build and verify coverage, and use a red team periodically to test whether the accumulated coverage actually holds against an opponent who is trying to avoid it.
The reason for that specificity is a failure I have seen: Three consecutive red team engagements each succeeded through a different path and produced 3 findings; a single purple team exercise across 40 techniques produced 29 concrete detection gaps.
Output per engagement.
| Exercise | Techniques exercised | Gaps identified |
|---|---|---|
| red team | 6 | 3 |
| purple team | 40 | 29 |
I would not consider it settled without evidence: Compare findings per unit of effort between the two, and use both for what each is good at.
Red teams find a path; purple teams find your coverage.
Curated: · Written: · Reviewed:
QA-33Which numbers tell you the security operation is improving?(show answer)
Where candidates lose the interview on measuring the security operation is reaching for user training first.
Alert counts and tickets closed measure activity. What indicates improvement is time to detect and contain, detection coverage confirmed by exercise, and the share of incidents found internally rather than reported from outside.
Concretely, track median time to detect and to contain, tested detection coverage, and internally-discovered share, and report them as trends with the caveat that a rising detection count can mean better detection rather than more intrusion.
The reason for that specificity is a failure I have seen: A team reported a 40 percent rise in alerts handled as evidence of improvement; median time to contain had risen from 4 to 11 hours over the same period because the additional alerts were noise.
Two views of the year.
| Measure | Q1 | Q4 | Reads as |
|---|---|---|---|
| alerts handled | 4200 | 5900 | busier |
| median time to contain | 4 h | 11 h | worse |
| internally discovered | 70% | 55% | worse |
I would not consider it settled without evidence: Report time to detect and contain alongside any volume metric.
Handling more alerts is not detecting more attacks.
Curated: · Written: · Reviewed:
QA-34You receive an alert that several endpoints across two business units are exhibiting rapid file renames with a common extension and shadow-copy deletion, and one host shows a scheduled task created by a service account. Walk me through your triage and containment steps in the first hour?(show answer)
I would answer ransomware triage in the first hour by separating what was detected from what was merely logged.
Confirm the alert at machine speed, then contain: the first hour exists to stop encryption spreading while preserving the evidence that reveals the entry point. Containment and evidence capture run in parallel, and affected hosts are isolated rather than powered off because the encryption keys, injected code and live sessions are in RAM and vanish with the power.
Concretely, validate the alert against EDR process lineage, the shadow-copy deletion command lines and file-creation rates before calling it ransomware, then bound the scope using the EDR console and asset inventory through the service account's logon events and lateral-movement indicators. Contain with EDR network isolation on confirmed hosts, disable and reset the service account, block the C2 addresses at the edge, and pull memory and triage images before anyone rebuilds anything.
The reason for that specificity is a failure I have seen: A team isolated six encrypting workstations within 20 minutes but left the service account active while they finished scoping; its scheduled task launched the encryptor against two file-server shares 40 minutes later, taking encrypted files from 9,000 to 310,000 and turning a one-day restore into a three-week recovery.
The first hour, in order.
| Elapsed | Action | What it buys |
|---|---|---|
| 0–10 min | Pull EDR lineage for the rename burst; confirm vssadmin delete shadows /all /quiet or wmic shadowcopy delete in the command line | Separates ransomware from a backup or sync agent |
| 10–15 min | EDR network-isolate the 4 confirmed hosts | Encryption stops spreading; disk and RAM stay intact for imaging |
| 15–25 min | Disable and reset the service account, pull its logon trail (4624 type 3 and 10, 4672) | Removes the mechanism that runs on every host it touched |
| 25–40 min | Block C2 IPs and domains at edge, proxy and DNS; hunt those indicators | Cuts second-stage pulls and exfiltration |
| 40–55 min | Memory capture plus KAPE triage on one host per business unit; raise the legal hold | Preserves keys and chain of custody before rebuild |
| 55–60 min | Sweep all hosts the account touched and all rename bursts in the estate | Turns 4 known hosts into a stated upper bound |
I would not consider it settled without evidence: Show the EDR process lineage behind the rename burst and the service account's authentication trail from the domain controllers for the preceding 48 hours, and state what each one tells you about scope.
Stop the spread in the first hour; the evidence you capture in that hour decides whether you ever find the door they came through.
Curated: · Written: · Reviewed:
QA-35Which response actions would you automate?(show answer)
The engineering content of security automation and its limits is the containment path and its rehearsal, not the framework name.
Automation is right where the action is reversible, the decision is unambiguous, and the cost of being wrong is low. Disabling a production account or isolating a critical server can be more damaging than the attack, so those belong behind a human decision.
Concretely, automate enrichment and reversible containment, require human authorisation for anything with material availability impact, make every automated action idempotent since the alerting path is at-least-once, and log the automation's decisions for review.
The reason for that specificity is a failure I have seen: An automation disabled accounts on a suspicious-login rule at 2 percent precision; it disabled 34 legitimate accounts in a week, including two during a customer incident.
Automation by action.
| Action | Reversible | Rule precision needed |
|---|---|---|
| enrich and tag | yes | any |
| block IP at edge | yes | moderate |
| disable account | partly | very high |
| isolate server | no, in effect | human decision |
I would not consider it settled without evidence: Check the precision of the triggering rule before automating a disruptive action, and require a threshold.
Automate what you would be comfortable doing wrong.
Curated: · Written: · Reviewed:
QA-36What makes a backup strategy survive ransomware?(show answer)
Before calling backup and recovery against ransomware covered I would write down the technique it does not stop.
Ransomware operators delete backups first, so a backup reachable with the credentials that reach production is not a recovery plan. Immutability, separation, and a tested restore are what make it one.
Concretely, keep at least one copy immutable and unreachable from production credentials, test restoration of a realistic system on a schedule and record the elapsed time, and know the recovery order across dependent systems before you need it.
The reason for that specificity is a failure I have seen: Backups were on a share reachable from the domain; they were encrypted along with everything else, and the only surviving copy was 5 weeks old and had never been restored.
Recovery position.
| Backup design | Survives | Restore tested | Recovery time |
|---|---|---|---|
| network share | no | no | none |
| offline copy | yes | no | unknown |
| immutable + tested | yes | monthly | 6 h |
I would not consider it settled without evidence: Restore a realistic system from the immutable copy on a schedule and record the elapsed time and data loss.
A backup reachable from production is part of production.
Curated: · Written: · Reviewed:
QA-37Who decides whether to pay a ransom, and what would you prepare in advance?(show answer)
The first thing I would establish about the ransom decision is what it costs an attacker to get past it.
It is a business decision with legal, regulatory, and insurance dimensions, not a technical one, and it must be made under time pressure by people who have thought about it before. The technical team's job is to make the decision informed and, ideally, unnecessary.
Concretely, agree the decision authority, legal counsel, and insurer contacts in advance, know your realistic recovery time so the alternative is quantified, and record the position before an incident rather than debating it during one.
The reason for that specificity is a failure I have seen: An organisation with no agreed position spent 14 hours in escalation while systems stayed down; recovery time was unknown, so the alternative to paying could not be evaluated.
Time lost to an unprepared decision.
| Preparation | Hours to decision | Recovery estimate available |
|---|---|---|
| none | 14 | no |
| authority agreed | 3 | no |
| authority + tested restore | 1 | yes |
I would not consider it settled without evidence: Rehearse the decision with the actual decision-makers and a measured recovery time estimate.
Decide who decides before the clock starts.
Curated: · Written: · Reviewed:
QA-38What makes an incident exercise worth the participants' time?(show answer)
I would start tabletop exercises with real systems from what the telemetry can actually show, not from the rule list.
An exercise finds gaps only where it requires people to actually do things. Describing what you would do confirms the plan is memorable; performing it finds the access nobody has and the command nobody knows.
Concretely, use an unrehearsed scenario, require each step to be executed against real systems where safe, inject a complication partway, and record every point where someone had to look something up or lacked access.
The reason for that specificity is a failure I have seen: An annual walkthrough concluded the plan was sound; the real incident found nobody could isolate a host out of hours, a step the walkthrough had only described.
Findings by exercise style.
| Style | Steps performed | Gaps found |
|---|---|---|
| walkthrough | 0 of 12 | 1 |
| hands-on | 12 of 12 | 9 |
I would not consider it settled without evidence: Count the steps completed unaided during the exercise, and treat each incomplete one as a finding.
Describing a step is not evidence anyone can perform it.
Curated: · Written: · Reviewed:
QA-39What belongs in a post-incident review beyond the timeline?(show answer)
This is an area where owning a tool for the post-incident review and detecting it are different things.
The timeline records what happened; the value is in the conditions that made it possible and the reasons the detection did not fire earlier. Those are separate questions and a review that conflates them produces one action item instead of several.
Concretely, separate the cause, the detection gap, and the response friction, give each its own actions with an owner and a date, and be blameless about people while being specific about systems and decisions.
The reason for that specificity is a failure I have seen: A review named a phishing email as the cause and produced one action, more training; the detection gap that allowed 9 days of dwell and the response friction that added 4 hours were never addressed.
Actions by question addressed.
| Question | Actions produced |
|---|---|
| cause | 1 |
| detection gap | 0 |
| response friction | 0 |
I would not consider it settled without evidence: Require each review to name a cause, a detection gap, and a response friction separately, with actions for each.
Why it happened and why nobody noticed are different questions.
Curated: · Written: · Reviewed:
QA-40How do you measure time to detect when you only know about the incidents you found?(show answer)
My answer to measuring time to detect honestly begins with the condition rather than the alert it produced.
The metric is computed over discovered incidents, so it systematically excludes the ones still undetected, and improving it can mean finding easier incidents faster rather than finding hard ones at all. It is useful as a trend and misleading as an absolute.
Concretely, report it alongside the share of incidents discovered internally rather than reported by a third party, since that share is the better indicator of whether the hard cases are being found.
The reason for that specificity is a failure I have seen: A team reported median detection falling from 14 days to 2; over the same period the share of incidents reported by outside parties rose from 20 to 45 percent, meaning the hard cases were increasingly found by someone else.
Both numbers together.
| Period | Median time to detect | Internally discovered |
|---|---|---|
| year 1 | 14 d | 80% |
| year 2 | 2 d | 55% |
I would not consider it settled without evidence: Report internally-discovered share alongside every time-to-detect figure.
The metric excludes what you have not found.
Curated: · Written: · Reviewed:
QA-41Where in the development lifecycle does security work pay best?(show answer)
I would treat secure software development lifecycle as a claim about an adversary that has to survive being attempted.
The cost of a defect rises with how late it is found, so design-stage work on the authorisation model and trust boundaries pays far more than late scanning. Late-stage tools are still worth having; they should not be the whole programme.
Concretely, engage while the trust boundaries are being chosen, provide secure defaults in the platform so most teams need no review, use automated checks as a backstop, and measure findings by the stage they were found.
The reason for that specificity is a failure I have seen: A programme consisting entirely of pre-release scanning found an authorisation flaw at implementation-complete; the fix took 6 weeks and the same flaw class recurred twice more because the design pattern was unchanged.
Cost by stage found.
| Stage | Findings | Median rework |
|---|---|---|
| design | 4 | 2 days |
| implementation | 3 | 3 weeks |
| pre-release | 2 | 6 weeks |
| production | 1 | incident |
I would not consider it settled without evidence: Report findings by lifecycle stage and rework cost, and shift effort toward the stage with the highest cost.
Late findings are the expensive ones by construction.
Curated: · Written: · Reviewed:
QA-42Static analysis reports 900 findings on a codebase. What now?(show answer)
The useful question for static analysis and its false-positive economy is what still holds on the hosts nobody has looked at.
A tool tuned for recall produces findings that are mostly not exploitable, and a queue like that is ignored within weeks. Precision matters more than coverage for a tool that developers are expected to act on.
Concretely, tune to the rules that produce actionable findings in your codebase, fail the build only on the high-precision set, run the broad set as a periodic review rather than a gate, and track the action rate to know whether the signal is trusted.
The reason for that specificity is a failure I have seen: A build gate on all 900 findings was disabled within a month; when it was re-enabled on 40 high-precision rules, the action rate rose from 3 percent to 78 percent.
Ruleset against action rate.
| Ruleset | Findings | Acted on | Gate survives |
|---|---|---|---|
| all rules | 900 | 3% | no |
| high-precision 40 | 61 | 78% | yes |
I would not consider it settled without evidence: Measure the share of findings developers act on, and cut the ruleset until that share is high.
A gate developers disable protects nothing.
Curated: · Written: · Reviewed:
QA-43How do you keep credentials out of source control?(show answer)
I would settle secrets scanning in repositories by generating the technique and seeing whether anything noticed.
Detection after the push is remediation, not prevention, because anything pushed to a shared remote must be treated as exposed. Prevention has to run before the commit leaves the developer's machine.
Concretely, run scanning as a pre-commit and pre-receive hook so the credential never reaches the remote, scan history for what is already there, and rotate anything found rather than only removing it.
The reason for that specificity is a failure I have seen: Post-push scanning found a key 40 minutes after the push; it had been cloned by CI and two forks by then, and rotation was still required despite the history rewrite.
Where the scan runs.
| Stage | Credential reaches remote | Rotation needed |
|---|---|---|
| post-push scan | yes | yes |
| pre-receive hook | no | no |
| pre-commit hook | no | no |
I would not consider it settled without evidence: Attempt to push a test credential and confirm it is rejected before it reaches the remote.
Once it reaches the remote, only rotation helps.
Curated: · Written: · Reviewed:
QA-44A dependency has no known vulnerabilities. Is it safe to add?(show answer)
The judgement in dependency risk beyond known vulnerabilities is which control removes the class, not which one closes the ticket.
A vulnerability database records what has been found and reported. An unmaintained package, one with a single maintainer, or one that recently changed hands carries risk that no advisory describes, and those are the conditions supply-chain attacks exploit.
Concretely, assess maintenance signals — release cadence, maintainer count, recent ownership changes — alongside advisories, prefer well-maintained alternatives, vendor or mirror what you depend on, and pin by digest so a compromised release does not arrive automatically.
The reason for that specificity is a failure I have seen: A widely used package with a single maintainer changed hands and shipped a malicious release; it had no advisories at the moment of adoption and none until after the compromise.
Signals at adoption.
| Signal | Package A | Package B |
|---|---|---|
| known vulnerabilities | 0 | 0 |
| maintainers | 1 | 14 |
| ownership changed | 2 months ago | no |
| releases last year | 1 | 22 |
I would not consider it settled without evidence: Record maintenance signals at adoption and re-check them periodically, not just advisory status.
An advisory database records the past.
Curated: · Written: · Reviewed:
QA-45What would you require from a supplier handling your data?(show answer)
Where candidates lose the interview on security requirements in procurement is reaching for user training first.
The leverage is at contract signature, and it disappears afterwards. The requirements worth insisting on are the ones you cannot add later: breach notification terms, audit or evidence rights, sub-processor disclosure, and a deletion and export path.
Concretely, require notification within a stated period, the right to evidence rather than only an attestation, disclosure of sub-processors with a change notification, and a tested export and deletion path. Verify the capabilities exist on the tier being purchased.
The reason for that specificity is a failure I have seen: A supplier contract required notification "without undue delay"; the eventual notification came 6 weeks after their discovery, which was defensible under the wording and useless to the response.
Terms against usefulness.
| Term | As written | Notification received |
|---|---|---|
| without undue delay | vague | 6 weeks |
| within 72 hours | specific | 3 days |
I would not consider it settled without evidence: Specify a numeric notification period and confirm the technical capabilities exist on the purchased tier before signing.
The only leverage is before signature.
Curated: · Written: · Reviewed:
QA-46A twenty-year-old system cannot be patched or replaced this year. What do you do?(show answer)
I would answer security architecture review of a legacy system by separating what was detected from what was merely logged.
Where the system cannot change, the controls have to move around it: restrict what can reach it, restrict what it can reach, and increase the monitoring on both. That is compensating control rather than acceptance, and it needs to be designed rather than declared.
Concretely, place it behind a controlled boundary with an explicit allowlist in both directions, put dedicated monitoring on those paths, restrict the credentials that can reach it, and set a review date rather than treating the arrangement as permanent.
The reason for that specificity is a failure I have seen: A legacy system was recorded as an accepted risk with no compensating controls; it sat on the flat network for 4 years and was the entry point for the eventual intrusion.
Position by treatment.
| Treatment | Reachable from | Monitored |
|---|---|---|
| accepted risk | whole network | no |
| boundary + allowlist | 3 hosts | yes |
| + credential restriction | 3 hosts | yes |
I would not consider it settled without evidence: Confirm the compensating controls exist and are monitored, rather than recording the risk as accepted.
Accepting a risk without compensating controls is deferring it.
Curated: · Written: · Reviewed:
QA-47What changes when the systems you defend control physical processes?(show answer)
The engineering content of operational technology and safety systems is the containment path and its rehearsal, not the framework name.
Availability and safety outrank confidentiality, patching windows may be annual, and a control that could stop the process can cause more harm than the attack. Defensive techniques from IT do not transfer unchanged.
Concretely, prefer passive monitoring to active scanning, enforce a controlled boundary between corporate and process networks rather than agents on process endpoints, keep safety instrumented systems independent, and design every control against the physical consequence of it failing.
The reason for that specificity is a failure I have seen: An active vulnerability scan against a process network caused a controller to fault and stopped a line for 6 hours, which was more disruption than the vulnerabilities it found represented.
Technique suitability.
| Technique | IT network | Process network |
|---|---|---|
| active scanning | yes | no |
| passive monitoring | yes | yes |
| endpoint agent | yes | rarely |
| boundary control | yes | yes |
I would not consider it settled without evidence: Test any active technique in a lab or on a spare controller before running it against a live process network.
The control must not be more dangerous than the threat.
Curated: · Written: · Reviewed:
QA-48Does physical access still matter when everything is encrypted?(show answer)
Before calling physical security and its digital consequences covered I would write down the technique it does not stop.
Physical access defeats most software controls: a running machine holds keys in memory, ports allow direct memory access on some hardware, and an unattended unlocked session needs no exploitation at all. Encryption at rest protects a powered-off device.
Concretely, require full-disk encryption with pre-boot authentication, enforce short screen locks, disable unnecessary direct-memory-access ports, and control physical access to areas holding infrastructure with the same seriousness as logical access.
The reason for that specificity is a failure I have seen: A laptop taken from a desk while unlocked yielded an authenticated session to internal systems; disk encryption was enabled and entirely irrelevant because the machine was running and unlocked.
What disk encryption covers.
| Device state | Data protected by disk encryption | Share of a working day |
|---|---|---|
| powered off | yes | 60% |
| suspended | partly | 15% |
| running, locked | no | 20% |
| running, unlocked | no | 5% |
I would not consider it settled without evidence: Test the actual scenario — a running, unlocked device — rather than the powered-off one the encryption addresses.
Encryption at rest protects a device at rest.
Curated: · Written: · Reviewed:
QA-49How do you address insider risk without treating staff as suspects?(show answer)
The first thing I would establish about insider risk proportionately is what it costs an attacker to get past it.
Most insider incidents are errors and departures rather than malice, so controls aimed at malice are poorly targeted. Reducing standing access and running a good leaver process addresses more of the actual risk than monitoring does, at lower cost and without the trust damage.
Concretely, minimise standing access, require elevation with a reason, separate duties for the highest-consequence actions, monitor a small set of high-signal events transparently, and involve legal and HR in the design rather than afterwards.
The reason for that specificity is a failure I have seen: Broad behavioural monitoring was deployed while 12 engineers held standing production administrative access; the incident that occurred was an accidental deletion by one of them, which minimisation would have prevented and monitoring only recorded.
Control against actual risk.
| Control | Addresses error | Addresses malice | Trust cost |
|---|---|---|---|
| access minimisation | yes | yes | none |
| separation of duties | yes | yes | low |
| broad monitoring | no | partly | high |
I would not consider it settled without evidence: Compare standing privileged access against monitoring investment, and address access first.
Reduce what an insider holds before watching what they do.
Curated: · Written: · Reviewed:
QA-50Where do privacy requirements change what a security team does?(show answer)
I would start privacy and security overlap from what the telemetry can actually show, not from the rule list.
Security wants to collect and retain telemetry; privacy constrains what may be collected, for how long, and for what purpose. The two agree on protecting data and disagree on how much to gather, and that disagreement has to be resolved deliberately.
Concretely, collect what has a stated detection purpose rather than everything available, set retention from the investigation window rather than indefinitely, minimise personal data in security telemetry, and document the basis for what is collected.
The reason for that specificity is a failure I have seen: A security data lake retained full web browsing history for every employee indefinitely with no stated purpose; a subject access request forced its disclosure and the retention had no defensible basis.
Telemetry with a stated purpose.
| Source | Detection purpose | Retention basis |
|---|---|---|
| authentication logs | yes | 1 year |
| DNS queries | yes | 90 days |
| full browsing history | no | none |
I would not consider it settled without evidence: Record the detection purpose and retention basis for each telemetry source, and remove sources with neither.
Telemetry with no stated purpose is a liability.
Curated: · Written: · Reviewed:
QA-51When does the notification clock start?(show answer)
This is an area where owning a tool for regulatory notification clocks and detecting it are different things.
Most regimes start the clock at awareness of the incident rather than at conclusion of the investigation, so waiting for certainty consumes the budget. The obligation is usually to notify with what is known and update, not to notify once complete.
Concretely, know the applicable clocks and their trigger before an incident, involve legal at declaration rather than at conclusion, and prepare a notification template that can be issued with partial information.
The reason for that specificity is a failure I have seen: A team investigated for 5 days before consulting legal; the applicable clock had started at initial awareness, and 120 of the available 72-hour budget had already been exceeded.
Budget against investigation.
| Consulted legal at | Hours elapsed | Budget remaining of 72 |
|---|---|---|
| declaration | 1 | 71 |
| after triage | 8 | 64 |
| after investigation | 120 | exceeded |
I would not consider it settled without evidence: Rehearse the clock decision with legal using an ambiguous scenario, and confirm everyone agrees when awareness occurred.
The clock starts at awareness, not at certainty.
Curated: · Written: · Reviewed:
QA-52What do you tell the business while an incident is still unfolding?(show answer)
My answer to communicating during an incident begins with the condition rather than the alert it produced.
Silence is read as either resolution or collapse, and both are wrong. Regular updates stating what is known, what is not, and when the next update comes keep the organisation from generating its own parallel response.
Concretely, publish on a stated cadence and send the update even when nothing has changed, describe impact in the audience's terms rather than technical ones, and separate confirmed facts from working hypotheses explicitly.
The reason for that specificity is a failure I have seen: A 4-hour silence during an incident led three business units to begin their own investigations and one to contact customers with an inaccurate description of the impact.
Effect of the silent interval.
| Interval without update | Parallel investigations | Inaccurate external statements |
|---|---|---|
| under 60 min | 0 | 0 |
| 4 h | 3 | 1 |
I would not consider it settled without evidence: Measure the interval between updates during exercises and treat a long gap as a finding.
An update saying nothing changed is still an update.
Curated: · Written: · Reviewed:
QA-53A candidate says data is secured because it is base64 encoded. What is the correction?(show answer)
I would treat hashing, encryption, and encoding as a claim about an adversary that has to survive being attempted.
Encoding is a reversible representation with no key and provides no confidentiality. Hashing is one-way and provides integrity or verification rather than confidentiality. Encryption provides confidentiality and requires key management. The three are not interchangeable.
Concretely, choose by the property required: encryption for confidentiality with a managed key, a keyed MAC or signature for integrity and authenticity, a slow salted hash for password verification, and encoding only for transport safety.
The reason for that specificity is a failure I have seen: Session data was base64 encoded and treated as opaque; an attacker decoded it, changed the role field, re-encoded it, and was granted administrative access.
Property by transformation.
| Transformation | Reversible | Needs key | Confidentiality |
|---|---|---|---|
| base64 | yes | no | no |
| SHA-256 | no | no | no |
| HMAC | no | yes | no |
| AES-GCM | yes | yes | yes |
I would not consider it settled without evidence: Ask what key is involved; if there is none, the transformation is not providing confidentiality.
No key means no confidentiality.
Curated: · Written: · Reviewed:
QA-54Why does nonce reuse matter in GCM?(show answer)
The useful question for authenticated encryption and nonce reuse is what still holds on the hosts nobody has looked at.
Reusing a nonce with the same key in a counter-based authenticated mode reveals the relationship between the two plaintexts and can expose the authentication key, which allows forgery rather than only decryption. The failure is more severe than a partial confidentiality loss.
Concretely, use a counter where the system can guarantee uniqueness, or random nonces with a message limit well below the birthday bound for the nonce size, and treat any design that might repeat a nonce after a restore or a fork as a defect.
The reason for that specificity is a failure I have seen: A service restored from a snapshot resumed its nonce counter from the snapshot's value; the reused nonces covered 40,000 messages and the authentication key was recoverable.
Consequence by reuse.
| Situation | Confidentiality | Authenticity |
|---|---|---|
| unique nonces | intact | intact |
| one reuse | partial loss | at risk |
| systematic reuse | lost | forgeable |
I would not consider it settled without evidence: Test the restore, restart, and fork paths specifically for nonce continuity rather than the steady-state path.
A repeated nonce breaks authenticity, not just secrecy.
Curated: · Written: · Reviewed:
QA-55A client disables certificate verification to make an integration work. What is the consequence?(show answer)
I would settle certificate validation by generating the technique and seeing whether anything noticed.
Without verification, TLS provides encryption to whoever answered, which is exactly the property an interception attack needs. The connection is encrypted and unauthenticated, which is worse than useless because it looks secure.
Concretely, fix the underlying trust problem — add the internal certificate authority, correct the hostname — rather than disabling verification, and lint for the disabling flags so a temporary workaround cannot ship.
The reason for that specificity is a failure I have seen: A verification flag disabled during development shipped to production; it went unnoticed for 2 years because everything worked exactly as it would have with verification enabled.
What verification provides.
| Setting | Encrypted | Authenticated | Interception detected |
|---|---|---|---|
| verification on | yes | yes | yes |
| verification off | yes | no | no |
I would not consider it settled without evidence: Scan for verification-disabling settings across the codebase and configuration, and fail the build on them.
Encrypted to whoever answered is not authenticated.
Curated: · Written: · Reviewed:
QA-56How do you rotate a signing key without breaking verification?(show answer)
The judgement in key rotation that works is which control removes the class, not which one closes the ticket.
Verifiers must accept both the old and the new key during the overlap, which means the design has to support multiple valid keys at once. A rotation that swaps a single key invalidates everything signed before it.
Concretely, publish key identifiers with each signature, allow verifiers to hold several valid keys, introduce the new key for verification before using it for signing, and retire the old one only after nothing signed with it remains valid.
The reason for that specificity is a failure I have seen: A single-key swap invalidated every token issued in the previous hour; every active session failed verification simultaneously and users were logged out mid-transaction.
Failures during rotation.
| Design | Verification failures |
|---|---|
| single key swap | all active sessions |
| overlap, key IDs | 0 |
I would not consider it settled without evidence: Rotate in a lower environment and confirm no verification failure occurs during the overlap.
Rotation needs an overlap, so the design needs two valid keys.
Curated: · Written: · Reviewed:
QA-57How do you generate a session token safely?(show answer)
Where candidates lose the interview on randomness and token generation is reaching for user training first.
A token must be unpredictable to anyone who has seen other tokens, which requires a cryptographic random source and enough length that guessing is infeasible. A general-purpose generator produces values whose state can often be recovered from a few outputs.
Concretely, use the platform's cryptographic source, give the token at least 128 bits of entropy, and never derive it from a timestamp, a counter, or a user identifier however hashed.
The reason for that specificity is a failure I have seen: Reset tokens were generated from a time-seeded general-purpose generator; observing a handful allowed the sequence to be reconstructed and another user's token predicted.
Guessability by source.
| Source | Bits of entropy | Outputs to predict next |
|---|---|---|
| time-seeded PRNG | ~32 | 1 |
| general PRNG | ~48 | a few hundred |
| cryptographic, 128-bit | 128 | infeasible |
I would not consider it settled without evidence: Confirm the call is to the cryptographic source by inspection, since output appearance says nothing.
Unpredictable is a property of the source, not the output.
Curated: · Written: · Reviewed:
QA-58What has to happen when a user logs out?(show answer)
I would answer session management by separating what was detected from what was merely logged.
Logging out must invalidate the session server-side, because clearing a cookie only affects a cooperating client. A session that remains valid after logout is usable by anyone who captured it.
Concretely, invalidate the session record on the server, rotate the session identifier on privilege change to prevent fixation, expire idle and absolute sessions, and invalidate all sessions on password change.
The reason for that specificity is a failure I have seen: Logout cleared the cookie and left the session valid for its remaining 12 hours; a session captured earlier remained usable long after the user believed they had logged out.
Session validity after logout.
| Implementation | Captured session usable |
|---|---|
| clear cookie only | up to 12 h |
| server-side invalidation | no |
I would not consider it settled without evidence: Replay a captured session identifier after logout and confirm it is rejected.
Clearing a cookie asks the client to forget.
Curated: · Written: · Reviewed:
QA-59What actually stops CSRF today?(show answer)
The engineering content of cross-site request forgery is the containment path and its rehearsal, not the framework name.
The attack relies on the browser attaching credentials to a cross-site request the user did not intend. SameSite cookie attributes address most of it by default, and an anti-forgery token remains necessary where cross-site requests are legitimately required.
Concretely, set SameSite appropriately, use anti-forgery tokens for state-changing requests where cross-site use is needed, verify the origin header, and never rely on the request method alone since a state-changing GET defeats all of it.
The reason for that specificity is a failure I have seen: A state-changing action was exposed over GET; SameSite in its default mode did not cover the top-level navigation case, and a link in an email performed the action.
Protection by request shape.
| Request | SameSite Lax | Anti-forgery token | Blocked of 2 controls |
|---|---|---|---|
| cross-site POST | blocked | blocked | 2 |
| top-level GET navigation | sent | blocked | 1 |
| same-site POST | sent | blocked if absent | 1 |
I would not consider it settled without evidence: Confirm no state-changing action is reachable by GET, then test the remaining paths cross-site.
A state-changing GET defeats every CSRF defence.
Curated: · Written: · Reviewed:
QA-60An application fetches a URL the user supplies. How do you make that safe?(show answer)
Before calling server-side request forgery covered I would write down the technique it does not stop.
The server can reach networks the user cannot, so a user-controlled fetch is a request issued from inside the perimeter. Blocklisting internal addresses fails to redirects, DNS rebinding, and alternative address encodings.
Concretely, allowlist permitted destinations rather than blocking internal ones, resolve and validate the address after every redirect rather than only the first URL, deny link-local and private ranges at the egress path, and consider a dedicated egress proxy for user-supplied fetches.
The reason for that specificity is a failure I have seen: A blocklist rejected the literal metadata address; a redirect to it from an allowed host was followed, and the instance credentials were returned in the response.
Bypasses by defence.
| Defence | Literal internal | Redirect | Alternative encoding | Blocked of 3 |
|---|---|---|---|---|
| blocklist | blocks | allows | allows | 1 |
| allowlist + per-hop check | blocks | blocks | blocks | 3 |
I would not consider it settled without evidence: Test with a redirect chain and an alternative encoding, not just the literal internal address.
Validate every hop, not the URL you were given.
Curated: · Written: · Reviewed:
QA-61Why is deserialising untrusted input dangerous?(show answer)
The first thing I would establish about deserialisation of untrusted data is what it costs an attacker to get past it.
Native deserialisation formats can instantiate arbitrary types and invoke their construction logic, which turns a data path into a code path. The vulnerability is in the mechanism rather than in any particular payload, so filtering payloads does not close it.
Concretely, use a data-only format such as JSON with an explicit schema and explicit type construction, never deserialise native formats from untrusted sources, and where legacy code must, restrict permitted types to an explicit allowlist.
The reason for that specificity is a failure I have seen: A cookie carried a native serialised object; a crafted value instantiated a type whose construction executed a command, and no payload filter would have covered the type space.
Risk by format.
| Format | Instantiates arbitrary types | Safe from untrusted source | Types reachable |
|---|---|---|---|
| native serialisation | yes | no | every loaded class |
| JSON to schema | no | yes | 1 |
| JSON to dynamic type | sometimes | no | many |
I would not consider it settled without evidence: Identify every native deserialisation of externally-supplied data and replace it, rather than filtering inputs.
Native deserialisation makes data into code.
Curated: · Written: · Reviewed:
QA-62Users can upload files. What are the controls?(show answer)
I would start file upload handling from what the telemetry can actually show, not from the rule list.
An upload is untrusted content that will later be stored, served, and possibly processed, and each of those is a separate risk. Checking the extension or the declared content type validates a label the uploader chose.
Concretely, validate by content inspection rather than by name, store outside the web root with generated names, serve from a separate origin with a forced download disposition where appropriate, and scan or sandbox anything that will be processed.
The reason for that specificity is a failure I have seen: An uploaded file with an image extension contained script and was served from the application's own origin; it executed in the context of authenticated users.
What each control stops.
| Control | Wrong extension | Script in image | Served executable | Stopped of 3 |
|---|---|---|---|---|
| extension check | stops | no | no | 1 |
| content inspection | stops | stops | no | 2 |
| separate origin + disposition | stops | stops | stops | 3 |
I would not consider it settled without evidence: Upload a file whose declared type and actual content disagree and confirm it is rejected and not served executable.
The extension is chosen by the uploader.
Curated: · Written: · Reviewed:
QA-63How much should an error tell the user?(show answer)
This is an area where owning a tool for error messages and information disclosure and detecting it are different things.
A detailed error helps the developer and helps the attacker equally, and the distinguishing factor is who sees it. Different messages for different failure causes also leak information — an authentication error that distinguishes unknown user from wrong password enumerates accounts.
Concretely, return a generic message with a correlation identifier to the user and the detail to the log, keep authentication failures indistinguishable including in timing, and check that error pages do not vary in a way that reveals internal state.
The reason for that specificity is a failure I have seen: A login endpoint returned distinct messages for unknown user and wrong password; an attacker enumerated 40,000 valid accounts before targeting them with credential stuffing.
Enumeration by response design.
| Design | Distinguishable | Accounts enumerable |
|---|---|---|
| distinct messages | yes | yes |
| same message, different timing | yes | yes |
| same message and timing | no | no |
I would not consider it settled without evidence: Compare responses and response times across failure causes and confirm they are indistinguishable.
Different errors enumerate what you did not intend to disclose.
Curated: · Written: · Reviewed:
QA-64How do you run a threat model that is not a documentation exercise?(show answer)
My answer to threat modelling that produces decisions begins with the condition rather than the alert it produced.
A model is useful when it changes a design decision. That requires it to be specific about the system's own trust boundaries and data flows rather than applying a generic checklist, and to happen while the design is still fluid.
Concretely, draw the actual data flows and trust boundaries, walk each crossing asking what an attacker at that point could do, record the assumptions so a later change that invalidates one is visible, and finish with decisions rather than a list of threats.
The reason for that specificity is a failure I have seen: A model produced 40 generic threats and no design change; the eventual intrusion used a path the model had drawn but never asked a question about.
Output by approach.
| Approach | Threats listed | Design changes |
|---|---|---|
| generic checklist | 40 | 0 |
| flows and boundaries | 14 | 6 |
I would not consider it settled without evidence: Count the design decisions the model changed, and treat zero as a failed model.
A model that changed nothing modelled nothing.
Curated: · Written: · Reviewed:
QA-65What is the cheapest security improvement available to most organisations?(show answer)
I would treat attack surface reduction as a claim about an adversary that has to survive being attempted.
Removing what is not needed — unused services, open ports, dormant accounts, unnecessary software, forgotten systems — costs little and eliminates whole categories of risk permanently, whereas defending them costs continuously.
Concretely, inventory exposed services and accounts, remove what has no owner or no use, disable rather than monitor where the function is not required, and repeat on a cadence because the surface grows back.
The reason for that specificity is a failure I have seen: An estate had 340 internet-exposed services; 61 had no owner, and the intrusion entered through an administrative interface on a system whose function had ended 2 years earlier.
Exposed services by ownership.
| Category | Services | Action |
|---|---|---|
| owned, needed | 210 | defend |
| owned, unneeded | 69 | remove |
| unowned | 61 | remove |
I would not consider it settled without evidence: Report exposed services with a named owner and a stated purpose, and remove those with neither.
What is not there cannot be exploited.
Curated: · Written: · Reviewed:
QA-66How do you check that two controls are actually independent?(show answer)
The useful question for defence in depth that is genuinely layered is what still holds on the hosts nobody has looked at.
Two controls are one control if they share an input, a configuration source, or a failure mode. Depth requires that a single failure cannot remove both, which is a property to be verified rather than assumed from having two names.
Concretely, for each pair, name what would have to fail for both to fail; prefer controls at different layers over two at the same one; and fault-inject the shared element to confirm at least one holds.
The reason for that specificity is a failure I have seen: A network rule and an application check both read one configuration file; a bad deploy removed both simultaneously, and the design had been described as two layers.
Independence by pair.
| Pair | Shared element | Failures to lose both |
|---|---|---|
| network rule + app check | config file | 1 |
| MFA + network restriction | none | 2 |
| WAF + input validation | none | 2 |
I would not consider it settled without evidence: Fault-inject the shared element and confirm a control remains.
Two controls with one shared input are one control.
Curated: · Written: · Reviewed:
QA-67How often should a control be tested?(show answer)
I would settle security control testing cadence by generating the technique and seeing whether anything noticed.
A control degrades silently through configuration drift, dependency change, and estate growth, so the interval between tests is the interval in which it may already have stopped working. The right cadence follows the rate of change rather than the calendar.
Concretely, test continuously where automation allows it, set the manual cadence from how quickly the environment changes, and treat any control that has never been tested since deployment as unverified rather than working.
The reason for that specificity is a failure I have seen: An annual test found that a control had stopped working 9 months earlier after an unrelated upgrade; the estate had been unprotected for three quarters and nothing had indicated it.
Unverified interval by cadence.
| Cadence | Worst unverified period |
|---|---|
| annual | 12 months |
| quarterly | 3 months |
| continuous automated | hours |
I would not consider it settled without evidence: Report the date each control was last verified working, and treat a stale date as a finding.
Between tests, the control is an assumption.
Curated: · Written: · Reviewed:
QA-68You are the only security person at a 60-person company. What do you do first?(show answer)
The judgement in security in a small organisation is which control removes the class, not which one closes the ticket.
With one person, leverage matters more than completeness. The controls that remove whole classes of attack for little ongoing cost — phishing-resistant authentication, backups that survive an intrusion, patching the internet-facing surface — come before anything requiring continuous attention.
Concretely, do the small number of high-leverage things properly, use managed services rather than operating tooling, and be explicit about what is not covered rather than claiming a programme you cannot staff.
The reason for that specificity is a failure I have seen: A sole practitioner deployed a SIEM and spent most of a year tuning it while multi-factor authentication was not enforced and backups had never been tested.
Leverage by control.
| Control | Setup | Ongoing/month | Classes removed |
|---|---|---|---|
| phishing-resistant MFA | 2 weeks | 2 h | several |
| tested immutable backup | 1 week | 4 h | ransomware recovery |
| self-run SIEM | 3 months | 60 h | detection, partial |
I would not consider it settled without evidence: Rank candidate work by attack classes removed per hour of ongoing effort, and start at the top.
One person should buy leverage, not coverage.
Curated: · Written: · Reviewed:
QA-69How do you evaluate a security product?(show answer)
Where candidates lose the interview on buying security tools is reaching for user training first.
The cost of a tool is dominated by the effort to operate it, not the licence, and a tool nobody has time to run produces alerts nobody reads. The right question is whether you can staff it, not whether it works.
Concretely, trial against your own environment and data rather than a vendor demonstration, estimate the ongoing operating effort honestly, check what it needs from you — agents, log sources, tuning — and prefer capabilities you will actually use over feature breadth.
The reason for that specificity is a failure I have seen: A product was bought on a demonstration; in the estate it needed a log source that did not exist, and after 8 months it was producing alerts nobody triaged.
Cost over the first year.
| Component | Estimated | Actual |
|---|---|---|
| licence | £40k | £40k |
| deployment | 2 weeks | 3 months |
| ongoing operation | 4 h/week | 20 h/week |
I would not consider it settled without evidence: Run a trial against your own data and measure the operating hours it consumes before purchase.
The licence is the cheap part.
Curated: · Written: · Reviewed:
QA-70A product claims to detect zero-day attacks. How do you assess that?(show answer)
I would answer reading a vendor security claim by separating what was detected from what was merely logged.
A claim about detecting unknown attacks is a claim about behavioural detection, which is testable. What matters is the detection rate on techniques you care about and the false-positive rate at that setting, both measured in your environment.
Concretely, test with techniques relevant to your threat model in your own environment, measure detection and false positives together since either alone is meaningless, and ask what telemetry the detection requires rather than accepting the claim.
The reason for that specificity is a failure I have seen: A product demonstrated 100 percent detection in a vendor lab; in the estate the default setting detected 6 of 40 techniques at 12 false positives a day, and the sensitive setting that reached 22 of 40 produced 200.
Vendor lab against your estate.
| Environment | Techniques detected | False positives/day |
|---|---|---|
| vendor lab | 40 of 40 | 0 |
| your estate, default | 6 of 40 | 12 |
| your estate, sensitive | 22 of 40 | 200 |
I would not consider it settled without evidence: Measure detection and false-positive rates together, on your own environment and threat model.
A detection rate without a false-positive rate is half a number.
Curated: · Written: · Reviewed:
QA-71You need to join 5 million authentication events against 200,000 known-bad indicators. How?(show answer)
The engineering content of hash tables in log correlation is the containment path and its rehearsal, not the framework name.
Checking each event against every indicator is a trillion comparisons. Loading the indicators into a hash set makes each lookup constant, turning the join into a single pass over the events.
Concretely, build the set once from the smaller side, stream the larger side through it, and where the indicator set does not fit in memory, use a Bloom filter as a pre-filter with an exact check on the candidates it passes.
The reason for that specificity is a failure I have seen: A nested-loop correlation over 5 million events and 200,000 indicators was still running after 14 hours; the same work with a hash set took 40 seconds.
Runtime by approach.
| Approach | Comparisons | Runtime |
|---|---|---|
| nested loop | 1e12 | 14 h, unfinished |
| hash set | 5e6 | 40 s |
| Bloom pre-filter | 5e6 | 25 s |
I would not consider it settled without evidence: Measure runtime against input size at two points; a fourfold rise for doubled input is the quadratic signature.
Index the smaller side and stream the larger.
Curated: · Written: · Reviewed:
QA-72How do you detect 5 failed logins in 10 minutes efficiently across a million accounts?(show answer)
Before calling sliding windows for rate-based detection covered I would write down the technique it does not stop.
Recomputing a count over a window for every event is expensive and unnecessary; a sliding window keeps the count incrementally, adding new events and expiring old ones so each event costs constant work.
Concretely, keep per-key timestamps in a bounded structure, expire on read rather than with a background sweep, and bound memory by capping tracked keys and shedding the least recently seen when the cap is reached.
The reason for that specificity is a failure I have seen: A detection recomputed a 10-minute count per event over the full log; at 4,000 events a second it fell 40 minutes behind within an hour and alerts arrived after the attack had finished.
Lag by implementation.
| Implementation | Cost per event | Lag at 4000/s |
|---|---|---|
| recompute window | O(window) | 40 min |
| incremental sliding window | O(1) | under 1 s |
I would not consider it settled without evidence: Measure detection lag under production event rate, not under a test rate.
A detection that falls behind is a detection that arrives late.
Curated: · Written: · Reviewed:
QA-73How do you decide which of a hundred findings actually matters?(show answer)
The first thing I would establish about graph analysis of attack paths is what it costs an attacker to get past it.
Findings are conditions; what matters is whether they compose into a path from something an attacker can reach to something they want. Three medium findings on one path can be worse than an isolated critical one.
Concretely, model identities, hosts, and reachability as a graph, run shortest-path from external entry points to sensitive assets, and prioritise the findings that lie on the shortest paths rather than by individual severity.
The reason for that specificity is a failure I have seen: A hundred findings were worked by severity; the intrusion used three medium findings that composed into a two-hop path from a public endpoint to the customer database.
Findings by path position.
| Path | Hops | Findings | Individual severity |
|---|---|---|---|
| public to customer DB | 2 | 3 | all medium |
| public to build system | 4 | 1 | critical |
| unreachable | — | 96 | mixed |
I would not consider it settled without evidence: Report the shortest attack paths to your crown-jewel assets and the findings that lie on them.
Attackers compose findings that scanners rank apart.
Curated: · Written: · Reviewed:
QA-74A detection rule's regular expression makes the pipeline stall. What happened?(show answer)
I would start regular expressions and catastrophic backtracking from what the telemetry can actually show, not from the rule list.
Certain patterns with nested quantifiers take exponential time on inputs that nearly match, so a crafted or merely unusual log line can consume a core indefinitely. It is a denial of service reachable from the data being analysed.
Concretely, avoid nested quantifiers over overlapping character classes, prefer a linear-time engine where available, set a match timeout, and fuzz new patterns with near-miss inputs before deploying them.
The reason for that specificity is a failure I have seen: A pattern with a nested quantifier took over 90 seconds on a 200-character log line; the pipeline stalled and 40 minutes of events queued behind it.
Match time by input length.
| Input length | Vulnerable pattern | Rewritten pattern |
|---|---|---|
| 20 chars | 1 ms | 0.01 ms |
| 40 chars | 900 ms | 0.02 ms |
| 200 chars | over 90 s | 0.1 ms |
I would not consider it settled without evidence: Test each pattern against long near-miss inputs with a timeout before deploying it.
The input that stalls it is the input that nearly matches.
Curated: · Written: · Reviewed:
QA-75You inherit 50 AWS accounts with no shared baseline. How do you find the exposures and fix them without breaking what is running?(show answer)
This is an area where owning a tool for cloud posture discovery and remediation and detecting it are different things.
Posture work is an inventory problem before a findings problem: you cannot rank what you have never enumerated. Fix what combines reachability, sensitive data and privilege, not what a benchmark rates highest.
Concretely, collect account, IAM, public exposure and logging state read-only from every account into one view, then rank by reachability, data sensitivity and privilege. Fix in alert-only mode with the owner's sign-off and a tested rollback; enforce only when the alert shows nothing legitimate depends on what you remove.
The reason for that specificity is a failure I have seen: A team applied an SCP denying iam:PassRole outside a named allowlist to all 42 accounts on a Friday evening, expecting quiet. ECS and CodeBuild in 13 accounts could no longer start task roles, 260 pipeline runs failed over the weekend, and rebuilding the exception list by hand took until Tuesday — the exposure was real, but the fix arrived as its own incident.
Fix order by exposure, not by scanner severity.
| Finding | Scanner severity | Internet reachable | Data class | Fix order |
|---|---|---|---|---|
| Public S3 bucket holding customer exports | High | Yes | Restricted | 1 |
| Deploy role assumable from any account | Critical | n/a | Admin | 2 |
| CloudTrail disabled in 3 production accounts | Medium | n/a | Restricted | 3 |
| Security group 0.0.0.0/0 on an isolated test host | Critical | Yes | Public | 4 |
| Unused IAM role, last used 400 days ago | Low | n/a | None | 5 |
Severity-sorted order would be 2, 4, 1, 3, 5. Exposure-sorted order puts the bucket with customer exports first despite being rated only High, and pushes the Critical test-host rule below two Medium-and-worse gaps because that host holds nothing and is not routed to anything that does.
I would not consider it settled without evidence: Name the query that returns every internet-reachable resource together with its data classification and owning team, and show the change rehearsed in one account in alert-only mode before it is enforced anywhere else.
Enumerate everything, fix the reachable-and-valuable first, and never let the remediation become the incident.
Curated: · Written: · Reviewed:
QA-76Where does Python help a security engineer and where does it get you into trouble?(show answer)
My answer to writing tooling in Python for security work begins with the condition rather than the alert it produced.
Python is well suited to log analysis, enrichment, and automation, where development speed matters more than execution speed. It becomes a problem in the streaming path at high event rates, where the interpreter's overhead and garbage-collection pauses cause the pipeline to fall behind.
Concretely, use it for analysis, orchestration, and one-off investigation; keep the high-rate streaming path in a compiled component or a purpose-built engine; and where it must process volume, batch and use vectorised libraries rather than per-event loops.
The reason for that specificity is a failure I have seen: A per-event Python enrichment step handled 900 events a second against a 4,000 per second stream; the queue grew until events were dropped.
Throughput by implementation.
| Implementation | Events/second |
|---|---|
| per-event Python loop | 900 |
| batched, vectorised | 12000 |
| compiled component | 90000 |
I would not consider it settled without evidence: Measure sustained throughput against the production event rate before putting any component in the streaming path.
Match the language to the rate the path must sustain.
Curated: · Written: · Reviewed:
QA-77Your analysis script parses attacker-controlled log content. What could go wrong?(show answer)
I would treat parsing untrusted input safely in tooling as a claim about an adversary that has to survive being attempted.
Security tooling processes hostile input by definition, so the tooling itself is an attack surface. A parser that evaluates, deserialises, or shells out on content from a log is a path from the attacker into the analyst's environment.
Concretely, never evaluate or deserialise log content, avoid shelling out with interpolated values, bound memory and time per record, and run investigation tooling in an isolated environment rather than on a privileged workstation.
The reason for that specificity is a failure I have seen: An enrichment script passed a log field into a shell command; a crafted field executed on the analyst workstation, which held credentials to the log platform and the estate.
Paths from log content to execution.
| Path | Present | Risk | Paths to remove |
|---|---|---|---|
| shell interpolation | yes | execution | 1 |
| native deserialisation | no | execution | 0 |
| bounded parse only | yes | none | 0 |
I would not consider it settled without evidence: Review every path where log content reaches a shell, an evaluator, or a deserialiser, and remove them.
Your tooling reads what the attacker wrote.
Curated: · Written: · Reviewed:
QA-78Why does layering still matter when you are investigating an intrusion?(show answer)
The useful question for the OSI model in an investigation is what still holds on the hosts nobody has looked at.
Layering tells you which evidence can answer which question and which cannot. A question about who talked to whom is answered at the network layer regardless of encryption above it; a question about what was said needs a layer that may be unavailable.
Concretely, choose the evidence source from the layer the question lives at, and where the needed layer is encrypted, decide whether an endpoint source can answer instead rather than pursuing decryption by default.
The reason for that specificity is a failure I have seen: An investigation pursued packet capture to determine what was exfiltrated; the traffic was encrypted, and the answer came from endpoint file-access telemetry that had been available throughout.
Question to source.
| Question | Layer | Source | Days spent on the wrong source |
|---|---|---|---|
| which hosts communicated | network | flow logs | 0 |
| which name was resolved | application | DNS logs | 0 |
| what data was read | endpoint | file access telemetry | 4 |
I would not consider it settled without evidence: Map each investigative question to the layer and source that can answer it before collecting.
Ask each question at the layer that can answer it.
Curated: · Written: · Reviewed:
QA-79What can connection metadata alone tell you about an intrusion?(show answer)
I would settle TCP behaviour in detection by generating the technique and seeing whether anything noticed.
Connection duration, direction, byte counts, and regularity carry a lot of signal even when the payload is encrypted. Beaconing in particular is visible as regularity in timing that normal traffic rarely shows.
Concretely, look for regular intervals with low jitter, small consistent request sizes with larger responses, long-lived connections to rare destinations, and outbound volume anomalies, all of which survive encryption.
The reason for that specificity is a failure I have seen: A beacon at a 60-second interval with 5 percent jitter ran for 11 weeks; payload inspection was impossible and nothing looked at inter-arrival regularity, which would have shown it in a day.
Beacon visibility by analysis.
| Analysis | Detects 60 s beacon | Needs payload |
|---|---|---|
| signature matching | no | yes |
| volume anomaly | no | no |
| inter-arrival regularity | yes | no |
I would not consider it settled without evidence: Run an inter-arrival regularity analysis over outbound connections and confirm it surfaces a synthetic beacon.
Encryption hides the content, not the rhythm.
Curated: · Written: · Reviewed:
QA-80A permission cache speeds up authorisation. What have you traded?(show answer)
The judgement in caching and stale authorisation decisions is which control removes the class, not which one closes the ticket.
A cached authorisation decision remains valid for its lifetime, so a revoked permission continues to work for that long. The cache TTL is the revocation delay, whether or not anyone has stated it that way.
Concretely, keep the TTL inside the revocation service level you have committed to, invalidate explicitly on permission change where the system allows it, and state the revocation delay in the design rather than leaving it implicit in a cache setting.
The reason for that specificity is a failure I have seen: A 15-minute permission cache meant an offboarded employee retained access for a quarter of an hour after removal; the offboarding process claimed immediate revocation.
Revocation delay by design.
| Design | Delay | Matches stated SLA |
|---|---|---|
| 15-minute cache | 15 min | no |
| 30-second cache | 30 s | yes |
| explicit invalidation | under 1 s | yes |
I would not consider it settled without evidence: Revoke a permission and measure how long the old decision continues to be honoured.
The cache TTL is your revocation delay.
Curated: · Written: · Reviewed:
QA-81Your detection pipeline falls behind during an attack. What should happen?(show answer)
Where candidates lose the interview on queueing and back-pressure in a detection pipeline is reaching for user training first.
The moment a pipeline is most likely to be overloaded is during an attack, which is when its output matters most. A pipeline that drops silently under load loses exactly the events the investigation will need.
Concretely, bound queues and shed deliberately with a recorded count rather than dropping silently, prioritise high-value sources over verbose ones when shedding, and alert on shedding as an incident rather than as an operational metric.
The reason for that specificity is a failure I have seen: A pipeline dropped 40 percent of events during a volumetric attack with no record; the investigation afterwards had gaps precisely across the period of interest.
Behaviour under 3x load.
| Design | Events lost | Recorded | Sources preserved |
|---|---|---|---|
| silent drop | 40% | no | random |
| bounded, prioritised shed | 40% | yes | high-value kept |
I would not consider it settled without evidence: Load-test above production peak and confirm shedding is recorded and prioritised rather than uniform and silent.
Losing events during the attack loses the investigation.
Curated: · Written: · Reviewed:
QA-82What does a good alert contain?(show answer)
I would answer designing an alert that an analyst can act on by separating what was detected from what was merely logged.
An alert exists to start a decision, so it needs what the decision requires: what fired, on which entity, what the surrounding activity was, and what to do next. An alert that requires the analyst to gather all of that has moved the work rather than done it.
Concretely, include the entity with its context, the triggering evidence, the recent related activity, the rule's own precision history, and a link to the response procedure, and enrich automatically so the analyst starts from a position rather than a string.
The reason for that specificity is a failure I have seen: Alerts contained a rule name and a hostname; median triage was 25 minutes, of which 22 were spent gathering context that the pipeline already had.
Triage time by alert content.
| Alert content | Median triage | Mechanical share |
|---|---|---|
| rule name, hostname | 25 min | 88% |
| + user, process tree | 9 min | 40% |
| + related activity, procedure | 4 min | 10% |
I would not consider it settled without evidence: Measure the share of triage time spent on mechanical enrichment, and automate that share.
Enrichment the pipeline can do should not be an analyst's job.
Curated: · Written: · Reviewed:
QA-83What makes a security runbook usable under pressure?(show answer)
The engineering content of runbooks for the security team is the containment path and its rehearsal, not the framework name.
It will be read by someone tired and possibly unfamiliar with the system, so prose describing an approach fails where exact commands with expected output succeed. The decision points and their authority have to be explicit.
Concretely, number the steps, give the exact command and the expected result, say what to do when the result differs, name who can authorise each decision, and validate by having someone unfamiliar execute it during a drill.
The reason for that specificity is a failure I have seen: A runbook said to isolate the affected host; during the incident nobody knew the command or who could approve it, and 40 minutes went to finding out.
Drill result by runbook style.
| Style | Steps completed unaided | Elapsed |
|---|---|---|
| prose | 4 of 12 | 40 min |
| exact commands | 12 of 12 | 6 min |
I would not consider it settled without evidence: Have an unfamiliar person execute the runbook in a drill and count the steps they could not complete.
Write it for someone who has never done it.
Curated: · Written: · Reviewed:
QA-84When should an analyst escalate rather than close?(show answer)
Before calling knowing what to escalate covered I would write down the technique it does not stop.
The cost of a late escalation is measured in dwell time and the cost of an early one in interruption, and the asymmetry favours escalating. What makes that workable is a defined threshold so the decision is not a judgement about looking foolish.
Concretely, define escalation criteria on observable conditions rather than on confidence, make escalation cheap and blameless, and review closed alerts periodically to find the ones that should have been escalated.
The reason for that specificity is a failure I have seen: An analyst closed an alert as a false positive because escalating had previously been met with irritation; it was the initial access, and the intrusion ran 9 further days.
Cost asymmetry.
| Decision | If wrong | Cost |
|---|---|---|
| escalate unnecessarily | interruption | 30 min |
| close incorrectly | dwell time | 9 days |
I would not consider it settled without evidence: Sample closed alerts for missed escalations and report the rate as a team metric rather than an individual one.
Escalation must be cheaper than being wrong about it.
Curated: · Written: · Reviewed:
QA-85What has to survive a shift handover?(show answer)
The first thing I would establish about handover between shifts is what it costs an attacker to get past it.
An investigation in progress carries state that lives mostly in the analyst's head — what has been checked, what was ruled out, and why. Without an explicit handover that state is lost and the next shift repeats the work or drops it.
Concretely, record in the case what has been checked and ruled out with the reasoning, not only the conclusion, hand over live rather than by document where the case is active, and flag anything time-sensitive explicitly.
The reason for that specificity is a failure I have seen: An active investigation was handed over with a one-line note; the incoming shift repeated 3 hours of work and missed the lead the previous analyst had been pursuing.
Handover quality.
| Handover | Work repeated | Leads lost |
|---|---|---|
| one-line note | 3 h | 1 |
| checked-and-ruled-out record | 20 min | 0 |
| live handover | 5 min | 0 |
I would not consider it settled without evidence: Sample handovers and check whether the incoming shift repeated work, and treat repetition as a process finding.
What was ruled out matters as much as what was found.
Curated: · Written: · Reviewed:
QA-86A team disagrees with a security finding. How do you handle it?(show answer)
I would start working with engineering teams from what the telemetry can actually show, not from the rule list.
A disagreement is usually about the exposure or the cost rather than about the facts, and treating it as non-compliance loses both the fix and the relationship. Establishing what would change either party's mind is what makes it resolvable.
Concretely, establish the specific technical disagreement, demonstrate the exposure where you can rather than asserting it, accept a different remediation that closes the same exposure, and record the decision including a genuine disagreement with its owner.
The reason for that specificity is a failure I have seen: A finding was escalated to management as non-compliance; the team was right that the exposure did not exist as described, and the next three findings were resisted before they were read.
Outcome by approach.
| Approach | Fixed | Later findings resisted |
|---|---|---|
| escalate as non-compliance | no | 3 |
| demonstrate exposure | yes | 0 |
I would not consider it settled without evidence: Establish the disagreement precisely and demonstrate the exposure before escalating.
Being right and being effective are separate achievements.
Curated: · Written: · Reviewed:
QA-87You have budget for two of five proposals. How do you choose?(show answer)
This is an area where owning a tool for prioritising with limited budget and detecting it are different things.
The right basis is attack classes removed per unit of ongoing cost, not severity of the thing each addresses. A control that removes a class permanently beats one that detects instances of it continuously, at equal cost.
Concretely, estimate the classes each proposal removes or detects, the setup and ongoing cost, and the dependency between them, and prefer the ones whose value does not depend on continuous attention you may not have.
The reason for that specificity is a failure I have seen: A budget went to a detection product and a scanning tool, both requiring continuous attention; phishing-resistant authentication, which would have removed the class that caused the eventual incident, was deferred as unglamorous.
Proposals scored.
| Proposal | Setup | Ongoing/month | Classes removed |
|---|---|---|---|
| phishing-resistant MFA | 2 weeks | 2 h | 2 |
| detection product | 3 months | 60 h | 0 |
| scanning tool | 3 weeks | 20 h | 0 |
I would not consider it settled without evidence: Score proposals by classes removed per ongoing hour and present that alongside the cost.
Prefer removing a class over detecting its instances.
Curated: · Written: · Reviewed:
QA-88How do you present security to a board?(show answer)
My answer to reporting to a board begins with the condition rather than the alert it produced.
A board decides between options with costs, so severity counts and technical detail do not translate. What does is a small number of plausible scenarios with their consequence, the current position, and a specific decision being asked for.
Concretely, present two or three scenarios with evidence for their plausibility, state what the organisation would lose, give the cost of the mitigation, and make a recommendation rather than presenting a menu.
The reason for that specificity is a failure I have seen: A quarterly deck of finding counts by severity produced no decision for 3 quarters; a single slide describing one plausible path to the customer database with its cost was funded that week.
Format against decisions.
| Format | Quarters | Decisions |
|---|---|---|
| findings by severity | 3 | 0 |
| scenario with cost | 1 | 1 |
I would not consider it settled without evidence: Judge the format by whether it produced a decision, and change it when it did not.
Boards fund scenarios, not severity counts.
Curated: · Written: · Reviewed:
QA-89Describe a security decision you got wrong.(show answer)
I would treat career-defining mistake as a claim about an adversary that has to survive being attempted.
A useful answer names the decision, the evidence available at the time, the signal that should have changed your mind sooner, and the practice you changed. An answer ending at a lesson learned describes an outcome rather than a change.
Concretely, say what you would measure earlier next time and what threshold would trigger a reversal, since that is the part that transfers.
The reason for that specificity is a failure I have seen: A team maintained a detection programme for 3 quarters at 2 percent precision because no threshold had been set at which the approach would be abandoned.
Precision against the reversal decision.
| Quarter | Precision | Threshold set | Action |
|---|---|---|---|
| Q1 | 2% | none | continue |
| Q2 | 2% | none | continue |
| Q3 | 3% | none | continue |
| Q4 | 3% | 25% | would have stopped |
I would not consider it settled without evidence: Set a reversal threshold at the start of any significant security investment and write it down.
Decide in advance what would tell you to stop.
Curated: · Written: · Reviewed:
QA-90How do you keep up without chasing every headline?(show answer)
The useful question for staying current as a defender is what still holds on the hosts nobody has looked at.
Relevance is determined by whether a technique applies to your estate and threat model, not by coverage. An inventory good enough to answer "do we run this" turns most advisories into a twenty-minute question rather than a multi-team investigation.
Concretely, maintain an inventory that answers version and configuration questions quickly, follow a small number of behaviour-focused sources, and convert each relevant item into a detection to test or a control to check rather than filing it.
The reason for that specificity is a failure I have seen: A widely covered vulnerability consumed 3 days across 12 teams; the affected configuration was not in use anywhere, and no inventory could answer that.
Time to assess.
| Inventory | Time | Teams involved |
|---|---|---|
| none | 3 days | 12 |
| versions | 2 h | 1 |
| versions + configuration | 20 min | 1 |
I would not consider it settled without evidence: Measure how long it takes to answer whether you run an affected version and configuration.
The inventory turns an advisory into an answer.
Curated: · Written: · Reviewed:
QA-91What mistake do experienced defenders make?(show answer)
I would settle the defender's own failure mode by generating the technique and seeing whether anything noticed.
The recurring one is optimising for what is measurable — rules written, findings closed, coverage claimed — rather than for what an attacker would actually do, because the first produces reports and the second produces arguments.
Concretely, reserve a fixed share of effort for adversary-perspective work, treat any exercise that finds nothing as evidence the exercise was too narrow rather than that the estate is secure, and report what that work finds separately from control coverage.
The reason for that specificity is a failure I have seen: A team spent a year raising claimed coverage from 60 to 92 percent while the unmanaged devices that carried the eventual intrusion were outside every measurement.
Where the year went.
| Activity | Effort | Real paths closed |
|---|---|---|
| coverage improvement | 80% | 3 |
| adversary-perspective work | 20% | 11 |
I would not consider it settled without evidence: Report adversary-perspective findings separately from coverage metrics, and protect the time for that work.
The attacker does not read your coverage matrix.
Curated: · Written: · Reviewed:
QA-92Is deception worth deploying?(show answer)
The judgement in honeypots and deception is which control removes the class, not which one closes the ticket.
A deception asset has no legitimate use, so any interaction with it is a high-precision signal in an environment where precision is the scarce resource. Its cost is deployment and maintenance rather than continuous triage.
Concretely, place credentials, shares, and hosts that nothing legitimate touches, make them plausible enough to be attractive, alert on any interaction, and maintain them so they do not become stale and obviously fake.
The reason for that specificity is a failure I have seen: An estate with 4,200 alerts a week had no high-precision signal; a handful of deception credentials later produced 3 alerts in a quarter, all of them genuine.
Precision by signal source.
| Source | Alerts/quarter | True positives |
|---|---|---|
| general rules | 54000 | 40 |
| deception credentials | 3 | 3 |
I would not consider it settled without evidence: Measure precision of deception alerts against the rest of the alert pipeline.
Nothing legitimate touches it, so anything that does is a finding.
Curated: · Written: · Reviewed:
QA-93What changes about detection when workloads move to the cloud?(show answer)
Where candidates lose the interview on cloud and on-premises detection differences is reaching for user training first.
The host-level telemetry that on-premises detection relies on is partly replaced by control-plane API logs, and the interesting events shift from process execution to identity and configuration change. Rules written for one do not transfer to the other.
Concretely, build detections on the control-plane audit log for identity, permission, and configuration events, keep workload telemetry where you still run workloads, and re-derive coverage for the cloud estate rather than assuming the on-premises matrix carries over.
The reason for that specificity is a failure I have seen: A mature on-premises detection set was assumed to cover the cloud migration; the cloud estate had 4 detections against 180 on-premises, and the first cloud intrusion was found by the provider.
Coverage after migration.
| Estate | Detections | Techniques exercised |
|---|---|---|
| on-premises | 180 | 34 |
| cloud | 4 | 0 |
I would not consider it settled without evidence: Build a separate coverage matrix for the cloud estate and exercise it independently.
A new platform is a new coverage matrix.
Curated: · Written: · Reviewed:
QA-94What does it mean in practice to say identity is the perimeter?(show answer)
I would answer identity as the new perimeter by separating what was detected from what was merely logged.
When workloads are distributed and users are remote, network location stops distinguishing trusted from untrusted, and the access decision rests on identity, device state, and context. That makes the identity system the highest-value target and the primary place to invest.
Concretely, invest in phishing-resistant authentication, conditional access on device and context, monitoring of identity-system configuration changes, and privileged access management, ahead of further network controls.
The reason for that specificity is a failure I have seen: An organisation invested heavily in perimeter controls while identity administration had no dedicated monitoring; the intrusion added a federation trust and nothing alerted for 5 weeks.
Monitoring by control plane.
| Control plane | Alerts on config change | Time to notice |
|---|---|---|
| firewall | yes | minutes |
| identity provider | no | 5 weeks |
I would not consider it settled without evidence: Confirm changes to identity-system trust and claim configuration produce an alert to a person.
Invest where the access decision is actually made.
Curated: · Written: · Reviewed:
QA-95How do you design conditional access without locking everyone out?(show answer)
The engineering content of conditional access design is the containment path and its rehearsal, not the framework name.
Conditional access evaluates every sign-in, so an error affects everyone at once and can lock out the administrators who would fix it. The design has to include the exclusion that survives its own mistake.
Concretely, exclude a monitored break-glass account from every policy, deploy in report-only mode to size the impact, roll out per group, and check that no policy applies to the account that would be needed to disable it.
The reason for that specificity is a failure I have seen: A policy requiring a compliant device applied to every account including administrators; nobody had a compliant device on the day of the rollout, and recovery required vendor support over 6 hours.
Rollout safety.
| Preparation | Lockout risk | Recovery |
|---|---|---|
| direct enforcement | high | 6 h, vendor |
| report-only first | low | n/a |
| + break-glass exclusion | none | minutes |
I would not consider it settled without evidence: Confirm a monitored break-glass account is excluded from every policy, and test that it works before each rollout.
The policy must not be able to lock out its own author.
Curated: · Written: · Reviewed:
QA-96Contractors use their own laptops. How do you handle that?(show answer)
Before calling device trust and unmanaged endpoints covered I would write down the technique it does not stop.
An unmanaged device cannot be attested, so any access it receives must assume the device may be compromised. The workable answers either bring the work to a managed environment or restrict what the unmanaged device can reach and take away.
Concretely, provide a managed virtual desktop or browser-isolated access for sensitive work, restrict unmanaged access to a defined low-sensitivity set, prevent download where the data justifies it, and require phishing-resistant authentication regardless.
The reason for that specificity is a failure I have seen: Contractors received the same access as employees from unmanaged devices; one device was already compromised and the session was used to access customer data, with no endpoint telemetry available.
Access by device state.
| Device | Data reachable | Download permitted | Telemetry | Controls of 3 |
|---|---|---|---|---|
| managed | full | yes | yes | 3 |
| unmanaged, isolated | limited | no | session only | 2 |
| unmanaged, direct | full | yes | none | 0 |
I would not consider it settled without evidence: Enumerate what an unmanaged device can currently reach and confirm it matches what you intended.
Unattested device means assume compromised.
Curated: · Written: · Reviewed:
QA-97Security logging costs more than the team's tooling budget. What do you cut?(show answer)
The first thing I would establish about logging cost against coverage is what it costs an attacker to get past it.
Log value varies enormously by source, and most volume comes from sources that appear in few investigations. Cutting by volume rather than by investigative value keeps the cost down without losing the sources that matter.
Concretely, rank sources by volume and by how often they appear in real investigations, cut from the bottom of that ranking, sample high-volume low-value sources rather than dropping them entirely, and archive cheaply rather than deleting.
The reason for that specificity is a failure I have seen: A uniform retention cut removed authentication logs beyond 30 days along with everything else; the next investigation needed 4 months of authentication history and had none.
Sources by value density.
| Source | Share of volume | Investigations using it |
|---|---|---|
| web access logs | 68% | 12% |
| authentication | 3% | 94% |
| DNS | 11% | 71% |
| endpoint process | 18% | 88% |
I would not consider it settled without evidence: Rank sources by appearance in past investigations per gigabyte, and cut from the bottom.
Cut by value per byte, not uniformly.
Curated: · Written: · Reviewed:
QA-98How do you keep detection content from decaying?(show answer)
I would start SIEM content lifecycle from what the telemetry can actually show, not from the rule list.
A rule written against a log format, a field name, or a product version stops working when any of those change, and it fails silently because a rule that matches nothing looks identical to a rule with nothing to match.
Concretely, test each rule on a schedule against generated data, alert on rules that have not fired within their expected period, version content alongside the platform, and retire rules that no longer have a source rather than leaving them in place.
The reason for that specificity is a failure I have seen: A field rename during a product upgrade silently broke 22 detections; nobody noticed for 7 months because no rule firing looks the same as a quiet environment.
Broken rules found by method.
| Method | Broken rules found | Time to find |
|---|---|---|
| waiting for an incident | 1 | 7 months |
| firing-interval alert | 22 | 2 weeks |
| scheduled generated tests | 22 | 1 day |
I would not consider it settled without evidence: Alert on rules that have not fired within their expected interval, and test with generated data on a cadence.
A rule matching nothing looks like a quiet environment.
Curated: · Written: · Reviewed:
QA-99How do you stop security debt becoming permanent?(show answer)
This is an area where owning a tool for security debt with an owner and detecting it are different things.
Security debt loses to work with a visible deadline because its cost lands on a future incident. Recording exposure, an owner, and a review date is what lets it compete, and reporting overdue items is what keeps it competing.
Concretely, record each item with the exposure it creates in operational terms, an owner with the authority to fix it, and a review date; report items past their date as overdue; and check after each incident which register items contributed.
The reason for that specificity is a failure I have seen: A register of 90 items had a named owner for only 22 and a review date for only 14; 6 of them appeared in a later incident as contributing conditions, all raised more than a year earlier.
Register health.
| Property | Count of 90 |
|---|---|
| named owner | 22 |
| review date | 14 |
| past review date | 61 |
| appeared in an incident | 6 |
I would not consider it settled without evidence: Report the register by age and by whether items later appeared in incidents.
Debt without an owner and a date is a note.
Curated: · Written: · Reviewed:
QA-100You join a security team with nothing in place. What does success look like after a year?(show answer)
My answer to what good looks like after a year begins with the condition rather than the alert it produced.
Success is measured by what an attacker's path costs, not by tools deployed. A year well spent removes the cheapest paths, establishes tested detection on the ones that remain, and leaves a team that can find and contain something.
Concretely, sequence by leverage: phishing-resistant authentication and privileged access first, tested backups second, asset inventory and internet-facing exposure third, then detection on the paths that remain, with each verified by exercise rather than by deployment.
The reason for that specificity is a failure I have seen: A year spent deploying six tools left multi-factor authentication unenforced and backups untested; the incident that followed used a reused password and encrypted the backups along with everything else.
Year one by sequence.
| Order | Work | Verified by |
|---|---|---|
| 1 | phishing-resistant MFA | relay attempt fails |
| 2 | immutable tested backup | restore in 6 h |
| 3 | inventory and exposure | 61 services removed |
| 4 | detection on remaining paths | 29 techniques exercised |
I would not consider it settled without evidence: Demonstrate each claim by exercise — attempt the phish, attempt the lateral movement, restore the backup — rather than by showing the tool.
Measure the year by what an attack now costs.
Curated: · Written: · Reviewed:
