Overview
Curated: · Written: · Reviewed:
Threat modeling turns system understanding into risk decisions
Threat modeling is a structured, repeatable way to ask what is being built, what can go wrong, what controls exist or are proposed, whether they work, and what remains. Risk assessment connects threat events, threat sources, vulnerabilities or predisposing conditions, likelihood, adverse impact, uncertainty and organizational risk tolerance. Neither process is a one-time diagram workshop or a numerical formula that produces objective truth. Their purpose is better design and accountable decisions.
NIST SP 800-154 remains an Initial Public Draft as of this edition; its data-centric methodology is useful context but is not presented here as final NIST guidance.
Scope and model the real system
State the decision, system boundary, version, environment, assumptions, owners, dependencies and exclusions. Identify mission and business processes, users and workloads, assets and data, security and privacy objectives, trust boundaries, entry and exit points, privileged operations, external services, administrative and recovery paths, and data lifecycle from collection through deletion. Diagrams must include identity, control planes, CI/CD, observability, backups, support tools, third parties and failure modes—not only the happy-path request flow.
Threat scenarios are concrete causal statements: a capable source uses a precondition or weakness through an attack path to affect an asset or mission outcome. STRIDE can prompt spoofing, tampering, repudiation, information disclosure, denial of service and elevation of privilege, but categories neither prove completeness nor calculate risk. ATT&CK and CAPEC provide behaviors and patterns; they are inputs, not checklists or evidence that every mapped scenario is relevant. Include mistakes, outages, dependency failure, malicious insiders, supply-chain compromise, privacy misuse, safety harm and authorized-feature abuse.
Assess and respond to risk
Assess likelihood from threat capability, intent and targeting; exposure and reachability; preconditions; vulnerability and control effectiveness; detection and recovery; and available evidence. Assess impact across safety, mission, customers, finance, legal/privacy, confidentiality, integrity, availability and systemic or downstream consequences. Record qualitative or quantitative scales, time horizon, evidence, confidence and assumptions. Multiplying ordinal labels can create false precision; rank ties with scenarios, consequence and uncertainty.
Distinguish inherent or pre-control risk from residual risk after verified controls. Choose responses: avoid the risky activity, mitigate likelihood or impact, transfer/share defined consequences, or accept residual risk within authority and tolerance. A control maps to scenario preconditions, path, detection, containment or recovery; generic control lists are weak evidence. Each action has owner, deadline, dependency, verification and residual risk. Risk acceptance states scope, rationale, assumptions, compensating controls, monitoring, expiration and reopening triggers; it is not a permanent waiver.
Model authority, abuse, and dependency failure
Architecture diagrams often omit the paths that dominate real incidents. Model administrative and support actions, CI/CD and update authority, secrets and signing roots, observability and backup control planes, partner federation, customer-managed configuration and emergency recovery. Show who can change policy or evidence, not only who can call the business API. For multi-tenant systems, trace tenant identity and object ownership through caches, queues, exports, analytics and operator tools. For safety- or mission-critical systems, include hazardous states, delayed decisions and loss of control even when no attacker is present.
Authorized features can be abused without violating a protocol: bulk search becomes enumeration, password reset becomes account takeover, export becomes exfiltration, automation becomes a confused deputy and a valid administrator session becomes destructive change. State business invariants and rate, volume, sequence, purpose or dual-control constraints around those actions. Third-party and platform threats include malicious updates, compromised maintainers, unavailable dependencies, control-plane takeover, correlated regional failure and misleading provider evidence. Transfer by contract may allocate financial responsibility, but it does not remove operational harm or the need for detection, containment, exit and recovery design.
Calibration and decision records
Risk scoring must be calibrated to the decision it supports. Define scale anchors with observable consequences and time horizons, distinguish uncertainty from low likelihood, and record whose evidence informed the estimate. Scenario-specific ranges or ordered categories are often more honest than a precise expected-loss number built from sparse data. Compare similar assessments for consistency, revisit estimates after tests and incidents, and expose sensitivity: if one unverified assumption moves the response from accept to mitigate, resolving that assumption is valuable work.
The final record connects scenario, affected objective, existing control and evidence, likelihood and impact rationale, uncertainty, chosen response, accountable authority and next verification. Accepted risk belongs to the person authorized to own the business consequence, not automatically to the engineering team implementing the system. Material changes reopen the decision, while superseded records remain available so reviewers can understand why a former design was reasonable and which assumption changed.
Validate and keep it alive
Test prevention, authorization boundaries, detection, fail behavior, containment and recovery through code/design review, unit and integration tests, adversarial simulation, fault injection, tabletop exercises and incident evidence. Control deployment is not proof of effectiveness. Track model coverage of critical flows and changes, action age, accepted-risk expiry, test results, incidents and recurring patterns—not threat count or diagram size.
Update when architecture, data, identity, exposure, dependencies, threats, controls, jurisdiction or mission changes, and after incidents or exercises. Automations and AI can inventory, suggest scenarios and find drift, but generated threats need provenance, deduplication, relevance review and security/privacy controls; they cannot own risk. A strong model exposes uncertainty and gives decision-makers options rather than claiming the system is secure.
Worked example: a STRIDE prompt is not a decision
The payments export API: STRIDE fills a board. Only one row changes the design this week.
| STRIDE | scenario | evidence | decision |
|---|---|---|---|
| Spoofing | stolen session on export | MFA + step-up exists | accept residual |
| Tampering | alter CSV in flight | TLS + checksum | accept residual |
| Repudiation | deny who exported | audit log untested | test the log first |
| Information disclosure | bulk export of all tenants | no object-level check | mitigate now |
| Denial of service | export fan-out | rate limit in prod | accept residual |
| Elevation | intern grants export | dual control missing | mitigate now |
Six stickers, two mitigations, one verification. The interview is the decision column, not the mnemonic.
