Overview
Curated: · Written: · Reviewed:
Zero trust is resource-centered access architecture
Zero trust is a set of design principles and an enterprise architecture strategy, not a product, protocol, or claim that no component is trusted. NIST SP 800-207 removes implicit trust based only on network location, ownership, or affiliation and focuses policy on access to individual enterprise resources. A subject and its device or workload are authenticated and authorized before a session is established; access is least-privilege, dynamically evaluated from multiple signals, monitored, and subject to reauthorization or termination. Encryption protects a channel but does not decide whether an action is allowed.
Logical components and policy
In NIST's logical model, the policy engine decides whether access should be granted using enterprise policy and inputs such as subject and asset identity, credential strength, device state, resource sensitivity, observed behavior, threat intelligence and legal or operational constraints. The policy administrator establishes or tears down the communication path, while a policy enforcement point enables, monitors and ends the connection. Implementations distribute or combine these roles, but separating decision, administration and enforcement clarifies authority, failure modes and audit evidence.
Policy must define actor, resource, action, context, obligations, duration, re-evaluation and denial reason. Prefer short-lived, scoped credentials and explicit service/workload identities over shared secrets, IP addresses or a network segment as identity. Device posture is one risk input, not proof the user or workload is benign. Signals have freshness, provenance and uncertainty; attackers can steal credentials, compromise managed devices, poison context or target the policy infrastructure itself.
Multiple enforcement layers
Identity governance, phishing-resistant authentication, conditional access, endpoint management, application authorization, API gateways, workload identity, service meshes, microsegmentation, data classification and monitoring can all contribute. Microsegmentation limits paths and lateral movement but does not replace resource-level authorization. NIST SP 800-207A emphasizes application and service identities and granular policies for cloud-native, multi-cloud workloads rather than relying only on IP addresses and subnets. Every application must still enforce object, tenant and business-action authorization at its authoritative boundary.
The control plane that issues identity, distributes policy and configures enforcement must be highly available, strongly authenticated, least-privileged, observable and recoverable. Define fail-open, fail-closed or cached-decision behavior per resource and action; a universal failure mode can either stop the mission or expose high-value data. Emergency access is intentionally designed, time-bound, monitored, independently recoverable and reviewed—not an undocumented bypass.
Migration and operations
Begin with assets, identities, data, flows, owners and business requirements. Choose a high-value but bounded use case, document its current trust assumptions, establish telemetry and success measures, then migrate incrementally with shadow evaluation, canaries and rollback. Legacy protocols, unmanaged devices, partners, operational technology and availability constraints require compensating controls and explicit residual risk; calling a network private does not resolve them. CISA's maturity model organizes progress across identity, devices, networks, applications/workloads and data with visibility/analytics, automation/orchestration and governance spanning them. It is a roadmap, not a certification checklist.
Decision semantics and degraded operation
A production design must say where the authoritative decision is made and what each enforcement point does with it. Policy inputs need a common vocabulary and explicit precedence: a broad identity grant must not silently override an object's tenant boundary, and an upstream allow must not replace the application's business-state checks. Carry obligations such as masking, step-up authentication, session limits and audit requirements with the decision. Version policy, signal enrichment and enforcement behavior together, retain an explainable trace, and test the same fixtures at gateways, services and data boundaries so implementation drift is visible.
Availability choices are resource-specific. A cached decision may be reasonable for a low-impact read during a brief control-plane outage but unacceptable for a privileged write after a revocation. Define cache lifetime, binding to subject/resource/action, invalidation, stale-signal behavior and the maximum exposure window. Exercise issuer failure, policy-store partition, delayed device posture, clock skew, regional isolation and loss of telemetry. Emergency access must use separately protected, individually attributable or quorum-controlled credentials, open an incident, alert out of band, expire automatically and be reconciled after normal authority returns.
Data, privacy, and proof
Continuous evaluation can itself become invasive. Collect only signals that can change a stated decision, establish purpose and retention, restrict raw telemetry, and prevent risk scores from becoming unexplained identity attributes reused elsewhere. A denial reason should be useful to the subject and responder without revealing a detection rule or another tenant's data. Protect policy logs from both the requestor and the administrator whose actions they record, and make correlation identifiers sufficient to reconstruct the decision without storing unnecessary payloads.
Acceptance evidence follows a representative request end to end: authoritative identity and device state, credential issuance, policy decision and version, enforcement at every reachable path, resource result, audit event, revocation and recovery. Negative tests cover direct endpoints, legacy protocols, administrator consoles, backup paths, service-to-service calls and stale sessions. Measure legitimate-user completion and false denial alongside unauthorized reachability, privilege and blast radius; otherwise teams can make the security metric look better simply by making the service unusable.
Measure reduced unauthorized reachability and standing privilege, policy coverage, credential and device assurance, segmentation and application-authorization tests, decision/enforcement availability and latency, stale signals, exception age, denied and revoked sessions, incident blast radius and recovery exercises. Do not use login friction, number of purchased tools, encrypted-traffic percentage or policy count as standalone proof of security. Zero trust complements prevention, detection, incident response, resilience and secure software design; it does not make compromise impossible or eliminate the need to understand data flows and business impact.
Worked example: VPN + TLS is not a resource decision
Stolen laptop on the office VPN. Request: GET /invoices/4419 for tenant Acme.
| enforcement | identity used | 4419 for Acme? |
|---|---|---|
| allow any RFC1918 source, TLS only | none | yes (any VPN client) |
ZTNA to the app; app checks if (user) | JWT present | yes, including other tenants |
PEP + object ACL invoice.tenant = session.tenant | subject + tenant | no for the stolen-other-tenant token |
The interview is the last row. A product that terminates TLS at a “zero-trust gateway” and then trusts the app’s session cookie has not moved the decision.
