Overview
Curated: · Written: · Reviewed:
Key takeaways
- Authentication proves who or what is calling; authorization evaluates whether that principal may perform this action on this resource in this context.
- Prefer federated, short-lived sessions for people and workload identities for software. Long-lived shared keys turn distribution and revocation into recurring incidents.
- Least privilege is a lifecycle: start narrow, observe legitimate use, refine, remove unused access, and re-evaluate when the workload or organization changes.
- Effective permission is produced by several policy layers. An allow in one document does not necessarily survive a boundary, session policy, organization guardrail, resource policy, or explicit deny.
- A role has two sides: who may assume it and what an assumed session may do. Secure both, especially across accounts and for systems that can pass roles.
- Network isolation, encryption, secrets management, audit logs, and detection complement IAM; none makes an over-privileged identity harmless.
1. Identities, credentials, and sessions
Human users should authenticate through a central identity provider with phishing-resistant MFA where possible, then obtain short-lived cloud sessions. Joiner, mover, and leaver events stay in one identity system, and access can be assigned to groups instead of copied across permanent cloud users.
Workloads need identities too. Attach a role or managed workload identity to compute rather than embedding an access key in source, images, environment files, or deployment variables. The platform issues temporary credentials that expire and can be refreshed by the SDK. Expiration limits the lifetime of a stolen token, but it does not make misuse during that window safe; permissions, session duration, audience, network context, and detection still matter.
Treat the root or tenant-owner identity as a recovery authority, not a daily administrator. Protect it with independent MFA, no routine access keys, tightly controlled recovery factors, monitored use, and an exercised break-glass process.
2. Policy anatomy and least privilege
An authorization request contains a principal, action, resource, and context. A policy statement combines an effect with actions, resources, and optional conditions. Prefer the narrow service actions and resource ARNs the task requires. Conditions can constrain source organization, network endpoint, region, resource tag, requested tag, authentication strength, or token claims, but only when the service and action actually support the chosen keys.
Least privilege is not achieved by replacing one wildcard with hundreds of unowned statements. Define a job or workload capability, grant its required operations, test normal and failure paths, observe access history, and remove permissions that are unused or no longer justified. Separate read, write, administrative, billing, security, and key-management duties. High-risk actions such as changing policies, passing roles, creating credentials, disabling logging, or decrypting broad data deserve explicit review.
3. Effective permission and guardrails
Requests begin implicitly denied and need an applicable allow. An explicit deny in an applicable policy wins over an allow. Identity and resource policies can contribute grants, while a permissions boundary limits the maximum identity permissions and an organization service-control policy limits what member accounts may use. A boundary or organization guardrail does not grant access by itself.
Session policies can further restrict an assumed session. Resource-policy and cross-account details vary by principal type, so troubleshoot the actual request context rather than memorizing 'all policies are unioned.' Capture principal ARN/session, action, resource, region, condition values, identity policy, resource policy, boundary, session policy, organization controls, and key policy. Use policy simulation as evidence, then verify the real service because resource state and service-specific behavior also participate.
4. Roles, trust, and delegation
A role's trust policy controls which principals may request a session. Its permissions policies control what that session can do. Cross-account access therefore needs cooperation: the target account trusts an external principal, and the caller is authorized to assume that role. Scope both sides and use conditions such as organization ID, source identity, audience, subject, or external ID where they address the threat model.
Protect against confused-deputy paths. A service provider acting for many customers should not let one customer name another customer's role and cause the provider to assume it. Bind each delegation to the expected tenant, resource, and request context. CI/CD federation should validate issuer, audience, repository/project, branch or environment claims; trusting every token from a large shared issuer is not least privilege.
Restrict role-passing independently from role-assumption. A principal that cannot perform an administrative action directly may still escalate if it can attach or pass a more privileged execution role to a service it controls.
5. Secrets, encryption, and key access
IAM credentials authorize cloud API calls. Application secrets such as database passwords, API tokens, and signing material belong in a managed secret store with scoped retrieval, rotation, versioning, and audit—not in source control or plaintext configuration. Rotation must update both the secret and every consumer safely; changing only the stored value can cause an outage.
Encryption limits what data is readable after storage or transport compromise, while IAM and key policies control who may invoke cryptographic operations. Envelope encryption uses a data key for the payload and a longer-lived managed key to protect that data key. Separate key administration from key use, scope decrypt permissions to the intended data path, and monitor unusual decrypt volume. Encryption does not repair an authorization rule that willingly returns plaintext to the wrong authorized principal.
6. Audit, detection, and recovery
Record sign-ins, session issuance, policy and trust changes, credential creation, role assumptions, denied requests where available, key operations, and logging changes in a protected central account. Preserve actor, session name or source identity, source, time, action, resource, outcome, and relevant request context. Alert on root use, new long-lived keys, public or cross-account grants, broad policy changes, disabled telemetry, unusual regions, and anomalous assumption or decrypt behavior.
Access Analyzer and policy validation can identify external access, unused permissions, and risky syntax, but findings require ownership and remediation deadlines. During a suspected credential leak, preserve evidence, disable or revoke the credential or trust path, contain active sessions where supported, identify what the principal could and did access, rotate dependent secrets, repair persistence, and learn why detection or least-privilege controls failed. Do not delete audit history to make the incident appear resolved.
7. Worked example: reading an effective permission
Effective access is the result of evaluating every policy that applies, and the order is what people get wrong. An explicit deny wins over any allow, no matter how specific the allow is:
Identity policy on role/DataAnalyst:
Allow s3:GetObject on arn:aws:s3:::reports/*
Bucket policy on reports:
Allow s3:GetObject to role/DataAnalyst
Deny s3:GetObject when aws:SourceVpce != vpce-0a41f9
Permission boundary on role/DataAnalyst:
Allow s3:*, athena:* (no kms:*)
Request: GetObject reports/q3.csv, from vpce-0a41f9, object encrypted with a customer-managed key
Result: DENIED - not by any of the three rules above, but by the missing kms:Decrypt
Every explicit rule passes. The request fails on a permission nobody wrote down: reading an SSE-KMS object
needs kms:Decrypt on the key, and the boundary silently caps that away. This is the most common
"the policy looks right" bug, and the answer to it is to read the evaluation as a pipeline rather than a list.
The order, and what can stop a request at each stage:
| Stage | Stops the request when |
|---|---|
| 1. Explicit deny, anywhere | Any policy denies. Nothing later can rescue it. |
| 2. SCP / org policy | The account is not permitted the action at all. |
| 3. Permission boundary | The action is outside the boundary, even if the identity policy allows it. |
| 4. Identity policy | No statement allows the action. |
| 5. Resource policy | Cross-account, and the resource does not allow the principal. |
| 6. Session policy | The assumed-role session narrowed itself below the need. |
| 7. Downstream key/service policy | The action needs a second permission, e.g. kms:Decrypt. |
The practical habit: when access fails, name which of these 7 stages stopped it before changing any policy. Adding an allow at stage 4 for a request that died at stage 1 or 3 widens the blast radius without fixing the symptom, and that widened policy usually outlives the incident by years.
