Skip to content
Tech Interview Prep home
Technical interview guide

Cloud IAM Policy Design

Writing least-privilege access policies for cloud resources, and avoiding the common over-permissioning failure modes.

Read
44 min
Practice MCQs
25
Interview QA
25
Edition
v4
Editorial status
Reviewed

Scope: AWS IAM, Google Cloud IAM, Azure RBAC/managed identity, and NIST SP 800-207 guidance current 2026-08-31.

Overview

Curated: · Written: · Reviewed:

Design effective permissions, not isolated policy documents

Cloud IAM questions show up in system design and security rounds at every level, and at senior/staff level the bar is not "knows what a policy is" — it's "can reason about effective permissions across layers and explain the failure modes." Interviewers here are usually probing for three things: whether you evaluate the whole request instead of one JSON document, whether you understand that identity administration is itself the crown-jewel privilege, and whether you've actually debugged an access denial in production. A weak answer recites vendor vocabulary ("we use least privilege") without being able to say what the cloud actually enforces when six policy types apply to one request.

The mental model: one request, every layer

Authorization answers whether a principal may perform an action on a resource in a specific request context. Authentication establishes an identity; authorization evaluates what that identity may do now. A network location, a company-owned device, or a successful sign-in is not itself authorization — a common weak answer conflates these.

Most cloud systems evaluate in the same shape:

  1. Start with implicit deny — no applicable policy means no access.
  2. Collect every applicable policy: identity policies, resource policies, role trust, session policies, permission boundaries, organization policies, deny assignments, service-specific controls.
  3. An explicit deny anywhere wins over any allow.
  4. Otherwise, an applicable allow grants the action — but caps like permission boundaries intersect rather than union.

The exact combination rules differ by vendor, especially for resource policies and session principals. The interview-grade point is not memorizing one vendor's matrix; it's knowing that the layers combine by both union and intersection, and that you never infer access by reading one document. When an interviewer asks "why does this principal have access you didn't expect?" the strong answer starts with "simulate or test the full request context," not with squinting at the role definition.

Likely follow-up: "walk me through what happens when an identity policy allows s3:GetObject and a resource policy denies it." If you can't say "explicit deny wins, request fails" without hesitation, that's the gap they're looking for.

The policy taxonomy, and what each type is for

Each policy type exists to answer a different question. Confusing them is the most common weak answer pattern:

  • Identity policies answer "what can this principal do?" — attach to users, groups, roles.
  • Resource policies answer "who can touch this resource?" — they grant to other principals, including cross-account, which identity policies cannot do.
  • Trust policies answer "who may assume this role?" — they are the front door, not the permissions inside.
  • Session policies answer "what does this particular session allow?" — a cap scoped to one assumption, useful for federation and just-in-time narrowing.
  • Permission boundaries / scopes answer "what is the ceiling on what this identity can ever be granted?" — they cap, they never grant. A boundary limiting a role to S3 read does not let the role read S3; it just means no attached policy can grant more.
  • Organization / management-group guardrails answer "what is forbidden everywhere, regardless of local admins?" — deny-only in most platforms, protecting invariants like "logging cannot be disabled."

Interviewers love the boundary question because it separates people who have used the feature from people who have read about it. "Does attaching a permission boundary grant anything?" No. "Can a boundary fully prevent escalation?" Also no — it doesn't cover every resource-policy or trust-policy path, and saying so with a concrete example (a trust-policy misconfiguration can bypass your mental model entirely) is a strong-answer move.

Least privilege is multidimensional — and fails in more ways than wildcards

"Avoid *" is the junior version of least privilege. The senior version limits along six dimensions: allowed actions, resource scope, time, session conditions, data boundary, and ability to delegate. And the failure modes interviewers probe are the ones that don't look like wildcards:

  • Usability collapse. A policy so brittle that operators route around it — sharing credentials, hoarding admin, building shadow bypasses — is a failed design. Least privilege that blocks legitimate work gets disabled.
  • Delegation leaks. The ability to pass a role to a service, attach a policy, or edit group membership is the target privilege. A principal that can modify its own policy has, effectively, every permission that policy could name.
  • Condition-key trust. A user who can retag a resource or principal can bypass a tag-based policy. Whoever controls the attribute controls the access.
  • Missing-attribute behavior. Deny conditions often fail closed when unevaluable; allow conditions generally need to evaluate true. If your whole design rests on a tag that's absent on new resources, you've built an accidental allow or an accidental outage, depending on direction.

Prefer vendor job-function roles when they fit, then reviewed custom roles for stable missing permission sets. When asked "how do you decide between a predefined role and a custom one?" the strong answer names the trade-off: predefined roles are maintained and reviewed by the vendor but over-grant; custom roles are exact but become unreviewed sprawl. Decide by blast radius and review cadence, not by convenience.

Condition keys: the real least-privilege lever

When an interviewer pushes past "what actions did you grant?" the follow-up is almost always conditions. Time, source resource, organization, network, device, authentication strength, resource tags and principal tags can all narrow a grant — and conditions are where real least privilege lives, because they bind the grant to context rather than to a static action list.

Two things make a condition trustworthy: the attribute is reliable (who sets it, can the grantee influence it?), and the platform supports it for that action. A tag condition on a resource is only as strong as who can write that tag. A KMS encryption-context condition is strong for a different reason: the caller does supply the encryption context in the request, but it is cryptographically bound to the ciphertext and must match exactly on decrypt — so a grant scoped to one context cannot be replayed against data encrypted under another.

Complex policy expressions need fixtures and negative tests. "How would you test a policy?" deserves more than "try it in the console": unit-test expected allows and expected denies, including boundary cases, inherited policy, anonymous access, cross-account principals, and privilege escalation. Validating only the happy path is the weak answer.

Cross-account access needs agreement on both sides

Cross-account grants fail in interviews the same way they fail in production: people configure one side. The trusting side constrains the role trust to an exact external principal; the trusted side still needs an identity policy allowing the assumption; and the role's permission policy is a third, separate document. All three must agree.

For third-party access: identify the exact external principal, constrain the trust, grant a narrow permission policy, use short sessions, and log assumption and use. In multi-tenant vendor scenarios, a customer-specific external identifier prevents confused-deputy attacks — and note it's a trust condition, not a secret; treating it as a secret is a misconception worth naming proactively.

Service-to-service resource policies should use supported source-account, source-resource or organization conditions, and be tested against unintended callers. A resource policy that says "allow from account X" is broader than it looks if account X is multi-tenant.

Humans, workloads, and the administration plane

Separate the two principal classes and say why:

  • Humans: workforce federation, strong authentication, short-lived sessions. Privileged actions just-in-time, approved or step-up protected when risk warrants.
  • Workloads: dedicated service or managed identities with temporary credentials bound to the runtime. Never embed access keys in code or images, never share one highly privileged identity across applications, never attach a powerful identity to an environment where less-trusted users can run code.

Then the part that separates senior answers: identity administration is itself privilege. Permissions to create roles, change policies, attach identities, pass or impersonate a service account, issue credentials, change group membership, or deploy onto a privileged runtime confer the target identity's authority. Model those escalation paths explicitly. When an interviewer asks "how could a principal with only IAM permissions escalate?" they're testing whether you see the administration plane as the real attack surface.

Guardrails, emergencies, and recovery

Organization guardrails protect a small set of high-value invariants — disabling logging, leaving the organization, creating long-lived keys — with carefully governed exceptions and impact analysis before deployment. Two properties interviewers probe: a guardrail that denies an action does not create a corresponding allow (deny-only means local admins still grant everything else), and broad denies can block recovery or break service principals — so test inheritance, break-glass paths, and propagation before enforcement.

Emergency access is designed before the emergency. A minimal number of strongly protected break-glass identities, outside ordinary dependency failure, with no routine use, tested credentials, narrow activation, alerting, and post-use review. Just-in-time access is different in kind: a normal governed elevation with approval, reason, duration, and automatic expiry. Neither should become a permanent administrator shortcut — and if you've seen "temporary" admin access survive two quarters, that anecdote is worth telling.

Operate it: policy as code, observability, decay

Manage policies as code: reusable definitions, stable identifiers, peer review, static validation, effective-permission analysis, public and cross-account findings, staged rollout. The test suite covers expected allows and expected denies.

Observe decisions without leaking secrets: record principal, session, action, resource, decision, relevant condition, and policy/change provenance where the platform supports it. Alert on privileged grants, public exposure, policy changes, break-glass use, credential creation, anomalous assumption. And the operational question interviewers actually ask — "a user reports access denied, walk me through it" — should be answered by explaining the evaluated context while preserving the guardrail. Fixing a denial by adding administrator access until it disappears is the weak answer; the strong one enumerates the layers and finds the explicit deny or missing allow.

Permissions decay. Inventory principals, owners, credentials, grants, and last-use evidence; review privileged, external, public and dormant access most frequently. Last-used data is evidence, not proof — seasonal recovery actions may legitimately fire once a year. Remove access through a staged, reversible process when uncertainty is material, and deprovision identity, sessions, keys, group membership, role assignments, resource policies and automation paths together.

The goal you should be able to state in one sentence at the end of an interview: a small, explainable set of effective permissions with accountable exceptions and a tested recovery path — compromise blast radius limited, escalation paths visible, and intended access continuously reconciled with what the cloud actually enforces.