Top 100 Identity and Access Management (IAM) Engineer Interview Questions and Answers
The questions most likely to actually come up in your Identity and Access Management (IAM) Engineer interview, ranked by likelihood — with detailed, senior-level answers covering what an interviewer is really listening for.
Curated: · Written: · Reviewed:
QA-1A team says they will use OAuth 2.0 to log users in. What do you tell them?(show answer)
The first thing I would establish about OAuth versus OIDC is which live path still works after the directory row looks right.
OAuth 2.0 is a delegation protocol: it produces an access token saying a client may call an API, and it says nothing about who is present or when they authenticated. OpenID Connect adds the identity layer — an ID token with an issuer, an audience, and an authentication time — which is what a sign-in decision actually needs.
Concretely, require the openid scope and validate the ID token for the sign-in decision, keep the access token for API calls only, and reject any design where a userinfo call on a bearer token stands in for proof that this user authenticated to this client.
The reason for that specificity is a failure I have seen: A partner portal treated a valid access token as a login. Tokens minted for a different application were replayed against the portal for 11 days, and 240 sessions were opened for people who had never authenticated there.
What each artifact answers.
| Question | Access token | ID token |
|---|---|---|
| names who authenticated | no | yes |
| carries an authentication time | no | yes |
| typical lifetime in minutes | 60 | 5 |
| audience checked by the app | rarely | required |
I would not consider it settled without evidence: attempt sign-in with a valid access token issued to another client and require the relying party to refuse it.
Delegation is not authentication.
Curated: · Written: · Reviewed:
QA-2Which OAuth flow do you pick for a public client that cannot hold a secret, and why not implicit?(show answer)
I would start authorization code flow with PKCE from the session, token, and grant that outlive the login, not from the login itself.
A public client cannot keep a secret, so the code has to be bound to the sender rather than to a credential. Authorization code with PKCE binds redemption to a verifier the attacker never saw; implicit returns the token in the URL fragment, where history, logs, and referrers can reach it.
Concretely, generate a 43-character random verifier per authorization request, send only its S256 challenge, redeem the code over the back channel with the verifier, and configure the authorization server to refuse plain challenges and to refuse a code redeemed twice.
The reason for that specificity is a failure I have seen: An internal application kept implicit flow through a migration window. Access tokens landed in the fragment, an embedded analytics script captured 1,900 URLs over 6 days, and 47 of the tokens were still inside their validity window when the capture was found.
Where the token can leak.
| Path | Implicit | Code with PKCE |
|---|---|---|
| token in URL fragment | yes | no |
| replay of a captured code | not applicable | 0 of 25 attempts |
| readable in browser history | yes | no |
I would not consider it settled without evidence: replay a captured authorization code without the verifier and require the token endpoint to return an invalid grant.
Bind the code to the sender, not to a secret the client cannot keep.
Curated: · Written: · Reviewed:
QA-3Your API accepts an ID token as a bearer credential. What breaks?(show answer)
This is an area where a successful federation and a current authorization decision are different events.
An ID token is a statement to one client that a user authenticated, and its audience is that client rather than the API. An API that accepts it is trusting an artifact minted for someone else, and it loses the scope and consent an access token carries.
Concretely, have the API compare the audience against its own resource identifier and refuse tokens whose audience is a client id, and have clients obtain an access token with the scopes the API publishes rather than forwarding whatever the login returned.
The reason for that specificity is a failure I have seen: A reporting API accepted ID tokens for convenience. A low-trust internal client's ID token, carrying no scopes at all, read 12,400 payroll rows in one afternoon because the scope check had nothing to evaluate.
What the API lost by accepting the wrong artifact.
| Property | ID token presented | Access token expected |
|---|---|---|
| audience value | a client id | the API resource id |
| scopes present | 0 | 3 |
| consent recorded | no | yes |
I would not consider it settled without evidence: present an ID token to the API and require a rejection that records an audience mismatch as the reason.
An ID token proves a login happened; it does not authorise a call.
Curated: · Written: · Reviewed:
QA-4Walk me through validating an ID token. What do you check beyond the signature?(show answer)
My answer to ID token validation across issuer, audience, nonce, and time begins with the artifact that is actually being trusted: token, assertion, group, or session.
A signature proves an issuer minted the token; it does not prove the token was minted for you, for this request, or recently. Validation is issuer match, audience match, authorised party when several audiences are present, nonce match against the stored request, and the expiry, issued-at, and authentication-time bounds.
Concretely, resolve the signing key from the issuer's published key set by key id, compare the issuer to the exact configured string, compare the audience to this client id, compare the nonce to the value held server-side for that authorization request, and refuse tokens whose authentication time is older than the maximum age the application requires.
The reason for that specificity is a failure I have seen: A relying party checked signature and expiry only. A correctly signed token from a second tenant of the same provider authenticated 63 accounts into the wrong customer's workspace, and the mismatch surfaced through a support ticket 4 days later.
Five checks, five separate failures.
| Check | Result when skipped | Rejections in the test |
|---|---|---|
| exact issuer match | cross-tenant sign-in | 3 of 3 |
| audience match | another app's token accepted | 2 of 2 |
| nonce match | replayed authorization response | 1 of 1 |
| authentication age | a 9-hour-old login reused | 4 of 4 |
I would not consider it settled without evidence: submit correctly signed tokens carrying the wrong issuer, the wrong audience, and a stale nonce, and require all three to be refused for distinct reasons.
A valid signature is the beginning of validation, not the end.
Curated: · Written: · Reviewed:
QA-5Are state, nonce, and PKCE not three names for the same protection?(show answer)
I would treat state versus nonce versus PKCE as a claim about effective access that has to survive a disablement.
They defend three different steps. State binds the authorization response to the browser that started it, the nonce binds the ID token to that same request, and PKCE binds code redemption to the client instance that asked; dropping one is not covered by keeping the other two.
Concretely, store state and nonce server-side against the pending authorization record, compare both on return, carry the PKCE verifier in the same record, and discard any response that arrives without all three before requesting a token.
The reason for that specificity is a failure I have seen: A team implemented PKCE and removed state as redundant. A crafted authorization response signed a victim's browser into the attacker's account on 3 of 5 tested browsers, and the account linkage survived the next sign-out.
Which attack each parameter stops.
| Parameter | Binds | Left open if removed |
|---|---|---|
| state | response to browser session | forced sign-in on 3 of 5 browsers |
| nonce | ID token to this request | token replay across 2 sessions |
| PKCE verifier | code to client instance | code interception on 1 device |
I would not consider it settled without evidence: drop each parameter in turn in a test harness and require the sign-in to fail for a different stated reason each time.
Three bindings, three attacks, and none substitutes for another.
Curated: · Written: · Reviewed:
QA-6A developer asks for a wildcard redirect URI on the staging client. What is your answer?(show answer)
The useful question for exact redirect URI matching is what still holds at the relying parties nobody has looked at this week.
The redirect URI is the only thing deciding where an authorization code is delivered, so it has to be compared as an exact string. Wildcards, prefix matching, and open path segments hand code delivery to whoever controls a subdomain or an open redirect on that host.
Concretely, register full absolute URIs including scheme, host, port, and path, compare byte for byte before issuing a code, and keep separate client registrations per environment so no production client ever holds a staging origin.
The reason for that specificity is a failure I have seen: A prefix-matched redirect on a marketing subdomain met an open redirect on the same host. Codes for 82 users were relayed to an external endpoint over 5 hours before the registration was corrected.
Registered pattern against what it accepts.
| Registered value | Accepts | Codes exposed in the drill |
|---|---|---|
| one absolute callback URI | that URI only | 0 |
| host with a path wildcard | any path, open redirect included | 82 |
| subdomain wildcard | any subdomain takeover | 14 |
I would not consider it settled without evidence: request authorization with an unregistered path on a registered host and require the authorization server to refuse before any code exists.
Comparison by prefix is delivery to a stranger.
Curated: · Written: · Reviewed:
QA-7Your authorization server still accepts a code challenge method of plain. Why does that matter?(show answer)
I would settle PKCE downgrade and verifier reuse by attempting the action after the identity event that was supposed to stop it.
PKCE is only as strong as the weakest method the server accepts and the freshest verifier the client generates. If plain is accepted, the challenge equals a verifier the attacker may already hold; if a verifier is reused, the binding is to a constant rather than to a request.
Concretely, refuse plain at the authorization server, require S256 in the client registration, generate a fresh verifier from a cryptographic source for every authorization request, and tie the verifier to the single-use request record so a second redemption finds nothing to compare against.
The reason for that specificity is a failure I have seen: A mobile SDK cached one verifier per install to save an allocation. An attacker who captured that verifier once redeemed 7 later codes over 19 days, and every redemption looked correct to the token endpoint.
Verifier lifecycle in the SDK before and after.
| Behaviour | Before | After |
|---|---|---|
| verifiers generated per install | 1 | 1 per request |
| codes redeemable with a stolen verifier | 7 | 0 |
| plain method accepted | yes | refused, 100 of 100 |
I would not consider it settled without evidence: attempt authorization with the plain method and with a previously used verifier, and require both attempts to be refused.
A constant verifier binds nothing.
Curated: · Written: · Reviewed:
QA-8Your application federates to four identity providers through one callback. What can go wrong?(show answer)
The judgement in multi-IdP authorization-server mix-up is which identifier is durable and which is only a convenient claim.
With one callback and several issuers, the response does not say which authorization server produced it unless you make it say so. Someone who controls one registered provider can have a code from their server redeemed at another, and the client will read the resulting identity as belonging to the honest provider.
Concretely, keep a per-request record naming the chosen issuer, require the response to carry an issuer parameter that matches it before redemption, and use a distinct redirect URI per provider where a provider will not return one.
The reason for that specificity is a failure I have seen: A marketplace let sellers register their own providers. A seller-controlled code was redeemed against the corporate issuer's token endpoint, and 5 corporate accounts were taken over before the mismatch was noticed on day 8.
Callback design against mix-up exposure.
| Design | Issuer identified | Accounts reached in the drill |
|---|---|---|
| shared callback, no issuer parameter | no | 5 |
| shared callback with an issuer check | yes | 0 |
| one callback per provider | yes | 0 |
I would not consider it settled without evidence: return an authorization response naming a different issuer than the one the request chose and require redemption to be refused.
One callback for many issuers needs the issuer inside the answer.
Curated: · Written: · Reviewed:
QA-9What do you store as the primary key for a federated user, and why not email?(show answer)
Where candidates lose the interview on issuer plus subject as the durable identity key is reading a green IdP dashboard as revocation.
The durable key is the pair of issuer and subject, because the issuer scopes the subject and only the issuer promises not to reassign it. Email, username, and display name are attributes that change with a marriage, a rebrand, or a reassignment, and any of those changes quietly becomes an account merge.
Concretely, store the issuer and subject as the account's identity row, keep email as a mutable attribute with its own verification state, and refuse to link an incoming assertion to an existing account on email alone without a deliberate and logged linking step.
The reason for that specificity is a failure I have seen: A mailbox address was reissued 90 days after the previous holder left. The new joiner's first sign-in matched on email and inherited 3 administrative group memberships along with 4 years of the previous holder's records.
Which attribute survived the year.
| Attribute | Changes in 12 months | Safe as a key |
|---|---|---|
| email address | 318 | no |
| display name | 502 | no |
| issuer with subject | 0 | yes |
I would not consider it settled without evidence: change a test user's email at the provider and confirm the application resolves the same internal account by subject.
Key on what the issuer promises not to reassign.
Curated: · Written: · Reviewed:
QA-10Your application grants roles straight from a groups claim. How do you make that safe?(show answer)
I would answer allowlisted claim-to-role mapping by separating authentication, authorization, and lifecycle into three control planes.
A claim is an assertion from the provider about the user, not a decision about your application. Mapping every group name straight through means whoever can create a group at the provider can create an administrator at the relying party.
Concretely, keep an explicit allowlist of group values that map to application roles, ignore unmapped values instead of passing them through, put mapping changes through the same review as a permission change, and log the claim values that were dropped.
The reason for that specificity is a failure I have seen: A self-service group feature let any employee create a group. Someone created one whose name matched a pass-through rule and held 6 privileged operations for 26 hours before the mapping was corrected.
Claim values seen in one synchronisation.
| Group value | In the allowlist | Role granted |
|---|---|---|
| finance readers | yes | reader |
| a self-service group named for admins | no | none |
| engineering on-call | yes | operator |
| 41 remaining values | no | none |
I would not consider it settled without evidence: create an unmapped group at the provider, sign in, and confirm the relying party grants nothing new.
The provider says what a user is in; you decide what that buys.
Curated: · Written: · Reviewed:
QA-11You disable a leaver at the identity provider. Which of their sessions are actually over?(show answer)
The engineering content of SSO disablement versus local session death is the revocation path and its measured latency, not the protocol name.
Disabling the source stops new authentications; it does not touch cookies already issued, refresh tokens already granted, or application sessions already established. Each of those is a separate object with its own lifetime held by a different party.
Concretely, pair disablement with provider session revocation, refresh-token revocation per client, and back-channel logout or a session-check obligation at each relying party, then measure how long each path takes to deny an action rather than assuming they end together.
The reason for that specificity is a failure I have seen: A dismissal disabled the account at 09:12. The browser session at the wiki lasted a further 7 hours, and a refresh token at the customer system kept minting access tokens for 13 days until it reached its absolute expiry.
Time to denial after disablement.
| Surface | Denied after | Mechanism |
|---|---|---|
| provider sign-in | 0 minutes | account disabled |
| wiki session cookie | 7 hours | cookie expiry |
| customer system refresh token | 13 days | absolute lifetime |
| remote access session | 42 minutes | reauthentication interval |
I would not consider it settled without evidence: disable a test identity, attempt a privileged action at each relying party, and record the minutes until each one denies.
The directory row is one object; the live grants are many.
Curated: · Written: · Reviewed:
QA-12Just-in-time provisioning created the accounts. What removes them?(show answer)
Before calling JIT provisioning versus SCIM deprovision done I would write down the grant that the source-of-truth change does not touch.
Just-in-time provisioning is a side effect of a sign-in, so it has a create path and no delete path. Without a provisioning protocol driving the other direction, an application accumulates accounts that only stop being used rather than stopping existing.
Concretely, run a provisioning protocol for lifecycle in both directions, keep just-in-time creation only where that is impossible, and reconcile the application's account list against the source weekly so accounts with no matching identity are surfaced and closed.
The reason for that specificity is a failure I have seen: An analytics tool created accounts on first sign-in for 3 years. A reconciliation found 214 accounts whose owners had left, 19 of them holding workspace administration, and the oldest had been unowned for 22 months.
Reconciliation of one application.
| Account state | Count | Action |
|---|---|---|
| active with a matching identity | 1,286 | keep |
| no matching identity | 214 | close |
| no matching identity, administrative | 19 | close and review |
I would not consider it settled without evidence: remove a test user at the source and confirm the application's own account list no longer contains them inside the stated window.
A create path without a delete path is an accumulation.
Curated: · Written: · Reviewed:
QA-13The SAML response signature verifies. What else must be true before you sign the user in?(show answer)
The first thing I would establish about SAML validation beyond XML signature is which live path still works after the directory row looks right.
Signature verification says the assertion was not altered by someone without the key. It does not say the assertion was meant for you, is still valid, or has not been used, and it does not say the signature covers the element your parser actually reads.
Concretely, verify that the signature reference covers the assertion you consume, compare the destination with your assertion consumer service URL, compare the audience with your entity id, enforce the condition window with a bounded clock skew, and match the in-response-to value against a request you issued.
The reason for that specificity is a failure I have seen: A service provider verified the signature and then read the first assertion in the document. A wrapped response carrying one signed but unrelated assertion plus an unsigned attacker assertion produced administrative sign-ins on 2 occasions before the parser was replaced.
Checks the service provider was missing.
| Check | Present before | Present after |
|---|---|---|
| signature covers the consumed assertion | no | yes |
| audience equals the entity id | no | yes |
| condition window with 120 seconds of skew | partial | yes |
| in-response-to matched | no | yes |
I would not consider it settled without evidence: submit a signature-wrapping response and an assertion addressed to another audience, and require both to be refused with distinct reasons.
Verify the signature, then verify it covers what you read.
Curated: · Written: · Reviewed:
QA-14How do you stop a captured SAML assertion from being used twice?(show answer)
I would start SAML assertion replay from the session, token, and grant that outlive the login, not from the login itself.
An assertion is a bearer artifact that stays valid for its whole condition window, so anyone who captures it can present it again until the window closes. Replay defence is a shared record of assertion identifiers plus a short window, not a stronger signature.
Concretely, record the assertion identifier with its expiry in a store every service-provider node reads, refuse a second presentation of the same identifier, keep the condition window to a few minutes, and make the store fail closed when it is unreachable.
The reason for that specificity is a failure I have seen: A load-balanced service provider kept its replay cache in each node's memory. One captured assertion was accepted 4 times across 6 nodes inside a 5-minute window, and each acceptance created a full session.
Replay attempts across the node pool.
| Cache design | Nodes | Successful replays |
|---|---|---|
| per-node memory | 6 | 4 |
| shared store, fail closed | 6 | 0 |
| no cache at all | 6 | 11 |
I would not consider it settled without evidence: present the same assertion twice against different nodes and require the second presentation to be refused.
One-time use needs shared memory of what was used.
Curated: · Written: · Reviewed:
QA-15You need to rotate the identity provider signing key. How do you do it without an outage?(show answer)
This is an area where a successful federation and a current authorization decision are different events.
Rotation is a two-key problem: relying parties must trust the new key before it signs anything, and must keep trusting the old key until nothing it signed is still in flight. A cut-over at a single instant is an outage for every party that caches metadata.
Concretely, publish the new key alongside the old in metadata and the key set, wait longer than the longest metadata cache plus the longest token lifetime, start signing with the new key, then withdraw the old one only after a second window and after confirming nobody still fetches the old key id.
The reason for that specificity is a failure I have seen: A key was swapped inside one change window. 38 relying parties that refreshed metadata daily failed signature validation for up to 26 hours, and the service desk took 1,150 calls before the change was rolled back.
Rotation timeline that held.
| Phase | Duration | Keys published |
|---|---|---|
| publish the new key | 48 hours | old and new |
| sign with the new key | 30 days | old and new |
| withdraw the old key | after 30 days | new only |
I would not consider it settled without evidence: confirm both keys resolve from published metadata and that a token signed by each validates at a sample relying party before the old key is withdrawn.
Overlap is the whole mechanism.
Curated: · Written: · Reviewed:
QA-16Your relying party fetches the key set at runtime. What are the failure modes?(show answer)
My answer to pinned OIDC discovery and JWKS begins with the artifact that is actually being trusted: token, assertion, group, or session.
Runtime discovery puts the provider's metadata endpoint inside your trust boundary and inside your availability path. If it can be redirected you accept attacker keys, and if it is briefly unreachable without a cache you stop authenticating anyone.
Concretely, fetch discovery and the key set over TLS to a pinned host, verify that the discovery document's issuer equals the configured issuer, cache keys with a bounded refresh, keep a last-known-good copy for the outage case, and refuse metadata that arrives through a redirect to another host.
The reason for that specificity is a failure I have seen: A misconfigured proxy served a discovery document from a lookalike host for 35 minutes. The relying party accepted tokens signed by an unrelated key and created 17 sessions before the certificate mismatch was investigated.
Discovery behaviour under three conditions.
| Condition | Without pinning | With pinning and cache |
|---|---|---|
| metadata host redirected | 17 bad sessions | request refused |
| key endpoint down 35 minutes | sign-ins fail | served from cache |
| a new key id appears | fetched on demand | at most 1 fetch per 5 minutes |
I would not consider it settled without evidence: point the relying party at a metadata document whose issuer differs from configuration and require startup or validation to fail.
Discovery is a trust decision made at runtime; treat it as one.
Curated: · Written: · Reviewed:
QA-17Why is reading the algorithm header from the token itself a vulnerability?(show answer)
I would treat JWT algorithm confusion as a claim about effective access that has to survive a disablement.
The token header is attacker-controlled input, so a library that selects its verification algorithm from it lets the attacker choose the verification. Two classic outcomes follow: the "none" algorithm, which verifies nothing, and an asymmetric algorithm downgraded to a symmetric one, where the published public key becomes the shared secret.
Concretely, configure the expected algorithm and key type in the relying party, pass both into the verifier explicitly, refuse tokens whose header disagrees, and resolve the key by key id from the issuer's published set rather than from anything the token carries inline.
The reason for that specificity is a failure I have seen: An internal gateway used a permissive verifier default. A forged symmetric token signed with the published public key granted an administrative scope, and it was accepted on 9 of 11 services before the library was pinned.
Forged tokens against the service fleet.
| Forgery | Services accepting before | After |
|---|---|---|
| none algorithm | 4 of 11 | 0 |
| symmetric signature over the public key | 9 of 11 | 0 |
| inline key in the header | 2 of 11 | 0 |
I would not consider it settled without evidence: submit tokens using the none algorithm, a symmetric signature over the public key, and an inline key header, and require all three to be refused.
The verifier chooses the algorithm; the token does not.
Curated: · Written: · Reviewed:
QA-18A user clicks sign out. What has to happen, and in what order?(show answer)
The useful question for local logout before federated logout is what still holds at the relying parties nobody has looked at this week.
The application's own session is the one the user is standing in, so it is destroyed first and unconditionally. Federated logout is a best-effort request to another party that can fail, redirect, or be blocked, and making it the first step leaves a live local session whenever it does.
Concretely, invalidate the server-side session record and clear the cookie before any redirect, revoke the refresh token tied to that session, then attempt the provider's end-session endpoint or back-channel logout, and record the outcome per relying party rather than assuming propagation.
The reason for that specificity is a failure I have seen: A sign-out button redirected to the provider first. A third-party cookie block ended the flow silently at the provider, and 1 in 6 sign-outs on shared kiosks left the local session alive for the remaining 45 minutes of its idle timeout.
Sign-out outcomes on kiosk browsers.
| Order | Local session ended | Kiosk sessions left live |
|---|---|---|
| provider first | sometimes | 1 in 6 |
| local first, then provider | always | 0 |
| local only | always | 0, provider session remains |
I would not consider it settled without evidence: sign out with the provider unreachable and confirm the application session is already gone when the browser returns.
End what you own before you ask anyone else.
Curated: · Written: · Reviewed:
QA-19How do you make sure a step-up challenge applies to the request that needed it?(show answer)
I would settle step-up authentication binding by attempting the action after the identity event that was supposed to stop it.
A step-up that only sets a flag on the session upgrades everything the session can do until it expires. The stronger authentication has to be bound to the specific operation, with its own freshness, so it authorises that transfer and not the next thirty.
Concretely, issue a short-lived single-use authorisation for the named operation, carry a reference to it on the request, check the authentication method and time at the point of decision, and record which challenge authorised which action.
The reason for that specificity is a failure I have seen: A payments console raised a step-up once per session. After a single challenge, a stolen session moved 26 payments in 12 minutes, and every one of them looked correctly authorised in the log.
Actions authorised per challenge.
| Binding | Challenges | Sensitive actions allowed |
|---|---|---|
| session flag | 1 | 26 |
| operation-scoped, 90 seconds | 1 | 1 |
| operation-scoped, reused | 1 | 0 after the first use |
I would not consider it settled without evidence: complete a step-up, then attempt a second sensitive action on the same session and require a fresh challenge.
Bind the challenge to the act, not to the session.
Curated: · Written: · Reviewed:
QA-20When would you issue pairwise subject identifiers instead of a public subject?(show answer)
The judgement in pairwise subject identifiers is which identifier is durable and which is only a convenient claim.
A public subject is the same string at every relying party, so any two applications can correlate the same person by comparing tokens. Pairwise identifiers give each relying party its own opaque subject, which keeps identity durable inside an application while removing the join key between applications.
Concretely, set pairwise as the subject type for external and multi-tenant relying parties, derive it from a sector identifier so several hosts of one product share a subject, and keep a provider-side mapping so support and audit can still resolve the person.
The reason for that specificity is a failure I have seen: A public subject shared across a consumer directory let two unrelated partners join their datasets. 74,000 accounts were correlated across the two products before the subject type was changed, and the join could not be undone.
Subject values for one person.
| Relying party | Public subject | Pairwise subject |
|---|---|---|
| partner A | identical string | distinct, 43 characters |
| partner B | identical string | distinct, 43 characters |
| correlation possible | yes | no |
I would not consider it settled without evidence: authenticate one test user at two relying parties and confirm the subject values differ while both resolve to the same internal account.
Durable inside an application, meaningless between applications.
Curated: · Written: · Reviewed:
QA-21Your service provider accepts identity-provider-initiated SAML. What do you have to add?(show answer)
Where candidates lose the interview on unsolicited IdP-initiated SAML is reading a green IdP dashboard as revocation.
An unsolicited response has no request to match, so the strongest anti-injection check in the protocol is unavailable. Accepting it means accepting a sign-in the service provider never asked for, which is also how a sign-in forgery drops a victim into the attacker's account.
Concretely, prefer service-provider-initiated flows, and where the business needs the other direction, enable it per connection only, carry a one-time relay value issued by the provider, keep the condition window tight, and require a confirmation step before an existing session is replaced.
The reason for that specificity is a failure I have seen: A finance portal accepted unsolicited responses on every connection. A crafted response signed the victim into an attacker-controlled tenant, and 3 invoices were uploaded to the wrong account over 2 days before the tenant mismatch was reported.
Connections allowed to send unsolicited responses.
| Connection | Unsolicited allowed | Confirmation required |
|---|---|---|
| corporate provider | yes | yes |
| 12 partner providers | no | not applicable |
| legacy portal | yes, until migration | yes |
I would not consider it settled without evidence: send an unsolicited response to a connection that has not enabled it and require the service provider to refuse.
A sign-in you did not request deserves a question before a session.
Curated: · Written: · Reviewed:
QA-22A user changes their surname and their email address. What must not change in your system?(show answer)
I would answer email change versus issuer-subject stability by separating authentication, authorization, and lifecycle into three control planes.
The internal account and every grant attached to it survive an attribute change, because the person did not change. Only the mutable attributes update, and the application's history stays attached to the subject the provider still asserts.
Concretely, treat email as a synchronised attribute with a verification state, update it in place on the account keyed by issuer and subject, keep an attribute history, and never create a second account because a new address arrived in an assertion.
The reason for that specificity is a failure I have seen: A rename created duplicate accounts at 4 applications. The user lost access to 61 saved reports and worked in the empty shell account for 9 days, while the original account kept its group memberships and stayed reachable through a live session.
What a rename touched.
| Item | Mishandled | Handled correctly |
|---|---|---|
| internal account id | 2 new ids created | 1 unchanged id |
| email attribute | stale on the old row | updated in place |
| duplicate accounts | 4 | 0 |
| saved reports retained | 0 | 61 |
I would not consider it settled without evidence: rename a test user at the provider and confirm the application resolves the same account id with the same permission set.
The person is the subject; the address is a label.
Curated: · Written: · Reviewed:
QA-23Provisioning set the user to inactive in the target application. Are their API clients dead?(show answer)
The engineering content of SCIM deactivation versus refresh tokens is the revocation path and its measured latency, not the protocol name.
Provisioning changes an attribute in the target's user store; it does not enumerate and destroy the grants that application issued. Refresh tokens, personal access tokens, and installed integrations each survive until something revokes them by name.
Concretely, make deactivation trigger a token revocation sweep for that subject at every authorization server and application, verify with an introspection call rather than by reading the user record, and treat a target with no revocation interface as an exception carrying a compensating control.
The reason for that specificity is a failure I have seen: A contractor was deactivated on a Friday. A refresh token inside a build agent kept exchanging for access tokens and pulled 3 repositories every 30 minutes for 16 days, while the user record read inactive the entire time.
State after deactivation, before the sweep.
| Object | State | Still usable |
|---|---|---|
| user record | inactive | no |
| refresh token in a build agent | 16 days remaining | yes |
| 2 personal access tokens | valid | yes |
I would not consider it settled without evidence: deactivate a test user, then introspect their outstanding refresh tokens and require every one to return inactive.
Deactivation writes a field; revocation ends a grant.
Curated: · Written: · Reviewed:
QA-24Your application grants an eight-hour session after SSO. What is wrong with that number?(show answer)
Before calling relying-party session lifetime after federation done I would write down the grant that the source-of-truth change does not touch.
A relying party's session lifetime is a promise it makes on its own, and after federation it usually outlives the provider's session and any change made there since. The lifetime should be chosen against how quickly this application must react to a revocation.
Concretely, set the session shorter than the revocation objective for that application's data, re-check the provider session or introspect the token on a fixed interval for sensitive applications, and force reauthentication for actions above a stated risk.
The reason for that specificity is a failure I have seen: A finance application issued eight-hour sessions with no re-check. An account suspended at 10:40 kept approving purchase orders until 18:05, and 6 approvals inside that window had to be reversed by hand.
Session policy against the revocation objective.
| Application | Session length | Worst-case stale access |
|---|---|---|
| finance approvals, old | 8 hours | 7 hours 25 minutes |
| finance approvals, revised | 15 minutes with re-check | 15 minutes |
| internal wiki | 12 hours | accepted in writing |
I would not consider it settled without evidence: suspend a test identity mid-session and measure the minutes until the application refuses the next sensitive action.
A session length is a revocation delay you have chosen.
Curated: · Written: · Reviewed:
QA-25A token arrives with a key id your relying party has never seen. What should happen?(show answer)
The first thing I would establish about unknown key id bounded refresh is which live path still works after the directory row looks right.
An unknown key identifier is either a legitimate rotation you have not observed yet or a probe. The safe behaviour is a rate-limited refresh of the key set followed by ordinary validation, never a decision to trust the token because the key was unfamiliar.
Concretely, refresh the key set at most once per fixed interval per issuer, keep a negative cache for identifiers that were not found, continue serving from last-known-good keys during the refresh, and alert when unknown identifiers arrive faster than rotations plausibly occur.
The reason for that specificity is a failure I have seen: A relying party refreshed keys on every unknown identifier. A loop of forged tokens produced 240 key-set requests per second, the provider rate-limited the caller, and legitimate sign-ins failed for 22 minutes.
Behaviour under a forged key-id burst.
| Policy | Outbound key requests | Valid sign-ins failing |
|---|---|---|
| refresh per unknown key id | 240 per second | 22 minutes of failures |
| at most 1 refresh per 5 minutes | 12 per hour | 0 |
| negative cache for 10 minutes | 4 per hour | 0 |
I would not consider it settled without evidence: send a burst of tokens carrying random key identifiers and confirm outbound key fetches stay within the configured rate while valid sign-ins continue.
Refresh on a schedule you control, not on an attacker's cue.
Curated: · Written: · Reviewed:
QA-26How do you make federation debuggable without storing tokens in your logs?(show answer)
I would start federation audit without raw tokens from the session, token, and grant that outlive the login, not from the login itself.
A raw assertion or token written to a log is a credential in a log, readable by anyone with log access and usable until it expires. The audit record needs the facts that let you reconstruct a decision — issuer, subject, audience, key id, timestamps, outcome, and failing check — not the bearer artifact itself.
Concretely, log a stable hash of the token identifier alongside issuer, audience, key id, authentication time, authentication method, and the validation outcome with its failing check, and redact signature and payload at the logging boundary so no code path can opt back in.
The reason for that specificity is a failure I have seen: A debug flag left assertion bodies in an application log for 6 weeks. 11,000 assertions were readable by every engineer with log access, and 90 of them were still inside their validity window when the flag was removed.
What the audit record keeps.
| Field | Stored | Purpose |
|---|---|---|
| hashed assertion identifier | yes | correlate 1 event across systems |
| issuer, audience, key id | yes | reconstruct validation |
| raw assertion body | no | credential material |
| failing check name | yes | 8 distinct reasons |
I would not consider it settled without evidence: search the log store for the token prefix pattern and require zero matches after the redaction change.
Log the decision, not the credential.
Curated: · Written: · Reviewed:
QA-27A mobile app sends the provider's access token to your backend as proof of sign-in. What is wrong?(show answer)
This is an area where a successful federation and a current authorization decision are different events.
An access token is opaque to the client and addressed to a resource, so a backend treating it as a sign-in accepts an artifact any other application could have obtained for the same person. With no audience the backend can check, nothing ties the token to your application.
Concretely, require an ID token whose audience is your client, or redeem the code at the backend yourself, and where a token must be accepted, introspect it and verify the client id and audience it was issued to before creating a session.
The reason for that specificity is a failure I have seen: A backend accepted provider access tokens for sign-in. A fitness application holding tokens for the same users could open sessions for 5,400 of them simply by pointing those tokens at the endpoint, and the report came from a researcher rather than from monitoring.
What the backend could verify.
| Artifact | Audience checkable | Sessions a third app could open |
|---|---|---|
| forwarded access token | no | 5,400 |
| introspected access token | yes | 0 |
| ID token with an audience | yes | 0 |
I would not consider it settled without evidence: present an access token issued to a different client and require the sign-in endpoint to refuse it.
Sign-in needs an artifact addressed to you.
Curated: · Written: · Reviewed:
QA-28You suspect the identity provider signing key has been exposed. What happens in the first hours?(show answer)
My answer to IdP signing-key compromise containment begins with the artifact that is actually being trusted: token, assertion, group, or session.
A leaked signing key means any assertion can be forged for any person at any relying party, so containment is rotation with distrust of the old key rather than rotation with overlap. Everything already signed is suspect, which makes session and token revocation part of the same operation.
Concretely, generate the replacement inside the hardware module, publish it, drive relying parties to refresh metadata, withdraw the old key immediately instead of after the usual overlap, revoke all sessions and refresh tokens, and reconcile accepted authentications against the provider's own issuance records.
The reason for that specificity is a failure I have seen: An exported key sat in a configuration repository. Forged assertions produced 12 administrative sessions across 4 relying parties over a weekend, and reconstruction took 3 weeks because relying-party logs held no key identifier.
Containment sequence and timing.
| Step | Target | Actual |
|---|---|---|
| new key published | 30 minutes | 41 minutes |
| old key distrusted everywhere | 2 hours | 5 hours |
| sessions and tokens revoked | 2 hours | 2 hours |
| issuance reconciliation complete | 5 days | 3 weeks |
I would not consider it settled without evidence: confirm the old key identifier is refused at every relying party and that the provider's issuance record matches each accepted assertion.
Compromise removes the overlap you would normally allow.
Curated: · Written: · Reviewed:
QA-29A user's second factor is a code sent to the phone they are logging in from. Is that two factors?(show answer)
I would treat independent MFA factor types as a claim about effective access that has to survive a disablement.
Two factors are independent only if compromising one does not deliver the other. A code delivered to the device already holding the session, or a prompt approved in the same browser, collapses into one compromise path however the policy labels it.
Concretely, classify factors by what an attacker must physically hold, require the second factor to be a separate possession or a hardware-backed key on a separate channel, and refuse combinations where a single device or a single consumer account controls both.
The reason for that specificity is a failure I have seen: A number was ported away in 40 minutes. The attacker already held the password from a public credential dump, received the delivered code on the ported number, and controlled the account before the handset showed no service.
Factor pairs under a single-device compromise.
| Pair | Survives device theft | Survives a number port |
|---|---|---|
| password with a delivered code | no | no, 40 minutes |
| password with a code app on the same phone | no | yes |
| password with a hardware key | yes | yes |
I would not consider it settled without evidence: model each factor pair against a single-device compromise and require at least one factor to survive it.
Independence is a property of the attack, not of the policy label.
Curated: · Written: · Reviewed:
QA-30Why can a phishing site not relay a WebAuthn assertion?(show answer)
The useful question for WebAuthn origin and RP-ID binding is what still holds at the relying parties nobody has looked at this week.
The browser puts the origin it is actually talking to into the client data, and the authenticator produces an assertion only for credentials scoped to the matching relying-party identifier. The signature therefore covers where the ceremony happened, which is the fact a relayed credential cannot fabricate.
Concretely, verify the origin string against your allowed origins, verify the relying-party identifier hash inside the authenticator data, check the challenge against a server-side value, and set the relying-party identifier at the registrable domain all your legitimate origins share rather than at one hostname.
The reason for that specificity is a failure I have seen: A team set the relying-party identifier to a single hostname and later moved sign-in to a new subdomain. 2,300 registered keys stopped working overnight, the fallback they enabled was a delivered code, and a relay page harvested codes from 34 of 210 targets in 6 days.
Same campaign, three factor types.
| Factor | Targets | Credentials relayed |
|---|---|---|
| delivered code | 210 | 34 |
| approval prompt | 210 | 12 |
| security key | 210 | 0 |
I would not consider it settled without evidence: run a relaying proxy against a test account and confirm no assertion validates while a delivered code would have.
The origin sits inside the signature, and that is the whole defence.
Curated: · Written: · Reviewed:
QA-31Your users have a code app. Why is a modern phishing page still a problem?(show answer)
I would settle TOTP real-time phishing relay by attempting the action after the identity event that was supposed to stop it.
A time-based code is the output of a shared secret that the user can read aloud, so a page that asks for it can forward it inside the validity window. Nothing in the code binds it to the site the user believes they are on.
Concretely, treat code apps as a mitigation for password reuse rather than for phishing, move privileged and high-value populations to origin-bound authenticators, and where codes remain, add device binding and a hard cap on concurrent verification attempts.
The reason for that specificity is a failure I have seen: A relay proxy captured code and session cookie together for 3 weeks. 88 accounts were taken over, and the median gap between the user typing a code and the attacker's first action was 9 seconds against a code lifetime of 30.
Relay window measured.
| Step | Seconds elapsed | Attacker holds |
|---|---|---|
| user enters the code | 0 | the code |
| code replayed to the real site | 9 | a session |
| cookie exported to another host | 11 | full access |
| code would have expired | 30 | already irrelevant |
I would not consider it settled without evidence: run a controlled relay against a test account and measure whether the code and the resulting session are usable from another network.
A code the user can read out is a code the page can forward.
Curated: · Written: · Reviewed:
QA-32You cannot roll out security keys everywhere at once. Who goes first?(show answer)
The judgement in phishing-resistant MFA for administrators is which identifier is durable and which is only a convenient claim.
Sequence by blast radius rather than by headcount. The accounts that can change identity itself — directory administrators, federation administrators, privileged-access operators, and break-glass holders — are the ones whose compromise makes every other control negotiable.
Concretely, enumerate the roles that can alter authentication policy, group membership, or federation trust, require hardware-backed origin-bound authenticators for those roles with no code fallback, and enforce it through a conditional policy that names the role rather than the person.
The reason for that specificity is a failure I have seen: A rollout ordered by department reached the directory administrators in month 5. In month 3 an administrator approved a fatigue prompt, and the attacker added a federation trust that survived the eventual key rollout by a further 2 weeks.
Rollout order by what the role can change.
| Population | Can alter federation trust | Wave and size |
|---|---|---|
| directory and provider administrators | yes | wave 1, 42 people |
| privileged-access operators | partly | wave 2, 130 people |
| general staff | no | wave 3, 6,400 people |
I would not consider it settled without evidence: attempt an administrative sign-in with a delivered code and require the policy to refuse before the rollout is called complete.
Start where a compromise rewrites the rules.
Curated: · Written: · Reviewed:
QA-33Does the fingerprint used on a passkey sign-in travel to your server?(show answer)
Where candidates lose the interview on local biometric as authenticator unlock is reading a green IdP dashboard as revocation.
The biometric unlocks the authenticator locally and never leaves the device, so it is not a factor your server evaluates. What the server learns is that a user-verification bit was set, which is the authenticator's assertion that some local check succeeded.
Concretely, require user verification in policy, check the verification flag in the authenticator data on every assertion, keep attestation only where the device model genuinely matters, and never design a flow that expects biometric material to reach the relying party.
The reason for that specificity is a failure I have seen: A procurement questionnaire was answered as though fingerprints were held server-side. The resulting data-protection review blocked a passkey rollout for 4 months, and 3,100 users stayed on delivered codes while the misunderstanding was unpicked.
What crosses the wire on a passkey sign-in.
| Item | Sent to the server | Size |
|---|---|---|
| signature over the challenge | yes | 71 bytes |
| user-verification flag | yes | 1 bit |
| credential identifier | yes | 32 bytes |
| fingerprint template | no | 0 bytes |
I would not consider it settled without evidence: inspect an assertion and confirm it carries flags and a signature with no biometric data anywhere in the payload.
The device checks the person; the server checks the device.
Curated: · Written: · Reviewed:
QA-34What do you verify when a user registers a new passkey?(show answer)
I would answer WebAuthn registration ceremony checks by separating authentication, authorization, and lifecycle into three control planes.
Registration is where a credential becomes trusted, so the checks made there decide what every later assertion means. The challenge must be one you issued, the origin and relying-party identifier must match, the credential identifier must be new, and the flags must show the verification level your policy claims.
Concretely, issue a single-use challenge with a short lifetime, confirm the ceremony type is a creation, compare origin and relying-party identifier hash, refuse a credential identifier already bound to any account, store the public key with its algorithm and counter, and require an existing strong factor to authorise the registration.
The reason for that specificity is a failure I have seen: A registration endpoint accepted any well-formed attestation without checking the challenge. A replayed registration bound the attacker's authenticator to 3 accounts, and those credentials passed every later assertion check because registration had already blessed them.
Registration checks and what they stop.
| Check | Stops | Caught in the test |
|---|---|---|
| server-issued unused challenge | replayed registration | 3 |
| origin and identifier match | cross-site enrolment | 2 |
| credential identifier unseen | credential grafting | 1 |
| verification flag set | silent enrolment | 5 |
I would not consider it settled without evidence: replay a captured registration payload and require the server to refuse on the challenge check.
Every later assertion inherits whatever registration accepted.
Curated: · Written: · Reviewed:
QA-35The signature counter in an assertion went backwards. What does that tell you?(show answer)
The engineering content of WebAuthn sign-counter anomalies is the revocation path and its measured latency, not the protocol name.
The counter is a cloning signal rather than an authentication decision. A value at or below the stored one is consistent with a duplicated authenticator, but many authenticators, including most synced passkeys, report zero permanently, so the check has to be applied per credential.
Concretely, record whether a credential has ever reported a non-zero counter, enforce monotonic increase only for those, treat a regression as a security event that suspends the credential and forces re-registration, and leave always-zero credentials to other signals.
The reason for that specificity is a failure I have seen: A regression check applied to every credential locked out 640 users on synced passkeys in one morning because their counters were permanently zero. The check was then switched off entirely, including for the 1,200 hardware keys where it had been working.
Counter behaviour by credential class.
| Credential class | Fleet size | Counter behaviour | Regression enforced |
|---|---|---|---|
| hardware key | 1,200 | increments | yes |
| synced passkey | 640 | always zero | no |
| platform key with a counter | 380 | increments | yes |
I would not consider it settled without evidence: replay an assertion with a stale counter for a counter-reporting credential and require that credential to be suspended.
Enforce the counter where the counter means something.
Curated: · Written: · Reviewed:
QA-36Should you accept synced passkeys for administrative access?(show answer)
Before calling synced passkeys versus device-bound keys done I would write down the grant that the source-of-truth change does not touch.
A synced passkey is phishing-resistant, but its private key exists wherever the vendor account is signed in, so the credential's security becomes the security of that consumer account and its recovery path. A device-bound key keeps the private key inside one piece of hardware you can inventory.
Concretely, allow synced passkeys broadly to win the phishing resistance, and require device-bound attested keys for administrative roles, checking the authenticator model against an allowed list and recording the backup-eligible and backup-state flags at registration.
The reason for that specificity is a failure I have seen: An administrator's personal cloud account was recovered by an attacker through a support call. The synced passkey followed that account onto a new device, and the attacker held phishing-resistant access for 5 days without ever touching a corporate machine.
Administrator credential inventory.
| Credential type | Count | Allowed for the role |
|---|---|---|
| device-bound and attested | 38 | yes |
| synced passkey | 11 | no, migrate |
| delivered-code fallback | 2 | no, remove |
I would not consider it settled without evidence: read the backup-eligible flag on every registered administrator credential and require the count of syncable ones to be zero.
Phishing-resistant is not the same as bound to hardware you control.
Curated: · Written: · Reviewed:
QA-37An attacker holding a stolen session enrols their own passkey. What stops that?(show answer)
The first thing I would establish about adding an authenticator after session theft is which live path still works after the directory row looks right.
Enrolling an authenticator changes who can be you, so it cannot be authorised by a session alone. The session may already belong to the attacker, and that is precisely the case the control has to survive.
Concretely, require a fresh phishing-resistant challenge with an existing credential at the moment of enrolment, notify every previously registered channel, hold the new credential back from privileged use for a cooling period, and let the owner revoke it straight from the notification.
The reason for that specificity is a failure I have seen: A stolen session added a passkey and removed the original key 40 seconds later. The owner was locked out for 3 days, and the only notification went to an address the attacker had already changed in the same session.
Enrolment controls against the theft window.
| Control | Present before | Attacker succeeded |
|---|---|---|
| fresh credential challenge | no | yes, within 40 seconds |
| notification to prior channels | partial | yes |
| 24-hour hold on privileged use | no | yes |
I would not consider it settled without evidence: attempt authenticator enrolment from a session that has not re-authenticated and require a challenge before it proceeds.
Changing who can be you is not a session-level action.
Curated: · Written: · Reviewed:
QA-38How do you store account recovery codes?(show answer)
I would start hashed one-time recovery codes from the session, token, and grant that outlive the login, not from the login itself.
Recovery codes bypass the second factor, so they get what a password gets and more: high entropy, single use, and storage only as a slow hash. A recoverable copy anywhere is a copy of the second factor.
Concretely, generate at least 128 bits of entropy per code, hash each with a memory-hard function, display them once, mark each used at first successful presentation, show the remaining count in the account view, and invalidate the whole set when a new set is issued.
The reason for that specificity is a failure I have seen: Codes were stored reversibly so support could read them back. A support-tool export covering 15,000 accounts contained every unused code, and 2 accounts were entered with codes generated 17 months earlier.
Recovery code handling.
| Property | Before | After |
|---|---|---|
| storage | reversible | memory-hard hash |
| entropy per code | 32 bits | 128 bits |
| single use enforced | no | yes, 10 codes per set |
I would not consider it settled without evidence: read the recovery-code column directly in the database and confirm no usable code can be recovered from it.
A code you can read back to a user is a second factor you have copied.
Curated: · Written: · Reviewed:
QA-39A user whose only factor is a security key loses it on a Friday. What is the recovery path?(show answer)
This is an area where a successful federation and a current authorization decision are different events.
Recovery is the weakest authentication path an account has, so it sets the account's real strength. A path that falls back to a mailed code or a persuasive phone call makes the key decorative for anyone willing to call.
Concretely, require a second registered authenticator at enrolment so loss is an ordinary event, run recovery through proofing of comparable strength such as a supervised check against a recorded document or a manager attestation plus a second registered device, and never issue a long-lived bypass.
The reason for that specificity is a failure I have seen: A single-key policy with service-desk recovery meant 1 in 9 support calls was a lost key. A caller impersonating an employee passed four knowledge questions in 4 minutes and held a working bypass for 48 hours.
Recovery paths by strength.
| Path | Bypass duration | Impersonation risk |
|---|---|---|
| second registered key | none | very low |
| supervised video proofing | 30 minutes | low |
| knowledge questions | 48 hours | 1 in 9 calls suspicious |
I would not consider it settled without evidence: run a recovery attempt using only publicly available facts about the user and require it to fail.
Recovery is the account's true floor.
Curated: · Written: · Reviewed:
QA-40Your risk engine rates this sign-in low risk. Can it skip the second factor?(show answer)
My answer to risk-adaptive MFA that cannot waive the baseline begins with the artifact that is actually being trusted: token, assertion, group, or session.
A risk score is a prior built from what the platform can observe, and everything it observes — network, device, geography, timing — is available to an attacker who has taken the endpoint. Adaptive policy may add friction; it must not remove the baseline.
Concretely, express policy as a floor plus escalations so a low score only chooses among acceptable strong factors while a high score adds a challenge or blocks, and require a named exception with an expiry for any request to drop below the floor.
The reason for that specificity is a failure I have seen: A trusted-network rule waived the second factor inside the office range. A compromised printer on that range was used with dumped passwords to reach 7 internal applications over 11 days, and every sign-in looked ordinary in the reports.
Policy shape before and after.
| Signal | Old outcome | New outcome |
|---|---|---|
| office address range | factor waived | key required |
| new device, 2 countries in 1 hour | factor required | key plus challenge |
| known device, low score | factor waived | key required |
I would not consider it settled without evidence: authenticate from the location the engine rates lowest risk and confirm a strong factor is still required.
Risk can add friction; it cannot remove the floor.
Curated: · Written: · Reviewed:
QA-41How do you decide which operations get a transaction-level confirmation?(show answer)
I would treat step-up scoped to one action as a claim about effective access that has to survive a disablement.
Scope a challenge to an action when the action is irreversible, moves value, or changes who can authenticate. The challenge must also show the parameters being authorised, because a user who approves an unnamed prompt has approved whatever the attacker requested.
Concretely, keep a register of operations with a required assurance level, render the concrete parameters — payee, amount, target account, role being granted — inside the challenge, and bind the resulting authorisation to a hash of those parameters so a changed request fails.
The reason for that specificity is a failure I have seen: A challenge displayed nothing but an approve button. A transfer was confirmed by a user who believed they were approving a sign-in, and 2 of the 4 payments in that hour were fraudulent before the prompt text was rewritten.
Operations register.
| Operation | Assurance required | Parameters shown |
|---|---|---|
| add a payee | key plus confirmation | payee and last 3 digits |
| grant an administrative role | key plus approval | role, scope, 90-day expiry |
| read a report | session only | none |
I would not consider it settled without evidence: alter one parameter after the challenge is issued and require the authorisation to be rejected.
A prompt that names nothing authorises anything.
Curated: · Written: · Reviewed:
QA-42Multi-factor authentication succeeded and the account was still taken over. How?(show answer)
The useful question for stolen session after successful MFA is what still holds at the relying parties nobody has looked at this week.
Authentication produces a session artifact, and that artifact is a bearer credential from the moment it exists. Theft through malware, a relaying proxy, or an exposed log bypasses every factor by starting after them.
Concretely, bind the session to the client with a device-bound proof, re-evaluate on network and device change, keep idle timeouts short for privileged applications, and treat replay from a new client fingerprint as an event that ends the session rather than as a warning.
The reason for that specificity is a failure I have seen: An information stealer exported cookies from one laptop. 14 applications accepted the exported session from another country the same night, and the earliest signal was a mailbox rule created 6 hours later.
Exported session accepted where.
| Application class | Accepted from a new client | After binding |
|---|---|---|
| mail and chat | 4 | 0 |
| internal tools | 8 | 0 |
| finance | 2 | 0, challenge forced |
I would not consider it settled without evidence: replay a captured session cookie from a different client and require the application to end the session rather than accept it.
Everything after the challenge is a bearer token.
Curated: · Written: · Reviewed:
QA-43Why can you not store code-app seeds the way you store passwords?(show answer)
I would settle TOTP seed encryption versus password hashing by attempting the action after the identity event that was supposed to stop it.
Verifying a time-based code requires the server to recompute it, so the seed has to be recoverable and hashing is not available. That makes the seed store a key-management problem rather than a hashing problem.
Concretely, encrypt seeds under a key held in a hardware module or a managed key service, keep that key out of the database and out of its backups, perform verification inside a service that returns only a yes or a no, and rotate the data key with a re-encryption job rather than by re-enrolling users.
The reason for that specificity is a failure I have seen: Seeds sat in plaintext in a table replicated to a reporting warehouse. A read-only analytics credential exposed 26,000 of them, and re-enrolment took 5 weeks during which the second factor was known to whoever held the export.
Seed store design.
| Property | Password hash | Code-app seed |
|---|---|---|
| recoverable by the server | no | required |
| storage | memory-hard hash | envelope encryption |
| impact of a table export | 0 codes usable | 26,000 accounts |
I would not consider it settled without evidence: read the seed column with a database credential and confirm the value cannot be used to generate a valid code.
A secret you must recompute with is a key-management problem.
Curated: · Written: · Reviewed:
QA-44How wide should the code acceptance window be, and what else does it need?(show answer)
The judgement in TOTP clock-window and replay is which identifier is durable and which is only a convenient claim.
A wider window buys tolerance for clock drift and buys the attacker time, so the window and the replay defence are one decision. One step either side of the current interval is usually enough, and a code already used must never be accepted again inside its window.
Concretely, accept one step of skew, record the last accepted step per user and refuse anything at or below it, track per-user drift so a consistently early device is corrected rather than accommodated by widening, and rate-limit verification per account and per source.
The reason for that specificity is a failure I have seen: A window of five steps was configured to stop drift complaints. The acceptance span became 330 seconds, one relayed code was reused successfully 3 times, and every guess had 11 valid targets instead of 3.
Window width against exposure.
| Window | Acceptance span | Valid intervals per guess |
|---|---|---|
| one step either side | 90 seconds | 3 |
| five steps either side | 330 seconds | 11 |
| one step with a replay record | 90 seconds | 1 |
I would not consider it settled without evidence: present the same code twice inside the window and require the second attempt to fail.
The window is an attacker's budget as much as a user's.
Curated: · Written: · Reviewed:
QA-45Users are approving prompts they did not start. What do you change?(show answer)
Where candidates lose the interview on MFA push fatigue is reading a green IdP dashboard as revocation.
A prompt that can be approved with one tap and no context turns the second factor into a doorbell the attacker can ring. The fix removes the blind approval rather than asking users to be careful at two in the morning.
Concretely, replace plain approvals with number matching, show the application, location, and requesting device, cap prompts per account per interval, lock after repeated denials, and give a one-tap route to report a prompt the user did not start.
The reason for that specificity is a failure I have seen: An attacker holding a valid password sent 63 prompts between 01:40 and 02:20. The 41st was approved, and mailbox rules were changed 4 minutes later.
Prompt behaviour after the change.
| Control | Before | After |
|---|---|---|
| prompts allowed per hour | unlimited, 63 observed | 3 |
| number matching | no | yes |
| account locked after denials | never | after 2 |
I would not consider it settled without evidence: drive repeated prompts at a test account and require throttling and a report path before the tenth arrives.
Remove the blind tap.
Curated: · Written: · Reviewed:
QA-46How do you throttle second-factor attempts without handing an attacker a lockout tool?(show answer)
I would answer layered MFA throttling by separating authentication, authorization, and lifecycle into three control planes.
Throttling has to be layered because the attacker chooses the axis: one code against many accounts, many codes against one account, or one account from many addresses. A single per-account counter is also a denial-of-service weapon anyone can point at a colleague.
Concretely, apply per-account, per-source, and per-value limits with different windows, prefer growing delay over hard lockout on the per-account axis, exempt a re-authenticated known device, and alert on the spread pattern rather than on any single counter.
The reason for that specificity is a failure I have seen: A per-account lockout after 5 failures was used to lock 300 staff out during a release weekend by an attacker who simply submitted wrong codes, and the service desk queue reached 4 hours.
Limits by axis.
| Axis | Limit | Window |
|---|---|---|
| per account | delay doubling from 2 seconds | 15 minutes |
| per source address | 20 attempts | 10 minutes |
| per code value across accounts | 6 attempts | 1 hour |
I would not consider it settled without evidence: spray one code value across many accounts and confirm the per-value limit stops it while no account is locked.
Lockout is a control an attacker can also operate.
Curated: · Written: · Reviewed:
QA-47What does your service desk need before it resets someone's second factor?(show answer)
The engineering content of helpdesk MFA reset social engineering is the revocation path and its measured latency, not the protocol name.
A factor reset carries the same privilege as issuing a new second factor, so it needs proofing of comparable strength to enrolment. A caller who knows a start date and a manager's name has demonstrated access to a staff directory, not identity.
Concretely, require a verifiable channel such as a supervised check against a recorded document or a manager approval delivered through an authenticated application rather than a phone call, attach the proofing evidence to the ticket, and separate the person who approves from the person who performs the reset.
The reason for that specificity is a failure I have seen: A caller using facts taken from a public profile passed proofing in 6 minutes and had a new authenticator on a senior finance account. 3 payment changes were attempted within the hour, and the ticket recorded no evidence beyond the caller's answers.
Reset proofing before and after.
| Requirement | Before | After |
|---|---|---|
| knowledge questions | 4 | 0 |
| supervised identity check | no | yes |
| approver separate from performer | no | yes |
| median reset time | 6 minutes | 34 minutes |
I would not consider it settled without evidence: run an unannounced impersonation attempt against the reset process and require it to fail at the proofing step.
A reset is enrolment; prove it like one.
Curated: · Written: · Reviewed:
QA-48A user cannot use the standard authenticator. What do you offer them?(show answer)
Before calling accessible MFA without sharing credentials done I would write down the grant that the source-of-truth change does not touch.
An accessibility need calls for a different strong factor, not for lower assurance and not for an assistant holding the credential. Shared credentials destroy attribution, and attribution is what every later investigation depends on.
Concretely, keep a menu of approved strong options covering several key form factors, platform authenticators, and a hardware code display for users who cannot use a phone, and route requests through an assessment that records which option was issued rather than through an exception that waives the factor.
The reason for that specificity is a failure I have seen: An assistant held the second factor for 3 executives as an informal accommodation. During a payment dispute nothing could establish which of 4 people had approved a transfer, and the investigation ran 7 weeks without an answer.
Approved options and uptake.
| Option | Users | Assurance |
|---|---|---|
| plug-in security key | 2,410 | phishing-resistant |
| contactless key with a tactile marker | 96 | phishing-resistant |
| hardware code display | 38 | code-based, restricted roles |
I would not consider it settled without evidence: confirm every issued authenticator in the inventory maps to exactly one named person.
Accommodate the person; do not share the credential.
Curated: · Written: · Reviewed:
QA-49Can you list every authenticator registered to a departing employee and remove them?(show answer)
The first thing I would establish about authenticator inventory and revocation is which live path still works after the directory row looks right.
Authenticators accumulate across enrolments, device replacements, and self-service, and an identity is only as revoked as its longest-lived credential. Leaver handling has to enumerate credentials rather than disable an account and assume enumeration happened.
Concretely, keep an inventory keyed by account holding credential identifier, type, model or serial, enrolment date, and last use, surface it in the leaver runbook, revoke each entry explicitly, and reconcile the inventory against the provider's own list on a schedule.
The reason for that specificity is a failure I have seen: A leaver's account was disabled while a hardware key registered against a second, older account object stayed active. That object was re-enabled during a mailbox migration 5 months later, and the key worked on the first attempt.
Credentials found for one leaver.
| Credential | Enrolled | Last used | Revoked |
|---|---|---|---|
| hardware key on a legacy object | 26 months ago | 5 months ago | yes |
| platform passkey | 8 months ago | 1 day ago | yes |
| code-app seed | 3 years ago | 18 months ago | yes |
I would not consider it settled without evidence: list credentials for a test identity, revoke them, and confirm the provider's own credential list then returns zero entries.
Revocation is per credential, not per account row.
Curated: · Written: · Reviewed:
QA-50When is role-based access the right model rather than an attribute expression?(show answer)
I would start RBAC for stable job functions from the session, token, and grant that outlive the login, not from the login itself.
Role-based access fits where the population clusters into stable job functions and the decision does not depend on the individual object. Its real strength is reviewability: a person can read a role and predict what it permits, which attribute expressions rarely allow.
Concretely, define roles from job functions with an owner and a written purpose, drive membership from source-of-truth attributes where possible, cap how many permissions one role may carry, and watch role count against distinct job functions so per-person permissions cannot hide as roles.
The reason for that specificity is a failure I have seen: A 900-role estate contained 610 roles with a single member. Reviewers approved 92 percent of entitlements without opening them, and a quarterly campaign consumed 3 weeks of manager time while removing 11 grants.
Role estate before and after consolidation.
| Measure | Before | After |
|---|---|---|
| roles defined | 900 | 214 |
| roles with one member | 610 | 7 |
| entitlements removed in review | 11 | 480 |
I would not consider it settled without evidence: count roles holding one member and require a reviewer to state each remaining role's purpose in one sentence.
A role nobody can describe is a permission with a name.
Curated: · Written: · Reviewed:
QA-51Your attribute rule needs a clearance value and it is absent. What happens?(show answer)
This is an area where a successful federation and a current authorization decision are different events.
Attribute-based access control (ABAC) turns policy into an expression over data, so the policy inherits that data's availability and freshness. An absent attribute has to deny, because treating missing as permissive means an upstream outage silently grants.
Concretely, evaluate against a schema that marks required attributes, deny with a distinct reason when one is missing, separate deny-for-missing from deny-by-rule in the decision log, cache attributes with an explicit staleness bound, and alert when the missing rate rises.
The reason for that specificity is a failure I have seen: A staff feed failed and 4,300 people lost their clearance attribute. A permissive default branch granted access to a restricted dataset for 90 minutes, and the log recorded those grants as ordinary allows because the reason was never separated.
Decision counts during the feed outage.
| Outcome | Old behaviour | New behaviour |
|---|---|---|
| allowed with the attribute missing | 4,300 | 0 |
| denied, missing attribute named | 0 | 4,300 |
| minutes of exposure | 90 | 0 |
I would not consider it settled without evidence: remove a required attribute in a test tenant and require the decision to be a denial naming the missing attribute.
Missing is not permissive.
Curated: · Written: · Reviewed:
QA-52Access to a resource depends on a tag. Who is allowed to change the tag?(show answer)
My answer to privilege-bearing tag mutation begins with the artifact that is actually being trusted: token, assertion, group, or session.
If a tag drives a decision, editing that tag is exercising the privilege, and permission to edit it belongs in the same review as the permission it controls. Otherwise the cheapest route to the data is to relabel it.
Concretely, restrict write access on decision-bearing tag keys to a named role, require approval for changes on sensitive resources, log old and new values with the actor, and run a detection comparing tag changes against subsequent access by the same principal.
The reason for that specificity is a failure I have seen: An engineer changed a classification tag from restricted to internal to unblock a job. 12 datasets became readable by 260 people for 6 days, and the change appeared in the trail as a routine metadata edit.
Tag keys and who may write them.
| Tag key | Drives access | Writers |
|---|---|---|
| data classification | yes | 6 named approvers |
| cost centre | no | any resource owner |
| environment | yes, for administrative scope | 11 platform engineers |
I would not consider it settled without evidence: attempt to edit a decision-bearing tag as a standard engineer and require the change to be refused.
The label is part of the lock.
Curated: · Written: · Reviewed:
QA-53One policy allows and another denies the same request. What should your engine do?(show answer)
I would treat explicit deny combining as a claim about effective access that has to survive a disablement.
The combining algorithm is itself a policy decision, and deny-overrides is the only one that lets a control owner state a prohibition another team cannot accidentally undo. It has to be declared, because engines differ and an unstated default becomes an argument during an incident.
Concretely, declare deny-overrides across the estate, keep denials narrow and owned so they do not become blanket blocks, put the deciding policy id and effect into every decision record, and provide a query that explains a denial so the answer is not found by trial.
The reason for that specificity is a failure I have seen: Two teams wrote overlapping policies on one storage bucket. A first-applicable engine let ordering decide, a deployment reordered the set, and 3 service accounts gained write access to an audit-log bucket for 8 days.
Same request under three algorithms.
| Algorithm | Result | Policies evaluated |
|---|---|---|
| deny-overrides | deny | 2 |
| permit-overrides | permit | 2 |
| first-applicable | depends on order | 2 |
I would not consider it settled without evidence: evaluate a principal covered by both an allow and a deny and require a denial that names both policy ids.
State the algorithm; do not inherit it.
Curated: · Written: · Reviewed:
QA-54Your endpoint checks that the caller holds a read scope. Is that enough?(show answer)
The useful question for object-level authorization is what still holds at the relying parties nobody has looked at this week.
A scope says what kind of thing a caller may do; it does not say which instance they may do it to. Object-level authorization compares the authenticated subject against the specific record, and it belongs on every path because the identifier arrives from the client.
Concretely, resolve the object first, evaluate ownership or relationship against the subject in shared middleware, refuse to treat unguessable identifiers as a control, and add a test per endpoint that requests another tenant's object and expects an empty response.
The reason for that specificity is a failure I have seen: A report endpoint checked scope only. Incrementing the identifier returned 4,700 reports belonging to 38 customers, and the crawl ran for 2 hours at 6 requests per second without tripping anything.
Endpoint audit for object checks.
| Endpoint group | Endpoints | Missing the object check |
|---|---|---|
| reports | 14 | 9 |
| exports | 6 | 2 |
| administration | 21 | 0 |
I would not consider it settled without evidence: request another tenant's object with a valid token and require the response to carry no data.
The scope is the verb; the object check is the noun.
Curated: · Written: · Reviewed:
QA-55What is the difference between static and dynamic separation of duty, and when do you need each?(show answer)
I would settle static versus dynamic separation of duty by attempting the action after the identity event that was supposed to stop it.
Static separation forbids holding two entitlements at all; dynamic separation allows holding both while forbidding their use on the same transaction. Static is the stronger and more expensive claim in a small team, and choosing the wrong one either blocks the business or leaves the conflict live.
Concretely, define conflict pairs at entitlement level rather than by role name, enforce static conflicts at grant time as a hard block, enforce dynamic conflicts at decision time by comparing the actor against the record's earlier actors, and record every override with an approver and an expiry.
The reason for that specificity is a failure I have seen: A finance team of 5 held both vendor creation and payment approval because static separation was unworkable, and no dynamic rule existed. One person created a vendor and approved 9 payments to it across 4 months before a reconciliation caught it.
Conflict pairs and enforcement.
| Pair | Static enforcement | Dynamic enforcement |
|---|---|---|
| vendor creation and payment approval | not feasible at 5 staff | enforced per record |
| user creation and access approval | enforced, 0 holders | not needed |
| merge to main and production deploy | enforced in 2 teams | enforced elsewhere |
I would not consider it settled without evidence: attempt to approve a record the same actor created and require the decision to refuse.
If you cannot separate the people, separate the transaction.
Curated: · Written: · Reviewed:
QA-56Your role hierarchy is six levels deep. What does that cost you?(show answer)
The judgement in shallow role hierarchy is which identifier is durable and which is only a convenient claim.
Every inherited level is permission a reviewer has to reconstruct in their head, and effective access stops being readable from the grant. Depth also turns a change at the top into an untraceable change everywhere below it.
Concretely, keep inheritance to two levels, display effective permissions rather than the grant path, require any hierarchy change to show the effective-access delta before it applies, and split a role rather than adding a level when a new need appears.
The reason for that specificity is a failure I have seen: A permission added to a base role reached 3,100 people through 4 inheritance levels. The change was approved on the belief that it affected 40, and the error surfaced 9 days later during an unrelated review.
Hierarchy change impact.
| Level touched | Roles below | People affected |
|---|---|---|
| base role | 27 | 3,100 |
| mid-tier role | 6 | 210 |
| leaf role | 0 | 40 |
I would not consider it settled without evidence: produce the effective permission difference for the affected population before approving a hierarchy change.
Review effective access, not the path that produced it.
Curated: · Written: · Reviewed:
QA-57An elevation was granted for two hours. What actually expires?(show answer)
Where candidates lose the interview on just-in-time elevation expiry is reading a green IdP dashboard as revocation.
An elevation expires only where something enforces the expiry, and the grant record, the issued credential, and any session opened under it are three separate clocks. If the target keeps the membership after the broker forgets it, the elevation never ended.
Concretely, issue elevation as a time-bound credential rather than a persistent membership, have the broker remove the grant at the target and confirm the removal, invalidate sessions created under the elevation, and reconcile target state against the broker's expectation on a schedule.
The reason for that specificity is a failure I have seen: A broker marked an elevation expired at 120 minutes while the group membership at the target remained. 14 accounts held privileged membership for an average of 26 days, and the broker's report showed zero active elevations throughout.
State 24 hours after expiry.
| Location | Expected | Found |
|---|---|---|
| broker grant record | removed | removed |
| target group membership | removed | present in 14 cases |
| sessions opened under elevation | ended | 3 still live |
I would not consider it settled without evidence: let an elevation expire and then query the target directly for the membership and for any live session.
The broker's clock is not the target's state.
Curated: · Written: · Reviewed:
QA-58You cache authorization decisions. What goes into the key?(show answer)
I would answer authorization cache key completeness by separating authentication, authorization, and lifecycle into three control planes.
A cached decision is reusable only for a request identical in every input the policy read. If the key omits one — tenant, object, action, or the version of the attributes the decision depended on — the cache answers a question that was not asked.
Concretely, derive the key from the full decision input including subject, tenant, object identifier, action, and a version stamp for policy and attribute snapshot, keep lifetimes short for privileged decisions, and invalidate by version rather than by waiting.
The reason for that specificity is a failure I have seen: A cache keyed on user and action alone served an allow from one tenant's object to the same person in another tenant. 78 cross-tenant reads happened in 12 minutes, and the log showed only the original decision.
Cache key contents.
| Component | Old key | New key |
|---|---|---|
| subject | yes | yes |
| action | yes | yes |
| tenant and object identifier | no | yes |
| policy version | no | yes, bumped 4 times this month |
I would not consider it settled without evidence: issue two requests differing in one policy input and confirm the cache produces two separate decisions.
A key missing an input is a wrong answer waiting for a hit.
Curated: · Written: · Reviewed:
QA-59The authorization service times out during a payment write. What does the caller do?(show answer)
The engineering content of policy-engine timeout on a write is the revocation path and its measured latency, not the protocol name.
An unavailable decision is not an allow. State-changing operations fail closed, and the availability of the decision path becomes part of the service's own availability budget rather than something traded away under load.
Concretely, set a short decision timeout, fail closed for writes and privileged reads, permit a bounded stale answer only for low-risk reads with a stated maximum age, run the engine close to the caller with a replicated policy bundle so the failure stays rare, and count fail-closed events as an incident signal.
The reason for that specificity is a failure I have seen: A permissive default added during a capacity incident was left in place. During a later 7-minute engine outage 1,900 writes were accepted with no authorization decision at all, and 5 of them were privileged changes nobody had permission to make.
Behaviour during a seven-minute outage.
| Path | Fail-open result | Fail-closed result |
|---|---|---|
| privileged write | 5 unauthorised changes | refused |
| ordinary write | 1,900 accepted | refused |
| low-risk read | served | served from cache, 60 seconds maximum |
I would not consider it settled without evidence: stop the decision service in a test environment and require every write path to refuse.
No decision means no.
Curated: · Written: · Reviewed:
QA-60How do you prove your multi-tenant authorization actually isolates tenants?(show answer)
Before calling cross-tenant resource substitution test done I would write down the grant that the source-of-truth change does not touch.
Isolation is a claim about what happens when a valid caller supplies another tenant's identifier, so it can only be demonstrated by doing exactly that. A design review cannot establish it, because the failure lives in whichever handler forgot the check.
Concretely, keep two seeded tenants in every environment, generate a substitution case per endpoint from the API description, run them in the pipeline, and fail the build on any response that returns data or discloses existence through a different status code or a timing difference.
The reason for that specificity is a failure I have seen: A new bulk endpoint shipped without a substitution case. It returned 30 records belonging to another tenant when identifiers were mixed, and the gap survived 21 days until a customer recognised a name they had never seen.
Substitution suite results.
| Endpoint set | Cases | Failures found |
|---|---|---|
| single-object reads | 62 | 0 |
| bulk operations | 18 | 1 |
| exports | 9 | 2 |
I would not consider it settled without evidence: run the substitution suite across every endpoint and require zero responses containing the other tenant's data.
Prove isolation by trying to cross it.
Curated: · Written: · Reviewed:
QA-61How do you review a policy change before it reaches production?(show answer)
The first thing I would establish about policy deployment decision diff is which live path still works after the directory row looks right.
A policy diff shows text, and what matters is the change in decisions. The reviewable artifact is the set of principals whose access moves, produced by replaying real decision inputs against both versions.
Concretely, capture a representative sample of decision inputs, evaluate old and new policy over it, present newly allowed and newly denied counts with examples, scale sign-off to the size of the newly allowed set, and block deployment when the difference cannot be produced.
The reason for that specificity is a failure I have seen: A refactor described as behaviour-preserving granted a wildcard action to 3 service principals. The text difference was 12 lines and looked tidy; the decision difference, produced afterwards, showed 41,000 newly allowed requests in the replay sample.
Decision difference for one change.
| Category | Requests | Handling |
|---|---|---|
| newly allowed | 41,000 | escalated |
| newly denied | 260 | accepted |
| unchanged | 1,240,000 | sampled |
I would not consider it settled without evidence: replay the recorded decision sample against both policy versions and require the newly allowed set to match the stated intent.
Review decisions, not text.
Curated: · Written: · Reviewed:
QA-62What does an authorization log need to be useful six months later?(show answer)
I would start authorization decision logging from the session, token, and grant that outlive the login, not from the login itself.
A log recording allow or deny answers the least interesting question. Reconstructing a past decision needs the subject, the object, the action, the policy version, the attribute values the decision read, and the rule that decided it.
Concretely, emit a structured record per decision with those fields and a correlation identifier, keep policy bundles addressable by version so an old decision can be re-evaluated, sample verbose attribute capture rather than dropping it, and retain decisions on privileged objects longer than ordinary ones.
The reason for that specificity is a failure I have seen: An access dispute six months old could not be settled because the log held only user, resource, and outcome. The policy had been rewritten twice since, and after 5 days of work the review closed on an assumption rather than a finding.
Fields needed to replay a decision.
| Field | Present before | Present after |
|---|---|---|
| subject and object | yes | yes |
| policy version | no | yes, 37 versions retained |
| attribute values read | no | yes, sampled at 5 percent |
| deciding rule identifier | no | yes |
I would not consider it settled without evidence: pick a decision from a past month and re-evaluate it from the log alone, reproducing the same outcome.
A decision you cannot replay is an opinion.
Curated: · Written: · Reviewed:
QA-63Where do you draw the line between a role and an attribute condition?(show answer)
This is an area where a successful federation and a current authorization decision are different events.
Roles carry the coarse, reviewable grant while attributes narrow it at decision time with context the grant cannot know. The workable split is that a person's job decides the role and the request's circumstances decide the condition.
Concretely, grant the role from job function, attach conditions on classification, location, device state, and time at the decision, keep conditions few and named so a reviewer can read the effective rule, and never encode an individual's identity inside a condition.
The reason for that specificity is a failure I have seen: Everything was pushed into attribute expressions. One rule reached 14 nested conditions, no reviewer could say who it permitted, and a mistake in operator precedence gave 220 contractors access to production data for 17 days.
Split for one entitlement.
| Element | Carried by | Reviewed by |
|---|---|---|
| read customer records | role with 340 holders | data owner, quarterly |
| classification is internal | condition | policy owner |
| device is managed | condition | 2 platform reviewers |
I would not consider it settled without evidence: state in one sentence who a rule permits, then check that sentence against an evaluation over the whole population.
Roles for who you are; conditions for where you are standing.
Curated: · Written: · Reviewed:
QA-64Someone moves from support into engineering. What does your process usually miss?(show answer)
My answer to department-move stale grants begins with the artifact that is actually being trusted: token, assertion, group, or session.
Access accumulates because the joining half of a move is urgent and the leaving half is not. The result is a person holding two jobs worth of entitlements who passes every review, since each grant was justified once.
Concretely, drive moves from the source-of-truth attribute change, compute old and new birthright sets, remove the difference on a stated clock rather than at the next campaign, and require the previous manager to confirm removal separately from the new manager confirming additions.
The reason for that specificity is a failure I have seen: A support engineer moved to engineering and kept a customer-impersonation entitlement for 11 months. It was used twice after the move, both times by the person's own account and both times legitimately, which is exactly why nothing fired.
Entitlements across one move.
| Set | Count | Handling |
|---|---|---|
| old birthright | 23 | removed within 3 days |
| new birthright | 19 | granted on day 1 |
| justified overlap | 7 | documented |
| retained without reason | 4 | present for 11 months |
I would not consider it settled without evidence: move a test identity between departments and compare effective access against the new birthright set inside the stated window.
A move is a leave and a join, and the leave is the half that slips.
Curated: · Written: · Reviewed:
QA-65A team needs an exception to a policy for a migration. How do you grant it?(show answer)
I would treat time-bound policy exceptions as a claim about effective access that has to survive a disablement.
An exception without an expiry is a policy change made by whoever was in the room. The grant carries a scope, an owner, a compensating control, and a date on which it stops working without anyone acting.
Concretely, implement exceptions as time-bound entries the engine enforces, scope them to the narrowest principal and resource that unblocks the work, require the owner to re-justify before renewal, cap the number of renewals, and publish the register with days remaining.
The reason for that specificity is a failure I have seen: A migration exception opened a broad write permission for ten working days. It was renewed by email 4 times, still existed 15 months later, and covered a system whose original owner had left the company.
Exception register.
| Exception | Age | Renewals |
|---|---|---|
| migration write access | 15 months | 4, owner gone |
| vendor read window | 6 days | 0 |
| legacy service bypass | 3 months | 1, expires in 12 days |
I would not consider it settled without evidence: list active exceptions and confirm each has an expiry within the cap and an owner who is still employed.
An exception without an expiry is the new policy.
Curated: · Written: · Reviewed:
QA-66A batch endpoint accepts five hundred identifiers. Where does authorization go?(show answer)
The useful question for bulk API per-object authorization is what still holds at the relying parties nobody has looked at this week.
A batch is a set of independent operations, so authorization is per element and the response must not reveal the elements the caller could not touch. Checking the first item, or the batch as a whole, turns one authorised object into a key for the rest.
Concretely, evaluate every element against the caller, return a per-element outcome in a uniform shape for refusals and for missing objects alike, cap batch size, account rate limits per element rather than per call, and make partial success explicit.
The reason for that specificity is a failure I have seen: A bulk delete was authorised on the first identifier only. One request removed 340 records across 12 tenants, restoration took 19 hours, and the trail recorded a single authorised call.
Batch of sixty mixed identifiers.
| Element class | Count | Acted on |
|---|---|---|
| the caller's own | 41 | 41 |
| another tenant's | 12 | 0 |
| non-existent | 7 | 0, identical response shape |
I would not consider it settled without evidence: submit a batch mixing permitted and forbidden identifiers and require exactly the permitted ones to be acted on.
A batch is many decisions wearing one request.
Curated: · Written: · Reviewed:
QA-67A user queues an export that runs an hour later. Whose authority does it run with?(show answer)
I would settle async job carrying user authority by attempting the action after the identity event that was supposed to stop it.
A queued job carries the requester's authority as a narrow, time-bound delegation, and it re-checks at execution because that authority may have been removed while the job waited. Running it as a service account with broad rights turns every queue into an escalation path.
Concretely, record the subject and the requested scope on the job, mint a short-lived delegated credential limited to that operation at execution time, re-evaluate the subject's current entitlement before producing output, and fail the job with a clear reason when the authority has gone.
The reason for that specificity is a failure I have seen: An export queue ran under a service identity holding tenant-wide read. A job queued 50 minutes before a dismissal completed afterwards and mailed 8,200 customer records to the leaver's personal address.
Job execution against current authority.
| Design | Re-checks at run time | Records released in the drill |
|---|---|---|
| service identity with tenant read | no | 8,200 |
| delegated credential, no re-check | no | 240 |
| delegated credential with re-check | yes | 0 |
I would not consider it settled without evidence: queue a job, remove the requester's access, and require the job to fail at execution.
The queue must not outlive the authority.
Curated: · Written: · Reviewed:
QA-68What is your revocation service level, and how do you know you meet it?(show answer)
The judgement in end-to-end access-removal latency is which identifier is durable and which is only a convenient claim.
Revocation latency runs from the source-of-truth event to the first denied action at the target, not to the ticket closing. Every hop adds to it — connector schedule, cache, session lifetime, token expiry — and the number is credible only when it is measured on real targets.
Concretely, keep a synthetic identity attempting a benign privileged action at each tier on a schedule, trigger a termination event, record the timestamp of the first denial per target, publish the distribution rather than a mean, and treat targets with no measurement as unmeasured rather than compliant.
The reason for that specificity is a failure I have seen: A stated one-hour objective was reported as met from ticket timestamps. Measured first denial at the third-party expense tool arrived 4 days and 6 hours later, and 2 of 9 tiers had no measurement at all.
Measured minutes to first denial.
| Target | Claimed | Measured |
|---|---|---|
| directory | 5 | 6 |
| core finance application | 60 | 74 |
| expense tool | 60 | 6,120 |
| data warehouse | 60 | never measured |
I would not consider it settled without evidence: run the synthetic termination monthly and publish the minutes to first denial for every target.
Measure at the target, not at the ticket.
Curated: · Written: · Reviewed:
QA-69Why remove standing administrative rights if the people holding them are trusted?(show answer)
Where candidates lose the interview on just-in-time privileged access is reading a green IdP dashboard as revocation.
Standing rights are exercised by whoever holds the session, not by the person who was trusted. Removing them turns a credential compromise into a request that must pass approval, and it leaves a record of every privileged interval.
Concretely, make privileged roles requestable with a justification, an approver, a maximum duration, and automatic removal, keep the everyday account without the role, and track standing privileged assignments as a number that must fall to a small named set.
The reason for that specificity is a failure I have seen: An administrator's workstation was compromised while the account held permanent domain rights. The attacker created 2 accounts and a scheduled task within 18 minutes, with no approval step to slow them and no elevation record to review afterwards.
Standing privilege reduction.
| Role | Standing holders before | After |
|---|---|---|
| domain administrator | 34 | 2 break-glass |
| cloud tenant owner | 12 | 0 |
| database owner | 46 | 3 |
I would not consider it settled without evidence: count accounts holding a privileged role outside an active elevation window and require that list to match the documented exceptions.
Trust the person; do not leave the rights lying on the desk.
Curated: · Written: · Reviewed:
QA-70A team needs to restart services in production. What role do you give them?(show answer)
I would answer just-enough administration scope by separating authentication, authorization, and lifecycle into three control planes.
Privilege should be scoped to the operations the job requires rather than to the container those operations live in. Handing over the whole administrative role because it happens to contain the needed action is how a restart permission becomes a permission to change authentication.
Concretely, identify the specific operations, assemble a role from those alone, scope it to the resource group or organisational unit involved, expose it through a constrained endpoint permitting the named commands, and review the operation list rather than the role name.
The reason for that specificity is a failure I have seen: An operations team was given a full administrative role to perform restarts. A scripted mistake used an unrelated capability of the same role, deleted 6 service principals, and broke authentication for 3 applications for 5 hours.
Role composition.
| Approach | Operations granted | Blast radius |
|---|---|---|
| full administrative role | 214 | the whole tenant |
| scoped custom role | 4 | 1 resource group |
| constrained endpoint | 4 commands | 9 named hosts |
I would not consider it settled without evidence: attempt an operation outside the named list with the scoped role and require it to be refused.
Grant the verb, not the container it arrived in.
Curated: · Written: · Reviewed:
QA-71Why should an administrator have two accounts?(show answer)
The engineering content of separate daily and privileged identities is the revocation path and its measured latency, not the protocol name.
Everyday work exposes an identity to mail, browsing, and documents, which is where compromise starts, and privileged work must not share that exposure. Two identities also make privileged activity legible, because everything the privileged account did was meant to be privileged.
Concretely, issue an administrative identity with no mailbox and no general browsing, require its own strong authenticator, restrict its sign-in to hardened administrative workstations through conditional policy, and alert on any interactive use from an ordinary endpoint.
The reason for that specificity is a failure I have seen: One account was used for both roles. A malicious document executed with rights that included directory administration, and the resulting persistence took 5 weeks to remove across 3 domains.
Identity separation in practice.
| Property | Daily account | Administrative account |
|---|---|---|
| mailbox | yes | none |
| general browsing | yes | blocked |
| permitted sign-in surfaces | 4,200 endpoints | 26 hardened workstations |
I would not consider it settled without evidence: attempt an administrative sign-in from an ordinary workstation and require the conditional policy to block it.
Separate the exposure from the privilege.
Curated: · Written: · Reviewed:
QA-72Who approves a privileged elevation request?(show answer)
Before calling independent elevation approval done I would write down the grant that the source-of-truth change does not touch.
Approval means nothing when the requester can supply it, directly or through a chain they control. Independence has to be structural: a different person, in a different reporting line for the highest tiers, with the record kept outside the requester's reach.
Concretely, define approver groups per privileged role, block self-approval and reciprocal pairs, require two approvers for the top tier, expire requests that go unapproved inside a short window, and audit approver-requester pairs for reciprocity patterns.
The reason for that specificity is a failure I have seen: Two engineers approved each other's elevations 61 times across 7 months. Both held permanent membership of the approver group for the very role they were requesting, and no rule prevented the pairing.
Approval rules by tier.
| Tier | Approvers required | Constraint |
|---|---|---|
| directory tier | 2 | different reporting line |
| application administration | 1 | not a report of the requester |
| reciprocal pairs | not applicable | blocked, 61 found in review |
I would not consider it settled without evidence: attempt self-approval and reciprocal approval in a test tenant and require both to be refused.
An approval the requester can produce is a formality.
Curated: · Written: · Reviewed:
QA-73Your privileged access tool checks out a password for four hours. What would you change?(show answer)
The first thing I would establish about ephemeral privileged credentials is which live path still works after the directory row looks right.
A checked-out password is a copyable secret whose lifetime is a policy setting rather than a fact. Ephemeral credentials — a short-lived certificate or a token minted per session — make the copy useless afterwards because nothing durable was handed over.
Concretely, issue certificates valid for the session with the principal and target encoded, keep the private key inside the broker or a hardware module, rotate the underlying account credential after every use where a password is unavoidable, and never display a credential the operator could paste elsewhere.
The reason for that specificity is a failure I have seen: A four-hour checkout was pasted into a personal notes application. The same password still worked 6 weeks later because rotation had failed silently on 40 of 380 accounts, and the failure was visible only in a log nobody read.
Credential model comparison.
| Model | Usable after the session | Depends on rotation |
|---|---|---|
| four-hour password checkout | yes, when rotation fails | 380 accounts |
| checkout with verified rotation | no | verified per use |
| twenty-minute session certificate | no | not at all |
I would not consider it settled without evidence: capture a credential issued for a session and confirm it fails at the target once the session ends.
Hand over a session, not a secret.
Curated: · Written: · Reviewed:
QA-74We have a password vault. Do we have privileged access management?(show answer)
I would start vaulting versus privileged access management from the session, token, and grant that outlive the login, not from the login itself.
A vault stores secrets; privileged access management governs a session — who may request it, who approved it, what was done, and what remains afterwards. A vault with broad read permission is a convenient copy of every privileged credential.
Concretely, put brokered sessions in front of the vault so operators never receive secrets, enforce request and approval, record the session, rotate after use, and reserve direct retrieval for a small number of documented cases carrying a compensating control.
The reason for that specificity is a failure I have seen: A vault held 2,100 privileged credentials with retrieval permitted to a group of 90 people. An audit could show who retrieved a credential and nothing about what was then done with it, leaving 3 incidents unattributable.
What each layer can answer.
| Question | Vault alone | Brokered session |
|---|---|---|
| who retrieved the secret | yes | yes |
| who approved it | no | yes |
| what commands ran | no | yes, 3 incidents resolved |
| credential rotated afterwards | sometimes | always |
I would not consider it settled without evidence: retrieve a privileged credential as a standard operator and require the request to be brokered rather than granted.
Storing the key is not governing the door.
Curated: · Written: · Reviewed:
QA-75A legacy system has one administrative account that four people use. What do you do before you can retire it?(show answer)
This is an area where a successful federation and a current authorization decision are different events.
Attribution has to be restored before the account can be governed, because every control after that depends on knowing who acted. Until individual identities exist there, the practical move is to make the shared credential unusable without a brokered and recorded session.
Concretely, put the account behind the broker so each use is requested by a named person, rotate the credential after every session, record the session, and run a parallel project to add individual accounts or a federation shim so the shared credential can be retired on a date.
The reason for that specificity is a failure I have seen: A shared administrative login on a manufacturing system was used by 4 engineers and 2 vendors. A configuration change halted a line for 9 hours, the investigation could not identify who made it, and the corrective action ended up being a policy reminder.
Shared account governance.
| Control | Before | After brokering |
|---|---|---|
| named user per session | no | yes, 38 sessions |
| credential rotation | annual | after every use |
| unattributable changes | one 9-hour outage | 0 |
I would not consider it settled without evidence: review the last month of sessions on the shared account and confirm every one names an individual.
Restore the name before you argue about the privilege.
Curated: · Written: · Reviewed:
QA-76You record privileged sessions. What must never end up in the recording?(show answer)
My answer to privileged session recording redaction begins with the artifact that is actually being trusted: token, assertion, group, or session.
A session recording is a high-value target because it captures administration verbatim, including everything typed. Secrets, personal data, and keystrokes revealing credentials have to be filtered at capture, and access to the archive is itself a privileged action.
Concretely, mask password prompts and known secret patterns during capture rather than in post-processing, store recordings under a separate encryption key, require dual authorisation to view, log every view, and set retention against the investigative need rather than keeping everything.
The reason for that specificity is a failure I have seen: Keystroke capture recorded a database password typed at a prompt. The archive was readable by an operations group of 22 people, and the credential was still valid 3 months later when the exposure was found.
Recording controls.
| Control | Change | Effect |
|---|---|---|
| capture-time masking | added | 0 secrets in 500 sampled sessions |
| archive readers | reduced from 22 | 4, under dual authorisation |
| retention | was indefinite | 180 days |
I would not consider it settled without evidence: search a sample of recordings for known secret patterns and require zero hits before the archive is trusted.
A recording of an administrator is an administrator's credential.
Curated: · Written: · Reviewed:
QA-77Your identity provider is unavailable. How does anyone get in to fix it?(show answer)
I would treat break-glass independent of the failed IdP as a claim about effective access that has to survive a disablement.
A break-glass path that authenticates through the system it exists to rescue is not a path. It has to be independent of the federation, of the multi-factor service, and of the network controls that could be the thing failing.
Concretely, keep a small number of non-federated accounts with long random credentials split between sealed physical copies, exclude them from conditional policies that depend on the provider, alert on any sign-in to several channels at once, exercise the path on a schedule, and rotate after every use and every test.
The reason for that specificity is a failure I have seen: Both emergency accounts were federated to the provider that failed. The tenant was unreachable for 5 hours and 40 minutes while support restored a trust relationship, and the runbook had not been exercised in 2 years.
Break-glass account properties.
| Property | Required | Found in review |
|---|---|---|
| not federated | yes | 0 of 2 |
| excluded from provider-dependent policy | yes | 0 of 2 |
| exercised in the last quarter | yes | last test 2 years ago |
| alerts on use | yes | 1 channel only |
I would not consider it settled without evidence: sign in with a break-glass account while the identity provider is blocked and confirm the path works and alerts.
The rescue path cannot depend on the thing being rescued.
Curated: · Written: · Reviewed:
QA-78Who administers the privileged access system itself?(show answer)
The useful question for PAM control-plane duty separation is what still holds at the relying parties nobody has looked at this week.
Whoever can change the privileged access platform can grant themselves anything it governs, so administering the platform is a separate duty from using it. Its audit trail also has to land somewhere its own administrators cannot edit.
Concretely, split platform administration, policy authorship, approval, and operation across different people, require two-person control for changes to approval rules and recording settings, stream the trail to a store outside the platform's control, and review platform-administrator actions monthly with a group that does not report to them.
The reason for that specificity is a failure I have seen: A platform administrator disabled recording for one target, used the account, and re-enabled it. The gap was found 4 months later only because the external store showed 26 minutes with no session events while the platform's own log had been trimmed.
Duties across the platform.
| Duty | Holder group | Overlap allowed |
|---|---|---|
| platform administration | 3 engineers | none with approval |
| policy authorship | 2 identity architects | none with operation |
| approval | 11 managers | none with requesting |
| audit review | 4 internal auditors | none |
I would not consider it settled without evidence: attempt to change recording settings with a single platform-administrator account and require a second approver.
The system that governs privilege needs its own separation.
Curated: · Written: · Reviewed:
QA-79A local administrators group on a server holds members the broker never granted. How did that happen, and what fixes it?(show answer)
I would settle privileged grant expiry at the target by attempting the action after the identity event that was supposed to stop it.
Targets accept changes through their own local administration path as well as through the broker, so the broker's record is a subset of effective privilege. Only reconciliation against the target's actual membership can say what is really there.
Concretely, enumerate effective privileged membership at each target on a schedule, compare it against the grants the broker believes are active, remove or convert anything unmatched, and disable direct local administration paths where the platform allows it so the broker becomes the only writer.
The reason for that specificity is a failure I have seen: A reconciliation across 610 servers found 143 local administrator entries the broker had never issued. 38 of them belonged to people who had changed teams, and the oldest dated back 3 years.
Reconciliation across the estate.
| Source of the entry | Count | Action |
|---|---|---|
| broker-issued and active | 402 | keep |
| broker-issued but expired | 27 | remove |
| never brokered | 143 | remove or convert |
I would not consider it settled without evidence: compare target-side privileged membership against broker-issued grants and require the unmatched set to be empty.
The target's membership is the fact; the broker's record is a claim.
Curated: · Written: · Reviewed:
QA-80A service needs to call another service. Why not give it an API key?(show answer)
The judgement in workload identity instead of static keys is which identifier is durable and which is only a convenient claim.
A static key is a bearer secret with no expiry, no binding to the caller, and no way to distinguish legitimate use from a copy. A workload identity issues a short-lived credential to an attested runtime, so the credential and the workload live and die together.
Concretely, use workload identity federation or a certificate issued to the attested workload, keep lifetimes in minutes, scope the audience to the specific callee, remove long-lived keys from the estate and block their creation by policy, and monitor for any credential older than the maximum lifetime.
The reason for that specificity is a failure I have seen: A static key baked into a container image was published to a public registry. It stayed valid for 17 months and was used from 3 unfamiliar networks before an inventory scan found it, and nothing about the calls looked different from legitimate traffic.
Service credential inventory.
| Credential type | Count | Maximum lifetime |
|---|---|---|
| static API keys | 1,480 | none |
| workload identity tokens | 96 services | 15 minutes |
| issued certificates | 210 | 24 hours |
I would not consider it settled without evidence: count credentials in the estate that carry no expiry and require the number to reach zero by a stated date.
A credential that never expires cannot be lost safely.
Curated: · Written: · Reviewed:
QA-81The broker holds credentials for every target. What stops a request from naming the wrong one?(show answer)
Where candidates lose the interview on PAM broker confused-deputy target bind is reading a green IdP dashboard as revocation.
The broker is a deputy holding authority over many systems, so the authorisation must bind to the exact target the approval named. If the target is taken from the request rather than from the approved grant, an operator approved for one host can reach another.
Concretely, carry the target identity inside the approved grant, resolve it server-side by a stable identifier rather than by a name the requester supplies, verify the target's host key or certificate before injecting the credential, and refuse any session whose resolved target differs from the grant.
The reason for that specificity is a failure I have seen: A request approved for a test database was redirected by editing a hostname parameter. The broker matched on display name and injected production credentials, and 4 tables were altered before the operator noticed the banner.
Target resolution before and after.
| Step | Before | After |
|---|---|---|
| target chosen from | the request parameter | the grant record |
| matching key | display name | stable identifier, 1 per host |
| host identity verified | no | yes, certificate pinned |
I would not consider it settled without evidence: approve a grant for one target, attempt a session against another, and require the broker to refuse.
The approval names the target; the request must not.
Curated: · Written: · Reviewed:
QA-82A vendor engineer needs privileged access for a support call. How do you set that up?(show answer)
I would answer named vendor privileged access by separating authentication, authorization, and lifecycle into three control planes.
Vendor access has to resolve to a named individual at the vendor with a start and an end, because the contract is with a company while accountability belongs to a person. A shared vendor account leaves you unable to say who was on the system.
Concretely, sponsor each vendor engineer as an individual identity whose expiry matches the engagement, require their own strong authenticator, broker the session with recording, scope it to the systems named in the ticket, and end access automatically when the ticket closes or the sponsorship lapses.
The reason for that specificity is a failure I have seen: A shared vendor login was used by 9 engineers in 2 countries across 3 years. When a configuration change caused a 12-hour outage, the vendor could not identify who had connected, and the contract contained no clause requiring them to.
Vendor access model.
| Attribute | Shared account | Sponsored identity |
|---|---|---|
| named individual | no | yes, 9 registered |
| expiry | none | ticket close or 30 days |
| session recording | no | yes |
I would not consider it settled without evidence: review vendor sessions from the last quarter and confirm each maps to a named individual with an unexpired sponsorship.
Contract with the company; authorise the person.
Curated: · Written: · Reviewed:
QA-83Managers approved every entitlement in the last campaign. What went wrong?(show answer)
The engineering content of privileged access recertification of effective rights is the revocation path and its measured latency, not the protocol name.
Recertification fails when the reviewer sees group names instead of what those groups permit, and when approving everything is the cheapest action available. The review has to present effective rights, recent use, and a default that removes whatever is not affirmatively kept.
Concretely, show the reviewer resolved permissions, last-use date, and a peer comparison, make no response mean removal after a grace period, route privileged entitlements to the resource owner rather than the line manager, and sample-test approvals afterwards by asking the approver to justify one.
The reason for that specificity is a failure I have seen: A campaign covering 12,000 entitlements closed with 98 percent approved at a median of 9 seconds per item. A follow-up sample found that 31 of 50 approvals could not be justified by the person who made them.
Campaign redesign results.
| Measure | Old campaign | New campaign |
|---|---|---|
| approval rate | 98 percent | 71 percent |
| median seconds per item | 9 | 48 |
| entitlements removed | 240 | 3,470 |
I would not consider it settled without evidence: ask a sample of approvers to justify a specific approval afterwards and measure how many can.
A review that costs nothing removes nothing.
Curated: · Written: · Reviewed:
QA-84What is your policy for the cloud root account?(show answer)
Before calling cloud root or tenant-owner rarity done I would write down the grant that the source-of-truth change does not touch.
The root or tenant-owner credential can undo every other control, including the logging that would show it being used. It should serve the small set of tasks that genuinely require it and be otherwise unusable without a deliberate, alerting, multi-party ceremony.
Concretely, remove programmatic keys from root, register a hardware authenticator held in a safe, split the password between two custodians, alert on any root event to security and to an executive channel, document the tasks that truly require it, and review each use within one working day.
The reason for that specificity is a failure I have seen: Root retained a long-lived access key from an early automation script. The key was used from an unfamiliar network to create an administrative user, and because the trail configuration was changed in the same session, reconstruction depended on a second account's logs and took 6 days.
Root account state.
| Item | Before | After |
|---|---|---|
| programmatic keys | 1, four years old | 0 |
| hardware authenticator | none | yes, held in a safe |
| uses in the last year | 11 | 2, both reviewed |
I would not consider it settled without evidence: list root credentials and confirm no programmatic key exists and that a sign-in raises an alert within minutes.
The account that can undo the controls should be almost unusable.
Curated: · Written: · Reviewed:
QA-85A manager asks you to give the new starter the same access as a colleague. What do you say?(show answer)
The first thing I would establish about joiner birthright from policy not a peer clone is which live path still works after the directory row looks right.
Copying a colleague copies their exceptions, their previous roles, and everything they collected over years, and it does so without anyone stating what the job needs. Birthright access should come from a policy keyed on job attributes that a reviewer can read and change once.
Concretely, define birthright sets per job code, location, and employment type, provision them automatically from the source-of-truth record, route anything beyond the set through a request with a justification, and report grants that came from cloning as a defect to be cleared.
The reason for that specificity is a failure I have seen: A clone of a nine-year veteran gave a graduate 62 entitlements, including production database write and a finance approval role. Nothing flagged it, because after the clone both accounts looked identical to every review.
Clone against policy birthright.
| Source | Entitlements granted | Justified by the job |
|---|---|---|
| clone of a colleague | 62 | 19 |
| birthright policy | 19 | 19 |
| requested additions | 3 | 3, each approved |
I would not consider it settled without evidence: compare a new joiner's effective access against the birthright set for their job code and require the difference to be zero or justified.
Copy the policy, not the person.
Curated: · Written: · Reviewed:
QA-86You create accounts a week before the start date. What keeps them safe until then?(show answer)
I would start pre-provisioned accounts remaining unusable from the session, token, and grant that outlive the login, not from the login itself.
An account that exists before its owner arrives is a credential nobody is watching. It has to be created in a state that cannot authenticate, with activation tied to a proofing event on the first day rather than to a date passing.
Concretely, create accounts disabled with no usable credential, hold the initial factor for delivery through a supervised enrolment, activate on a confirmed start event from the personnel system, and expire the pre-provisioned record automatically when the start does not happen.
The reason for that specificity is a failure I have seen: Accounts were created 10 days early with a predictable initial password and enabled immediately. Two offers were withdrawn, the accounts stayed live for 5 weeks, and one of them was used to read an internal document library.
Pre-provisioning states.
| State | Can authenticate | Trigger to the next state |
|---|---|---|
| created and disabled | no | confirmed start on day 1 |
| activated at enrolment | yes | supervised authenticator setup |
| start cancelled | no | automatic deletion after 7 days |
I would not consider it settled without evidence: attempt to authenticate as a pre-provisioned account before its activation event and require the attempt to fail.
Exists is not the same as usable.
Curated: · Written: · Reviewed:
QA-87Should a mover's new access land before or after the old access is removed?(show answer)
This is an area where a successful federation and a current authorization decision are different events.
The order matters because the two halves fail differently: adding first and removing later reliably produces accumulation, while removing first produces a short and visible gap somebody will fix. Where a handover genuinely needs overlap, that overlap should be explicit and dated.
Concretely, sequence removal ahead of addition by default, allow a named handover window with an expiry for the specific entitlements that need it, notify both managers, and report any mover whose old and new sets are both live past the window.
The reason for that specificity is a failure I have seen: Add-then-remove was the standard sequence and removal depended on a manual step. Across 340 moves in one year, 51 people held both sets for more than 30 days and 6 held them for over 6 months.
Move sequencing outcomes.
| Sequence | Moves | Both sets live past 30 days |
|---|---|---|
| add then remove | 340 | 51 |
| remove then add | 288 | 2 |
| explicit handover window | 46 | 0, expiry enforced |
I would not consider it settled without evidence: sample recent movers and confirm the old birthright set has gone inside the stated window.
Overlap should be a decision, not a leftover.
Curated: · Written: · Reviewed:
QA-88How does access granted for a project come to an end?(show answer)
My answer to automatic expiry of project access begins with the artifact that is actually being trusted: token, assertion, group, or session.
Project access should be granted with the project's end date attached, because the reason for the grant is known at the moment it is made and will never be clearer. Anything granted without an expiry becomes permanent by default.
Concretely, require an end date on every project grant with a maximum length, tie membership to the project record so closing the project removes access, warn owners before expiry with a one-click extension that records a reason, and report grants approaching the cap.
The reason for that specificity is a failure I have seen: A data-migration project ended and its access group survived it. 47 people kept read access to a customer dataset for 8 months, and 3 of them had left the company but retained access through a contractor identity.
Project group lifecycle.
| Grant type | Expiry set | Members after project close |
|---|---|---|
| project group, no expiry | no | 47 |
| project group, dated | yes, 90 days | 0 |
| extended with a reason | yes, plus 30 days | 5 |
I would not consider it settled without evidence: close a test project and confirm every membership derived from it disappears.
Attach the end date while the reason is still fresh.
Curated: · Written: · Reviewed:
QA-89A dismissal is happening at two in the afternoon. What is your sequence?(show answer)
I would treat immediate leaver kill chain as a claim about effective access that has to survive a disablement.
An immediate termination is an ordered sequence, because the steps have different latencies and one of them warns the person. Ending authentication and live sessions comes first, and the tidying of records comes afterwards.
Concretely, disable the account and revoke sessions and refresh tokens at the source, revoke device certificates and remote-access sessions, disable the authenticators, sweep tokens at the major relying parties, then handle mailbox delegation, data ownership transfer, and physical access, with a named owner and a timestamp per step.
The reason for that specificity is a failure I have seen: A termination began with a mailbox transfer and reached session revocation 40 minutes later. Inside that window the person cleared a shared folder and forwarded 260 messages from an active session on a personal device.
Termination sequence and timing.
| Step | Target elapsed | Owner |
|---|---|---|
| disable account, revoke sessions | 0 to 2 minutes | identity on-call |
| revoke tokens at 6 major relying parties | 5 minutes | automation |
| device and remote-access certificates | 10 minutes | endpoint team |
| mailbox and data transfer | same day | the manager |
I would not consider it settled without evidence: run the sequence against a test identity and record the elapsed seconds to first denial at each tier.
End the live paths first; the paperwork does not run away.
Curated: · Written: · Reviewed:
QA-90Your provisioning create timed out but the user may already exist. What now?(show answer)
The useful question for idempotent SCIM create after timeout is what still holds at the relying parties nobody has looked at this week.
A timeout is an unknown outcome rather than a failure, so a blind retry creates a second account for the same person. Idempotency needs a stable external identifier that the target enforces as unique, so a repeated create resolves to the same object.
Concretely, send the source-of-truth immutable identifier as the external identifier, look the user up by it before retrying, use conditional requests where the target supports them, treat a conflict response as success, and reconcile daily for objects sharing an external identifier.
The reason for that specificity is a failure I have seen: A connector retried after a 30-second timeout and created duplicates for 220 people during one slow afternoon. Group memberships split across the pairs, and 14 of them could not open files their own duplicate account owned.
Retry behaviour.
| Strategy | Duplicates created | Reconciliation needed |
|---|---|---|
| blind retry | 220 | 3 days |
| lookup by external identifier | 0 | none |
| conditional create | 0 | none |
I would not consider it settled without evidence: force a timeout on a create and confirm the retry resolves to one object rather than two.
Retry against an identifier, not against hope.
Curated: · Written: · Reviewed:
QA-91One malformed user record is blocking your provisioning queue. What is the design fix?(show answer)
I would settle poison provisioning record isolation by attempting the action after the identity event that was supposed to stop it.
A queue that halts on a bad record turns one data problem into an outage for every identity change behind it, terminations included. Isolation means the bad record stops while everything else continues.
Concretely, process records independently with per-record error handling, move repeated failures into a quarantine holding the reason and the payload, keep terminations on a priority path that creates cannot block, alert on quarantine depth, and require a person to clear the quarantine rather than a blanket retry.
The reason for that specificity is a failure I have seen: A user whose display name contained an unsupported character failed the connector, which then retried the whole batch. 1,100 changes queued behind it, including 3 terminations that completed 26 hours late.
Queue behaviour with a poison record.
| Design | Records blocked | Late terminations |
|---|---|---|
| batch retry on failure | 1,100 | 3 |
| per-record isolation | 1 | 0 |
| priority lane for leavers | 1 | 0 |
I would not consider it settled without evidence: inject a malformed record and confirm the queue continues while that record lands in quarantine.
One bad record should cost one record.
Curated: · Written: · Reviewed:
QA-92Someone rejoins after two years away. Do you reactivate the old account?(show answer)
The judgement in rehire without inheriting old sessions is which identifier is durable and which is only a convenient claim.
Reactivating an old object restores whatever was attached to it — group memberships, application grants, delegations, sometimes credentials — none of which was reviewed against the new job. The identity may be reused for continuity, but the entitlements are rebuilt from birthright.
Concretely, keep the identity record for history, strip entitlements and credentials at departure rather than at rehire, require fresh enrolment of authenticators, provision the new birthright set from the new job code, and use the pre-departure set as a comparison rather than as a source.
The reason for that specificity is a failure I have seen: A rehire reactivated an account dormant for 26 months. It carried 18 stale group memberships including a payroll approval role, and an old delegation kept forwarding a former colleague's mail for 3 weeks.
Rehire access reconstruction.
| Source | Entitlements | Kept |
|---|---|---|
| dormant account, 26 months | 18 | 0 |
| new birthright set | 14 | 14 |
| requested additions | 2 | 2 |
I would not consider it settled without evidence: reactivate a test identity and confirm its effective access equals the new birthright set exactly.
Keep the person's history; rebuild the person's access.
Curated: · Written: · Reviewed:
QA-93A sign-in form passes the username straight into a directory search filter. What is the risk?(show answer)
Where candidates lose the interview on LDAP filter injection escaping is reading a green IdP dashboard as revocation.
Filter syntax is a language, and unescaped input rewrites the query rather than parameterising it. Injected wildcards and boolean clauses turn a lookup into an enumeration or into a condition that is always true.
Concretely, escape the reserved characters per the filter encoding rules, build filters with a library that separates structure from values, bind against the returned distinguished name rather than concatenating credentials into a filter, restrict the search base and returned attributes, and reject input failing a length and character check first.
The reason for that specificity is a failure I have seen: An unescaped wildcard in the username field returned the first matching entry for any prefix. An attacker enumerated 5,600 accounts in 40 minutes and authenticated as one whose password appeared in a public breach list.
Input handling in the search filter.
| Input | Unescaped result | Escaped result |
|---|---|---|
| a single wildcard character | 5,600 entries returned | 0 entries |
| a closing clause plus a new one | always-true condition | literal search |
| an ordinary username | 1 entry | 1 entry |
I would not consider it settled without evidence: submit filter metacharacters in a sign-in field and confirm the resulting query is a literal search for that string.
Escape the syntax; do not trust the shape of the input.
Curated: · Written: · Reviewed:
QA-94Why should you not key application accounts on the distinguished name?(show answer)
I would answer DN versus immutable directory object ID by separating authentication, authorization, and lifecycle into three control planes.
A distinguished name encodes the object's position in the tree, so it changes whenever the object moves between organisational units or is renamed. The immutable object identifier is the value that survives the move.
Concretely, store the immutable identifier as the join key, keep the distinguished name as a display and search convenience refreshed on each synchronisation, handle rename and move events by updating the name against the stable key, and refuse a synchronisation that would create a new account because a name changed.
The reason for that specificity is a failure I have seen: A restructure moved 1,900 people between organisational units. Applications keyed on distinguished name treated every one as a new person, creating duplicates and orphaning 12 years of records, and reconciliation took 4 weeks.
Identifier behaviour during the restructure.
| Identifier | Changed for 1,900 people | Safe as a join key |
|---|---|---|
| distinguished name | yes, all of them | no |
| account name | 40 changes | no |
| immutable object identifier | 0 | yes |
I would not consider it settled without evidence: move a test object between organisational units and confirm the application still resolves the same account.
Position changes; identity should not.
Curated: · Written: · Reviewed:
QA-95How do you answer who can reach this share when groups are nested?(show answer)
The engineering content of nested group effective membership is the revocation path and its measured latency, not the protocol name.
Effective membership is the transitive closure of the nesting graph, and nobody can compute it by reading a group's direct members. Nesting also crosses ownership boundaries, so a group owner can add people to an access set they cannot see.
Concretely, resolve membership transitively and store the closure, expose an effective-members view to the resource owner, cap nesting depth, alert when a group used for access gains a nested group with a different owner, and check for cycles.
The reason for that specificity is a failure I have seen: A share granted to one group was reachable through 4 levels of nesting by 2,700 people rather than the 60 the owner believed. The path ran through a self-service group anyone could join, and the depth surfaced only during a data-loss investigation.
Effective membership resolution.
| Depth resolved | Groups included | People reachable |
|---|---|---|
| direct only | 1 | 60 |
| two levels | 7 | 480 |
| full closure at four levels | 23 | 2,700 |
I would not consider it settled without evidence: compute the transitive membership of the access group and compare it against the owner's stated expectation.
Direct members are a fragment of the answer.
Curated: · Written: · Reviewed:
QA-96A user presents a valid service ticket. Are they authorised?(show answer)
Before calling Kerberos service ticket versus authorization done I would write down the grant that the source-of-truth change does not touch.
A service ticket proves the domain controller vouched for the principal to that service, and the authorisation data it carries is group membership captured when the ticket was issued. The service still applies its own access check, and the group data can be older than the ticket's freshness suggests.
Concretely, have the service evaluate its own access control list against the presented identity, keep ticket lifetimes short enough that stale group data cannot outlive the removal objective, and force reauthentication for privileged operations rather than trusting the authorisation data inside the ticket.
The reason for that specificity is a failure I have seen: A removal from a privileged group did not take effect while existing tickets lived. The person reached a file service for the remaining 10 hours of the service ticket, and every access was logged as authorised.
Ticket lifetimes and staleness.
| Artifact | Lifetime | Stale group data window |
|---|---|---|
| service ticket | 10 hours | up to 10 hours |
| ticket-granting ticket | renewable for 7 days | up to 7 days |
| revised policy | 4 hours, no renewal | up to 4 hours |
I would not consider it settled without evidence: remove a group membership, then attempt the action at the file service once a minute and record the elapsed time until it first refuses, rather than reading the directory row and calling it done.
The ticket says who; the service must still decide what.
Curated: · Written: · Reviewed:
QA-97A web tier needs to reach a database as the user. Which delegation do you configure?(show answer)
The first thing I would establish about constrained Kerberos delegation is which live path still works after the directory row looks right.
Unconstrained delegation hands the front-end service the user's ticket-granting ticket, which lets it impersonate that user anywhere in the domain. Constrained delegation limits impersonation to named services, and the resource-based form moves the trust decision to the owner of the target.
Concretely, use resource-based constrained delegation configured on the target so the back end names who may impersonate to it, restrict the permitted service principal names, mark sensitive accounts as not delegatable, and inventory any remaining unconstrained delegation as a finding with a removal date.
The reason for that specificity is a failure I have seen: One unconstrained web server cached the ticket of a domain administrator who browsed to it. The cached ticket was extracted and used against 3 other services, and the exposure had existed for the 2 years the server had been configured that way.
Delegation configuration in the domain.
| Mode | Hosts | Impersonation reach |
|---|---|---|
| unconstrained | 1 | any service in the domain |
| constrained | 26 | listed service principal names |
| resource-based constrained | 74 | decided by each target |
I would not consider it settled without evidence: list computer objects trusted for unconstrained delegation and require the count to be zero outside documented exceptions.
Name the services delegation may reach.
Curated: · Written: · Reviewed:
QA-98You disabled an Active Directory account. How long can they still reach file shares?(show answer)
I would start disabling AD versus ticket lifetime from the session, token, and grant that outlive the login, not from the login itself.
Disabling stops new tickets being issued; it does not invalidate tickets already in the client's cache. Access continues until those tickets expire, which is a property of policy rather than of the disable action.
Concretely, pair disablement with a password reset that blocks renewal, shorten ticket lifetimes for privileged accounts, force logoff of active sessions on the relevant hosts, and where the risk warrants it act on the endpoint to clear the cache rather than waiting for expiry.
The reason for that specificity is a failure I have seen: An urgent disablement at 16:20 was followed by continued access to two file servers for 7 hours and 30 minutes on cached service tickets, because the session was never terminated and the password was never reset.
Access after disablement.
| Action taken | First denial | Mechanism |
|---|---|---|
| disable only | 7 hours 30 minutes | ticket expiry |
| disable plus password reset | 4 hours 10 minutes | renewal blocked |
| disable, reset, and force logoff | 3 minutes | session terminated |
I would not consider it settled without evidence: disable a test account and attempt a share access every minute, recording when it first fails.
Disable stops issuance; expiry stops use.
Curated: · Written: · Reviewed:
QA-99You want to require directory signing and channel binding. How do you get there safely?(show answer)
This is an area where a successful federation and a current authorization decision are different events.
Enforcement breaks every client still binding without integrity protection, and those clients are usually appliances and scripts nobody owns. The work is the inventory and the migration; the setting change is the last five minutes.
Concretely, turn on the diagnostic logging that records unsigned and simple binds with their source addresses, build an owner list from those addresses, migrate clients in waves against a published date, enforce first on one domain controller as a canary, then across the fleet while keeping the log to catch stragglers.
The reason for that specificity is a failure I have seen: Enforcement was applied fleet-wide from a baseline with no inventory. 26 appliances including a badge system and a printing service lost directory authentication for 5 hours, and the rollback also reverted 4 unrelated settings.
Unsigned binds by source over one week.
| Source class | Binds observed | Owner identified |
|---|---|---|
| appliances | 41,000 | 19 of 26 |
| in-house scripts | 8,300 | all of them |
| unknown addresses | 640 | 3 |
I would not consider it settled without evidence: count unsigned binds per source in the directory event log and require zero for a full week before enforcing.
The inventory is the project; the toggle is the ceremony.
Curated: · Written: · Reviewed:
QA-100Your forest is compromised at domain controller level. What does recovery look like?(show answer)
My answer to Active Directory forest recovery isolation begins with the artifact that is actually being trusted: token, assertion, group, or session.
Once the directory database is untrusted, every credential inside it is untrusted, including machine accounts and the key that signs tickets. Recovery is a rebuild in isolation from known-good media with credential resets, not a restore that carries the compromise forward.
Concretely, restore the first controller from a backup predating the compromise onto an isolated network, seize the roles, reset the ticket-signing account twice with a gap longer than the maximum ticket lifetime, reset service and administrative credentials, rebuild the remaining controllers rather than restoring them, and reconnect only after validation.
The reason for that specificity is a failure I have seen: A restore reconnected to the production network before the ticket-signing account was reset. Forged tickets created before the incident were accepted again within 20 minutes, and starting over added 4 days to what became a 9-day outage.
Recovery sequence.
| Step | Network isolated | Elapsed |
|---|---|---|
| first controller from a clean backup | yes | day 1 |
| ticket-signing reset, first pass | yes | day 1 |
| ticket-signing reset, second pass | yes | day 2, after 11 hours |
| reconnect to the production network | no | day 5 |
I would not consider it settled without evidence: confirm the ticket-signing account has been reset twice with the required interval and that tickets issued before the incident no longer validate.
Restore into isolation, or restore the intruder.
Curated: · Written: · Reviewed:
