Overview
Curated: · Written: · Reviewed:
Design the full authenticator lifecycle around the threats that matter
When an interviewer asks you to design MFA, they are not asking for a list of methods. They are testing whether you start from the threat model, rank methods by the attacks they actually stop, and treat enrollment, recovery, and step-up as part of the same system. The single most common failure in these interviews is designing a beautiful login flow and then letting a helpdesk call or a recovery email quietly bypass it.
Start from the threat model, not the factor count
Name the attacks MFA is supposed to stop before you name a method:
- Password theft and credential stuffing. Any second factor stops these, which is why even SMS beats password-only.
- Phishing that captures one factor. A phishing page collects the password; whether the second factor survives depends entirely on the method.
- SIM-swap account takeover. The attacker moves the victim's number to their own SIM and receives the OTP directly.
Then name what MFA does not stop, because interviewers probe for this:
- Post-authentication token theft. A stolen session cookie works regardless of how strong the login was. MFA is a login control, not a session control.
- Consent phishing. A user legitimately authenticates to a malicious OAuth app and grants it access. No factor was defeated.
- Insider abuse and a compromised endpoint. Malware on the user's device observes or initiates actions after authentication succeeds.
A weak answer jumps straight to "we'll use TOTP" without this list. A strong answer opens with "what are we defending against?" and lets the threat model pick the method.
Factor independence is the real unit of strength
Multi-factor means proving control of independent factor types: something known, possessed, or inherent. Two passwords are not two factors. A device PIN or biometric that locally unlocks a cryptographic key can activate a multi-factor authenticator, but the server never receives a biometric and should not store one.
Independence is where designs quietly fail. If one malware event, email compromise, or support workflow defeats every factor, the label "MFA" overstates the protection. The classic break: the account's recovery email is the same email that could be the second factor, or the phone number receives both the SMS OTP and the password-reset link. The attacker compromises one channel and owns the account. Interviewers love this question — "walk me through how you'd take over this account" — and the answer is almost always a shared channel, not a broken algorithm.
When you present a design, audit each factor's compromise path and show they don't overlap.
Rank methods by real-world resistance
Phishing resistance is the most important distinction among common methods, and being able to rank them with reasons is a core interview skill.
SMS and voice codes depend on the public switched telephone network and face SS7 interception, SIM swap, carrier social engineering, number reassignment, and availability problems. They stop credential stuffing; they do not stop a determined attacker who targets the number. Treat them as a restricted fallback where risk permits, never the strongest factor.
TOTP (RFC 6238) is better — the secret never travels at login — but the code is typed by a human, so a phishing proxy can relay it to the real service in real time within the 30-second window. Screen capture and code capture work too. Useful as a transitional or fallback method; not phishing-resistant.
Push approval fails through fatigue: an attacker triggers prompt after prompt at 2 a.m. until the user taps approve. Number matching reduces blind approval, but a phishing proxy can relay the number as well, so it is not equivalent to verifier-bound WebAuthn. Push needs transaction context, rate limits, and explicit user intent — and after denials or repeated prompts, stop the flow, notify the account owner, and investigate instead of sending more pushes.
WebAuthn/FIDO2 and passkeys are the top of the ranking. The authenticator signs a fresh challenge scoped to the relying-party ID and authenticated origin, so an impostor site cannot reuse the response at the legitimate site. Prefer phishing-resistant cryptographic authenticators for administrators, high-impact actions, and as an option for all users.
Certificate-based and platform authenticators (smart cards, device TPMs) sit alongside WebAuthn at the high end; device-bound non-exportable hardware supports the highest assurance but creates enrollment and break-glass obligations.
Passkeys may be device-bound or synced through a provider. Synced credentials improve usability and recovery but expand trust to the sync fabric and its account recovery. Select policy from risk and assurance, not from product names.
A weak answer says "TOTP is secure because the secret is encrypted." A strong answer says "TOTP is relayable in real time; here's the attack, and here's when I'd still accept it."
Map assurance levels to what each buys
Anchor the design in NIST SP 800-63B's authenticator assurance levels rather than adjectives:
- AAL1 — single-factor authentication; a memorized secret qualifies. Buys almost nothing against phishing.
- AAL2 — two different factors, or a single factor plus a verifier-impersonation-resistant property. TOTP, push, and WebAuthn with user presence all reach AAL2. This is the sensible floor for most consumer and workforce accounts.
- AAL3 — requires a hardware-based authenticator with verifier-impersonation resistance: the attacker who collects every transmitted value still cannot authenticate. WebAuthn with a hardware-backed, user-verification-capable authenticator is the standard path.
The interview-ready move is assigning levels to actions, not just logins. Key rotation, recovery flows, wire transfers, privilege changes, and admin console access demand phishing-resistant factors at the highest level, because those are exactly the actions an attacker who has partially compromised the account wants to perform. If a stolen password plus a relayed SMS can rotate the MFA credentials, your AAL3 login is decoration.
WebAuthn mechanics: what the server actually validates
Registration begins only inside a recent authorized session. The server creates a fresh unpredictable challenge and sends the intended RP ID, user identity, supported algorithms, authenticator selection, and attestation policy. It validates challenge, origin, RP ID hash, type, user presence, user verification when required, algorithm, credential ID, and public key before storing the credential. Attestation should be required only when a concrete device-management or assurance policy justifies the privacy and operational cost.
Authentication repeats the contextual checks: one-time challenge, expected origin and RP ID, credential ownership, signature, flags, and sign counter where meaningful. Counters can be absent or unreliable for multi-device credentials, so a counter anomaly is a risk signal rather than automatic proof of cloning. Backup eligibility and backup state describe whether a credential may be or has been synced; they do not replace verification of user presence or user verification.
Expect the follow-up: "what breaks if the RP ID is misconfigured?" If the RP ID is broader than the service (say, the parent domain), a sibling site can request credentials for it. If it's narrower, login breaks across subdomains. Knowing that the RP ID is the security boundary — not the URL — is what separates a candidate who has shipped WebAuthn from one who has read about it.
Step-up and adaptive authentication as a design decision
Risk-adaptive authentication requests step-up for a new device, impossible travel, an impossible-time login (a successful authentication at 3 a.m. from a user who never works those hours), an unmanaged device, a sensitive action, or unusual behavior. Interviewers want the architecture, not the buzzword:
- Where the policy engine sits. Typically an authorization/risk service consulted at login and at sensitive-action time, returning a decision: allow, challenge, or deny.
- The evaluate-then-challenge flow. Evaluate risk signals against a declared baseline, then issue a step-up bound to the authenticated user, session, intended action, and a short validity window. The challenge must be bound to the transaction — an unbound "re-enter your password" adds little.
- The cost trade-off. Challenging too often trains users to approve mechanically (the push-fatigue problem, recreated by your own policy); challenging too rarely leaves the baseline unenforced. Measure step-up success and abandonment by cohort and tune.
Risk signals are noisy and spoofable. They supplement a declared baseline; they should not silently waive required MFA. Avoid permanent device trust that becomes an unrevoked possession factor stored in a weak cookie.
Enrollment, recovery, and lifecycle are authentication operations
Recovery is an alternate authentication path, not an exception to the policy — this is the sentence most candidates never say, and the one interviewers are waiting for. Adding a factor should require a recent session at the appropriate assurance, explicit confirmation, and notification over an independently established channel. Removing or replacing the last strong authenticator may require step-up, a delay, or human proofing. Recovery cannot be a secret question or helpdesk conversation that silently bypasses the stronger login. Recovery codes should be high entropy, hashed like credentials, one-time, downloadable once, rate-limited, and regenerated after use or suspected exposure.
For OTP specifically: generate a high-entropy per-account secret, display it only during protected enrollment, store it encrypted with narrowly controlled key access, never log it, validate with a small clock window, reject reuse within the accepted interval, throttle attempts, and offer a clear clock-recovery path.
Operate MFA as inventory and lifecycle. Let users view credential type, nickname, creation, last use, and revoke safely. Provide accessible alternatives and support shared or replaced devices without forcing insecure sharing. Provide multiple authenticators so one lost device does not force an intentionally weak recovery path. Notify on enrollment, reset, recovery, and high-risk changes. Audit outcomes and policy versions without storing secrets, OTP values, biometric data, or full sensitive device fingerprints.
What interviewers probe, and how to close
Likely follow-ups, in rough order of frequency:
- "How would you take over this account?" — Walk every path: login, recovery, helpdesk, sync provider, session theft. The weak answer only audits the login.
- "Why is TOTP not phishing-resistant if the secret never leaves the phone?" — Real-time relay through a proxy; the user types the code into the impostor site.
- "What happens when a user loses their only device?" — Break-glass design: recovery codes, a second enrolled authenticator, or a delayed, proofed recovery — never a silent downgrade.
- "How do you roll this out without locking out half the users?" — Phased enrollment, accessible alternatives, per-cohort abandonment metrics.
- "Does biometric MFA mean the server stores my fingerprint?" — No; the biometric locally unlocks the private-key operation.
Measure effective protection, not enrollment counts: phishing-resistant enrollment and use, accounts with two independent recovery-capable authenticators, fallback and recovery rates, time to revoke lost devices, suspicious prompt denials, reset takeover, support overrides, authentication failure and abandonment by cohort, and successful high-risk step-up. Treat a high MFA prompt-success rate as a usability signal, not as phishing resistance: typed TOTP can succeed at 99 percent while remaining relayable.
Test replay, relay, origin/RP mismatch, challenge reuse, user-verification downgrade, OTP reuse/window abuse, push flooding, SIM change, lost device, recovery-code guessing, session theft, concurrent factor changes, and notification failure. The strongest login is irrelevant if enrollment, recovery, or session management offers an easier path around it — and closing that sentence with a concrete example from your own design is what turns a good interview answer into a staff-level one.
