Skip to content
Tech Interview Prep home
Technical interview guide

DNS & DHCP

How devices get an IP address automatically, and how domain names get resolved into one.

Read
45 min
Practice MCQs
25
Interview QA
25
Edition
v4
Editorial status
Reviewed
Relevant for
Network Engineer

Scope: Current IETF DNS, negative caching, DNSSEC, DNS over TCP/TLS/HTTPS, DHCPv4/options/relay, DHCPv6, SLAAC, and mDNS references reviewed 2026-09-04.

Overview

Curated: · Written: · Reviewed:

Operate naming and addressing as cached distributed state

DNS and DHCP are the two systems where a change you make is not the state your users see. DNS maps names to typed records through a delegated hierarchy; DHCP hands out time-bounded host configuration. Both are distributed control systems whose results are cached everywhere and outlive the edit that produced them. When an interviewer asks you about either one, they are usually probing whether you understand that gap: an authoritative update or a scope edit is not complete when the server accepts it — clients, recursive resolvers, relays, secondaries, and applications all have to converge.

The mental model: resolution is a chain of caches

The question interviewers open with is "walk me through what happens when you type a URL." A weak answer defines DNS. A strong answer walks the chain with a decision at each hop:

  1. The application calls the OS stub resolver. If the OS or browser cache has a live entry, resolution ends here — no packet leaves the machine.
  2. The stub resolver sends a recursive query to its configured recursive resolver (typically over UDP/53, falling back to TCP on truncation).
  3. If the recursive resolver has a cached answer within TTL — positive or negative — it returns it. This is where most lookups actually terminate.
  4. On a miss, the recursive resolver does the iterative work: it asks a root server, gets a referral to the TLD nameservers, asks a TLD server, gets a referral to the zone's authoritative servers, asks one of those, and gets the answer.
  5. The resolver caches the answer for its TTL and returns it to the stub, which caches it again.

Every hop caches, so the answer a client gets can be up to TTL old at each layer, and the layers stack. This is also why "I changed the record and it's still wrong" is the single most common DNS incident: the authoritative zone was updated in a second; the caches take minutes to hours to agree.

Who does the work: recursion, forwarding, and where to place each

The roles differ by who performs the iterative walk:

  • Stub resolver: asks a recursive resolver, does no work itself.
  • Recursive resolver: performs the iterative walk (or forwards) and caches results.
  • Forwarder: a recursive resolver that, instead of walking from the root, passes selected or all queries to an upstream resolver. A conditional forwarder only forwards queries for a specific zone — e.g., send corp.example.internal to the internal resolver, everything else to the root. A stub zone is a variant that periodically fetches the zone's NS records and resolves from there.

The design question interviewers actually ask is "where would you place recursion in your design and why?" The trade-off: recursion at the edge (every site runs its own resolver) gives local cache hit rates and survives WAN loss but multiplies the places you must run, secure, and monitor. Centralized recursion with forwarders is easier to govern and monitor but makes the central resolver a single point of failure and adds RTT to every cache miss. Conditional forwarding is how you stitch internal and external namespaces together without making internal names leak to public resolvers.

Two rules that come up in follow-ups: never run an open recursive resolver (it will be used for amplification attacks), and remember that a resolver returning an address proves nothing about the destination — not that it's reachable, listening, authenticated, or authorized.

Records are typed contracts, not a glossary

Interviewers rarely ask "what is a CNAME." They ask "why does this lookup fail" or "why is this answer stale," and the answer is almost always a record-type or TTL interaction:

  • A / AAAA: name to address. A lookup fails mysteriously when the client only asks for A and the name only has AAAA (or vice versa) — address-family selection is the application's job.
  • CNAME: aliases one owner name to another. It cannot coexist with other data at the same owner — a CNAME at the zone apex is illegal because the apex must hold SOA and NS. This is why ALIAS/ANAME-style provider-specific records exist for apex pointing; they're a provider feature, not a standard.
  • PTR / reverse zones: in-addr.arpa and ip6.arpa delegations. Forward and reverse are independent zones; a correct A record with no matching PTR breaks protocols that verify it (some mail servers, Kerberos, logging tools).
  • NS / SOA: delegation and zone authority. SOA also carries the negative-caching TTL (the MINIMUM field, per RFC 2308) — a detail that decides the next question.
  • MX / SRV / TXT: service discovery and policy. TXT is where SPF, DKIM, and DMARC glue live; SRV adds priority, weight, and port for protocols like SIP and LDAP.

TTL and negative caching are where staleness questions land. TTL bounds how long a cache may reuse a record — it is not a refresh deadline and not a health signal. A migration done right: lower the TTL early enough for existing entries to age out (e.g., drop 3600s to 60s a full hour before the change), verify resolvers are serving the short TTL, make the change, keep old capacity for the tail of old-TTL entries, then raise it back. Negative caching is the trap: an NXDOMAIN or no-data answer is also cached, for the SOA MINIMUM time. Create a record that a resolver recently answered NXDOMAIN for, and clients keep getting "name does not exist" until that negative entry expires. There is no globally enforceable cache flush — flush commands exist per-resolver, not for the Internet.

Zone mechanics: delegation, glue, and transfers

The namespace is hierarchical: registrars establish delegations at a parent, and NS records identify the authoritative servers for each cut. Glue records break the dependency cycle when a nameserver's own name lives under the zone it serves — ns1.example.com can't be resolved by first asking example.com's nameservers, so the parent must carry the A record as glue. A lame delegation — an NS record pointing at a server that isn't actually authoritative for the zone — is a classic "why does resolution fail for some resolvers but not others" answer: resolvers with cached good glue keep working; others hit the lame path. Validate delegation from the parent's referral through each authoritative endpoint, over both UDP and TCP, and from multiple networks — a lame or partially broken delegation often only shows up from one vantage point, because different resolvers hold different cached glue.

Zone serving has its own mechanics:

  • Primary/secondary: secondaries load the zone by transfer — AXFR (full) or IXFR (incremental) — triggered by the SOA serial in a notify, authenticated with TSIG. Restrict AXFR to trusted secondaries; an open AXFR leaks your entire zone including internal-naming patterns.
  • Stub zones at the resolver side track a zone's NS set to shortcut referrals.
  • Forward vs. reverse zones are configured and delegated separately; owning example.com gives you no automatic control of 3.2.1.in-addr.arpa.
  • Split-horizon/views return different data by client subnet — useful for private endpoints, but it multiplies cache, VPN, resolver-selection, and troubleshooting complexity, and "view leak" (internal answers escaping to external clients) is a real security failure mode.

DNS commonly starts over UDP, but responses that don't fit must fall back to TCP — a resolver or firewall that blocks TCP/53 breaks large answers and DNSSEC-signed responses. Keep responses under the path MTU; the widely recommended UDP payload size is 1232 bytes, because it fits inside a typical 1500-byte MTU without fragmentation. EDNS advertises a buffer size but does not change the real path MTU.

On transport security: DoT and DoH protect the client-to-resolver hop from network observers; they do not make an unsigned answer authentic. DNSSEC does that — it signs DNS data and denial-of-existence proofs via a chain of trust from the root, giving origin authentication and integrity, not confidentiality or authorization. Key rollovers that skip the double-signature or pre-publish window are a standard outage story.

Name semantics at the application edge

Before a query even reaches the hierarchy, application-level naming semantics decide whether two parties are talking about the same name:

  • Search suffixes: a bare hostname db may expand to db.prod.example.com on one host and db.dev.example.com on another. Same query, different answer — and a classic source of "works on my machine."
  • Trailing dots: example.com. is a fully qualified name; example.com may be expanded by search suffixes. A missing or extra dot changes which name is actually queried.
  • IDNA and normalization: internationalized names must be punycode-encoded (xn--…) before hitting DNS; two visually identical names can be different code points. Case-insensitivity applies to ASCII labels, and different libraries normalize differently.
  • DNS rebinding controls: a public name can resolve to a public IP on first lookup and a private RFC 1918 address on the second, letting a browser treat an attacker-controlled page as trusted internal origin. Rebinding protections (pinning resolved addresses, blocking public names that resolve to private ranges) live in browsers and resolvers, not in DNS itself.
  • mDNS: link-local multicast resolution, especially for .local, for printer and device discovery on one segment. It is not a substitute for governed unicast DNS across routed enterprise domains — .local collisions between mDNS and an internal unicast zone are a recurring support headache.

DHCP: the lease lifecycle and its failure modes

DHCPv4 follows DORA: Discover, Offer, Request, Acknowledge. The client broadcasts because it has no configuration yet; the server offers from a scope; the client formally requests one offer; the server acknowledges and records the lease.

The lifecycle thresholds are fixed fractions of the lease time (RFC 2131): at T1 = 50% of the lease the client tries to renew with its original server (unicast); at T2 = 87.5% it enters rebinding, broadcasting to any server that will extend it; at 100% the lease expires and the address must be released. These are lifecycle thresholds, not arbitrary retry timers.

The failure modes interviewers probe:

  • No response at all → the client self-assigns an APIPA address in 169.254.0.0/16. Seeing a 169.254 address on a host means "DHCP is broken," not "the network assigned this."
  • NAK → the server refuses a renewal (wrong scope, lease database lost, client moved subnets) and the client restarts from Discover.
  • Lease exhaustion → the pool is empty; new clients get nothing while existing leases ride until expiry. Pool utilization monitoring is the fix.
  • Relay across subnets → broadcasts don't cross routers, so a DHCP relay (Cisco's ip helper-address) forwards requests to the server and stamps gateway/interface context (option 82 territory) so the server picks the right scope. Without a relay, a server only serves its own subnet.
  • Options → router, DNS servers, domain name, MTU, and vendor options like 66 (TFTP server for PXE) must be consistent and fit the packet; option overload exists precisely because option space is tight.

Allocation can be dynamic, automatic, or reserved. A reservation is an operational assignment keyed on client identity — and a MAC address is spoofable, so a reservation is not authentication.

The server-side failure modes are where designs actually break:

  • Conflict detection: before offering, the server should probe the candidate address (ARP/ICMP echo) and clients should probe too; a client that finds the address in use sends DHCPDECLINE and the server must mark the address conflicted and exclude it. Without this, a statically configured host silently collides with a leased one.
  • Lease database durability: the lease database must survive restart and be backed up. A server that loses its database and re-offers from scratch hands out addresses still in use — the NAK storm and duplicate-address fallout follow.
  • High-availability ownership and clocks: failover pairs must agree on who owns the pool and on time. Lease expiry and T1/T2 math depend on clocks; unsynchronized servers can both believe they own the same pool and allocate the same address to two clients. During failover or restore, preserve lease state and prevent duplicate allocation before optimizing availability — a duplicate address is a far worse failure than a slow one.

The relay is a trust boundary: accept relay-agent information only on trusted ports, and use DHCP snooping on switches to restrict server responses to trusted ports and build binding tables — but validate its fail-open/fail-closed behavior and its interaction with topology changes.

IPv6 changes the shape of the problem

IPv6 splits configuration across two mechanisms: Router Advertisements supply prefixes and default-router info for SLAAC, while DHCPv6 can supply stateful addresses, prefix delegation, or other configuration. The RA M/O flags hint at what clients should do but don't eliminate platform differences — and DHCPv6 does not normally supply the default gateway; the RA does. Duplicate Address Detection, preferred/valid lifetimes, temporary vs. stable addresses, and prefix renumbering all behave differently across client OSes, so test SLAAC-only, DHCPv6, and mixed modes with the actual platforms you run.

What interviewers probe, and what a weak answer sounds like

The follow-ups cluster around convergence and trust:

  • "I changed the DNS record an hour ago and some clients still hit the old server — why?" Weak answer: "DNS is slow." Strong answer: which caches hold the old entry, what the TTL was before and after the change, whether a negative cache entry predates the record, and how you'd verify from a specific resolver.
  • "Where does recursion live in your design?" Weak answer: "on the DNS server." Strong answer: edge vs. central trade-offs, forwarders and conditional forwarders for internal namespaces, and why open recursion is unacceptable.
  • "Walk me through a DHCP failure where a client has a 169.254 address." Weak answer: "DHCP is broken." Strong answer: relay path, scope exhaustion, rogue server, snooping config, and how you'd isolate which hop dropped the DORA exchange.
  • "What breaks when DNSSEC keys roll over?" Weak answer: "nothing if you're careful." Strong answer: the pre-publish/double-sign window, validators caching the old chain, and why you test with a validating resolver before the cut.

The operational checklist underneath all of it: restrict zone transfers and dynamic updates, protect registrar and provider accounts, rate-limit abuse, monitor for unexpected records and delegations, watch DHCP pool utilization and rogue offers, and test the failure scenarios before they test you — delegation loss, stale positive and negative caches, TCP fallback, view leaks, relay loss, duplicate addresses, and lease database restore.

Worked trace: one lookup, end to end

A host with a cold cache resolves api.example.com (assume a fresh zone, TTL 300s, no DNSSEC on this path):

  1. Stub resolver misses locally, sends a recursive query for api.example.com A to the recursive resolver.
  2. Resolver cache misses for example.com too, so it queries a root server. Root doesn't know the answer but returns a referral: .com NS records plus their glue A records. Resolver caches the .com NS set for the TTL the root zone sets on those records (at the time of writing, the root servers set a 172800-second — 2-day — TTL on .com NS/glue; check with dig +norecurse @a.root-servers.net com NS rather than assuming, since these values do change).
  3. Resolver asks a .com server, which returns a referral: example.com NS records and glue for those nameservers. Cached again.
  4. Resolver asks example.com's authoritative server, which returns the A record with TTL 300.
  5. Resolver caches the A record for 300 seconds and returns it to the stub, which caches it for its own policy (often capped by the OS).

Now delete the record at the authoritative server. For up to 300 seconds, resolvers that already hold the A record keep answering it — the record is gone authoritatively but alive in caches. A resolver that queries after the deletion gets NXDOMAIN, cached for the SOA MINIMUM time. Both directions of staleness come from the same mechanism, and neither can be flushed globally. That asymmetry — server state changes instantly, world state converges on TTL — is the one idea to carry into every DNS and DHCP answer you give.