Skip to content
Tech Interview Prep home

Top 100 Network Engineer Interview Questions and Answers

The questions most likely to actually come up in your Network Engineer interview, ranked by likelihood — with detailed, senior-level answers covering what an interviewer is really listening for.

Curated: · Written: · Reviewed:

Reviewed 64Review pending 36
QA-1A routing table shows 10.0.0.0/8 via OSPF and 10.4.0.0/16 via BGP. Where does 10.4.1.7 go, and why is administrative distance the wrong first answer?(show answer)

The first thing I would establish about longest-prefix match is which forwarding entry the packet actually hits.

Forwarding selects the longest matching destination prefix first. Preference, metric, and administrative distance decide only among candidates for the same prefix length, so a /16 wins over a /8 even when the /8 is from a more-trusted source.

Concretely, list every destination match for the packet, pick the greatest prefix length, then inspect that winner's next hop, recursion, and FIB adjacency, and refuse any diagnosis that compared OSPF distance with BGP distance across different lengths.

The reason for that specificity is a failure I have seen: An operator withdrew the OSPF default after seeing 10.0.0.0/8 in the table and left a stale BGP /16 in place. Traffic to 10.4.1.7 followed the /16 for 18 minutes, and 47,200 packets hit a dead next hop while the /8 remained installed.

Lookup for 10.4.1.7 against two overlapping prefixes.

PrefixSourceAdministrative distanceMatch lengthPackets forwarded
10.0.0.0/8OSPF11080
10.4.0.0/16BGP201647,200
time until the /16 was withdrawn———18 minutes

I would not consider it settled without evidence: run an exact-prefix lookup for 10.4.1.7 on the forwarding table and record the winning length and next hop, not the protocol of the less-specific.

Specificity is evaluated before preference.

Curated: · Written: · Reviewed:

QA-2show ip route lists a prefix. Why is that not proof the packet will be forwarded there?(show answer)

I would start RIB versus FIB from the packet and the return path, not from the protocol adjacency.

The RIB holds selected control-plane candidates. The FIB is the programmed forwarding structure the ASIC or forwarding process actually uses, and it can lag, omit, or differ after recursion, hardware capacity, or line-card programming fails.

Concretely, compare the selected RIB prefix against a platform forwarding lookup and hardware adjacency, count RIB versus FIB entries, and treat a missing FIB hit as a programming fault rather than as a routing-protocol adjacency fault.

The reason for that specificity is a failure I have seen: A line card accepted 14,112 BGP prefixes into the RIB and programmed 12,004 into the FIB. The remaining 2,108 prefixes black-holed for 3 hours 20 minutes while every BGP session stayed Established.

Control-plane count versus programmed forwarding.

TableEntriesDestinations that forwarded
RIB (selected BGP)14,112not measured here
FIB (hardware)12,00412,004
missing from FIB2,1080 for 3 hours 20 minutes

I would not consider it settled without evidence: issue a hardware exact-route lookup for the failing destination and require a programmed adjacency, not only a RIB row.

A RIB row is a candidate; a FIB hit is forwarding.

Curated: · Written: · Reviewed:

QA-3The VLAN is saturated and traceroute looks clean. How do you tell a Layer 2 loop from an IP forwarding loop?(show answer)

This is an area where a green session and a working path are different observations.

An Ethernet frame has no hop limit, so a Layer 2 loop multiplies frames until the broadcast domain saturates. An IP forwarding loop decrements TTL and dies at a finite hop count, which shows as alternating traceroute hops rather than as a storm.

Concretely, measure unknown-unicast and broadcast rates per VLAN, inspect MAC move counters, and look for TTL-expired ICMP; treat a storm with zero TTL expiry as Layer 2, and treat alternating L3 hops with decaying TTL as a routing loop.

The reason for that specificity is a failure I have seen: A redundant access cable without spanning tree filled VLAN 40 to 98 percent of a 1.2 Gbps uplink in 4 minutes. Traceroute to the gateway still completed in 1 hop, and 0 TTL-expired messages were logged because the frames never left Layer 2.

Two loops, two counters.

ObservationLayer 2 stormIP TTL loop
VLAN utilisation98 percent of 1.2 Gbps4 percent
minutes to saturation4not applicable
TTL-expired ICMP01,900 per minute
traceroute1 hop to gatewayalternating 2 hops

I would not consider it settled without evidence: compare VLAN flood counters with ICMP time-exceeded counters on the same interval and require the diagnosis to name which counter actually moved.

Ethernet does not age a looping frame; IP TTL does.

Curated: · Written: · Reviewed:

QA-4Spanning tree elected a closet switch as root. What traffic path did you just accept?(show answer)

My answer to STP root placement begins with which plane is making the decision: control, forwarding, or policy.

The root is the reference for every path cost in the VLAN. A root in an access closet forces core-to-core traffic through that closet, adding hops and concentrating failure on a device that was never sized for it.

Concretely, set bridge priority so the intended core pair are primary and secondary root, verify root port and designated roles after convergence, and refuse any topology whose shortest path to the root leaves the distribution layer.

The reason for that specificity is a failure I have seen: A new access switch kept the default priority 32768 matching the core, won on MAC address, and increased the path from 2 hops to 11 over a 1 Gbps uplink. East-west latency rose from 6 ms to 41 ms and stayed there for 2 days until someone listed the root.

Path after the closet won root.

PathHopsLatencyDuration
designed, core as root26 msintended
after closet election1141 ms2 days
uplink speed at the accidental root1 Gbps——

I would not consider it settled without evidence: show spanning-tree root and the hop count from each distribution switch, and require the core pair to own root and secondary.

Root placement is the traffic map, not a cosmetic priority.

Curated: · Written: · Reviewed:

QA-5Both ends of a trunk are up and in trunk mode. Why can a VLAN still be silent across it?(show answer)

I would treat VLAN trunk allowed list as a claim about a specific five-tuple, VRF, and direction.

A trunk forwards only VLANs that exist, are allowed, and are in forwarding state on both sides. Line protocol up does not mean the allowed list, pruning, or STP state carried the VLAN.

Concretely, compare operational allowed VLANs, VLAN database membership, and per-VLAN STP state in both directions, then capture a tagged frame of that VLAN on the trunk rather than inferring carriage from the interface being up.

The reason for that specificity is a failure I have seen: VLAN 47 was pruned on the far end of an otherwise healthy LAG. 1,840 endpoints on that VLAN lost the gateway for 7 hours while the trunk stayed up and VLAN 10 continued to pass.

Two VLANs on one trunk.

VLANAllowed both endsSTP forwardingSilent endpointsHours until noticed
10yesyes0—
47no, pruned far endblocked1,8407

I would not consider it settled without evidence: capture an 802.1Q frame with VLAN 47 on the trunk and confirm the same VLAN is forwarding on both member switches.

An up trunk is not an allowed-VLAN list.

Curated: · Written: · Reviewed:

QA-6Untagged frames on a trunk are being learned in different VLANs at each end. What is leaking?(show answer)

The useful question for native VLAN mismatch is what a hash-diverse probe would show on the other members.

The native VLAN is the untagged broadcast domain on a trunk. If the two ends disagree, untagged control and data frames cross into a foreign VLAN, which is a segmentation failure, not a cosmetic numbering issue.

Concretely, set the native VLAN explicitly and identically, tag the native VLAN where the platform allows it, and refuse trunks whose untagged VLAN numbers differ even by one.

The reason for that specificity is a failure I have seen: One side used native VLAN 1 and the other native VLAN 99. 312 CDP and untagged user frames crossed the boundary over 14 days, and a workstation in VLAN 99 obtained a VLAN 1 address.

Untagged frames on a mismatched trunk.

EndNative VLANUntagged frames learnedDays until found
switch A131214
switch B9931214
workstations that received the wrong subnet—1—

I would not consider it settled without evidence: send a known untagged test frame and require it to be learned in the same VLAN identifier on both switches.

Disagreeing native VLANs are a VLAN merge.

Curated: · Written: · Reviewed:

QA-7A MAC address oscillates between two ports. What plane is sick?(show answer)

I would settle MAC flapping by capturing both sides of the encapsulation, not by reading the neighbour table.

A switch learns a source MAC on an ingress VLAN and port. Rapid moves mean either a Layer 2 loop, a duplicated adapter, or an unexpected active-active attachment, not a routing metric problem.

Concretely, record move rate, VLAN, and both ports, trace those ports through trunks, LAGs, and hypervisors, and contain flood first only if the domain is already saturating; then remove the second source or the loop.

The reason for that specificity is a failure I have seen: A mis-bundled dual-homed server presented one MAC on two access ports. The table recorded 8,400 moves per minute for 22 minutes, unknown-unicast flooded the VLAN, and the first ticket was filed as a gateway ARP fault.

One MAC, two ports.

PortMoves per minuteMinutes observedFirst ticket category
Gi1/0/124,20022gateway ARP
Gi1/0/184,20022gateway ARP
total8,40022wrong plane

I would not consider it settled without evidence: export MAC-move timestamps for that address and require a single stable port before declaring the VLAN healthy.

A flapping MAC is a Layer 2 topology statement.

Curated: · Written: · Reviewed:

QA-8The port-channel is up with four collecting-distributing members. Why is a quarter of the traffic still black-holing?(show answer)

The judgement in LACP member data-plane failure is which state the packet uses, not which state the operator intended.

A member that LACP suspends is normally removed from the aggregate's distribution set. But LACP can remain healthy when its control frames pass and ordinary payload traffic does not, so collecting-distributing state is not proof of the member's data path.

Concretely, compare actor and partner state, key, speed, and VLAN on every member, then probe with hash-diverse five-tuples and verify payload counters and loss per member rather than relying on LACP state or a single ping through the bundle.

The reason for that specificity is a failure I have seen: One member remained collecting and distributing while a data-plane fault dropped ordinary payloads. Hashing sent 25 percent of flows onto it for 90 minutes, and 1 in 4 application sessions failed while the port-channel showed up/up.

Four LACP members, one payload path failed.

MemberLACP stateHash shareSessions that failed
1collecting, distributing25 percent0 of 100
2collecting, distributing25 percent0 of 100
3collecting, distributing25 percent0 of 100
4collecting, distributing; payload fault25 percent100 of 100
minutes until the failed member was removed from hash——90

I would not consider it settled without evidence: send at least one flow per member hash bucket and require each member's TX/RX counters to move only on collecting-distributing ports.

Bundle line protocol is not the member distribution set.

Curated: · Written: · Reviewed:

QA-9The neighbour is FULL. Why are the prefixes you expected still missing?(show answer)

Where candidates lose the interview on OSPF adjacency versus advertised prefix is calling the tunnel or the BGP session the path.

FULL means the link-state database for that adjacency synchronized. It does not mean the desired prefix was originated, passed area policy, won preference, or was programmed. Adjacency is a transport for LSAs, not a list of routes.

Concretely, check the expected LSA type and advertising router, confirm area and redistribution policy, then look up the prefix in RIB and FIB; treat a missing Type-5 external, Type-3 summary, or intra-area prefix advertisement as a generation or policy fault.

The reason for that specificity is a failure I have seen: Two cores stayed FULL while redistribution of 64 static prefixes was left in a deny-all route-map. The prefixes were absent for 5 hours, and every health check that stopped at neighbour state reported green.

Adjacency versus origin.

CheckResultHours until the route-map was found
neighbour stateFULL5
Type-5 LSAs for the statics05
prefixes missing from RIB645

I would not consider it settled without evidence: show the specific LSA for the missing prefix and a FIB lookup, and require both, not the neighbour state machine.

FULL is database sync, not prefix presence.

Curated: · Written: · Reviewed:

QA-10You have three OSPF areas and no contiguous backbone. What paths are illegal?(show answer)

I would answer OSPF area 0 by separating reachability, policy, and observed delivery.

Non-backbone areas exchange inter-area routing through area 0. A disjoint backbone, or areas that only touch each other, cannot compute a correct inter-area path without a virtual link or a redesigned backbone.

Concretely, draw area 0 so every ABR has a resilient backbone link, refuse area-to-area shortcuts that bypass it, and test ABR and backbone-link loss before calling the area design done.

The reason for that specificity is a failure I have seen: Sites A and B were placed in area 1 and 2 with an extra link between them and no area-0 path after a core failure. Inter-area prefixes disappeared for 40 minutes, and 3 areas partitioned even though intra-area neighbours stayed FULL.

Partition after the backbone link failed.

PathAreas involvedPrefixes retainedMinutes of partition
designed via area 03all inter-area0
leftover A-to-B extra link1 and 2 onlyintra-area only40

I would not consider it settled without evidence: fail the backbone path in a change window and require inter-area prefixes to remain installed through the remaining area-0 ABR.

Area 0 is the inter-area spine, not a label.

Curated: · Written: · Reviewed:

QA-11The eBGP session has been Established for two hours. Why might you still be advertising nothing useful?(show answer)

The engineering content of BGP Established versus advertised is the return path and the more-specific prefix, not the protocol name.

Established means the TCP session and OPEN completed. Import, export, address-family activation, and outbound policy decide which NLRI actually leave. A quiet session can be a policy deny, not a healthy default.

Concretely, compare advertised and received prefix counts per address family with the intended policy, and require an explicit outbound permit rather than treating session state as an advertisement.

The reason for that specificity is a failure I have seen: A new peering stayed Established while the IPv4 unicast family was never activated. 0 prefixes were advertised for 2 hours, and 1,100 customer prefixes stayed inside the AS until someone counted Adj-RIB-Out.

Session versus NLRI.

ObservationCountDuration
BGP stateEstablished2 hours
address families activated02 hours
prefixes in Adj-RIB-Out02 hours
customer prefixes held in1,1002 hours

I would not consider it settled without evidence: show advertised routes to that neighbour and require a non-zero, policy-expected count.

Established is transport; Adj-RIB-Out is the advertisement.

Curated: · Written: · Reviewed:

QA-12An implementation still accepts and advertises all eBGP routes by default. What does RFC 8212 say you should do instead?(show answer)

Before calling RFC 8212 default-reject done I would write down the ECMP member, VLAN, or VRF nobody hashed onto.

RFC 8212 requires default-reject import and export on eBGP sessions. Without explicit policy, a new neighbour is a full-table leak and a full-table advertisement.

Concretely, configure inbound and outbound policy that permits only the intended prefixes and AS paths, set a max-prefix, and refuse any eBGP neighbour whose policy is implicit-allow.

The reason for that specificity is a failure I have seen: A lab router with implicit-allow eBGP accepted 84,000 prefixes from a transit test peer and advertised them internally for 11 minutes, attracting production traffic onto a 100 Mbps lab link.

Implicit-allow versus default-reject.

PolicyPrefixes acceptedMinutes until withdrawnLink
implicit allow84,00011100 Mbps lab
RFC 8212 default-reject0——

I would not consider it settled without evidence: bring up an eBGP session with no policy and require 0 prefixes accepted and 0 advertised.

eBGP without policy is a leak waiting for a neighbour.

Curated: · Written: · Reviewed:

QA-13You need your AS's outbound traffic to prefer one exit. Why is MED the wrong first knob, and what does local preference actually decide?(show answer)

The first thing I would establish about BGP local preference versus MED is which forwarding entry the packet actually hits.

Local preference is compared inside the AS and typically wins early in the decision process, so it selects the exit for outbound traffic. MED is a hint to an external neighbour and is compared later, often only among routes from the same AS, so it does not steer your own outbound path.

Concretely, set local preference on inbound advertisements to choose your outbound exit, use MED or communities only as a request to a neighbour, and never treat a lower MED as beating a higher local preference inside your AS.

The reason for that specificity is a failure I have seen: A team lowered MED on the primary exit and left local preference equal. Internal routers kept the secondary exit because of IGP metric, adding 180 ms for 6 hours, while the neighbour never compared the MED at all.

Same prefix, two attributes.

Attribute set on primaryInternal winnerExtra latencyHours
MED 50, local pref 100secondary (IGP)180 ms6
local pref 200, MED unchangedprimary0 ms—

I would not consider it settled without evidence: show the BGP decision trace for a prefix and require local preference, not MED, to be the first differing attribute among internal candidates.

Local preference chooses your exit; MED asks a neighbour.

Curated: · Written: · Reviewed:

QA-14One path has two AS hops and another has four. Why might the four-hop path still be the right one?(show answer)

I would start AS_PATH is not cable distance from the packet and the return path, not from the protocol adjacency.

AS_PATH length is a policy loop-prevention and preference signal, not geographic or latency distance. A short AS_PATH can traverse a congested or distant cable; a longer path can be the close peering.

Concretely, measure RTT and loss on the forwarding path you actually install, and if latency matters, set local preference or communities from that measurement rather than from AS hop count.

The reason for that specificity is a failure I have seen: Outbound policy preferred a 2-AS transit over a 4-AS peering. The 2-AS path added 91 ms versus 12 ms on the peering, and a voice cluster stayed on the long path for 3 days because the attribute looked shorter.

Path length versus delay.

PathAS hopsRTTDays selected
transit291 ms3
peering412 ms0 until policy change

I would not consider it settled without evidence: compare RTT of hash-pinned probes on each candidate next hop and install from that, not from AS_PATH length.

AS hops are not milliseconds.

Curated: · Written: · Reviewed:

QA-15A prefix is RPKI invalid. Do you drop it, and what happens if you still forward it?(show answer)

This is an area where a green session and a working path are different observations.

Route origin validation can mark a route valid, not found, or invalid against ROAs. Invalid means either the origin AS is unauthorized or the route is more specific than the ROA's maxLength permits; accepting it can admit a hijack or a legitimate route broken by stale ROA data. RFC 6811 leaves treatment to local policy, and not found is not the same as valid.

Concretely, make invalid routes ineligible under the documented import policy, keep not-found routes unless policy says otherwise, and monitor the invalid count so a stale ROA does not silently drop a legitimate more-specific.

The reason for that specificity is a failure I have seen: Invalids were accepted to "avoid risk." One hijacked /24 stayed in the FIB for 6 hours and 19,000 packets followed it, while the valid covering /22 remained unused because of longest match.

Hijack that won on length.

PrefixRPKIInstalledPackets
203.0.112.0/22validno, less specific0
203.0.113.0/24invalidyes for 6 hours19,000

I would not consider it settled without evidence: inject a test invalid in a lab VRF and require it to be discarded before RIB installation.

Invalid is not a softer unknown.

Curated: · Written: · Reviewed:

QA-16A customer re-advertises your transit table to another provider. What did they just become, and how do you limit the blast?(show answer)

My answer to route leak begins with which plane is making the decision: control, forwarding, or policy.

A route leak is an advertisement of prefixes beyond the intended customer/provider/peer relationship, often visible as an unexpected AS_PATH or as more-specifics that attract a continent. Peering policy and prefix filters exist to stop that role change.

Concretely, apply role-based export (customer may send only their prefixes), max-prefix, RPKI, and IRR filters, and treat a sudden more-specific from a customer as a leak until proven otherwise.

The reason for that specificity is a failure I have seen: A customer advertised 0.0.0.0/1 and 128.0.0.0/1 learned from another ISP. Four of our peers preferred those more-specifics for 37 minutes, and 2.1 Gbps of foreign traffic entered a 1 Gbps access circuit.

Leak of default-like more-specifics.

Prefix from customerIntendedPeers that preferred itMinutesTraffic on 1 Gbps access
0.0.0.0/1no4372.1 Gbps offered
128.0.0.0/1no437included above

I would not consider it settled without evidence: alarm on prefix count and unexpected default-like more-specifics from customers, and require filters that would have dropped both /1s.

A customer session is not a transit settlement.

Curated: · Written: · Reviewed:

QA-17OSPF and BGP redistribute into each other at two border routers. What can the prefixes do?(show answer)

I would treat mutual redistribution loop as a claim about a specific five-tuple, VRF, and direction.

Mutual redistribution without tagging or filters re-injects the same prefix with a new source, which can loop, inflate the table, and oscillate SPF or BGP churn. Each redistribution point is a new origin unless you mark and deny re-entry.

Concretely, tag or community-stamp on export, deny tagged routes on the reverse redistribution, prefer a single redistribution point, and count prefixes before and after.

The reason for that specificity is a failure I have seen: Two ABRs redistributed BGP into OSPF and OSPF into BGP with no tags. The LSDB gained 2,400 extra external prefixes, SPF ran 9 times per minute, and the loop lasted 50 minutes until one redistribution was shut.

Table growth during the loop.

MetricBeforeDuring the 50 minutes
OSPF external prefixes1202,520 (120 + 2,400)
SPF runs per minute0.19
redistribution points with tags0 of 20 of 2

I would not consider it settled without evidence: trace one prefix's path attributes and LSA advertising-router across both borders and require a tag that stops re-entry.

Untagged mutual redistribution is a prefix photocopier.

Curated: · Written: · Reviewed:

QA-18One ping across an eight-next-hop ECMP succeeds. What have you not measured?(show answer)

The useful question for ECMP one-ping fallacy is what a hash-diverse probe would show on the other members.

ECMP hashes a flow onto one member. A single five-tuple samples one bucket. A failed member, a polarized hash, or a one-way path on another member stays invisible until a different tuple is used.

Concretely, install the full next-hop set, generate many controlled tuples covering each member, and measure loss and latency per member rather than averaging one success.

The reason for that specificity is a failure I have seen: Operators signed off a new core after one ping. Three of eight members dropped 100 percent of hashed flows because of a unicast RPF failure. The outage lasted until 8 minutes of user traffic filled the other buckets.

Eight members, one ping.

MemberPing usedLossUser flows after 8 minutes
1–5not sampled0 percenthealthy
6–8not sampled100 percentall hashed flows failed
member 1 (the ping)1 tuple0 percentproved only itself

I would not consider it settled without evidence: map each next hop to at least one probe five-tuple and require all eight members to forward.

One flow is one hash bucket.

Curated: · Written: · Reviewed:

QA-19The SYN leaves through firewall A and the SYN-ACK returns through firewall B. What happens to the session?(show answer)

I would settle asymmetric routing through a stateful firewall by capturing both sides of the encapsulation, not by reading the neighbour table.

A stateful firewall forwards only packets that match created state. Asymmetry that splits the handshake across two unsynchronized devices looks like an unsolicited SYN-ACK and is dropped, even though IP routing itself allows asymmetry.

Concretely, keep both directions of a flow on the same firewall cluster, enable state sync only where it actually shares the table, or use external routing that pins the pair; never assume routing symmetry from a successful traceroute in one direction.

The reason for that specificity is a failure I have seen: After an ECMP change, SYN-ACKs hashed to the other firewall. 100 percent of new TCP sessions died for 8 minutes, and both firewalls showed healthy CPU while dropping 22,000 unsolicited SYN-ACKs.

Handshake split across two firewalls.

PacketFirewallState presentActionMinutes
SYNAcreatedallow8
SYN-ACKBnodrop 22,0008
new TCP sessions completing——0 percent8

I would not consider it settled without evidence: capture the SYN and SYN-ACK at both firewalls and require them to hit the same state table.

Statefulness makes asymmetry a drop, not a curiosity.

Curated: · Written: · Reviewed:

QA-20You enable strict uRPF on an edge with two ISPs. Whose legitimate inbound packets die?(show answer)

The judgement in uRPF strict versus loose is which state the packet uses, not which state the operator intended.

Strict uRPF requires the source to be reachable via the ingress interface. Loose uRPF requires the source to be reachable via any interface. Asymmetric multihoming fails strict and usually passes loose.

Concretely, use loose or feasible-paths uRPF on asymmetric edges, reserve strict for single-homed access where spoofing risk dominates, and test inbound from each provider with a source that returns via the other.

The reason for that specificity is a failure I have seen: Strict uRPF on both ISP interfaces dropped 31 percent of inbound customer replies for 4 hours because return routing preferred the other ISP, while spoofed packets were not the traffic that failed.

Two ISPs, strict mode.

ModeInbound droppedHoursCause
strict31 percent4source not via ingress
loose (after change)0 percent of legitimate—any-interface reachability

I would not consider it settled without evidence: send a packet whose source prefix is installed only on the other ISP interface and require strict to drop it and loose to accept it.

Strict uRPF assumes symmetric routing.

Curated: · Written: · Reviewed:

QA-21VRRP fails over in three seconds. Why are still-open firewall sessions dead?(show answer)

Where candidates lose the interview on VRRP versus session state is calling the tunnel or the BGP session the path.

VRRP moves a virtual IP and MAC. It does not copy NAT or firewall session tables. The new master answers ARP for the VIP and then drops packets that belong to state it never saw.

Concretely, place stateful devices in an active/standby pair that syncs sessions, or accept that failover resets sessions; do not quote VRRP advertisement timers as a session-preservation SLA.

The reason for that specificity is a failure I have seen: The VIP moved in 3 seconds. 12,600 NAT translations existed only on the old master, and every translated TCP session reset. Users saw a 3-second ICMP recovery and a 12-minute application rebuild.

VIP move versus NAT table.

ObjectOn old masterOn new master at +3 sApplication rebuild
virtual IPreleasedowned—
NAT sessions12,600012 minutes
ICMP to VIPfails 3 ssucceedsnot the app

I would not consider it settled without evidence: count session-table entries on the new master immediately after failover and require either sync or an explicit reset budget.

VRRP moves the address, not the flow table.

Curated: · Written: · Reviewed:

QA-22BFD is set to 50 ms times three. The fibre is fine. Why is OSPF still bouncing?(show answer)

I would answer BFD false positives by separating reachability, policy, and observed delivery.

BFD is a forwarding-path liveness detector, not an application health check. Aggressive timers trip on queueing, CoPP, or control-plane scheduling, tearing down clients that were still forwarding user packets.

Concretely, size interval and multiplier to the platform's control-plane budget, protect BFD in CoPP, and correlate flaps with CPU and drop counters before blaming the optical path.

The reason for that specificity is a failure I have seen: Timers of 50 ms × 3 produced 14 flaps per hour for 2 hours during a control-plane spike at 88 percent CPU. User packet loss on the same interfaces stayed at 0 percent; only protocol adjacencies died.

Liveness versus user frames.

SignalValue during the 2 hours
BFD interval × multiplier50 ms × 3
adjacency flaps per hour14
user packet loss0 percent
control-plane CPU88 percent

I would not consider it settled without evidence: show BFD drop and CoPP counters alongside interface error counters, and require optical errors before calling the path dead.

BFD can kill a healthy adjacency.

Curated: · Written: · Reviewed:

QA-23You advertise 10.8.0.0/16 while 10.8.4.0/24 is down. Where do packets for 10.8.4.10 go?(show answer)

The engineering content of summary black hole is the return path and the more-specific prefix, not the protocol name.

A summary asserts reachability for its entire range. If a covered more-specific fails and no discard or withdrawal policy exists, the aggregate keeps attracting traffic into a hole or a loop.

Concretely, install a discard route for the summary, advertise only when components justify it, and leak more-specifics or withdraw the aggregate when the last component dies.

The reason for that specificity is a failure I have seen: An ABR kept advertising the /16 after the site's only component /24 failed. All 256 addresses covered by that /24, including its live hosts, were black-holed for 1 hour 15 minutes because the aggregate was not withdrawn.

Covered prefix down, aggregate still out.

PrefixStateAttracts 10.8.4.10Duration
10.8.4.0/24downno, absent1 hour 15 minutes
10.8.0.0/16still advertisedyes1 hour 15 minutes
discard for 10.8.0.0/16missingloop or hole1 hour 15 minutes

I would not consider it settled without evidence: fail the component prefix in a lab and require either aggregate withdrawal or a discard that does not loop.

An aggregate is a reachability promise.

Curated: · Written: · Reviewed:

QA-24Hosts send 1500-byte inner frames into a VXLAN overlay. The underlay is 1500. What breaks, and where?(show answer)

Before calling VXLAN underlay MTU done I would write down the ECMP member, VLAN, or VRF nobody hashed onto.

VXLAN adds outer Ethernet, IP, UDP, and VXLAN headers (commonly 50 bytes, more with options). If the underlay MTU cannot carry inner 1500 plus overhead, large inner frames drop. Overlay control plane can stay up.

Concretely, raise underlay MTU or clamp inner MSS, account for the actual header stack, and test 1500-byte inner frames on every underlay ECMP member.

The reason for that specificity is a failure I have seen: Underlay MTU stayed 1500. Inner 1500-byte frames needed 1550 on the wire and dropped. Small pings worked for 3 hours while file copies failed; EVPN sessions remained Established.

Inner 1500 on a 1500 underlay.

FrameSize on underlayForwardedHours of silent large-drop
ICMP 64-byte inner~114yes3
inner 15001550no3
EVPN BGP state—Established3

I would not consider it settled without evidence: send a 1500-byte inner frame and capture the encapsulated size on the underlay, requiring the underlay MTU to be at least that size.

Overlay adjacency does not enlarge underlay MTU.

Curated: · Written: · Reviewed:

QA-25A Type-2 MAC/IP route is in the EVPN table. Why can the inner frame still die?(show answer)

The first thing I would establish about EVPN control plane versus data plane is which forwarding entry the packet actually hits.

EVPN advertises reachability. Encapsulation, VNI mapping, underlay next hop, and the actual VXLAN data path are separate. A Type-2 without a working VTEP data plane is a directory entry, not a delivered frame.

Concretely, resolve the next-hop VTEP, confirm VNI and encapsulation counters increment, and ping or capture inner traffic rather than stopping at the BGP EVPN table.

The reason for that specificity is a failure I have seen: Type-2 routes were present while a mis-set VNI dropped data-plane decapsulation. Inner frames produced 0 encap hits for 45 minutes; the control plane stayed full.

Type-2 present, data plane silent.

PlaneObservationMinutes
EVPN Type-2installed45
VXLAN encap hits for the VNI045
inner frames delivered045

I would not consider it settled without evidence: require VXLAN encap/decap counters to move for that VNI, not only a Type-2 show command.

EVPN is the directory; VXLAN counters are delivery.

Curated: · Written: · Reviewed:

QA-26EVPN looks healthy and one underlay ECMP member is lossy. Which users fail?(show answer)

I would start overlay versus underlay failure from the packet and the return path, not from the protocol adjacency.

Overlay sessions ride hashed underlay paths. Overlay control-plane health does not prove every underlay member. A single bad underlay next hop takes the overlay flows that hash onto it.

Concretely, measure underlay member loss with diverse hashes, map overlay VTEP pairs onto those members, and repair the underlay rather than restarting BGP EVPN.

The reason for that specificity is a failure I have seen: One of four underlay members dropped 12 percent of large packets. Overlay BGP stayed up. Flows that hashed onto that member failed for 6 hours while a VTEP ping that hashed elsewhere succeeded.

One underlay member, overlay still Established.

PathLoss on large packetsOverlay BGPHours
underlay members 1–30 percentEstablished6
underlay member 412 percentEstablished6
VTEP ping on members 1–30 percent—6

I would not consider it settled without evidence: probe each underlay member between VTEPs with large packets and require all members below a stated loss bound.

Overlay green is not underlay ECMP green.

Curated: · Written: · Reviewed:

QA-27A shared-services prefix appears in a tenant VRF you did not intend to join. What leaked?(show answer)

This is an area where a green session and a working path are different observations.

VRF isolation is the import and export of route targets, plus the routing table the packet is classified into. An extra import RT, a leaked static, or a global leak undoes the tenant boundary even when interfaces stay in the right VRF.

Concretely, inventory route targets per VRF, deny unexpected imports, and traceroute inside the tenant VRF to see whether the next hop belongs to another tenant.

The reason for that specificity is a failure I have seen: An extra import RT copied 17 prefixes from tenant A into tenant B. The leak lasted 9 days, and 1 billing query from B reached A's database because longest match preferred the leaked /32.

Unintended import.

VRFExtra import RTForeign prefixesDaysPackets that crossed
tenant Bone1791 documented query plus unknown

I would not consider it settled without evidence: show the tenant VRF RIB for foreign prefixes and require zero unless a documented leak with policy exists.

The VRF name on the interface is not the import policy.

Curated: · Written: · Reviewed:

QA-28The IPsec SA is up. Why is the interesting traffic still leaving in the clear?(show answer)

My answer to IPsec SA up versus selector begins with which plane is making the decision: control, forwarding, or policy.

An SA is a keyed channel for packets that match its traffic selectors. If the ACL or TS is wrong, packets miss the SA and follow the normal route, while the SA idles in the up state.

Concretely, compare the selector to the actual inner five-tuple, confirm encapsulating counters increment for that flow, and never treat IKE or SA up as proof of match.

The reason for that specificity is a failure I have seen: The SA covered 10.8.0.0/24 while the application used 10.8.1.0/24 inside 10.8.0.0/16. The SA stayed up with 0 matching bytes for 2 hours, and 880 MB left unencrypted.

Idle SA beside cleartext.

PrefixIn selectorApplication trafficSA bytesHours
10.8.0.0/24yesno02
10.8.1.0/24noyes, 880 MB02

I would not consider it settled without evidence: generate the application five-tuple and require IPsec encapsulating byte counters to rise by the payload size.

SA up is a channel; selectors are admission.

Curated: · Written: · Reviewed:

QA-29A GRE tunnel is up. What confidentiality did you just claim?(show answer)

I would treat tunnel versus encryption as a claim about a specific five-tuple, VRF, and direction.

GRE encapsulates; it does not encrypt. Confidentiality requires a crypto transform such as IPsec. A tunnel icon is reachability of the outer header, not secrecy of the inner.

Concretely, place GRE inside IPsec or use a crypto tunnel, and verify ESP or equivalent byte counters rather than GRE keepalive state.

The reason for that specificity is a failure I have seen: A "VPN" built as GRE-only carried 4.1 GB of inner cleartext for 1 day on a shared underlay; the tunnel stayed up and the ticket closed as encrypted.

GRE without a crypto transform.

StateValue
GRE tunnelup for 1 day
IPsec transformnone
inner bytes on the underlay4.1 GB readable

I would not consider it settled without evidence: capture the underlay and require the inner payload to be unreadable, not merely GRE-encapsulated.

Encapsulation is not confidentiality.

Curated: · Written: · Reviewed:

QA-30IKE comes up on UDP 500, then the path has a NAT. Why does ESP die until you move to 4500?(show answer)

The useful question for NAT-T is what a hash-diverse probe would show on the other members.

Ordinary port-mapped NAT cannot distinguish naked ESP flows because ESP has no transport ports; IPsec-aware pass-through is special handling. NAT-Traversal wraps ESP in UDP 4500 so a normal mapping can exist. Blocking 4500 after allowing 500 yields a control plane with no data plane.

Concretely, allow UDP 500 and 4500 plus the return mapping, detect NAT via the NAT-D payloads, and refuse designs that permit IKE only.

The reason for that specificity is a failure I have seen: A firewall allowed UDP 500 and denied 4500. 1,200 remote SAs reached IKE then passed 0 ESP for 70 minutes until NAT-T was unblocked.

Control allowed, data blocked.

Port or protocolPolicySAsData bytesMinutes
UDP 500allow1,200 IKE070
UDP 4500deny—070
ESP protocol 50n/a behind NAT—070

I would not consider it settled without evidence: place a NAT in the path and require ESP-in-UDP 4500 counters to increment.

IKE on 500 is not ESP through NAT.

Curated: · Written: · Reviewed:

QA-31Small packets through the VPN work. 1500-byte packets vanish. Where is the hole?(show answer)

I would settle VPN MTU black hole by capturing both sides of the encapsulation, not by reading the neighbour table.

IPsec and tunnel headers reduce effective inner MTU. If PMTUD is blocked and MSS is not clamped, large inner packets drop. The SA remaining up is expected.

Concretely, compute inner MTU from outer MTU minus headers, clamp TCP MSS on the VPN path, and test large inner packets in both directions.

The reason for that specificity is a failure I have seen: Inner 1500-byte packets needed a 1400-byte inner MTU. 22 percent of application writes failed for 5 hours; keepalives at 64 bytes succeeded and the SA stayed up.

Size-dependent drop on an up SA.

Inner sizeForwardedApplication writes failedHours
64yes—5
1400yes—5
1500no22 percent5

I would not consider it settled without evidence: send inner packets from 64 to 1500 bytes and record the first size that dies, then set MSS and MTU from that size.

An up SA still has an inner MTU.

Curated: · Written: · Reviewed:

QA-32You exclude RFC1918 from the tunnel to save bandwidth. What inspection did you just skip?(show answer)

The judgement in split tunnel versus full tunnel is which state the packet uses, not which state the operator intended.

Split tunnel sends listed prefixes in the clear to the local network. Full tunnel sends all user traffic to the concentrator. Exclusion is a trust decision about those destinations, not a performance toggle.

Concretely, enumerate excluded prefixes, treat them as off-VPN, and put host firewall and DNS policy on that path; do not exclude "private ranges" that include attacker-controlled space.

The reason for that specificity is a failure I have seen: Eight RFC1918 exclusions included a hotel LAN. Three malware C2 hosts on 10.0.0.0/8 were reachable in the clear for 14 days while the VPN icon stayed connected.

Exclusions that leaked.

Excluded prefixesC2 hosts reached in the clearDays
8 RFC1918 summaries314
full-tunnel (after)0—

I would not consider it settled without evidence: from a tunneled client, traceroute an excluded prefix and require it to be either intended or removed from the split list.

Split tunnel is a bypass list.

Curated: · Written: · Reviewed:

QA-33You point 0.0.0.0/0 at a tunnel whose destination is learned via the default. What happens at the next lookup?(show answer)

Where candidates lose the interview on default-route tunnel recursion is calling the tunnel or the BGP session the path.

If the tunnel destination recurses through the tunnel itself, the next hop is undefined and the tunnel collapses. The tunnel destination must resolve via a more-specific underlay route that does not use the tunnel.

Concretely, install a more-specific for the tunnel endpoint via the underlay, never via 0.0.0.0/0 through the tunnel, and test by withdrawing the more-specific in a lab.

The reason for that specificity is a failure I have seen: A static default into the tunnel removed the only route to the tunnel destination. The interface stayed administratively up with 0 next hops for 11 minutes, taking the site offline.

Recursive default.

RouteNext hopMinutes of 0 forwarding
0.0.0.0/0tunnel11
tunnel destination /320.0.0.0/0 (recursive)11
underlay more-specificmissing11

I would not consider it settled without evidence: lookup the tunnel destination and require a non-tunnel next hop.

The tunnel endpoint cannot live inside the tunnel.

Curated: · Written: · Reviewed:

QA-34Phase 1 is up. Phase 2 is up. A middlebox allows UDP 500 and drops protocol 50. What still fails?(show answer)

I would answer IKE versus ESP data path by separating reachability, policy, and observed delivery.

IKE negotiates keys; ESP carries the data. Filtering ESP or UDP 4500 after IKE succeeds is a data-path drop with a green control plane. Phase 1 is not the payload path.

Concretely, permit the negotiated data protocol or NAT-T port end to end, and confirm ESP or ESP-in-UDP counters, not only IKE state.

The reason for that specificity is a failure I have seen: IKE and IPsec SAs showed up while ESP was filtered. 0 data bytes passed for 4 hours; both vendors' "VPN up" tiles stayed green.

Keys without payload.

ChannelStateBytes in 4 hours
IKE (UDP 500)upnegotiation only
ESP (protocol 50)filtered0
application through the VPN—0

I would not consider it settled without evidence: capture beyond the middlebox and require ESP (or UDP 4500) packets, not only ISAKMP.

IKE is the handshake; ESP is the path.

Curated: · Written: · Reviewed:

QA-35A reordering path drops ESP after a burst. Why does increasing the replay window matter, and what does a window of 64 actually mean?(show answer)

The engineering content of ESP anti-replay is the return path and the more-specific prefix, not the protocol name.

ESP anti-replay rejects packets whose sequence number falls too far behind the highest seen. A window of 64 means 64 packets of slack. ECMP reordering or QoS can look like a replay attack.

Concretely, size the window to observed reorder depth, keep a flow on one path where possible, and distinguish replay drops from integrity failures.

The reason for that specificity is a failure I have seen: Window 64 met a two-member ECMP that reordered by 80 packets. 2,200 ESP packets dropped in 18 minutes after the reorder began; IKE stayed up.

Reorder deeper than the window.

Replay windowObserved reorderESP dropsMinutes
6480 packets2,20018
128 (after)80 packets0—

I would not consider it settled without evidence: count anti-replay drops versus integrity failures and measure reorder depth on the underlay.

Anti-replay is a sequence window, not a virus scanner.

Curated: · Written: · Reviewed:

QA-36Four backends have 2, 8, 8, and 32 cores. Round-robin is "fair." Who saturates first?(show answer)

Before calling round robin unequal work done I would write down the ECMP member, VLAN, or VRF nobody hashed onto.

Round-robin equalises request count, not work. Unequal capacity means the smallest member reaches 100 percent while others idle, which looks like a random latency spike.

Concretely, weight by capacity or use a load signal, and measure per-member utilisation rather than request share.

The reason for that specificity is a failure I have seen: Equal round-robin sent 25 percent of requests to a 2-core node. That node hit 100 percent CPU and 4× latency for 35 minutes while the 32-core node sat at 12 percent.

Four members, round-robin.

CoresRequest shareCPULatency vs baselineMinutes
225 percent100 percent4×35
825 percent40 percent1.1×35
825 percent41 percent1.1×35
3225 percent12 percent1.0×35

I would not consider it settled without evidence: compare request share with CPU per member and require weights that keep utilisation within a stated band.

Equal requests are not equal work.

Curated: · Written: · Reviewed:

QA-37Least-connections shows one backend with one connection and 80 percent of the work. How?(show answer)

The first thing I would establish about HTTP/2 least-connections is which forwarding entry the packet actually hits.

HTTP/2 multiplexes many streams on one TCP connection. A least-connections balancer that counts TCP sessions, not streams, parks work on whichever connection opened first.

Concretely, balance on stream count, request rate, or HTTP/1.1 where session count equals work, and do not read TCP connection count as concurrency.

The reason for that specificity is a failure I have seen: One HTTP/2 connection to a warm backend counted as 1. 80 percent of streams landed there for 26 minutes; three other backends each showed 1 idle connection.

Connections versus streams.

BackendTCP connectionsStreamsShare of workMinutes
A180080 percent26
B170~7 percent26
C170~7 percent26
D160~6 percent26

I would not consider it settled without evidence: compare stream or request counters with TCP connection counters per member.

One HTTP/2 connection is many requests.

Curated: · Written: · Reviewed:

QA-38The load balancer marks the pool up because port 443 accepts SYN-ACK. The app returns 503. Who is served?(show answer)

I would start TCP health check shallowness from the packet and the return path, not from the protocol adjacency.

A TCP connect check proves the handshake, not that the application will complete a request. Shallow health hides 503, TLS failures, and dependency errors behind an open port.

Concretely, check an HTTP path that exercises the dependency, require a success status, and fail the member when that path fails even if TCP still completes.

The reason for that specificity is a failure I have seen: TCP checks stayed green while the app returned 503. The pool served errors for 14 minutes to 9,400 requests; the balancer never marked a member down.

Open port, failing handler.

CheckResultRequests served 503Minutes
TCP 443up9,40014
HTTP GET /health503—14
members marked down0—14

I would not consider it settled without evidence: fail the application handler with the port still open and require the member to be removed.

SYN-ACK is not an application response.

Curated: · Written: · Reviewed:

QA-39The liveness probe fails on a still-ready process and Kubernetes kills it. What traffic pattern did you just create?(show answer)

This is an area where a green session and a working path are different observations.

Liveness means the process should be restarted. Readiness means it should receive traffic. Using liveness for a busy or slow-but-alive process causes kill-restart loops and 502s.

Concretely, point liveness at a cheap alive check, point readiness at dependency-complete, and never kill a process for being overloaded if the fix is to stop sending it work.

The reason for that specificity is a failure I have seen: A slow handler tripped liveness. 3,200 requests returned 502 over 9 minutes while pods restarted; readiness had been passing until the kill.

Wrong probe kills a live worker.

ProbeShould doDid502sMinutes
livenessrestart if deadkilled slow pods3,2009
readinessstop trafficstill passing until kill—9

I would not consider it settled without evidence: induce slowness without deadlock and require readiness to drop while liveness stays passing.

Not ready is not dead.

Curated: · Written: · Reviewed:

QA-40Each layer makes up to three total attempts, and three layers independently repeat work. What load did one user click become?(show answer)

My answer to retry amplification begins with which plane is making the decision: control, forwarding, or policy.

Independent attempts multiply. Three total attempts at each of three layers can produce 3 × 3 × 3 = 27 downstream attempts, which turns a timeout into a self-DDoS. If "three retries" means three retries after the initial try, the corresponding maximum is 4³ = 64.

Concretely, budget attempts end to end, define whether a count includes the initial attempt, use idempotency keys, and fail fast on a known-bad member rather than stacking timeouts.

The reason for that specificity is a failure I have seen: A 1-second dependency timeout with up to 3 total attempts at each of 3 layers produced 27 attempts per click. The dependency's 8-minute brownout ran at 27× offered load and took a second pool with it.

One click, three layers.

LayerTotal attempts per layerDownstream attempts per clickOffered load vs baselineMinutes
edge33—8
mid39—8
data32727×8

I would not consider it settled without evidence: count attempts per user request at each layer during a fault injection and require the product to stay inside a stated budget.

Retries are load.

Curated: · Written: · Reviewed:

QA-41You pull a pool member with 940 open connections and no drain. What do those connections see?(show answer)

I would treat graceful drain as a claim about a specific five-tuple, VRF, and direction.

A hard remove resets or black-holes existing connections. Graceful drain stops new work, waits for in-flight requests, then removes the member. Connection count at pull time is the blast radius.

Concretely, enable connection draining with a timeout greater than the longest in-flight request, and refuse changes that set the member down with a non-zero connection count.

The reason for that specificity is a failure I have seen: A member was set down with 940 connections and 0 drain. Clients saw 940 RSTs over 4 minutes; in-flight checkouts rolled back.

Hard pull versus drain.

MethodOpen connections at changeRSTsMinutes
hard down9409404
drain then down940 then 00—

I would not consider it settled without evidence: remove a member under load and require in-flight requests to finish or be handed off, not RST.

Down is not drain.

Curated: · Written: · Reviewed:

QA-42The app trusts X-Forwarded-For from the internet. Who just became the client IP?(show answer)

The useful question for Forwarded header trust is what a hash-diverse probe would show on the other members.

Forwarded and X-Forwarded-For are hop-by-hop claims. Only a trusted proxy should be allowed to set them; an edge that appends without stripping attacker values lets the attacker pick the client identity.

Concretely, strip or overwrite incoming Forwarded headers at the first trusted proxy, append the true socket IP, and ignore those headers from untrusted sources.

The reason for that specificity is a failure I have seen: The edge appended without stripping. 1 spoofed X-Forwarded-For value passed 6,100 admin-panel requests as an internal IP over 2 days.

Forged client identity.

Header handlingAdmin requests as internal IPDays
append without strip6,1002
strip then append socket IP0—

I would not consider it settled without evidence: send a request with a forged X-Forwarded-For from the internet and require the app to see the real socket address.

A header is not a source address.

Curated: · Written: · Reviewed:

QA-43A member dies. 28 percent of clients stay hashed to it by cookie. What did stickiness preserve?(show answer)

I would settle sticky sessions by capturing both sides of the encapsulation, not by reading the neighbour table.

Stickiness pins a client to a member. If that member is gone and the cookie is still honoured, the client is pinned to a hole until the cookie expires or the balancer fails over.

Concretely, fail sticky when the member is down, keep a shared session store, and treat stickiness as a cache, not as a hard location.

The reason for that specificity is a failure I have seen: Cookie stickiness kept 28 percent of clients on a dead member for 20 minutes. New clients hashed to live members; returning clients got errors.

Dead member, live cookie.

Client classShareOutcomeMinutes
new72 percentlive member20
returning sticky28 percentdead member20

I would not consider it settled without evidence: kill a member and require sticky clients to be remapped before the cookie TTL.

A sticky cookie can outlive the member.

Curated: · Written: · Reviewed:

QA-44You lowered a record's TTL from 3600 to 60 and changed the address. Why are clients still on the old IP 47 minutes later?(show answer)

The judgement in DNS TTL is not a global flush is which state the packet uses, not which state the operator intended.

TTL is the maximum cache reuse interval for resolvers that fetched the previous answer. It is not a push flush. Resolvers that cached at the old TTL 3600 keep the old answer until that remaining lifetime elapses.

Concretely, lower TTL well before the change, wait the previous TTL, then cut over, and measure remaining cache with diverse resolvers rather than with a local lookup.

The reason for that specificity is a failure I have seen: The cutover happened immediately after setting TTL 3600 to 60. Caches that had 47 minutes remaining on the old 3600-second answer kept the old IP; 41 percent of queries missed the new address for that remaining window.

Cutover without waiting out the old TTL.

Resolver classRemaining cacheQueries on old IPWindow
cold after TTL 6000 percent—
cached at TTL 3600up to 47 minutes41 percent47 minutes

I would not consider it settled without evidence: query resolvers that cached before the TTL change and record remaining lifetime, not only a recursive lookup from a cold cache.

TTL is cache reuse, not a global flush.

Curated: · Written: · Reviewed:

QA-45You create a record after its name has been cached as NXDOMAIN for 12 minutes. Why do some clients still get NXDOMAIN for 18 more minutes?(show answer)

Where candidates lose the interview on negative DNS cache is calling the tunnel or the BGP session the path.

Negative answers are cached for the smaller of the SOA MINIMUM field and the SOA record's TTL, not the new record's positive TTL. Creating the record does not erase remote negative cache.

Concretely, keep negative TTL short on zones you edit live, and wait out the SOA negative cache or use a different name rather than expecting instant presence.

The reason for that specificity is a failure I have seen: SOA negative TTL was 1800 seconds. After creating the A record, 18 minutes of remaining negative cache served NXDOMAIN to 1,200 clients; the authoritative server already had the record.

Positive record, leftover negative cache.

SourceAnswerClientsRemaining negative TTL
authoritativeA record——
recursive that cached NXDOMAINNXDOMAIN1,20018 minutes of an 1800-second cache

I would not consider it settled without evidence: query a resolver that received NXDOMAIN before the create and require it to keep NXDOMAIN until its negative TTL, matching the SOA.

NXDOMAIN is cached too.

Curated: · Written: · Reviewed:

QA-46The authoritative servers have the new record. A laptop still sees the old one. Whose cache is that?(show answer)

I would answer recursive versus authoritative by separating reachability, policy, and observed delivery.

Authoritative servers answer from zone data. Recursive resolvers answer from cache. Debugging the wrong one "fixes" nothing. Stub clients rarely query authority directly.

Concretely, query the authoritative name servers by address, then the recursive the client actually uses, and compare TTLs; fix the recursive cache or wait, not the zone that is already right.

The reason for that specificity is a failure I have seen: Operators edited the zone three times in 11 hours because a corporate recursive still held 2 stale answers. Authority had been correct after the first edit.

Authority right, recursion stale.

ServerAnswerStale recordsHours of extra edits
authoritative NSnew00 after first edit
corporate recursiveold211

I would not consider it settled without evidence: dig the record at NS addresses and at the client's resolver and require the two answers to be labelled as such.

Fix the cache that answered the client.

Curated: · Written: · Reviewed:

QA-47Parent NS records point at hosts that do not answer authoritatively for the child. What do resolvers return?(show answer)

The engineering content of lame delegation is the return path and the more-specific prefix, not the protocol name.

A lame delegation is an NS set that names servers which refuse or are not authoritative for the zone. Resolvers then SERVFAIL or serve stale, and the parent looking "correct" in whois is not a working delegation.

Concretely, query each parent-listed NS for the child SOA with AA bit set, and replace any that do not answer as authority.

The reason for that specificity is a failure I have seen: One of two NS hosts was a web box not running DNS. 35 percent of queries hit SERVFAIL for 3 days depending on which NS the resolver tried first.

One lame NS in a set of two.

NSAA SOAQueries that SERVFAILDays
ns1yes0 percent when chosen3
ns2 (web host)no35 percent overall3

I would not consider it settled without evidence: require AA SOA from every NS in the parent delegation.

An NS name is not a running authority.

Curated: · Written: · Reviewed:

QA-48A name has an A record and you add a CNAME at the same owner. What must happen?(show answer)

Before calling CNAME coexistence done I would write down the ECMP member, VLAN, or VRF nobody hashed onto.

A CNAME cannot coexist with A, MX, NS, or other ordinary data at the same owner name; DNSSEC proof and signature records are the narrow exception. The CNAME says the name is an alias, so pairing it with address or service data creates an inconsistent zone that conforming authorities must reject.

Concretely, put the CNAME on a name that has no other records, or use ALIAS/ANAME at the apex only if the product is explicit, and never pair CNAME with A at the same owner.

The reason for that specificity is a failure I have seen: Apex CNAME plus A was loaded on 1 of 3 nameservers and rejected on 2. Resolvers saw a mix for 2 hours; 1 of 3 authorities answered CNAME, the others kept the A.

Split authorities after an illegal pair.

AuthorityLoadedAnswerHours
1CNAME+A acceptedCNAME2
2rejectedold A2
3rejectedold A2

I would not consider it settled without evidence: load the zone on every authority and require either CNAME-only or A-only at that owner.

CNAME occupies the owner.

Curated: · Written: · Reviewed:

QA-49The browser uses DNS over HTTPS. Does that sign the zone?(show answer)

The first thing I would establish about DNSSEC versus DoH is which forwarding entry the packet actually hits.

DoH encrypts the stub-to-resolver channel. DNSSEC authenticates zone data with signatures and a chain to a trust anchor. DoH without DNSSEC still accepts a lying resolver; DNSSEC without DoH is visible on the path but authenticable.

Concretely, validate DNSSEC at a resolver you control, and treat DoH as privacy of the query path, not as origin authentication.

The reason for that specificity is a failure I have seen: A team enabled DoH and skipped DS records. For 9 days the zone had 0 DNSSEC authentication; the selected DoH resolver could still return injected answers over its encrypted channel.

Encrypted channel, unsigned zone.

ControlPresentDaysInjected answers possible
DoH to the resolveryes9yes, by that resolver
DS at parentno9—
AD bit on answersno9—

I would not consider it settled without evidence: check DS at the parent and AD bit on a validating resolver, separately from whether the stub used DoH.

DoH hides the path; DNSSEC proves the data.

Curated: · Written: · Reviewed:

QA-50A MAC has a reservation for 10.20.20.8. Why can another device still use that address?(show answer)

I would start DHCP reservation is not authentication from the packet and the return path, not from the protocol adjacency.

A reservation is a DHCP server policy to offer that address to that chaddr. It is not 802.1X, not DHCP snooping binding enforcement by itself, and not a check that the MAC is authentic. MAC addresses are spoofable.

Concretely, bind DHCP snooping and port security or 802.1X if the address must be exclusive, and treat reservations as convenience, not as access control.

The reason for that specificity is a failure I have seen: A spoofed MAC obtained the reserved 10.20.20.8. The legitimate printer was offline for 40 minutes; the server log showed a normal ACK to the reservation.

Reservation honoured for a spoof.

ClientMACAddressMinutes printer offline
printerrealnone, offline40
laptopspoofed10.20.20.8 ACK40

I would not consider it settled without evidence: spoof the reserved MAC from a second port and require the offer to be denied unless port-level control is in place.

A reservation is an offer policy.

Curated: · Written: · Reviewed:

QA-51Two access circuits share the same relayed subnet but require different address policies. How can the server distinguish them?(show answer)

This is an area where a green session and a working path are different observations.

DORA is Discover, Offer, Request, Ack. A relay sets giaddr so the server can select the relayed subnet. Option 82 can additionally identify the access circuit or remote device when clients sharing that subnet need circuit-specific policy; it is not a substitute for a correct giaddr.

Concretely, set giaddr to the client-facing relayed subnet, insert trusted option 82 circuit-id or remote-id when circuit-specific policy is required, and match that policy without accepting forged option 82 from clients.

The reason for that specificity is a failure I have seen: Two access circuits used the same giaddr and required different gateway policy, but the trusted relays omitted option 82. For 2 hours, 1,100 clients on the second circuit received the first circuit's policy even though DORA completed successfully.

Two circuits, one relayed subnet, no option 82.

BuildingOption 82Clients with wrong gatewayHours
1absent02
2absent1,1002

I would not consider it settled without evidence: capture Discover at the server and require option 82 plus a giaddr that maps to one scope.

A completed DORA can still be the wrong subnet.

Curated: · Written: · Reviewed:

QA-52You offer DHCPv6 addresses and no Router Advertisements. How do hosts get a default gateway?(show answer)

My answer to IPv6 RA versus DHCPv6 gateway begins with which plane is making the decision: control, forwarding, or policy.

IPv6 default routers come from Router Advertisements, not from DHCPv6. DHCPv6 can assign addresses and options; it does not replace RA for on-link routers. RA is ICMPv6 from the router, not a DHCPv6 message.

Concretely, send RAs with the intended prefix and router lifetime, use DHCPv6 if you need extra options, and never expect a default route from DHCPv6 alone.

The reason for that specificity is a failure I have seen: DHCPv6 assigned addresses with 0 default routes because RA was disabled. Dual-stack hosts sat with IPv6 addresses and no gateway for 26 minutes; IPv4 kept working.

Address without a router.

SourceIPv6 addressDefault routeMinutes
DHCPv6 onlyyes026
RA enabledyes1—

I would not consider it settled without evidence: show a host's IPv6 route table after DHCPv6-only and require a default only after RA.

DHCPv6 is not the IPv6 default-gateway protocol.

Curated: · Written: · Reviewed:

QA-53AAAA is faster to resolve than A, and IPv6 is a black hole. What does the client do?(show answer)

I would treat dual-stack Happy Eyeballs IPv6 black hole as a claim about a specific five-tuple, VRF, and direction.

Happy Eyeballs gives IPv6 an initial preference but races a later IPv4 attempt after a short connection-attempt delay; it should not wait for the IPv6 connect timeout before trying working IPv4. A broken or disabled fallback implementation can still turn an IPv6 black hole into user-visible delay or failure.

Concretely, do not publish AAAA until the IPv6 path forwards, and test that real clients launch and complete the IPv4 fallback within the configured Happy Eyeballs delay when IPv6 is silently filtered.

The reason for that specificity is a failure I have seen: AAAA returned 200 ms faster than A and IPv6 was filtered. A client implementation with IPv4 fallback disabled waited for the IPv6 connect timeout and failed 100 percent of new sessions for 7 minutes despite healthy IPv4.

Faster AAAA into a hole.

RecordResolve timePathNew sessionsMinutes
AAAA200 ms faster than Ablack hole0 percent success7
Aslowerworkingfallback disabled7

I would not consider it settled without evidence: publish AAAA toward a filtered path in a lab and require the client either to fail over within the Happy Eyeballs timer or not to receive the AAAA.

A published AAAA is a promise of a working IPv6 path.

Curated: · Written: · Reviewed:

QA-54The router shows incomplete ARP for the next hop. What has not happened, and what do packets do?(show answer)

The useful question for incomplete ARP is what a hash-diverse probe would show on the other members.

ARP incomplete means no neighbour Ethernet address is known. The router cannot build an L2 rewrite and drops or queues the packet. IP reachability of a prefix does not populate ARP; a reply to who-has does.

Concretely, capture ARP request and reply on the VLAN, fix L2 connectivity or filtering, and refuse to troubleshoot IGP until the adjacency MAC exists.

The reason for that specificity is a failure I have seen: 1,024 incomplete entries sat behind a VLAN ACL that dropped ARP. Packets to those next hops dropped for 15 minutes while ICMP to the router's own interface address still worked.

Incomplete neighbour table.

ARP stateEntriesForwarded to those next hopsMinutes
incomplete1,024015
ICMP to router interface—yes15

I would not consider it settled without evidence: require a complete ARP entry with a MAC before blaming the routing protocol.

No MAC, no rewrite.

Curated: · Written: · Reviewed:

QA-55IPv4 works and IPv6 neighbour solicitation is filtered. Why is that not "ARP for v6"?(show answer)

I would settle NDP versus ARP by capturing both sides of the encapsulation, not by reading the neighbour table.

NDP uses ICMPv6 neighbour solicitation and advertisement, multicast to solicited-node groups, and router advertisements. It is not ARP with a different name; filters that allow ARP ethertype 0x0806 can still drop ICMPv6 type 135/136.

Concretely, permit ICMPv6 NS/NA and RA on the segment, and diagnose IPv6 neighbour with ndp tables, not with arp -a.

The reason for that specificity is a failure I have seen: A VLAN ACL allowed ARP and dropped ICMPv6 NS. IPv4 forwarded; 44 percent of dual-stack sessions failed IPv6 neighbour for 1 hour 10 minutes.

ARP allowed, NS dropped.

ProtocolAdjacencyDual-stack sessions failing IPv6Duration
IPv4 ARPcomplete—1 hour 10 minutes
IPv6 NS/NAfiltered44 percent1 hour 10 minutes

I would not consider it settled without evidence: capture NS/NA and require them, not ARP, for IPv6 adjacency.

IPv6 neighbours are ICMPv6, not ARP.

Curated: · Written: · Reviewed:

QA-56DF is set, a hop has MTU 1400, and ICMP destination unreachable fragmentation-needed is filtered. What size dies?(show answer)

The judgement in PMTUD black hole is which state the packet uses, not which state the operator intended.

PMTUD relies on ICMP type 3 code 4 (or IPv6 packet too big) to shrink the path MTU. Without that signal, 1500-byte DF packets drop at the 1400 hop and never retry smaller.

Concretely, permit the needed ICMP, clamp MSS, or raise MTU, and test with DF set at 1500.

The reason for that specificity is a failure I have seen: A GRE hop of 1400 with filtered ICMP 3/4 black-holed 1500-byte DF packets for 4 hours. Small packets and traceroute (often smaller) succeeded.

DF 1500 into MTU 1400 without ICMP.

PacketDFForwardedHours
1400yesyes4
1500yesno4
ICMP 3/4 returned—04

I would not consider it settled without evidence: send DF 1500 and DF 1400; require 1500 to fail and ICMP 3/4 to be visible, or MSS clamped.

PMTUD is an ICMP protocol, not magic.

Curated: · Written: · Reviewed:

QA-57You allow echo and block all other ICMP. Which breakage did you keep?(show answer)

Where candidates lose the interview on filtering ICMP is calling the tunnel or the BGP session the path.

Echo is a convenience. Destination unreachable, packet too big, and time exceeded are control messages for fragmentation, PMTUD, and traceroute. Filtering "ICMP except ping" breaks those.

Concretely, allow the ICMP types the path needs, rate-limit rather than blanket-deny, and test large DF packets and traceroute after the ACL.

The reason for that specificity is a failure I have seen: Echo was allowed; dest-unreach was dropped. 19 percent of large transfers failed for 6 hours while ping stayed green.

Echo versus dest-unreach.

ICMPACLLarge transfer lossHours
echoallow—6
dest-unreachdeny19 percent6

I would not consider it settled without evidence: block dest-unreach in a lab and require large DF transfers to fail even as echo succeeds.

Ping is not the ICMP you needed.

Curated: · Written: · Reviewed:

QA-58You see SYNs out and no SYN-ACK. How do you decide whether the return path or the listener is at fault?(show answer)

I would answer incomplete TCP handshake by separating reachability, policy, and observed delivery.

A missing SYN-ACK is either the SYN not arriving, the listener not answering, or the SYN-ACK not returning. One capture on the client cannot distinguish those. You need both ends and the middle.

Concretely, capture at client, at server, and at the stateful midpoint, and classify by where the SYN is last seen and whether a SYN-ACK is generated.

The reason for that specificity is a failure I have seen: Client captures showed 8,400 embryonic SYNs. The server had generated SYN-ACKs that a return ACL dropped. The ticket stayed on "application down" for 21 minutes.

SYNs without completing handshake.

LocationSYNsSYN-ACKsMinutes misdiagnosed
client8,400021
server8,4008,40021
return ACL—0 passed21

I would not consider it settled without evidence: show the SYN-ACK leaving the server and dying on the return ACL, or show the SYN never arriving.

Embryonic on one side is not a closed listener.

Curated: · Written: · Reviewed:

QA-59The client receives a RST. Why might the server log show no such RST?(show answer)

The engineering content of TCP RST provenance is the return path and the more-specific prefix, not the protocol name.

A RST can come from the server, from a middlebox, or from an old stack sending on a closed 4-tuple. Sequence validity, TTL/hop-limit evidence, and captures at multiple points can localize it; a TTL alone does not prove the sender. Blaming the server process without a capture is a guess.

Concretely, compare sequence validity and captures before and after each suspected middlebox, using TTL only as supporting evidence; a RST absent at the server-side capture but present downstream was injected on that segment.

The reason for that specificity is a failure I have seen: A middlebox injected 2,200 RSTs over 16 minutes. Server logs showed accepted handshakes; the app team was paged for a crash that did not occur.

Injected RSTs.

Claimed senderRSTs capturedServer crash logsMinutes
application2,200 at client016
middlebox (TTL too high)2,200—16

I would not consider it settled without evidence: capture the RST at two points and require its first appearance and sequence validity to match the claimed sender.

A RST has a source; name it.

Curated: · Written: · Reviewed:

QA-60Wireshark on the sender shows 64 KB TCP segments. The network MTU is 1500. What was on the wire?(show answer)

Before calling TSO/GSO capture illusion done I would write down the ECMP member, VLAN, or VRF nobody hashed onto.

TCP segmentation offload and GSO let the stack hand large buffers to the NIC, which segments them. A capture before segmentation shows huge segments that never existed on the Ethernet. Jumbo is a different configuration.

Concretely, capture after the NIC or on a TAP, or disable TSO on the capture host, before concluding jumbo or MSS failure.

The reason for that specificity is a failure I have seen: A 64 KB capture led to a 3-hour jumbo hunt. On-wire frames were 1,460-byte payload; 0 jumbo frames existed.

Two capture points, one flow.

PointSegment sizeJumbo framesHours of wrong hunt
before TSO64 KBn/a3
TAP1,460-byte payload03

I would not consider it settled without evidence: compare host capture size with TAP capture size for the same flow.

TSO is a host buffer, not a wire frame.

Curated: · Written: · Reviewed:

QA-61UDP traceroute dies at hop 6 while the TCP application works. What did traceroute actually test?(show answer)

The first thing I would establish about traceroute caveats is which forwarding entry the packet actually hits.

Classic traceroute sends UDP to high ports or ICMP echo with incrementing TTL. Firewalls treat that unlike the application's TCP/TLS path. Missing hops are often filtered TTL-exceeded, not a forwarding hole.

Concretely, use TCP traceroute to the application port, compare with a working transaction, and do not change routing because UDP 33434 was dropped.

The reason for that specificity is a failure I have seen: UDP traceroute to 33434 was blocked at hop 6. The HTTP path of 6 hops worked. A 90-minute routing change was proposed for a filter that was working as designed.

UDP 33434 versus TCP 443.

ProbeHops shownApplicationMinutes of proposed change
UDP 33434dies at 6not HTTP90
TCP 4436 workingsuccess0 needed

I would not consider it settled without evidence: run traceroute with the application's protocol and port and require agreement with a successful transaction capture.

Traceroute is not the application five-tuple.

Curated: · Written: · Reviewed:

QA-62You see one TCP retransmission. Which hop dropped the packet?(show answer)

I would start one retransmission is not a location from the packet and the return path, not from the protocol adjacency.

A retransmission says a segment was not ACKed in time. It does not name the hop. Loss could be anywhere on the forward or reverse path, or be delay that was not loss.

Concretely, sample loss with captures at multiple points or with instrumented hops, and refuse a "the problem is hop 3" conclusion from one retransmit counter.

The reason for that specificity is a failure I have seen: One retransmission on a 40 ms path was blamed on the WAN. Three possible hops were later instrumented; the drop was the access switch buffer for 90 minutes of wrong WAN tickets.

One loss, three suspects.

Hop blamed firstDrops actually countedMinutes of WAN tickets
WAN090
access switch bufferthe loss90
other two hops090

I would not consider it settled without evidence: place counters or captures on each candidate hop during the loss and require the drop to increment in one place.

A retransmit is a symptom, not a coordinate.

Curated: · Written: · Reviewed:

QA-63You connect by IP to a name-based HTTPS virtual host. Which certificate do you get?(show answer)

This is an area where a green session and a working path are different observations.

TLS Server Name Indication tells the terminator which certificate to present. Connecting by IP omits SNI or sends a mismatch, so you get the default cert, which will not match the name you intended.

Concretely, connect with the hostname in SNI, and do not validate a service by IP unless that IP has its own cert.

The reason for that specificity is a failure I have seen: Health checks used the IP. 12 name-based hosts presented the default cert; checks failed for 1 hour even though clients using SNI succeeded.

Default cert without SNI.

HandshakeHostsCertHours of failed IP checks
no SNI (by IP)12default1
SNI = hostname12matching0 client impact

I would not consider it settled without evidence: compare a handshake with SNI to one without and require different certificates where vhosts differ.

The IP is not the name the cert binds.

Curated: · Written: · Reviewed:

QA-64The interface is administratively up with 0 pps and no light. What have you not got?(show answer)

My answer to interface UP is not carrier begins with which plane is making the decision: control, forwarding, or policy.

Admin up is the configured desire. Carrier/line protocol is photons or a negotiated link. Protocol up/up is still not proof of the intended VLAN or IPv6 ND. Empty pps with no optics is a physical or negotiation fault.

Concretely, read optics, line protocol, and speed/duplex, then pps, and do not start IGP debug on an interface with no carrier.

The reason for that specificity is a failure I have seen: A shutdown was removed but the SFP sat unseated. The interface stayed admin up, protocol down, 0 pps for 25 minutes while OSPF was restarted twice.

Admin up, no photons.

StateValueMinutes of OSPF restarts
adminup25
line protocoldown25
pps025
OSPF restarts225

I would not consider it settled without evidence: require carrier and non-zero light levels before protocol debug.

Admin up is a command, not a link.

Curated: · Written: · Reviewed:

QA-65Forward traceroute succeeds. The application still fails. What have you not mapped?(show answer)

I would treat bidirectional path mapping as a claim about a specific five-tuple, VRF, and direction.

Forward and return paths can differ. A filter, NAT, or firewall on the return path drops the response while the forward traceroute looks clean. Mapping one direction is half a path.

Concretely, trace or capture both directions, include the stateful midpoint, and treat one-way success as a return-path incident until proven otherwise.

The reason for that specificity is a failure I have seen: Forward ACL allowed the app; return ACL dropped it. 100 percent of sessions failed one-way for 16 minutes; forward traceroute was clean.

Clean forward, dead return.

DirectionTracerouteApplication packetsMinutes
forwardsuccessSYN arrives16
returnnot mapped100 percent drop16

I would not consider it settled without evidence: capture the response leaving the server and not arriving at the client, and name the return hop that counters it.

A path has two directions.

Curated: · Written: · Reviewed:

QA-66The receiver window is 2 segments and cwnd is large. Which control is limiting, and what is not "the network"?(show answer)

The useful question for TCP flow control versus congestion control is what a hash-diverse probe would show on the other members.

Flow control is the receiver's window. Congestion control is the sender's cwnd from loss or ECN. A tiny window is an application or socket issue, not WAN congestion.

Concretely, read rwnd and cwnd from a capture, and only tune the path when cwnd is the limiter.

The reason for that specificity is a failure I have seen: A 40 ms RTT path was blamed for roughly 0.6 Mbps throughput. rwnd was 2 × 1460-byte segments; cwnd never collapsed. The advertised receive window was about 2.9 KB for 2 hours of WAN tickets.

Window versus cwnd.

SignalValueWAN ticketsHours
rwnd2 segmentsopened2
cwndlarge—2
throughput~0.6 Mbps—2
RTT40 msblamed2

I would not consider it settled without evidence: require a capture that shows whether rwnd or cwnd is the min.

A closed window is not a congested path.

Curated: · Written: · Reviewed:

QA-67ICMP port unreachable comes back and the UDP application does not retry. Whose job is that?(show answer)

I would settle UDP application semantics by capturing both sides of the encapsulation, not by reading the neighbour table.

UDP has no transport retry. ICMP unreach is a hint the stack may surface. If the application ignores it, the user sees silence. That is application semantics, not a missing TCP feature on the router.

Concretely, document whether the app retries, and do not "fix UDP" in the network when the payload protocol has no retransmission.

The reason for that specificity is a failure I have seen: A syslog over UDP was sent once. ICMP unreach from a closed port was ignored. Operators spent 22 minutes looking for packet loss that was a closed collector and 0 app retries.

One datagram, no retry.

EventCountMinutes of loss hunt
UDP datagrams sent122
ICMP unreach received122
application retries022

I would not consider it settled without evidence: show the ICMP unreach at the sender and 0 retransmission from the application.

UDP will not save an application that does not retry.

Curated: · Written: · Reviewed:

QA-68UDP keepalives are 45 seconds and the NAT mapping idle timer is 30. What happens to the flow?(show answer)

The judgement in NAT mapping expiry is which state the packet uses, not which state the operator intended.

A NAT mapping exists only while traffic refreshes it. If the idle timer is shorter than the keepalive, the mapping dies and return traffic hits a foreign 5-tuple.

Concretely, set idle timers longer than keepalives, or send keepalives inside the timer, and test after silence longer than the timer.

The reason for that specificity is a failure I have seen: 30-second UDP mappings met 45-second keepalives. 1,800 flows per hour dropped after silence; the VPN looked up while voice media died.

Keepalive longer than idle.

TimerValueDrops per hour
NAT UDP idle30 s1,800
keepalive45 s1,800

I would not consider it settled without evidence: silence a UDP flow for longer than the NAT timer and require the mapping to be gone, then fix the timer.

The mapping timer is the session.

Curated: · Written: · Reviewed:

QA-69A TCP connection starts at one anycast PoP and a later packet arrives at another. What state does the second PoP have?(show answer)

Where candidates lose the interview on anycast TCP state is calling the tunnel or the BGP session the path.

Anycast shares an address, not a TCP table. A mid-flow move to another PoP meets a stack with no TCB and typically a RST. Long TCP therefore needs route stability, flow pinning, shared state, or an explicit reset budget.

Concretely, pin a TCP flow to one PoP with ECMP stability or use a connection-aware front, and do not announce anycast from PoPs that cannot share state.

The reason for that specificity is a failure I have seen: Two PoPs advertised the same VIP. 14 percent of long TCP flows moved after a routing change and were RST for 6 minutes.

Mid-flow PoP move.

PoPsFlows that movedResultMinutes
2 anycast14 percentRST6

I would not consider it settled without evidence: change IGP cost so a flow's next hop moves and require either pinning or an accepted reset budget.

Anycast is an address, not a shared TCB.

Curated: · Written: · Reviewed:

QA-70You drain a PoP in DNS. TTL is 3600. How long do clients keep steering there?(show answer)

I would answer DNS steering cache by separating reachability, policy, and observed delivery.

Geo-DNS steering is cached like any A/AAAA. Draining a PoP in the geo map does not evict caches. Remaining TTL is the drain tail.

Concretely, lower TTL before drain, wait the old TTL, then withdraw, and measure application requests still arriving at the drained PoP.

The reason for that specificity is a failure I have seen: PoP drain was immediate with TTL 3600. After 8 minutes, 36 percent of application requests still arrived at the drained PoP because clients retained cached steering answers.

Drain versus TTL 3600.

Time after geo changeApplication requests still at drained PoP
8 minutes36 percent
designed wait of 3600 snear 0

I would not consider it settled without evidence: count application requests at the drained PoP versus time since the geo change and match the remaining TTL curve.

Steering is DNS cache, not an instant map.

Curated: · Written: · Reviewed:

QA-71The load balancer presents the cert and talks HTTP to the backend. Where is the cleartext?(show answer)

The engineering content of TLS termination at a load balancer is the return path and the more-specific prefix, not the protocol name.

TLS terminated at the balancer decrypts there. The backend hop is whatever you configured: cleartext HTTP, or a new TLS session. Users seeing a padlock have not encrypted the backend hop.

Concretely, encrypt the backend if that hop is untrusted, and inventory hop-by-hop TLS rather than quoting the browser padlock.

The reason for that specificity is a failure I have seen: The balancer terminated TLS and forwarded HTTP. 2.4 GB of backend cleartext crossed 5 hops of a shared VLAN for 3 days.

Browser TLS, backend HTTP.

HopEncryptedBytes capturedDays
client to LBTLS—3
LB to backendHTTP2.4 GB3
L2 hops on that VLAN5—3

I would not consider it settled without evidence: capture behind the balancer and require TLS or an isolated segment.

The padlock stops at the terminator.

Curated: · Written: · Reviewed:

QA-72A dependency is 100 percent failing. The client keeps sending at full rate. What is the breaker supposed to do?(show answer)

Before calling circuit breaker done I would write down the ECMP member, VLAN, or VRF nobody hashed onto.

A circuit breaker stops sending when error rate crosses a threshold, then probes half-open. Without it, retries and timeouts pile onto a dead dependency.

Concretely, open on a measured error ratio, cap half-open probes, and apply it per dependency, not per process.

The reason for that specificity is a failure I have seen: No breaker opened. 100 percent of traffic hit a dead pool for 11 minutes; timeouts occupied every worker.

No open state.

BreakerTraffic to dead poolMinutesWorkers blocked on timeout
never opens100 percent11all
opens at 50 percent errorsprobe only—few

I would not consider it settled without evidence: fail a dependency and require sending to drop to probe rate, not stay at 100 percent.

A closed breaker on a dead pool is a flood.

Curated: · Written: · Reviewed:

QA-73The proxy never rejects; it queues 48,000 requests. What did p99 become?(show answer)

The first thing I would establish about unbounded proxy queue is which forwarding entry the packet actually hits.

An unbounded queue trades drop for latency. Requests sit until they are stale, then fail anyway, after occupying memory. A bounded queue with 503 is a faster signal.

Concretely, cap queue depth and time, reject overflow, and alert on queue age, not only on CPU.

The reason for that specificity is a failure I have seen: The queue hit 48,000. p99 rose to 9 seconds with 0 rejects for 13 minutes, then clients timed out anyway.

Queue instead of 503.

Queue depthp99RejectsMinutes
48,0009 s013
bounded 2,000200 msoverflow 503—

I would not consider it settled without evidence: load past capacity and require either a bound with rejects or a stated latency budget that was met.

Infinite queue is delayed failure.

Curated: · Written: · Reviewed:

QA-74Twelve nodes, consistent hashing, one key is 38 percent of traffic. Where is the CPU?(show answer)

I would start consistent hashing hot key from the packet and the return path, not from the protocol adjacency.

Consistent hashing maps a key to a node. A hot key is not spread by adding nodes unless you shard the key. The node that owns it takes that share plus its fair keys.

Concretely, split hot keys, add a local cache, or use bounded loads, and measure per-key rate before scaling the ring.

The reason for that specificity is a failure I have seen: One key was 38 percent of requests. That node ran at 91 percent CPU for 17 minutes while 11 others sat near 20 percent.

Hot key on a 12-node ring.

NodesHottest key shareThat node's CPUMinutes
1238 percent91 percent17
other 11remainder~20 percent17

I would not consider it settled without evidence: rank keys by rate and require the hottest key's node utilisation to be explained by that key.

The ring does not shard a single key.

Curated: · Written: · Reviewed:

QA-75You carry tenant traffic in GRE across a shared underlay. Who can read it?(show answer)

This is an area where a green session and a working path are different observations.

GRE adds a header. Anyone who can see the underlay packet can see the inner IP unless a crypto transform is applied. Isolation of VNI or VRF at the edge does not encrypt the underlay.

Concretely, add IPsec or equivalent, and treat GRE as topology, not as a confidentiality boundary.

The reason for that specificity is a failure I have seen: GRE-only tenant links leaked 880 MB of inner traffic to a SPAN on the underlay for 1 afternoon; VRFs at the PE stayed correctly separated.

VRF correct, underlay readable.

PlaneIsolatedInner bytes on SPANDuration
PE VRFyes—1 afternoon
GRE underlayno crypto880 MB1 afternoon

I would not consider it settled without evidence: SPAN the underlay and require inner payloads to be unreadable.

GRE is a pipe, not a vault.

Curated: · Written: · Reviewed:

QA-76Three tenants share an underlay, each with a VNI. What stops a curious VTEP from reading VNI 10010?(show answer)

My answer to VXLAN is not confidential begins with which plane is making the decision: control, forwarding, or policy.

VXLAN identifies a VNI; it does not encrypt. A VTEP or tap that receives the UDP packet can parse the inner frame. Confidentiality needs crypto (MACSec, IPsec, or encrypted overlays), plus control-plane auth.

Concretely, restrict who can be a VTEP, encrypt the underlay, and do not quote VNI as encryption.

The reason for that specificity is a failure I have seen: VNI 10010 was visible to 3 tenants' debug taps on a shared underlay with 0 IPsec for 5 days.

Shared underlay, tagged only.

VNIIPsecTenants that could read innerDays
10010035

I would not consider it settled without evidence: decap a captured VXLAN packet without keys and require the inner Ethernet to be readable unless crypto is present.

A VNI is a tag.

Curated: · Written: · Reviewed:

QA-77A policy-based VPN misses 17 prefixes that a route-based tunnel would have taken from the routing table. Why?(show answer)

I would treat route-based versus policy-based VPN as a claim about a specific five-tuple, VRF, and direction.

Policy-based VPN encrypts what the ACL/selector lists. Route-based VPN encrypts what is routed into a tunnel interface. A missed ACL entry is cleartext or unrouted, while the SA for other prefixes stays up.

Concretely, prefer a tunnel interface with routes for changing prefixes, and audit selectors when policy-based is required.

The reason for that specificity is a failure I have seen: PBR/ACL missed 17 prefixes. Those flows left in the clear for 3 hours; the SA for the listed prefixes stayed up.

Missed ACL entries.

PrefixesIn selectorPathHours
intended setyesencrypted3
17 missednoclear default3

I would not consider it settled without evidence: traceroute a prefix that should be encrypted and require the next hop to be the tunnel, not the clear default.

Selectors are a static list; a tunnel interface follows the RIB.

Curated: · Written: · Reviewed:

QA-78Both sides of a VPN use 10.0.0.0/8. What must you do before interesting traffic can be unique?(show answer)

The useful question for overlapping RFC1918 is what a hash-diverse probe would show on the other members.

Overlapping RFC1918 means the same address is two different hosts. Without NAT or re-addressing, selectors and return routes cannot name a unique endpoint.

Concretely, nAT one or both sides to unique ranges, or renumber, and put the translated prefixes in the selectors.

The reason for that specificity is a failure I have seen: Both campuses used 10.0.0.0/8. 64 connections hit the local host instead of the remote for 2 hours; the SA was up.

Two 10/8s, one SA.

SidePrefixConnections to local-by-mistakeHours
A10.0.0.0/8included in 642
B10.0.0.0/8included in 642
SAup0 unique remote2

I would not consider it settled without evidence: show a destination 10.1.1.1 routing to the tunnel translated prefix, not to the local LAN.

The same 10/8 is not a unique locator.

Curated: · Written: · Reviewed:

QA-79Two VRFs use the same RD and different RTs. What collided, and what did not import?(show answer)

I would settle L3VPN route distinguisher versus route target by capturing both sides of the encapsulation, not by reading the neighbour table.

The route distinguisher combines with an IPv4 prefix to form the VPN-IPv4 NLRI. Reusing the same RD for the same overlapping IPv4 prefix makes those routes the same NLRI, so BGP path selection can hide one; different IPv4 prefixes do not collide merely because the RD matches. The route target is the import/export community that selects VRFs, so different RTs still mean no import.

Concretely, use unique RDs per VRF (or per PE+VRF), set RTs for the intended topology, and never treat RD as the import policy.

The reason for that specificity is a failure I have seen: Identical RDs collided two tenants' 10.0.0.0/24 in VPNv4. Different RTs meant 0 import into the other VRF, and the collision hid one prefix for 3 hours.

Same RD, different RT.

VRFRDRT import10.0.0.0/24 visibleHours
tenant A65000:165000:10yes, won collision3
tenant B65000:165000:20hidden in VPNv43
cross import—none03

I would not consider it settled without evidence: show VPNv4 for that RD and the import RT list, and require unique RD plus intended RT.

RD uniqueness is not RT policy.

Curated: · Written: · Reviewed:

QA-80The VPN drops and the kill switch is off. Where does the next packet go?(show answer)

The judgement in always-on VPN kill switch is which state the packet uses, not which state the operator intended.

A kill switch blocks traffic when the tunnel is down. Without it, always-on VPN fails open to the local default, which is a leak. Split tunnel already defined some leaks; this is the rest.

Concretely, enable fail-closed, test by killing IKE, and require 0 bytes on the physical default except the tunnel outer.

The reason for that specificity is a failure I have seen: IKE died. With no kill switch, 100 percent of user traffic leaked to the local LAN for 12 minutes while the client icon showed reconnecting.

Tunnel down, switch off.

Kill switchInner traffic on LANMinutes
off100 percent12
on0—

I would not consider it settled without evidence: kill the tunnel and capture the physical interface for inner destinations.

Always-on without fail-closed is sometimes-on.

Curated: · Written: · Reviewed:

QA-81VRRP master is elected and has no default route. What do hosts with a VIP gateway get?(show answer)

Where candidates lose the interview on first-hop redundancy versus upstream routing is calling the tunnel or the BGP session the path.

First-hop redundancy shares a gateway address. Upstream routing on that master is separate. A master without a default still answers ARP and then drops.

Concretely, track the upstream default in VRRP, preempt to a master that can forward, and ping past the VIP, not only the VIP.

The reason for that specificity is a failure I have seen: The master had 0 default after an upstream BGP drop. Hosts ARPed the VIP successfully and black-holed for 7 minutes.

VIP up, default gone.

ObjectStateMinutes of host black hole
VRRP masterelected7
default route on master07
ARP for VIPsuccess7

I would not consider it settled without evidence: fail upstream on the master and require VIP ownership to move, or tracking to decrement priority.

Owning the VIP is not owning a default.

Curated: · Written: · Reviewed:

QA-82OSPF offers 10.0.0.0/16 at AD 110 and BGP offers 10.4.1.0/24 at AD 20. Who wins for 10.4.1.9, and who would win if they were the same length?(show answer)

I would answer administrative distance versus longest match by separating reachability, policy, and observed delivery.

Longest match is first, so the /24 wins regardless of AD, and AD is compared only among same-prefix-length candidates.

Concretely, state the match length first, then AD among ties, and use a forwarding lookup to prove it rather than quoting AD as the reason the /16 lost.

The reason for that specificity is a failure I have seen: An operator deleted the BGP /24 because "OSPF AD is worse so it cannot be used," leaving the /16. Traffic followed the remaining /16 into a hole for 27 minutes.

Different lengths, AD is irrelevant.

PrefixADWinner for 10.4.1.9Minutes after /24 deleted
10.4.1.0/24 BGP20yes while present—
10.0.0.0/16 OSPF110after deletion27

I would not consider it settled without evidence: lookup 10.4.1.9 with both routes present and require the /24.

AD does not beat a longer prefix.

Curated: · Written: · Reviewed:

QA-83A /32 static to a dead next hop remains installed, and 0.0.0.0/0 is healthy. Where do packets to that /32 go?(show answer)

The engineering content of more-specific broken path beating default is the return path and the more-specific prefix, not the protocol name.

A more-specific that is still installed wins forever against a default. Withdrawal, tracking, or a worse AD on failure is required; otherwise the default never sees the packet.

Concretely, track the next hop, remove or degrade the /32 on failure, and test by failing the next hop.

The reason for that specificity is a failure I have seen: A /32 static stayed up with an unresolved next hop for 55 minutes. The default would have worked; 100 percent of that host's traffic died.

Stuck /32 versus healthy default.

RouteNext hopPackets to the hostMinutes
host /32unresolved100 percent of them55
0.0.0.0/0healthy055

I would not consider it settled without evidence: fail the /32 next hop and require either removal or a working tracked backup, not a stuck more-specific.

A dead more-specific beats a live default.

Curated: · Written: · Reviewed:

QA-84Forwarding is at 0 percent drop and the routing process is at 94 percent CPU. What fails next?(show answer)

Before calling control-plane CPU done I would write down the ECMP member, VLAN, or VRF nobody hashed onto.

The forwarding plane can keep switching while the control plane starves. Hellos, BGP keepalives, and BFD then time out, which looks like a link failure. CoPP exists to protect that CPU.

Concretely, read control-plane CPU and CoPP drops separately from interface drops, and protect protocol traffic before adding more iBGP sessions.

The reason for that specificity is a failure I have seen: Data-plane drop stayed 0 percent. Control-plane CPU at 94 percent missed hellos; OSPF flapped for 20 minutes on otherwise clean links.

Data plane fine, hellos dying.

PlaneDrop or CPUOSPF flapsMinutes
forwarding0 percent drop—20
control94 percent CPUyes20

I would not consider it settled without evidence: show CoPP and process CPU during the flap and require interface error counters to stay flat.

Fast switching does not keep hellos alive.

Curated: · Written: · Reviewed:

QA-85The RIB holds 518,000 prefixes and the TCAM holds 512,000. What happens to the extra 6,000?(show answer)

The first thing I would establish about FIB/TCAM exhaustion is which forwarding entry the packet actually hits.

Hardware FIB/TCAM is finite. Prefixes that do not program punt or black-hole. Software RIB can still list them. Scale is a hardware number, not a protocol success.

Concretely, monitor RIB versus FIB counts, filter or summarize before the limit, and alarm before installation fails.

The reason for that specificity is a failure I have seen: 6,000 prefixes over a 512,000 TCAM limit were not programmed for 2 hours. Those destinations black-holed while BGP was Established with 518,000 in the RIB.

RIB larger than TCAM.

TablePrefixesForwardedHours
RIB518,000—2
TCAM limit512,000512,0002
not programmed6,00002

I would not consider it settled without evidence: compare RIB and hardware FIB counts and require them equal for forwarded prefixes.

The TCAM size is the forwarding table.

Curated: · Written: · Reviewed:

QA-86Storm control drops 1 percent of broadcast and the loop is still there. What did you treat, and what remains?(show answer)

I would start storm control versus a loop from the packet and the return path, not from the protocol adjacency.

Storm control rate-limits flood to protect uplinks. It does not break the loop. Frames still circulate; MAC tables still flap. STP or unplugging the loop is the cure.

Concretely, use storm control as a shock absorber, find the loop with MAC moves and disabled STP, and do not close the incident on a 1 percent drop rate.

The reason for that specificity is a failure I have seen: Storm control at 1 percent drop ran for 14 minutes on a loop with STP off. The VLAN stayed unhealthy; flood continued at the limiter.

Limited flood, loop intact.

ControlDropSTPMinutes still looping
storm control1 percentoff14
after STP enabled—blocking a port0

I would not consider it settled without evidence: require STP blocking or a removed cable, not only storm-control drops, before closing.

A limiter is not a loop-free topology.

Curated: · Written: · Reviewed:

QA-87You enable BPDU guard on an uplink. What happens when the upstream sends a BPDU?(show answer)

This is an area where a green session and a working path are different observations.

BPDU guard errdisables a port that receives a BPDU. It belongs on edge ports that should never see BPDUs. On an uplink it takes down the path the first time the upstream speaks STP.

Concretely, place BPDU guard only on PortFast/edge ports, use root guard toward untrusted switches, and test with a BPDU on an edge.

The reason for that specificity is a failure I have seen: BPDU guard on an uplink errdisabled the port for 9 minutes at the first hello from the core, isolating an IDF.

Guard on the wrong port.

PortBPDU guardResultMinutes
uplinkonerrdisable9
edge (correct)onerrdisable on rogue—

I would not consider it settled without evidence: send a BPDU to an edge port and require errdisable, and send one on an uplink and require normal STP processing without errdisable.

BPDU guard is an edge weapon.

Curated: · Written: · Reviewed:

QA-88PortFast is enabled on a trunk to another switch. What window do you open?(show answer)

My answer to PortFast on a trunk begins with which plane is making the decision: control, forwarding, or policy.

PortFast skips listening/learning on edges. On a switch-to-switch trunk it forwards immediately and can loop until STP blocks, which on a trunk is many VLANs of flood.

Concretely, enable PortFast only on true edges, keep trunks in normal STP, and use BPDU guard on those edges.

The reason for that specificity is a failure I have seen: PortFast on a trunk looped for 8 seconds across 2 VLANs before STP blocked; both VLANs stormed.

Trunk with PortFast.

LinkPortFastLoop durationVLANs stormed
switch trunkyes8 s2

I would not consider it settled without evidence: require PortFast operational only on access edges, not on trunks to bridges.

PortFast assumes nothing behind the port bridges.

Curated: · Written: · Reviewed:

QA-89The classifier matches DSCP EF internally and the WAN still shows CS0. What was missing?(show answer)

I would treat QoS classification versus marking as a claim about a specific five-tuple, VRF, and direction.

Classification selects a class. Marking writes the DSCP/CoS the next hop will see. Matching without marking leaves the original mark, often CS0 on the WAN.

Concretely, mark explicitly on ingress or at the WAN edge, and verify with a capture on the WAN, not with a class-map hit counter alone.

The reason for that specificity is a failure I have seen: Class-maps hit for 3 hours. WAN captures showed CS0 on 100 percent of those packets because no set-dscp was configured.

Matched, not marked.

LocationEF markHours
class-map hitscounted3
WAN DSCPCS0 on 100 percent3

I would not consider it settled without evidence: capture DSCP on the WAN for a known EF flow and require EF, not only a policy-map hit.

A class-map hit is not a DSCP on the wire.

Curated: · Written: · Reviewed:

QA-90NetFlow uses random 1:1024 packet sampling. A 3-second 400-packet burst may not appear. What can you still claim?(show answer)

The useful question for sampled NetFlow is what a hash-diverse probe would show on the other members.

Sampled NetFlow estimates. Short bursts smaller than the sampling chance can be invisible. It is not a packet-complete audit log.

Concretely, use packet capture or unsampled counters for the burst, and quote NetFlow as a sample with an error bar.

The reason for that specificity is a failure I have seen: A 3-second burst of 400 packets left 0 sampled flow records at random 1:1024 packet sampling, where only about 0.39 sampled packets were expected. The incident was closed as "no traffic" for 28 minutes.

Burst versus 1:1024.

SourcePackets in 3 sRecordsMinutes of "no traffic"
interface counters400—28
random NetFlow 1:1024expected ~0.39 sampled packets028

I would not consider it settled without evidence: compare interface counters for the 3 seconds with NetFlow and require the counter, not the sample, as proof of absence.

A sample can miss a burst.

Curated: · Written: · Reviewed:

QA-91Syslog shows 0 ACL denies and a TAP shows 12,000 hits. Which one forwards?(show answer)

I would settle syslog versus packet truth by capturing both sides of the encapsulation, not by reading the neighbour table.

Syslog is a sampled, buffered, filtered management plane. ACL hit counters and captures are the forwarding record. Missing syslog is not a missing drop.

Concretely, read ACL counters and captures first, then syslog, and never close a deny with an empty log buffer.

The reason for that specificity is a failure I have seen: Logging was rate-limited to 0 visible denies. The TAP showed 12,000 ACL drops in 10 minutes; the change was thought to have no effect.

Silent syslog, busy TAP.

SourceDeniesMinutes
syslog010
TAP / ACL counter12,00010

I would not consider it settled without evidence: compare ACL counters with syslog counts for the same rule.

The packet counter is the policy.

Curated: · Written: · Reviewed:

QA-92An ACL change locks you out with no rollback. How long is the restore, and what should have been in the window?(show answer)

The judgement in change without rollback is which state the packet uses, not which state the operator intended.

A change without a known-good restore is a bet that the new ACL is perfect. Out-of-band or commit-rollback exists because the in-band path is what you just filtered.

Concretely, keep console or a confirmed commit, save the prior ACL, and test a session that would be denied before dropping the old permit.

The reason for that specificity is a failure I have seen: One ACL line blocked management. Restore from memory took 40 minutes because the only copy was on the now-unreachable SSH path.

Lockout without rollback.

ChangeRollback readyMinutes to restore
one ACL lineno40
same with confirmed commityesauto

I would not consider it settled without evidence: require a rollback method that does not use the traffic you might deny.

The management path is in the packet filter.

Curated: · Written: · Reviewed:

QA-93You commit a routing change with a 10-minute confirm. What happens if you are locked out?(show answer)

Where candidates lose the interview on commit-confirmed is calling the tunnel or the BGP session the path.

Commit-confirmed applies the change and reverts unless a confirm arrives. It is the difference between a 10-minute automatic restore and a 2-hour drive to the console.

Concretely, use confirmed commit on remote boxes, set the timer longer than the test and shorter than the outage budget, and confirm only after forwarding checks.

The reason for that specificity is a failure I have seen: A bad static was committed without confirm. The site needed 2 hours to reach console. The same change with a 10-minute confirmed commit would have rolled back inside the timer.

Timer versus drive time.

MethodRestoreHours or minutes
no confirmconsole2 hours
commit-confirmedautomatic10 minutes

I would not consider it settled without evidence: fail to confirm in a lab and require the prior config to return at the timer.

Confirm is the rollback clock.

Curated: · Written: · Reviewed:

QA-94The IPv4 ACL is tight and the IPv6 ACL is absent. What did dual-stack just open?(show answer)

I would answer IPv4/IPv6 policy parity by separating reachability, policy, and observed delivery.

Address families are separate filters. Dual-stack hosts choose IPv6 under Happy Eyeballs. An IPv4-only ACL is not an IPv6 ACL. Parity is two policies, not a slogan.

Concretely, write matching IPv6 filters, tests, and logging, and probe with IPv6 as well as IPv4.

The reason for that specificity is a failure I have seen: IPv4 ACL blocked admin; IPv6 was open. 1,400 probes reached the service over IPv6 in 1 day.

IPv4 closed, IPv6 open.

FamilyACLProbes in 1 day
IPv4tight0
IPv6absent1,400

I would not consider it settled without evidence: run the same test over AAAA and require the same deny or permit.

One family filtered is not the host filtered.

Curated: · Written: · Reviewed:

QA-95Four anycast DNS PoPs share an address. One is overloaded and still advertised. What do some clients get?(show answer)

The engineering content of anycast DNS is the return path and the more-specific prefix, not the protocol name.

Anycast DNS steers by routing, not by a DNS health protocol. An unhealthy PoP that still announces the service address receives a share of clients who will SERVFAIL.

Concretely, withdraw the anycast prefix when local health fails, and monitor SERVFAIL per PoP, not only a global query success average.

The reason for that specificity is a failure I have seen: One of four PoPs stayed advertised while broken. 28 percent of clients hit SERVFAIL for 22 minutes; the global success average looked "mostly fine."

One sick PoP still announced.

PoPsUnhealthy still advertisedClient SERVFAILMinutes
4128 percent22

I would not consider it settled without evidence: fail a PoP's nameserver and require its anycast prefix to be withdrawn from the routing domain.

Anycast advertisement is the health signal.

Curated: · Written: · Reviewed:

QA-96One side is 9000 and the other is 1500, and ICMP too-big is filtered. Which frames die?(show answer)

Before calling jumbo MTU mismatch done I would write down the ECMP member, VLAN, or VRF nobody hashed onto.

A jumbo sender toward a 1500 hop with DF set and no ICMP too-big black-holes large frames. NFS and other large UDP/TCP then fail while ping works.

Concretely, match MTU end to end, or clamp, and permit too-big; test with 9000-byte frames.

The reason for that specificity is a failure I have seen: NFS at jumbo met a 1500 hop. 33 percent of NFS ops failed for 4 hours; ICMP echo of 64 bytes succeeded.

9000 into 1500.

FrameForwardedNFS ops failedHours
64-byte echoyes—4
jumbo NFSno33 percent4

I would not consider it settled without evidence: send 9000-byte DF frames and require either delivery or a too-big message.

Jumbo is a path property, not a NIC checkbox.

Curated: · Written: · Reviewed:

QA-97You SPAN 10 Gbps of sources to a 1 Gbps analyser port. How much of the evidence is missing?(show answer)

The first thing I would establish about SPAN oversubscription is which forwarding entry the packet actually hits.

SPAN is not lossless when the destination is slower than the source. Oversubscription drops span frames silently. The analyser then lies about loss on the production path.

Concretely, size the SPAN destination, use a TAP, or filter the session, and treat analyser drops as analyser drops.

The reason for that specificity is a failure I have seen: 10 Gbps offered to a 1 Gbps destination dropped at least 90 percent of span traffic. A 2-hour hunt chased production loss that was SPAN oversubscription.

10G source, 1G dest.

PortSpeedSpan frames droppedHours of wrong hunt
sources10 Gbps—2
dest1 Gbpsat least 90 percent2

I would not consider it settled without evidence: compare production interface counters with analyser received frames.

A SPAN port has an MTU and a speed.

Curated: · Written: · Reviewed:

QA-98The pool has 254 addresses and 0 free. New hosts get what?(show answer)

I would start DHCP pool exhaustion from the packet and the return path, not from the protocol adjacency.

A DHCP pool that is full offers nothing. New hosts fail DORA. This is capacity, not a relay bug, and reservations still consume pool space unless excluded.

Concretely, count used versus free, expand or reclaim, and alert before 0 free.

The reason for that specificity is a failure I have seen: The /24 pool hit 0 free. New hosts failed for 46 minutes; existing leases kept working, which looked like "DHCP is up."

254 used, 0 free.

Pool sizeFreeNew DORAMinutes
2540no Offer46
existing leases—still valid46

I would not consider it settled without evidence: show free addresses and a Discover with no Offer.

A leased-up pool is a silent deny.

Curated: · Written: · Reviewed:

QA-99You clamp MSS on the VPN but not on the internet edge. Which sessions still black-hole on PMTUD failure?(show answer)

This is an area where a green session and a working path are different observations.

MSS clamping is per rewrite point. A clamp on the tunnel does not rewrite sessions that never enter it. Scope is the interfaces that see the SYN.

Concretely, clamp on every path that has reduced MTU, and test SYNs on VPN and on clear internet separately.

The reason for that specificity is a failure I have seen: VPN sessions were clamped; internet sessions were not. 18 percent of internet large transfers failed for 2 hours on a 1400-MTU middlebox.

Clamp on one path only.

PathMSS clampedLarge-transfer failHours
VPNyes0 percent2
internet edgeno18 percent2

I would not consider it settled without evidence: read MSS on SYN for both paths and require the clamp only where the SYN actually passed.

MSS clamp is an interface action, not a campus spell.

Curated: · Written: · Reviewed:

QA-100Twelve hops look fine from one ping. What does a complete validation still have to prove?(show answer)

My answer to end-to-end path validation begins with which plane is making the decision: control, forwarding, or policy.

End-to-end validation is both directions, every ECMP member, the reduced-MTU size, the application five-tuple, and the VRF the user is in. A single ping proves one hash, one size, one direction, one protocol.

Concretely, run a matrix of hashes, sizes, directions, and the real port, and record which member each probe used.

The reason for that specificity is a failure I have seen: A 12-hop ping succeeded. One of four hashes dropped large UDP on the return path. The application failed until a diverse probe found it 33 minutes later.

One ping, four hashes.

ProbeHopsLarge UDP returnMinutes until found
one ping12not tested33
hash members 1–312ok—
hash member 412drop33

I would not consider it settled without evidence: require a recorded matrix covering members, directions, and sizes before signing the path.

The path is a matrix, not a ping.

Curated: · Written: · Reviewed: