Skip to content
Tech Interview Prep home
Technical interview guide

Network Troubleshooting

Systematically isolating where in the stack a connectivity problem actually lives, instead of guessing.

Read
45 min
Practice MCQs
25
Interview QA
25
Edition
v4
Editorial status
Reviewed
Relevant for
Network Engineer

Scope: Current Wireshark, tcpdump, Linux iproute2/ss/tc/ethtool, IETF host/TCP/ICMP/PMTUD/DNS/TLS references reviewed 2026-09-04.

Overview

Curated: · Written: · Reviewed:

Narrow the failing transition with comparable evidence — and say so out loud in the interview

Interviewers grade network troubleshooting answers on process, not on whether you name the right fault first try. What they probe for:

  • Do you scope before you touch anything? Weak answer: "I'd check the switch." Strong answer: "First, is this one user, one subnet, or the whole site? That decides my first check."
  • Do you isolate along a path, or guess-and-reboot? Weak answer: "I'd restart the service and see." Strong answer: "I'll test the gateway, then hop by hop, and compare against a known-good control."
  • Do you know what each tool actually proves? Weak answer: "ping the server to check it's up." Strong answer: "ping proves ICMP reachability and RTT; it says nothing about the TCP listener or the app."
  • Can you name signature faults from symptoms? Interviewers love "small packets pass, large ones stall" → PMTUD/MTU black hole, and "ping works but the app is dead" → filtered ICMP vs dead listener.

Expect follow-ups like: "What does a retransmission in your capture actually prove?" (missing ACK at that observer, not where loss happened), "Why is connecting directly to the IP not an equivalent test?" (bypasses DNS, changes SNI/Host/LB routing), and "What would you never do on a production path?" (disable firewall or TLS validation globally to "see if it works").

Scope triage before touching a device

Before any command, classify the blast radius: one user, one subnet/VLAN, one site, one application, or everything. The classification picks your first check:

  • One user → local machine: address, mask, gateway, DNS settings, local firewall.
  • One subnet/VLAN → that segment's gateway interface, ARP/NDP table, switch port counters, DHCP pool.
  • One site → site exit: WAN link counters, routing adjacency, NAT state, upstream path.
  • One application → listener, load balancer, backend health, DNS answer for that name.
  • Everything → shared dependencies: DHCP, DNS resolver, core switching (STP event?), authentication.

Begin with a precise failing transaction: client and server identities, source and destination addresses, protocol and port, time window, expected outcome, observed outcome, frequency, affected cohorts, topology, and recent change. Preserve state before restarting or flushing it — a cleared ARP cache or restarted process destroys the evidence. "The network is slow" is not a testable problem; "new IPv6 clients in one site time out during TLS after TCP connects, while IPv4 succeeds" is.

Three isolation strategies, and why the middle wins under time pressure

  • Bottom-up (OSI 1 upward): carrier, errors, ARP, route, transport, app. Thorough, but slow when the fault is at layer 7.
  • Top-down (app downward): start at the application and logs. Fast when the app layer is clearly implicated, wasteful when it's a bad optic.
  • Divide-and-conquer: test the middle of the path first — can the client reach its default gateway? Then does the traffic get past hop N? Each answer halves the path. This is the strategy to name in an interview when asked "how would you start, you have ten minutes": gateway ping, then hop-by-hop traceroute/mtr against a healthy control, and the branch point tells you which half to descend into.

Draw the packet and control path in both directions: client namespace/VRF, resolver, interface/VLAN, neighbor and gateway, routing and ECMP, firewall/NAT, tunnel, load balancer, server listener, downstream app. For each boundary, predict the tuple, encapsulation, MTU, state and counters. Compare a healthy control that differs in one dimension. Never change several layers at once — improvement then destroys causal evidence.

Each tool answers exactly one question

ToolThe one question it answersWhat it does not prove
pingICMP reachability, RTT, loss patternThat the TCP service or app works; ICMP may be filtered
traceroute / mtrWhere in the path latency or loss beginsWhy it begins there; per-hop response policies vary
ARP / neighbor tableL2 resolution state for the next hopAnything beyond that segment
dig / resolvectlName resolution: answer, TTL, resolver, validationTransport reachability to the returned address
ss / netstatIs the app listening, on which address, queue depthThat the path to it works
Interface countersErrors, CRC, duplex, drops, queue saturationIntermittent faults since last clear (check deltas, not lifetime totals)
Route table / FIBControl-plane (and forwarding) lookup for exact src/dst/VRFThe return path, or that stateful middleboxes allow the flow
Packet captureFlags, retransmissions, alerts, payload, timingGround truth at that observer only

Start at local facts: expected address family and destination, socket/listener state, carrier, address/prefix, route and policy lookup for the exact source/destination, next-hop neighbor state, local firewall. An interface marked UP does not prove link carrier, VLAN carriage, error-free transmission, or a usable gateway. A route in the control table does not prove the forwarding path or the return route.

The splits that decide the branch

Every fork in the diagnosis is a split; name the split, then take the branch:

  • L2 vs L3: does the client resolve its gateway's MAC? ARP/NDP state, duplicate addresses, VLAN membership. Neighbor resolution is a prerequisite for local delivery — and don't paper over it with a static neighbor entry; that hides a failed broadcast/multicast or trust boundary.
  • Data plane vs control plane: is the route present but not forwarded (FIB drift, ACL, hardware table limits), or absent entirely (adjacency down)?
  • Name resolution vs transport reachability: dig succeeds but the connection times out → path problem. Connection to the IP works but the name fails → resolver problem. And connecting to an IP bypasses more than DNS: it changes TLS SNI, HTTP Host, LB routing, certificate validation, and virtual hosting — so it is never an equivalent application test.
  • One flow vs all flows: only new IPv6 clients? Only large uploads? Only one LAG member? A duplex mismatch, bad optic, overloaded queue, or one-way link can affect only some flows — check every selected member and correlate both switch and host ends.
  • ICMP reachability vs TCP service reachability: ping works but the app is dead → listener, firewall, or LB health. Ping fails but the app works → ICMP filtered. Never treat either as proof of the other.

Signature faults and their symptoms

Recognizing these from the symptom alone is what makes a senior answer sound senior:

  • Duplex mismatch / bad optic / dirty fiber: CRC and error counters climbing as rates under traffic, not lifetime totals; often after a speed change or one side hard-coded.
  • MTU black hole / broken PMTUD: handshakes and small packets pass, large payloads stall mid-transfer. Look for tunnels, DF set, blocked ICMP Fragmentation Needed / Packet Too Big, and check TCP MSS plus offloads. Probe packet-size boundaries and capture inner and outer packets on both sides of encapsulation. Blocking ICMP or blindly lowering MSS hides only part of the problem.
  • Missing default gateway / wrong mask vs DNS failure: can you reach IPs by address but not by name? Then it's resolution, not routing. Wrong mask shows up as "works to nearby hosts, not to far ones."
  • Duplicate IP / ARP conflict: MAC flapping in the neighbor table, intermittent reachability that follows the duplicate's traffic pattern.
  • Missing VLAN / trunk or native-VLAN mismatch: port up, carrier up, no reachability; the frame never arrives tagged the way the far side expects.
  • STP loop or reconvergence: broadcast storm, switch CPU spikes, intermittent loss, MAC tables churning.
  • Missing route or asymmetric return path: return traffic can follow different routing and state boundaries — a stateful firewall that sees only one direction drops the flow. Routing analysis must use the exact source, destination, VRF, policy rule, longest-prefix result, recursive nexthop, ECMP hash, and reverse path.
  • DHCP pool exhaustion: new clients fail while existing leases keep working; check utilization before anything else.
  • ACL dropping traffic: silent, consistent, cohort-specific; check hit counters and log drops rather than guessing.
  • Congestion / buffer / QoS: latency and loss that scale with load; queue and drop counters move with traffic.

Reading the transport and TLS layers

TCP diagnosis follows the state machine and byte sequence: does the SYN reach the listener, does the SYN-ACK return, does the final ACK arrive, does TLS begin, are request bytes acknowledged, does the response arrive? Inspect sequence ranges, cumulative/SACK acknowledgments, retransmission timing, RTT, receive window, congestion evidence, reset/FIN direction, socket queues, application timing. A retransmission at one capture point proves missing acknowledgment at that observer, not where or why loss occurred.

TLS failures occur after or alongside transport success. Preserve SNI, ALPN, protocol version, certificate chain and name, trust, validity, alert direction, client auth, and proxy termination boundaries — without logging secrets. A TCP connect test cannot validate TLS or application authorization; an HTTPS success through one proxy cannot validate a direct backend path with a different identity.

Latency accumulates across queueing, propagation, serialization, processing, retransmission, setup, application work, and client consumption. Synchronize clocks or use single-observer deltas; separate network RTT from server response time from end-to-end time.

Captures, safety, and validation

Packet capture is an observation with limitations. Record host/interface/direction, filter, snap length, promiscuous mode, timestamp source, dropped packets, tool/version, and offload state — host transmit captures may show pre-offload large segments or incomplete checksums that never appear that way on the wire. A capture at a proxy sees a different connection than the client. Preserve raw pcap, command output, clocks, configs and hashes; restrict sensitive payload access; redact only derived copies.

Troubleshooting must be safe. Prefer read-only observations, then reversible, narrowly scoped experiments with an owner, expected signal, timeout, abort, and rollback. Never disable firewalls, certificate validation, redundancy, or rate limits globally to "see whether it works." During incidents, keep a timeline of hypotheses, evidence, changes, and outcomes. After mitigation: reproduce the fault where safe, identify the failed mechanism and the detection gap, add regression coverage or monitoring, verify rollback left no residue, and document what remains uncertain.

Strong validation injects one controlled fault at a time — loss, delay, reorder, bandwidth, MTU, route, DNS, certificate, process, NAT timeout, queue saturation — in an isolated environment, checks both protocol evidence and user outcome, and removes the impairment with verified cleanup. Segment every measurement by site, path, address family, interface/member, client/server version, payload size, and time.