Overview
Curated: · Written: · Reviewed:
Separate forwarding state, control intent, and observed path
Switching moves frames within a link-layer domain; routing moves packets between IP prefixes. Both devices have a data or forwarding plane that handles traffic at speed and a control plane that learns or computes state. The management plane configures and observes them. Diagnose the installed forwarding entry actually used for a packet, not merely the route or VLAN an operator intended to configure.
An Ethernet switch learns a source MAC address and ingress VLAN/port, then forwards a known unicast toward the learned destination. Unknown unicast, broadcast, and relevant multicast may be flooded within the VLAN. Entries age and move; loops or unstable endpoints cause MAC flapping and storms. A VLAN is a broadcast-domain and policy boundary carried untagged on an access port or tagged over a trunk according to explicit allowed and native VLAN semantics. A matching VLAN number on two switches does not prove end-to-end carriage.
Redundant Layer 2 links create loops because Ethernet frames lack a hop limit. Spanning Tree elects a root and blocks selected paths to create a loop-free active topology. Root choice, port roles, costs, protection features, and convergence must be designed rather than accepted from incidental MAC priorities. Edge/portfast behavior belongs only on true endpoints; BPDU guard, root guard, loop guard, storm control, and broadcast containment address different faults. Link aggregation uses a negotiated or static bundle, but member speed, VLAN, system ID, key, hashing, and failure behavior must be consistent.
Routers choose the most specific matching destination prefix—longest-prefix match—then use next-hop resolution and an egress adjacency to rewrite the link envelope. A routing information base can contain candidates from connected, static, OSPF, BGP, and other sources; policy, preference, metric, and recursion determine the selected route. The forwarding information base is programmed from that result and can lag or fail independently. Equal-cost multipath commonly hashes flows across nexthops, so one probe may not represent every path.
Route aggregation reduces state and limits churn, but a summary asserts reachability for the covered address space. Install a discard route or equivalent and advertise only when the intended components and failure policy justify it, or traffic can loop or black-hole. Default routes are broad summaries and should have clear ownership, tracking, and fail behavior. Validate return routing: stateful firewalls, NAT, policy routing, and ECMP make asymmetry operationally important even though IP forwarding itself can be asymmetric.
OSPF is a link-state interior routing protocol. Neighbors form adjacencies after compatible parameters and exchange a link-state database; each router runs shortest-path computation. Areas limit flooding and calculation scope, with area 0 providing backbone structure. Network type, timers, authentication, MTU, router ID, area, stub flags, and duplicate addressing can prevent or destabilize adjacency. A FULL neighbor does not prove that the desired prefix is originated, preferred, installed, programmed, or reachable in both directions.
BGP is a policy-driven path-vector protocol. Sessions exchange NLRI with attributes such as AS_PATH, NEXT_HOP, origin, local preference, MED, and communities. Route selection and export are implementation and policy controlled; BGP does not automatically choose the physically shortest or lowest-latency route. Treat import and export policy as deny-by-default, validate prefixes and AS paths, cap accepted routes, and use routing registries or RPKI where applicable. A session can be Established while silently advertising or accepting the wrong routes.
First-hop redundancy protocols such as VRRP provide a virtual router address with master election and failover. They do not synchronize arbitrary firewall/NAT sessions or prove the new master has upstream reachability. Track critical dependencies, coordinate advertisement timers and preemption, update neighbor state, and test partial failures. BFD can rapidly detect forwarding-path failure for clients such as routing protocols, but aggressive timers consume resources and may flap under congestion or control-plane pressure.
VXLAN creates Layer 2 overlays over an IP underlay using VTEPs and VNIs. Flood-and-learn or EVPN can distribute endpoint information. Underlay reachability and MTU must cover encapsulation overhead. Separate tenant/VRF/VNI identities, control unknown traffic, authenticate control-plane peers, prevent route leaks, and validate endpoint movement and multihoming. An overlay can be healthy at BGP control plane while one underlay ECMP member drops large encapsulated packets.
Changes to routing and switching are distributed state transitions. Before change, capture topology, device and software versions, interface/VLAN/VRF, neighbors, RIB/FIB, MAC/ARP/NDP, STP/LAG, policy counters, traffic baseline, and rollback. Stage or canary where topology permits, use commit-confirmed or automated rollback, check both directions and diverse ECMP hashes, and watch convergence rather than only configuration acceptance.
Measure interface carrier/errors/discards, utilization and queues, MAC move and unknown flood, STP topology change and blocked roles, LAG members/hash imbalance, neighbor state, route count/churn, RIB-to-FIB programming, prefix/next-hop recursion, OSPF adjacency/LSA/SPF, BGP session/update/policy/RPKI, BFD/VRRP transitions, ECMP path outcomes, MTU/ICMP, packet loss/latency, CPU/memory, and end-to-end application transactions. Test link/member/device loss, one-way fiber, VLAN mismatch, loop, storm, MAC move, ARP/NDP failure, route withdrawal, policy error, summary black hole, BFD false positive, first-hop failover, control-plane restart, FIB exhaustion, MTU reduction, and rollback.
Longest-prefix match is the rule that decides which entry wins, and reading a routing table without applying it is the most common source of a wrong diagnosis. A table holding 10.0.0.0/8 via one next hop and 10.4.0.0/16 via another sends traffic for 10.4.1.7 down the /16 regardless of metric, administrative distance, or the order the entries appear in the output, because specificity is evaluated before any preference between sources. Preference decides only among candidates for the same prefix. A default route is simply the least specific entry, 0.0.0.0/0, so it is used when nothing else matches rather than as a fallback applied after a failure; a more specific route that remains installed while its path is broken will keep attracting traffic that the default would otherwise have carried.
Layer 2 and Layer 3 fail differently, and the difference decides where to look. An Ethernet frame carries no hop limit, so a Layer 2 loop does not decay: frames multiply through every redundant path until the broadcast domain saturates, which is why spanning tree blocks ports rather than relying on any per-frame counter. An IP forwarding loop is bounded by TTL and shows as a traceroute alternating between two hops and traffic that stops at a fixed distance, costing latency and loss rather than the whole segment. A MAC address moving between ports faster than its ageing timer points at a loop or a duplicated endpoint, not at a routing fault, and reading it as one sends the investigation to the wrong plane.
Convergence is a budget, not an event, and every protocol timer spends part of it. Detection is the largest term: waiting for OSPF dead intervals is typically tens of seconds, while BFD with sub-second intervals detects a failure in a fraction of that, which is why it is deployed alongside the routing protocol rather than instead of it. After detection come propagation, recomputation, and FIB programming, and the last of these scales with table size, so a device carrying a full Internet table converges more slowly than the protocol timers alone imply. State the target end to end, measure each term under a real link failure, and treat aggressive timers with care: intervals short enough to trip on ordinary jitter turn a stable network into one that reconverges for no reason.
Worked example: which route wins, and why
A router with several sources for the same destination does not pick the shortest path. It picks by administrative distance first, and only compares metrics within one protocol:
Destination 10.20.0.0/16, four candidates:
Static route via 192.0.2.1 AD 1 metric -
eBGP via 198.51.100.7 AD 20 AS-path length 2
OSPF intra-area via 203.0.113.4 AD 110 cost 12
RIP via 203.0.113.9 AD 120 hop count 3
Installed: the static route, AD 1.
The OSPF path may genuinely be faster; it is never compared, because administrative distance is evaluated before any metric. That is what makes a hand-typed static route such a durable outage: it wins against every dynamic protocol and does not withdraw when its next hop degrades.
Longest-prefix match is evaluated before all of this, which is the rule people invert:
| Route | Source | AD | Matches 10.20.30.40? | Wins? |
|---|---|---|---|---|
| 0.0.0.0/0 | static | 1 | yes | no - shortest prefix |
| 10.0.0.0/8 | OSPF | 110 | yes | no |
| 10.20.0.0/16 | static | 1 | yes | no |
| 10.20.30.0/24 | eBGP | 20 | yes | yes - longest prefix |
| 10.20.30.40/32 | - | - | - | would win if it existed |
The /24 learned by eBGP at AD 20 beats the /16 static at AD 1, because prefix length is compared first and AD only breaks ties among routes of equal length. An interviewer asking "does a static route always win" is usually testing exactly this: it wins at equal prefix length, and loses to any longer match.
The operational consequence: advertising a more specific prefix is how traffic gets hijacked or redirected, intentionally or not. A /24 leaked into BGP pulls traffic away from a correctly configured /16 anywhere it propagates, and no local configuration on the /16's owner prevents it.
