Overview
Curated: · Written: · Reviewed:
Balance admitted work, not just connection counts
A load balancer selects a backend or path for traffic so a service can scale, tolerate failures, and enforce a consistent edge contract. Distribution is only one responsibility. Production behavior also depends on discovery, health, capacity, connection reuse, protocol semantics, retries, timeouts, draining, overload, security, topology, and observability. A healthy load-balancer process can still amplify an outage by sending accepted work to an overloaded or semantically broken dependency.
In an interview, the strong answer is rarely "round robin vs least connections." The interviewer is probing whether you understand that a balancer balances work, not connections or selections, and whether you can reason about what happens when a backend is slow, a member leaves, or a retry storm starts. A weak answer recites algorithms and stops; a strong answer picks an algorithm, states the workload assumption it encodes, and names the failure mode when that assumption breaks.
Layer 4 vs Layer 7: what the balancer terminates
Layer 4 balancing uses transport and network tuples — source/destination IP and port — without interpreting application requests. Two deployment shapes dominate:
- Proxy/NAT mode: the balancer terminates the TCP connection, translates the client address to its own (SNAT), and opens a second connection upstream. Every byte crosses the balancer twice (in and out), and the balancer owns connection state: sequence numbers, retransmission, and half-open detection.
- DSR (direct server return): the balancer forwards the request to a backend but the backend replies directly to the client, bypassing the balancer. This exists because return traffic is usually the bulk of the bytes (a small HTTP request, a large response), so DSR roughly halves the balancer's bandwidth requirement and removes it from the response path's latency. The cost is that the backend must be able to reach the client directly and the balancer sees only half the conversation, which breaks L7 features and per-request observability.
Layer 7 proxies terminate or understand HTTP or another application protocol. That buys host/path/header routing, authentication integration, retries, and richer metrics, and costs additional connection, trust, parsing, and resource boundaries: the proxy now owns two connection lifetimes per request, must parse untrusted input, and becomes a TLS termination point. TLS passthrough retains end-server termination; TLS termination at the balancer exposes plaintext there and may require authenticated re-encryption upstream.
The follow-up interviewers use here: "What breaks if you move from L4 to L7?" Weak answers say "L7 is slower." Strong answers name the concrete changes: client IP visibility now depends on forwarded headers, connection counts double, the balancer becomes a parsing attack surface (request smuggling), and idle/keep-alive semantics now belong to the proxy.
VIP plumbing: how the address actually reaches the balancer
Before any algorithm runs, something has to make traffic arrive at the balancer at all. The common mechanisms:
- ARP announcement / floating IP: a virtual IP is announced via ARP (or NDP for IPv6) by exactly one active node at a time; failover tools like keepalived (VRRP) move the announcement. Convergence is fast on a LAN but clients and adjacent routers cache ARP entries, so the failover timeline includes ARP cache expiry unless gratuitous ARP is honored.
- Anycast: the same IP is announced from multiple sites; routing steers the client to the nearest announcer. Failover is a route withdrawal, which propagates on BGP timescales (seconds to minutes) and depends on the client's resolver/router behavior.
- DNS steering: the name resolves to different IPs per region. Cached answers mean you cannot instantly revoke anything — a TTL of 60s means up to 60s of stale traffic after a withdrawal.
Know which layer your failover actually lives at. A weak answer says "we failed over" without saying whether that meant an ARP move, a route withdrawal, or a DNS TTL expiring — three very different timelines.
Algorithms and the assumptions they encode
- Round robin distributes selections, not equal work. If request cost varies (some endpoints do 1ms, others do 500ms), round robin produces skewed load. Its failure mode is uneven session or request duration.
- Weighted round robin fixes known capacity differences, but weights must be tied to measured sustainable capacity, not guesses. Stale weights after a backend is resized are a classic silent imbalance.
- Least connections helps when connection duration tracks cost. It fails when HTTP/2 or pooled connections carry many unequal streams — one connection holding 200 active streams counts as "1."
- IP hash / consistent hashing provides affinity (and cache friendliness). Consistent/ring hashing limits how many keys remap when membership changes — adding a node moves only roughly 1/N of keys, not everything. Its failure modes: hot keys concentrate on one node, and hashing cannot repair a session design that breaks when the session's home node dies.
- Randomized choice scales well and avoids shared-state coordination; over enough requests it approximates uniform, but short windows can be lumpy.
The question to be ready for: "Given our workload, which would you pick and what would you measure to know it's wrong?" Measure actual requests, bytes, concurrency, latency, errors, and resource saturation per backend — not selection counts.
Health checks and the failover timeline they produce
Health has layers. A transport connect proves a listener accepted; an HTTP success may prove one shallow dependency; neither proves the backend can serve representative work within budget. Active checks probe predefined paths on an interval; passive health uses real request failures and catches what synthetic checks miss.
Do the math out loud in the interview — interviewers probe whether you can derive the failover timeline rather than quote a default:
- Detection time ≈ failure threshold × check interval. A check every 2s with a 3-failure threshold means a dead backend keeps receiving traffic for ~6s plus one in-flight check timeout.
- Add ejection propagation (control-plane convergence) and connection drain before the last in-flight request is done.
- So end-to-end failover = detection + ejection + drain. If your SLO budget is 5s, a 2s interval with threshold 3 already blows it.
Use consecutive thresholds, intervals, timeouts, jitter, and recovery hysteresis to prevent flapping and synchronized probes. Readiness should reject new work before a process is fully initialized or when it cannot serve; liveness should avoid restarting a merely overloaded dependency and causing a restart loop.
Draining, retries, and overload: the failure modes interviewers dig into
Draining is a state transition: remove a backend from new selection, allow bounded in-flight requests to finish, then force close at a documented deadline. Account for HTTP/1 keep-alive, HTTP/2 GOAWAY, WebSockets, QUIC, long polling, and background work. Deregistration delay does not prove applications stopped accepting jobs. Deployments must coordinate endpoint readiness, balancer convergence, process shutdown, and rollback.
Retries create new load and can duplicate effects. Retry only when the method and operation are idempotent or protected by an idempotency key and durable outcome record. Respect an end-to-end deadline and remaining attempt budget, use bounded exponential backoff with jitter, avoid retrying at every layer, and distinguish connect failure, reset before request, partial response, timeout, overload, and explicit retry-after. Hedging reduces tails but intentionally adds parallel work and must be tightly budgeted.
Timeouts form a budget hierarchy: DNS, connect, TLS, headers/body, upstream queue, response, idle, and total deadline. A proxy timeout is not automatically safe to retry — the backend may already have committed. Circuit breakers cap pending requests, connections, and retries; load shedding rejects early when admitted work would violate objectives. Queueing hides overload while increasing age and memory. Scale-out helps only if discovery, warm-up, downstream capacity, and control-loop delay are modeled.
Connection pooling amortizes handshakes but changes balancing granularity and failure lifetime. HTTP/2 multiplexes streams on one connection, so least-connections misreads load and one connection failure affects many streams. QUIC connection IDs and migration change tuple assumptions. Test backend restart, stale DNS, half-close, GOAWAY, certificate rotation, and asymmetric packet paths.
Session persistence: why it exists and what it costs
Sticky sessions exist because state lives somewhere: TLS session resumption state, in-memory shopping carts, long-lived WebSocket subscriptions. If the state is in the process, the request must land on the process.
Mechanisms, roughly in order of preference:
- Application-level: a shared session store (Redis, a database) keyed by a cookie — no balancer involvement, any backend can serve.
- Cookie-based affinity: the balancer (or the app) sets a cookie naming the chosen backend; subsequent requests route by it. Survives balancer restarts better than IP-based methods.
- IP hash: no cookie needed, but breaks for clients behind NAT/CGNAT (one IP = many users, all pinned to one backend) and for mobile clients whose IP changes mid-session.
The costs interviewers want named: persistence fights the balancing algorithm (a hot user pins one backend regardless of load), and it makes failover worse — when the sticky backend dies, its sessions either break or must be re-homed, and re-homing is only safe if the state was external anyway. The strong answer: "Sticky sessions are a patch for state in the wrong place; the durable fix is externalizing session state." That's also the likely follow-up: *"How would you remove the need for stickiness?"
Identity, security, and the edge contract
Client identity changes at a proxy. The source IP a backend sees is usually the proxy's unless transparent modes are used. Forwarded or X-Forwarded-* headers are trustworthy only when the edge strips untrusted incoming values and writes a canonical chain from an authenticated network boundary. Never use client IP alone as user authorization. Bound header size and hop count, protect PROXY protocol listeners from direct untrusted access, and preserve request correlation without exposing sensitive data.
Secure the administrative and data planes: authenticate configuration and discovery sources, apply least privilege, validate endpoint identity, patch parsers, rate-limit handshakes and requests, defend request smuggling through consistent HTTP framing, set header/body limits, manage certificates, and isolate tenants. A balancer is not a substitute for backend authorization or application input validation.
What to measure and what to test
Measure accepted/rejected requests, connection and TLS handshakes, active connections/streams, queue depth and oldest age, retries/hedges, timeout stage, response codes, bytes, latency by stage and backend, health transitions with ejection reason, endpoint membership/version, algorithm decisions and load skew, saturation, circuit-breaker and load-shed events, config propagation, DNS/global steering, certificate expiry, and downstream application outcomes.
Validate nominal distribution plus the failure modes: slow/failed/partial backends, flapping checks, hot keys, long connections, HTTP/2 multiplexing, retry storms, queue saturation, large/slow bodies, malformed framing, cert failure, discovery staleness, zone/region loss, drain/restart, overload recovery, config rollback, and negative authorization.
