Skip to content
Tech Interview Prep home
Technical interview guide

Capacity Planning & Load Testing

Knowing how much traffic a system can take before it does, through modeling and deliberate load testing.

Read
28 min
Practice MCQs
25
Interview QA
25
Edition
v4
Editorial status
Reviewed

Scope: Google SRE, Kubernetes, OpenTelemetry, Prometheus, k6, JMeter, PostgreSQL, and IETF guidance current 2026-08-31.

Overview

Curated: · Written: · Reviewed:

Capacity planning turns demand and failure assumptions into tested operating limits

Capacity is the amount of useful work a system can sustain while meeting explicit user, correctness, reliability and cost objectives. Planning begins with demand, not CPU: identify critical journeys and workload units; forecast arrival rate, concurrency, payload and dataset size; segment by tenant, region and request class; model peaks, bursts, seasonality, launches and organic growth; and state uncertainty. Translate demand through the dependency graph into compute, memory, connections, storage IOPS, network, quotas, queues, workers and third-party capacity. Include failover and maintenance states: a fleet that serves normal traffic only when every zone is healthy has no defensible resilience headroom.

The usable limit is rarely one advertised hardware maximum. It is the first constraint that causes an objective or safety boundary to fail: tail latency, error rate, stale or lost work, connection exhaustion, queue age, throttling, memory pressure, downstream quota, recovery time, or unacceptable unit cost. A utilization target is therefore an operating policy tied to response time, scaling delay, forecast error and failure reserve—not a universal percentage. Record assumptions, owners, measurement provenance, confidence ranges, lead time and an action threshold. Revisit the model when workload shape, code, data, topology, dependencies or objectives change.

Load testing is an experiment that evaluates those assumptions under controlled demand. A smoke test checks the scenario and instrumentation at low load. An average-load test checks expected steady operation. A stress or breakpoint test explores saturation and failure behavior. A spike test tests sudden arrival changes and autoscaling lag. A soak test exposes leaks, fragmentation, queue drift, compaction and thermal or quota effects. A resilience-capacity test removes a zone, dependency or capacity slice while load continues. Each test needs a written hypothesis, representative environment and data, workload model, ramp, steady period, safety guardrails, abort authority, expected bottleneck, success criteria and reproducible versioned configuration.

Representativeness is multidimensional. Match request mix, arrival process, concurrency, payload and response sizes, cache warmth, data cardinality and skew, authentication, think time, background jobs, retries, writes, hot keys, tenant distribution and dependency latency. Closed-loop virtual users can understate overload because slow responses reduce offered load; an open arrival-rate model better represents demand that continues while the service slows. Coordinated omission can hide the delay users experience when a generator waits behind the system. A test that bypasses TLS, authorization, edge controls or a production bottleneck measures a different system and must say so.

Observe from the client through every constrained dependency. Measure achieved and attempted rate, concurrency, success and rejection classes, latency distributions by journey, queue depth and age, saturation, garbage collection, connections, locks, I/O, network, autoscaler decisions, cold starts, cache behavior, database plans, downstream quotas and cost per useful result. Preserve versions and traces for diagnosis, but control telemetry cardinality and load. Averages conceal tails and mixed populations; p99 alone can also mislead without sample count, interval, histogram design and segment context. Validate correctness and durability under load—fast incorrect, duplicated or silently dropped work is not capacity.

Run tests progressively. Start below the expected limit, confirm measurement, then change one principal variable at a time where feasible. Use isolated or production-like environments by default; production tests require explicit approval, tenant/user protections, bounded blast radius, rate ceilings, abort automation, synthetic data, cleanup and incident readiness. Prevent the generator from becoming the bottleneck by measuring its CPU, network, scheduling lag and connection limits, and distribute it only with clock and result aggregation controls. Repeat runs to distinguish a stable limit from noise and compare against a controlled baseline.

Overload behavior is part of the capacity contract. Bound queues; apply admission control and fair rate limits before scarce work; propagate deadlines and cancellation; use retry budgets, exponential backoff and jitter; shed optional work; protect health/control paths; and degrade intentionally. Rejecting excess work quickly can preserve successful capacity better than accepting everything into a growing queue. Validate recovery after load falls: retries, backlog and autoscaling oscillation can keep a system overloaded. Test fairness so one tenant or expensive request class cannot consume the whole reserve.

Turn evidence into decisions. State the sustainable envelope, bottleneck sequence, objective margin, failure-mode capacity, scaling/recovery time, forecast scenarios, cost curve and residual uncertainty. Define trigger thresholds far enough ahead to cover procurement, quota, sharding or engineering lead time. Prefer removing the limiting constraint or demand rather than scaling an unrelated tier. Re-run after material change and before high-risk events; trend forecast error and tested headroom. A benchmark number is not a permanent certification, and a passing test at one traffic mix does not prove every mix, failure state or production environment.

Reading a capacity result honestly

A load test produces a number, and the number is only as good as the assumptions behind the traffic that generated it. Synthetic load that reuses a small key set, skips authentication, or holds a fixed think time will exercise caches and connection pools in a way real traffic does not, and the resulting headroom estimate can be several times too optimistic. Before quoting a capacity figure, state what the generated workload looked like — request mix, key cardinality, payload sizes, concurrency ramp, and client behaviour on error — and say which of those differ from production.

The second honesty question is what limit was actually found. A test that stops at the first error found the point where something broke, which may be a client-side limit, a generator bottleneck, or a downstream quota rather than the service under test. A test that ramps until latency crosses the objective found a service-level limit, which is the useful one. Reporting throughput without the latency objective it was measured against, or without saying whether the system was still stable at that point, produces a figure that cannot be used for planning.

Finally, capacity is a moving figure rather than a property. Traffic mix, retry behaviour, dependency latency, dataset size and code changes all move it, so a capacity number carries an expiry the same way a benchmark does. The operational form of this is a headroom metric derived continuously from production signals, with the load test used to calibrate it rather than to replace it.