Skip to content
Tech Interview Prep home
Technical interview guide

Scalability Fundamentals

Production scalability fundamentals for technical interviews: bottlenecks, scaling, load balancing, autoscaling, capacity, overload control, and failure behavior.

Read
18 min
Practice MCQs
25
Interview QA
25
Edition
v7
Editorial status
Reviewed

Scope: Vendor-neutral fundamentals; Google SRE, AWS Well-Architected, Google Cloud, and IETF references accessed 2026-08-30..

Interview QA

Treat each question like a live interview question: answer out loud first (structure, assumptions, tradeoffs), then open the model answer to spot gaps and rehearse a tighter follow-up.

Curated: · Written: · Reviewed:

QA-1

Why is horizontal scaling generally preferred over vertical scaling for large-scale systems, despite vertical scaling being simpler?

QA-2

Explain the difference between throughput and latency using a concrete example, and describe a scenario where optimizing one hurts the other.

QA-3

Why does statelessness matter for horizontally scaling an application tier?

QA-4

How do you identify the bottleneck in a system before deciding what to scale?

QA-5

How would you design a load balancer's health check to avoid both false positives (marking a healthy server unhealthy) and false negatives (missing a genuinely unhealthy server)?

QA-6

Explain how auto-scaling based on CPU utilization can sometimes react too slowly to a sudden traffic spike, and what mitigation strategies address this.

QA-7

Why can scaling out the application tier alone fail to improve overall system throughput?

QA-8

How does a CDN reduce load on an origin server, and what kind of content is it best suited for versus poorly suited for?

QA-9

Explain graceful degradation with a concrete example of a feature you'd disable under heavy load to keep a system's core functionality responsive.

QA-10

How would you implement backpressure in a system where a fast producer service is overwhelming a slower downstream consumer?

QA-11

Why is capacity planning based on peak load (not average load) essential for systems with highly variable traffic patterns, like e-commerce during a holiday sale?

QA-12

How would you identify whether a system's bottleneck is CPU-bound, I/O-bound, or network-bound, and why does the answer change what scaling approach helps?

QA-13

Explain why stateless application servers alone don't guarantee a horizontally-scaled system is fully scalable, using the database as a counter-example.

QA-14

How would you use a combination of vertical and horizontal scaling pragmatically, rather than treating them as mutually exclusive choices?

QA-15

How would you use a load test to validate a capacity plan before a known high-traffic event, and what specifically would you look for in the results?

QA-16

Explain why round-robin load balancing can perform poorly when backend servers have uneven processing times for different requests.

QA-17

How would you design a system to gracefully handle a sudden 10x traffic spike without any advance warning, combining several of the techniques discussed?

QA-18

Why does a stateless application tier still need to think carefully about where session-related state (like a shopping cart or login session) actually lives?

QA-19

How would you determine whether a system's bottleneck under load is the application code itself versus the infrastructure it's running on?

QA-20

Explain the relationship between Amdahl's Law and horizontal scaling — why adding more parallel workers eventually yields diminishing returns.

QA-21

How would you use rate limiting to protect a system from being overwhelmed, and what's the tradeoff of setting the limit too conservatively versus too permissively?

QA-22

Why might a system choose to shed load (deliberately reject some requests) during an overload event rather than trying to serve every request slowly?

QA-23

How would you decide between scaling a stateless application tier reactively (auto-scaling based on live metrics) versus proactively (scheduled scaling for known traffic patterns)?

QA-24

Explain why a system's overall availability can be lower than any individual component's availability, when multiple components are chained together sequentially.

QA-25

Why does over-provisioning capacity beyond the historically observed peak matter for resilience, using a partial-failure scenario as an example?