Skip to content
Tech Interview Prep home
Technical interview guide

Model Serving & Inference Infrastructure

The infrastructure choices behind getting predictions out of a trained model at production latency and scale.

Read
45 min
Practice MCQs
25
Interview QA
25
Edition
v6
Editorial status
Reviewed

Scope: NVIDIA Triton, KServe, Kubernetes autoscaling, Amazon SageMaker inference, and NIST AI RMF guidance current 2026-09-04.

Overview

Curated: · Written: · Reviewed:

Serve a versioned prediction contract, not merely a model file

Model serving turns a trained artifact into a dependable decision capability. The production contract includes input and output schemas, preprocessing and postprocessing, model and dependency versions, hardware assumptions, latency and availability objectives, authorization, observability, fallback behavior, and rollback. A server that can load weights is only one component.

Choose the inference pattern from the user need. Synchronous online inference fits interactive decisions with strict end-to-end deadlines. Asynchronous inference accepts a request, queues longer work, and returns a result later. Batch inference processes bounded datasets efficiently without a continuously running endpoint. Streaming inference is appropriate when events must be scored continuously. Do not operate an expensive low-latency endpoint when freshness permits scheduled batch output.

Budget latency end to end: client and network time, admission and queueing, serialization, feature retrieval, preprocessing, model execution, postprocessing, and downstream work. Report percentiles and deadline misses by model version and request class; an average hides tail pain. Throughput, concurrency, batch size, sequence length, payload size, accelerator type, precision, and memory pressure must be measured together under a representative arrival pattern.

Dynamic batching combines compatible stateless requests to raise accelerator utilization and throughput. Its queue delay consumes part of the latency budget, so tune maximum batch size and wait time using load tests rather than folklore. Stateful sequence workloads require correlation and ordered routing instead of arbitrary dynamic batching. Bound every queue: overload should produce explicit backpressure, rejection, or asynchronous admission rather than unbounded latency and memory growth.

Autoscale on a signal that leads demand. CPU may be weak for GPU-bound inference; useful signals include in-flight requests, queue depth, concurrency, request rate, accelerator utilization, or estimated work such as input tokens. Account for cold-start and model-load time, retain safe warm capacity for strict SLOs, cap scale-out against quotas, and use stabilization to avoid oscillation. Scale-to-zero trades idle cost for cold-start latency and belongs only where that trade is acceptable.

Package immutable artifacts and pin the full serving environment. Validate checksums and provenance before load. A versioned repository or registry should connect the model to code, tokenizer or feature contract, runtime image, dependencies, evaluation evidence, approval, and rollback target. Readiness must remain false until the correct version is loaded and warmed; liveness should detect a stuck process without causing destructive restart loops during ordinary overload.

Optimize only against a measured quality boundary. Quantization, compilation, pruning, distillation, caching, smaller models, and specialized hardware can reduce cost or latency, but must be evaluated on representative slices. A numerically valid response can still be a quality regression. Cache keys must include every behavior-changing input and version, and sensitive predictions need deliberate retention and tenant isolation.

Design failure before launch. Set request deadlines, cancellation propagation, bounded retries with jitter, load shedding, circuit breaking, and idempotency where requests can be replayed. Decide whether each use case should fail closed, use a previous model, return a rules-based fallback, serve a stale cached result, or route to a person. Never silently turn missing features or inference errors into a plausible default prediction.

Observe service and model behavior together: traffic, errors, saturation, queue time, batch formation, execution time, cold starts, model loads, memory, accelerator utilization, cost, input/output validity, prediction distributions, feature freshness, slice outcomes, and delayed ground truth. Preserve traceability from a prediction to the exact model and serving contract while minimizing sensitive payload capture.

Capacity tests should include realistic payload and sequence distributions, bursts, dependency slowdown, node loss, model reloads, cold starts, quota exhaustion, malformed inputs, cancellation, and regional failure. Determine sustainable throughput at the latency objective, not the largest number produced by a saturated benchmark. Production readiness means predictable behavior at and beyond the operating envelope, with a rehearsed rollback and evidence that recovery restores both technical health and decision quality.

Dynamic batching is a deliberate trade of latency for utilization, and the numbers should be measured rather than assumed. A GPU endpoint serving single requests may run at 12 percent accelerator utilization with a 24 ms p99; the same endpoint with a maximum batch of 16 and a 10 ms wait window can reach 70 percent utilization and roughly four times the throughput, at a p99 near 45 ms. Whether that is a win depends entirely on the deadline: a synchronous ranking call inside a 50 ms page budget cannot absorb it, while an asynchronous scoring job can absorb far more. Tune the pair of knobs — maximum batch size and maximum wait — against a load test using the real payload and sequence-length distribution, because a benchmark of uniformly sized inputs will overstate achievable batching for any workload whose inputs vary.

Cost per prediction is a first-class serving metric, and it usually falls out of utilization rather than of model choice. An accelerator held for a service that averages 8 percent utilization costs the same as one at 80 percent, so the practical levers are consolidating models onto shared endpoints, right-sizing the instance to the memory the model actually needs, and moving work that tolerates delay to batch or asynchronous paths. Compression helps when its quality cost is measured: 8-bit quantization commonly cuts memory roughly in half and improves throughput materially, but the acceptance test is per-slice quality against the full-precision baseline rather than an aggregate metric, since the loss usually concentrates in exactly the rare inputs that matter. Publish cost per thousand predictions next to latency and quality, so an optimization that halves cost while degrading a minority slice is visible as the trade it is.

Worked example: batching, and what it actually costs

A GPU serving a 7B model processes one request in 180 ms alone. Batching amortizes the weight loading across the batch, so throughput rises far faster than per-request latency:

Batch sizeService latencyCapacity (batch / latency)Time to accumulate a full batch at 40 req/s
1180 ms5.6 req/s25 ms
8220 ms36 req/s200 ms
16260 ms61 req/s400 ms
32390 ms82 req/s800 ms
64710 ms90 req/s1,600 ms

Latency per request rises about 4x from batch 1 to batch 64 while capacity rises 16x. That is the trade everyone quotes. The fourth column is the part that gets left out, and it runs the other way: the larger the batch, the longer you wait to fill it.

There is a floor under that cost, and it is worth deriving because it is not obvious. A stable server needs arrivals at or below capacity, so lambda <= batch / latency. Filling a batch takes batch / lambda, and substituting gives batch / lambda >= latency. Accumulating a batch always takes at least as long as processing one. Batching cannot buy latency; it can only buy throughput, and it always charges waiting for it.

At 40 req/s that floor bites immediately:

  Batch 1   capacity  5.6 req/s  <  40  -> cannot keep up, queue grows without bound
  Batch 8   capacity   36 req/s  <  40  -> cannot keep up, queue grows without bound
  Batch 16  capacity   61 req/s  >  40  -> stable. 400 ms to fill + 260 ms to serve = 660 ms
  Batch 32  capacity   82 req/s  >  40  -> stable. 800 ms to fill + 390 ms to serve = 1,190 ms
  Batch 64  capacity   90 req/s  >  40  -> stable. 1,600 ms to fill + 710 ms to serve = 2,310 ms

So the utilization graph and the user's experience point in opposite directions: batch 64 has the best hardware efficiency and by far the worst latency. Batch 16 is the smallest batch that keeps up, and at this arrival rate that makes it the right answer.

Real servers do not actually wait 1,600 ms. They set a maximum batch size and a maximum formation window and dispatch whatever has arrived when either is hit — which means at 40 req/s with a 20 ms window you gather 0.8 requests, run at an effective batch of 1, and none of the throughput in that table is reachable. Large batches need dense arrivals, not just a large setting.

Two follow-ups worth having ready. Continuous batching changes the shape for generation: a static batch waits for its slowest member, so one 800-token completion holds 15 finished 40-token ones hostage, while continuous batching evicts and admits per step. And the SLO picks the batch size, not the utilization graph — a 300 ms p99 against 260 ms of service leaves 40 ms of formation budget, which at 40 req/s buys 1.6 requests, so that SLO and batch 16 cannot both hold at this load.