Top 100 Backend Engineer Interview Questions and Answers
The questions most likely to actually come up in your Backend Engineer interview, ranked by likelihood — with detailed, senior-level answers covering what an interviewer is really listening for.
Curated: · Written: · Reviewed:
QA-1How would you design the endpoints for an order management API?(show answer)
I would start designing a REST resource from the failure a caller would see, not from the framework feature.
Resources are nouns with stable identifiers, and the verb belongs in the HTTP method rather than the path. An action that is not naturally a resource, such as cancelling, is modelled as a sub-resource or a state transition rather than a remote procedure bolted onto a URL.
Concretely, name collections and items, use the methods for read, create, replace, and delete, model transitions as explicit sub-resources, and keep response shapes consistent across endpoints.
The reason for that specificity is a failure I have seen: An API exposed 34 verb-shaped endpoints such as one that created an order and sent an email together, and 3 clients each called them in a different order, which left 2 percent of orders with no confirmation.
Three operations, two designs.
| Operation | Verb-shaped path | Resource design |
|---|---|---|
| create an order | POST /createOrder | POST /orders |
| cancel an order | POST /cancelOrder | POST /orders/42/cancellation |
| list 50 orders | GET /getOrders | GET /orders with limit 50 |
I would not consider it settled without evidence: List every operation as a method plus a resource and confirm each maps to a single state change.
Put the verb in the method, not in the path.
Curated: · Written: · Reviewed:
QA-2What should an API return when the request is valid JSON but names a customer that does not exist?(show answer)
The first question I ask about status codes and error contracts is what happens on the second attempt.
The status code carries the class of outcome and the body carries the detail, so a missing resource is 404, an unprocessable request is 422, and a conflicting state is 409. Returning 200 with an error field forces every client to parse prose to learn whether the call worked.
Concretely, choose the code from the outcome, return a machine-readable error code with a human message, keep one error schema across endpoints, and never leak internal exception text.
The reason for that specificity is a failure I have seen: A service returned 200 with a success flag, one client ignored the flag, and 1,900 failed payments were recorded as complete.
Four outcomes.
| Case | Status | Error code |
|---|---|---|
| missing customer | 404 | customer_not_found |
| invalid field | 422 | validation_failed |
| duplicate submit | 409 | already_exists |
| dependency down | 503 | upstream_unavailable |
I would not consider it settled without evidence: Assert the status code and the error code for every failure mode in the contract tests.
The status line is part of the contract.
Curated: · Written: · Reviewed:
QA-3A client retries a payment request after a timeout. How do you prevent a double charge?(show answer)
A retry is only distinguishable from a new request if the client supplies an idempotency key and the server remembers what it did with it. So the fix is: client generates a UUID per logical operation, sends it as an Idempotency-Key header, and the server stores key → (status, response) before returning. A repeat with the same key returns the stored response without touching the payment provider again.
The part people get wrong is the race: the timeout happens, the client retries, and the second request arrives while the first is still in flight. If the server just checks "does this key exist," both proceed. You need to insert the key atomically — INSERT INTO idempotency_keys (key, status) VALUES (k1, 'in_progress') with a unique constraint — and treat a duplicate-key violation as "someone is already working." At that point either block until the first finishes (poll the row, or SELECT ... FOR UPDATE on it) or return 409 Conflict and let the client back off. Stripe's idempotency layer does the first; returning the in-flight response is friendlier to clients with short timeouts.
The other failure mode is the timeout itself: the server may have charged the card and lost the response, or never reached the provider. The stored record has to capture which. If the first attempt genuinely failed before the provider call, the key should be reusable or the client needs a fresh one — otherwise a transient network blip permanently fails a valid payment. So the state machine is in_progress → completed | failed, and only completed results get replayed verbatim.
| Attempt | Key | Server behavior | Charges |
|---|---|---|---|
| first request | k1 | insert key, call provider, store response | 1 |
| retry after timeout | k1 | key exists, return stored response | 0 |
| concurrent retry (race) | k1 | insert fails, wait on in-flight row | 0 |
| new purchase | k2 | new key, new charge | 1 |
Expiry matters too: keys live long enough to cover the client's maximum retry window plus provider settlement time — 24 hours is Stripe's default, and it's a reasonable floor. Replaying a stored response after expiry is a real double-charge bug, so expiry has to be enforced, not just documented.
How I'd verify it: send the same request twice with one key and confirm exactly one charge exists and both responses are byte-identical; then fire two concurrent requests with one key and confirm the same. The concurrent case is the one staging tests usually miss, and it's the one that bites in production.
One caveat worth stating: idempotency keys deduplicate retries, not duplicate submissions. A user clicking "pay" twice generates two logical operations with two keys. Catching that needs a higher-level guard — an order ID or a uniqueness constraint on the order's payment — keyed on business identity rather than transport identity.
Curated: · Written: · Reviewed:
QA-4How would you paginate a list endpoint that clients poll while the data changes?(show answer)
First I'd pin down what the client actually needs, because it decides the design: is the poll a "what's new since last time" feed, or a full walk of the list? Those have different answers.
Full walk over changing data: cursor pagination on a stable key.
Offset pagination breaks in two ways here. Mechanically, LIMIT 40 OFFSET 40000 forces the database to scan and discard 40,000 rows on every request — at page 800 of a 50-row page size that's real work each poll, and it gets linearly worse with depth. Correctness-wise, if rows are inserted at the head of the sort order between polls, the window shifts: every page after the insert silently skips rows, and rows deleted mid-walk show up as repeats. I've debugged exactly this — an export job using offset paging against a table taking live inserts missed 1,100 of 50,000 records and emitted 700 duplicates. The client never knew; the counts only surfaced when a downstream reconciliation failed.
The fix is a cursor over a sort key that's immutable and unique. In practice that means (created_at, id) — created_at alone collides when rows share a timestamp, and id alone isn't monotonic if you use random UUIDs. The cursor is the last row's key, encoded opaquely (base64 of the tuple, or a signed token if you don't want clients tampering with it):
-- cursor = (created_at, id) of the last row returned
SELECT created_at, id, ...
FROM records
WHERE (created_at, id) < (:last_created_at, :last_id)
ORDER BY created_at DESC, id DESC
LIMIT 50;
This is an index seek into (created_at, id) regardless of depth — page 800 costs the same as page 2 — and inserts at the head don't move the window because the cursor anchors to a key, not a position.
"What's new" polling: a since-cursor, not pagination at all. If the client polls for changes, the right interface is ?since=<cursor> returning everything after that point, not numbered pages. Same keyset mechanism, but the contract is "here's what you missed," which is what a poller wants. If rows can be updated (not just inserted), keyset on created_at misses mutations — you need an updated_at index and cursor on that instead, and accept that a row may appear twice in the stream (once on create, once on update). Dedupe is the client's job; document it.
Trade-offs to state out loud:
| Offset | Cursor | |
|---|---|---|
| Deep page cost | scans N rows, linear | index seek, constant |
| During inserts | skips/repeats | stable window |
| Jump to page 40 | trivial | not supported |
| Implementation | one line | cursor encode/decode, composite index |
You lose random page access and total counts get awkward (a cheap COUNT on a moving table is a lie anyway). If the product genuinely needs "page 40 of 90," offset on a snapshot is defensible — e.g. Postgres REPEATABLE READ transaction or a temp-table materialization — but that's the wrong tool for a polling client.
Failure modes I'd watch for: the composite index missing and the keyset predicate degrading to a sort of the whole table (check EXPLAIN for an index scan, not a sort); cursors that leak internal IDs if you want them opaque; and page-size caps — cap at something like 100 and enforce it, or a client asking for 50,000 rows reintroduces the deep-scan problem through the back door.
How I'd prove it works: insert rows continuously during a full paged walk in a test environment, then assert the union of all pages contains every record exactly once. That test catches the skip/repeat bug class directly — it's the test the export job above never had.
Curated: · Written: · Reviewed:
QA-5You must change a response field that clients depend on. How do you ship it?(show answer)
First I'd classify the change, because that decides everything downstream. If I'm adding a field or changing an internal representation clients never see, there's no breaking change — ship it. If I'm removing, renaming, or retyping a field clients parse, that's breaking, and the only safe way to ship it is an additive path: both fields live side by side for a window, and the old one is retired on evidence, not on a calendar guess. A version boundary (v2 endpoint, new media type) is the fallback when the change is too large to express additively — a field that moves from a scalar to a nested object, say — but for one field, dual-writing is cheaper than maintaining two API versions for years.
The rollout I'd actually run:
- Ship the new field alongside the old one. Both populated from the same source, so they can't drift. If the response is built from a DB row, that means the serializer emits both until cutover — not a copy in the handler, which is where drift bugs come from.
- Mark the old field deprecated in the contract (OpenAPI
deprecated: true, plus aDeprecation/Sunsetheader if the framework supports it — RFC 8594 definesSunset). Announce a removal date, but treat it as a target, not a trigger. - Measure who still reads it. Server logs can't tell you directly — the field is in a response — so I'd instrument it: emit a per-client counter when the old field is serialized, or better, check whether clients send the old field back in requests (echoing behavior is a strong signal they still consume it). Client-side, ask integration owners directly; for first-party clients, add a telemetry ping on the old parse path.
- Remove when the per-client counter hits zero and stays there for a full business cycle — a month, not a week, because batch jobs and quarterly cron integrations exist.
A retirement trace from a change I shipped this way:
| Stage | Old field | New field | Clients still reading old | | --- | --- | | Release | present | present | 9 | | +2 weeks | deprecated, Sunset header set | present | 5 | | +6 weeks | deprecated | present | 1 (batch job, monthly) | | +10 weeks | removed | present | 0 |
The batch job at week 6 is the reason for the long tail: it ran monthly, so a two-week "all clear" window would have broken it.
Failure modes I'd watch for:
- Strict parsers. Clients using strict deserialization (Rust serde with
deny_unknown_fields, or generated protobuf/avro clients) break on added fields too, not just removed ones. If any client is codegen'd from a strict schema, "additive is safe" isn't automatically true — I'd check with the integration owners before the first release, not after. - Silent drift. If old and new fields disagree for some rows (a migration that backfilled half the table), clients reading either one see different truths. Detect with a reconciliation job: sample N responses, assert old == new, alert on mismatch. I'd run that from day one of dual-writing.
- Caches and CDNs. If responses are cached, the new field doesn't reach clients until cache entries expire. Purge or version the cache key at release, or the "nobody reads the new field" measurement lies to you.
- The removal itself breaking someone you never saw. Unknown third-party consumers (scrapers, internal scripts) don't appear in your client registry. That's why the removal ships behind a flag for one release before the code is deleted — if something screams, the flag flips back in minutes.
What I'd avoid: renaming in place and bumping the API version in the same release. I've seen that exact mistake — a field renamed with no transition window, and 4 of 9 client integrations broke within the hour because they parsed responses strictly. The version bump was on a changelog nobody read. Additive dual-write with measured retirement would have caught all four before removal.
Curated: · Written: · Reviewed:
QA-6When would you choose gRPC or GraphQL over a REST API?(show answer)
I start with the caller, not the style. The question that decides it: who calls this API, over what network, and how often?
REST when the interface is resource-shaped and the callers are outside my control — public APIs, third parties, anything that needs to survive without a coordinated deploy. You get plain HTTP semantics, CDN and Cache-Control caching for free, and every ecosystem tool works. The cost is over- and under-fetching: a mobile client that needs 3 fields gets a 40-field JSON blob, or makes 4 round trips to assemble a screen.
gRPC for internal service-to-service calls. HTTP/2 multiplexing, a typed .proto contract with generated stubs in every language we run, and Protobuf on the wire — typically 3-10x smaller than the equivalent JSON and no hand-written serialization code. Streaming (server, client, bidirectional) is first-class. The costs are real: browser support requires grpc-web plus a proxy, debugging means reading binary frames unless you set up reflection with grpcurl, and load-balancing over HTTP/2 long-lived streams needs care — one TCP connection can pin all traffic to one pod unless the LB understands it. Deadline and retry semantics are built into the contract, which is why it wins for high-volume internal traffic.
GraphQL when many different clients need to shape their own queries over the same data — a classic case is one product with web, iOS, and Android teams shipping on different schedules. One endpoint, client-specified fields, no versioned REST sprawl. The costs land on the server: you own query-cost control, because a client can ask for arbitrarily deep nesting. And HTTP caching mostly breaks — POST bodies, per-client queries — so you're caching in the resolver layer with DataLoader-style batching instead of a CDN.
The failure mode I've seen that makes me insist on cost limits before adopting GraphQL: a team replaced six internal REST calls with one GraphQL endpoint, and a single unbounded query — a nested list field with no first: argument — pulled 240,000 rows in one request and exhausted the database connection pool. Every other request timed out. The fix was graphql-cost-analysis with a max query depth of 5 and a complexity budget per query, plus a per-client rate limit on cost, not request count. That's the standing tax: without persisted queries or a cost analyzer, any client can write a query your server can't afford.
| REST | gRPC | GraphQL | |
|---|---|---|---|
| Wire format | JSON | Protobuf (binary) | JSON |
| Transport | HTTP/1.1 or 2 | HTTP/2 | HTTP POST, usually 1.1 |
| Caching | HTTP/CDN, free | application-level | resolver-level, must build it |
| Contract | OpenAPI, optional | .proto, enforced | schema, enforced |
| Streaming | SSE/WebSocket hacks | native, 3 modes | subscriptions over WebSocket |
| Browser | native | needs grpc-web proxy | native |
Concrete decision points I'd give:
- Public API, unknown callers → REST. Nobody outside can generate your stubs or trust your schema introspection.
- 5,000 calls/sec between two services I own, with streaming or strict latency budgets → gRPC. Measured on a past project: an internal inventory lookup went from 8 ms p99 with JSON/HTTP to 3 ms with gRPC, mostly from payload (2.1 KB JSON → 340 bytes Protobuf) and connection reuse. Labelled as one measurement, not a law — Protobuf's win shrinks when payloads are small and TLS handshake dominates.
- Three client platforms, rapidly changing screens, one backend team → GraphQL, with cost limits and persisted queries from day one.
Mixed is normal and fine: gRPC inside, REST at the edge, GraphQL as a BFF layer for a specific client family. What I'd refuse is choosing by fashion — the evidence I'd ask for before a migration is payload size, p99 latency, and worst-case query cost for the same call in each style, measured on our traffic, not a benchmark.
Curated: · Written: · Reviewed:
QA-7Where should request validation live in a service?(show answer)
Short version: at the boundary, before any business logic — but the interesting part is what "validation" means there versus deeper in the stack, because each layer catches a different class of mistake.
The boundary layer. The handler (or the deserialization step just before it) should parse the request into typed objects and reject anything that doesn't fit the declared shape: unknown fields, wrong types, out-of-range values, malformed dates. This is structural validation, and it belongs at the edge for two reasons. First, invalid data never reaches the domain, so the domain code can assume well-formed input and stay readable. Second, the failure is a client error, so you return a 422 with a stable error code — not a 500 from a constraint violation three layers down, which turns the client's mistake into your incident.
The domain layer. Invariants that need business context — "you can't refund more than was charged", "this state transition isn't allowed from status=shipped" — live in the domain, not the handler. The handler can't know these; it doesn't have the loaded entity. Splitting it this way keeps the boundary fast and dumb and the domain authoritative.
The database. Constraints (NOT NULL, CHECK, FK, uniqueness) are the second line, not the first. They exist for the case the boundary can't cover: concurrent writers racing on the same invariant. Two requests can both pass handler validation for "username must be unique"; only the database sees the interleaving. So the constraint fires, and your job is to catch that specific error and translate it into a 4xx rather than letting it surface as a 500.
Why this actually matters — a failure I've seen. A list endpoint accepted a limit parameter with no upper bound. A client asked for 1,000,000 records. The query held a connection for about 90 seconds; the pool (20 connections) saturated, and every other request on that service timed out until it finished. One unvalidated integer took the whole service down. That's the failure mode of boundary validation done late or not at all: it's not just bad data in the database, it's resource exhaustion.
| Input | No boundary validation | With boundary validation |
|---|---|---|
limit=1000000 | ~90 s query, pool exhaustion | 422, max 200 |
unknown field is_admin | silently ignored (or worse, bound by a framework) | 422 |
negative amount | CHECK constraint fires → 500 | 422 |
The unknown-field row is worth dwelling on. Silently ignoring extra fields is how mass-assignment bugs happen: a framework auto-binds request fields onto a model, someone adds a role column, and a client can now set it. Rejecting unknown fields at the boundary closes that class entirely.
How I'd verify it's actually working. Fuzz each endpoint with missing, extra, and out-of-range fields and assert every rejection is a 4xx with a stable, documented error code — in CI, not as a one-off. Then a negative test for the race case: two concurrent requests violating the same unique constraint, asserting the loser gets a 409, not a 500.
When not to put it all at the boundary: if validation needs a database lookup ("does this SKU exist"), don't do it in the handler — that's a domain concern, and doing it at the edge couples your transport layer to your data layer. Keep the boundary to checks that need only the request itself.
So: shape at the edge, invariants in the domain, constraints in the database as the concurrency backstop — and each layer's rejection translated into the right status code.
Curated: · Written: · Reviewed:
QA-8How do you handle identity in a service-to-service call chain?(show answer)
The part of identity in a service call chain interviewers probe is the partial failure, not the happy path.
Authentication proves who is calling and authorization decides what they may do, and both must be checked by the service that owns the data rather than only at the gateway. A downstream service that trusts a header can be called directly.
Concretely, verify a signed token at every service, carry the end-user identity separately from the calling service's identity, check authorization against the resource being touched, and never trust a client-supplied user id.
The reason for that specificity is a failure I have seen: An internal service trusted a user-id header set by the gateway, and a request sent directly to the pod returned another tenant's records.
Where each check belongs.
| Layer | Checks | Bypassable |
|---|---|---|
| gateway | token signature, rate limit | yes, by direct call |
| service | token, tenant, permission | no |
| database | row-level constraint | no |
| 3 layers | 2 enforce on every request | 1 bypassable |
I would not consider it settled without evidence: Call each internal endpoint directly with a forged identity header and confirm it is rejected.
Every service checks, not just the gateway.
Curated: · Written: · Reviewed:
QA-9Would you use signed tokens or server-side sessions for an API?(show answer)
I would anchor tokens or server-side sessions in a number from a load test or a trace rather than a preference.
A signed token avoids a lookup per request but cannot be revoked before it expires, while a server-side session is revocable at the cost of a store read. The choice follows how quickly access must be withdrawn.
Concretely, keep tokens short-lived with a refresh path, hold a revocation list for high-risk actions, or use sessions with a fast store when immediate revocation matters, and never put secrets in a token payload.
The reason for that specificity is a failure I have seen: Access tokens were valid for 24 hours with no revocation, so a dismissed employee's token kept working for 19 hours after their account was disabled.
Revocation against cost.
| Scheme | Revocation delay | Store read per request |
|---|---|---|
| 24-hour token | up to 24 h | none |
| 5-minute token with refresh | 5 min | on refresh |
| server session | immediate | every request |
I would not consider it settled without evidence: Disable an account and measure how long its existing credential still works.
Decide how fast you must be able to say no.
Curated: · Written: · Reviewed:
QA-10How would you protect an API from a client that sends too many requests?(show answer)
What separates a strong answer on rate limiting an API is knowing how the naive version fails under load.
A limiter sets a quota per identity and window, and the algorithm decides how bursts are treated. A fixed window admits up to twice the quota across its boundary, while a token bucket allows a burst you chose deliberately.
Concretely, limit per client and per endpoint class, keep counters in a shared store so limits hold across instances, return 429 with a retry-after header, and raise limits for known partners deliberately.
The reason for that specificity is a failure I have seen: A fixed-window limiter of 100 requests a minute admitted 200 calls in 2 seconds across the window boundary, which saturated the database pool.
Three algorithms at a 100 per minute quota.
| Algorithm | Worst burst | Boundary problem |
|---|---|---|
| fixed window | 200 in 2 s | yes |
| sliding window | 100 per minute | no |
| token bucket | the configured burst | no |
I would not consider it settled without evidence: Replay a burst across a window boundary and confirm the limiter never admits more than the intended rate.
Pick the algorithm by how it handles bursts.
Curated: · Written: · Reviewed:
QA-11How do you configure timeouts and retries for a call to a slow dependency?(show answer)
I would treat timeouts and retries for a dependency as something that must survive retries, restarts, and a slow dependency.
Every remote call needs a deadline shorter than its caller's, or one slow dependency exhausts the threads and connections above it. Retries without backoff and jitter turn a slow dependency into a self-inflicted flood.
Concretely, set each timeout from the dependency's measured tail latency, propagate a deadline through the chain, retry only idempotent calls with exponential backoff and jitter, and cap the total attempts.
The reason for that specificity is a failure I have seen: A service with no client timeout waited on a dependency for 120 seconds, and within 4 minutes all 200 worker threads were blocked and it stopped answering health checks.
Settings derived from measurement.
| Setting | Value | Reason |
|---|---|---|
| dependency p99 | 400 ms | measured |
| call timeout | 1 s | tail plus margin |
| retries | 2 with jitter | idempotent calls only |
| caller deadline | 3 s | below the gateway's 5 s |
I would not consider it settled without evidence: Inject a hung dependency in a test and confirm the caller fails fast and keeps serving other traffic.
A call without a deadline is an outage waiting for a slow day.
Curated: · Written: · Reviewed:
QA-12A downstream service is failing. How do you stop it taking your service down?(show answer)
Before changing anything for circuit breakers and dependency isolation, I would write down how the result will be measured.
A circuit breaker stops calling a dependency that is clearly failing, which frees resources and gives it room to recover. Isolating each dependency into its own pool keeps one slow call path from consuming every worker.
Concretely, trip the breaker on an error rate over a window, serve a degraded response while it is open, probe with a few calls before closing, and give each dependency its own connection and thread budget.
The reason for that specificity is a failure I have seen: One failing recommendation dependency consumed every request thread, so checkout latency rose from 120 milliseconds to 9 seconds although checkout did not need recommendations.
Budgets and fallbacks.
| Dependency | Pool | Degraded response |
|---|---|---|
| payments | 40 connections | fail the request |
| recommendations | 10 connections | empty list |
| search | 20 connections | cached results |
I would not consider it settled without evidence: Fail one dependency in a test and confirm unrelated endpoints keep their normal latency.
Isolate dependencies so one failure stays one failure.
Curated: · Written: · Reviewed:
QA-13Your queue consumer cannot keep up with producers. What do you do?(show answer)
I would start backpressure when consumers fall behind from the failure a caller would see, not from the framework feature.
When arrivals exceed the service rate, something must slow down, shed load, or buffer within a bound. An unbounded buffer converts a throughput problem into a memory outage and hides the real capacity gap.
Concretely, bound the queue, reject or shed low-priority work when it is full, push back on producers with 429 or flow control, and scale consumers on queue depth rather than on processor use.
The reason for that specificity is a failure I have seen: An unbounded in-memory queue grew to 4.2 million messages during a spike, the process was killed by its memory limit, and every buffered message was lost.
At twice capacity.
| Strategy | Behaviour | Message loss |
|---|---|---|
| unbounded buffer | memory climbs to the limit | all buffered |
| bounded and reject | 429 to producers | none |
| bounded and shed low priority | best effort below the cut | low priority only |
I would not consider it settled without evidence: Load test above capacity and confirm the service rejects or sheds rather than growing without bound.
Bound the buffer and push back.
Curated: · Written: · Reviewed:
QA-14How do you keep two services from breaking each other's expectations?(show answer)
The first question I ask about contract testing between services is what happens on the second attempt.
Unit tests on each side pass while their assumptions drift apart, so the contract needs a test both sides run. A consumer-driven contract records what the caller actually relies on.
Concretely, record each consumer's expectations, verify the provider against every consumer contract in its own pipeline, and fail the provider build on a breaking change.
The reason for that specificity is a failure I have seen: A provider tightened a field from optional to required, both test suites stayed green, and 2 consumers failed in production within 20 minutes of the deploy.
Which check catches it.
| Check | Catches the breaking change |
|---|---|
| provider unit tests | no |
| consumer tests against a mock | no |
| consumer contracts in the provider pipeline | yes, 2 of 2 |
I would not consider it settled without evidence: Run every consumer contract against the provider in the provider's pipeline and confirm a breaking change fails the build.
Test the contract where the change happens.
Curated: · Written: · Reviewed:
QA-15You must notify customers of events. What does a reliable webhook system need?(show answer)
With delivering webhooks reliably, a passing staging run is where I start checking rather than stop.
The receiver may be slow, down, or intolerant of duplicates, so delivery needs retries with backoff, a signature so origin can be verified, an event id for deduplication, and a way to replay. At-least-once is the realistic guarantee.
Concretely, sign each payload, include an event id and timestamp, retry with exponential backoff into a dead-letter state, expose a delivery log and a replay endpoint, and document that receivers must deduplicate.
The reason for that specificity is a failure I have seen: A webhook system retried every failure immediately and without a cap, which sent 90,000 duplicate calls to one customer in an hour and got the sending address blocked.
One failing receiver.
| Attempt | Delay | Outcome |
|---|---|---|
| 1 | immediate | 500 from receiver |
| 2 | 1 min | 500 |
| 3 to 8 | doubling to 8 h | delivered at attempt 6 |
| beyond 8 | dead letter | replayable |
I would not consider it settled without evidence: Take a receiver offline for an hour and confirm deliveries resume with no loss and no unbounded retry.
Assume at-least-once and help the receiver deduplicate.
Curated: · Written: · Reviewed:
QA-16A request needs 40 seconds of processing. How should the API handle it?(show answer)
I would answer long-running work behind an API by following one request through the service and its dependencies.
Holding a connection for a long job wastes resources and dies with any network hiccup, so the request should accept the work and return a handle. The client then polls a status URL or receives a callback.
Concretely, return 202 with a job id and status URL, process on a worker, expose progress and a terminal state, make submission idempotent, and keep results retrievable for a documented window.
The reason for that specificity is a failure I have seen: A synchronous report endpoint held connections for 40 seconds, and in a busy hour 70 percent of requests hit the proxy's 30 second timeout and were retried, doubling the load.
Against a 30 second proxy timeout.
| Design | Response time | Behaviour at the timeout |
|---|---|---|
| synchronous | 40 s | timeout and retry storm |
| 202 with polling | 30 ms | unaffected |
I would not consider it settled without evidence: Submit the job and confirm the request returns in milliseconds with a status URL that reaches a terminal state.
Accept the work rather than holding the caller.
Curated: · Written: · Reviewed:
QA-17Users upload 500 megabyte files. How do you design that?(show answer)
My approach to accepting large file uploads separates what the database guarantees from what the code assumes.
Streaming large bodies through the application costs memory and ties up workers, so the upload should go directly to object storage with a short-lived signed URL. The service records metadata and validates afterwards.
Concretely, issue a presigned URL scoped to one key with a short expiry, let the client upload directly, verify size and type after the fact, and use multipart upload so a failure can resume.
The reason for that specificity is a failure I have seen: Uploads proxied through the application buffered 500 megabytes per request, and 6 concurrent uploads exhausted the container's 3 gigabyte limit and restarted it.
Two upload paths.
| Path | Service memory per upload | Worker held for |
|---|---|---|
| proxied through the app | 500 MB | the whole upload |
| presigned direct upload | negligible | milliseconds |
I would not consider it settled without evidence: Upload the largest supported file and confirm the service's memory use stays flat.
Let storage take the bytes.
Curated: · Written: · Reviewed:
QA-18A client needs to create ten thousand records. What do you offer?(show answer)
For bulk write endpoints, I would name the invariant that must hold before choosing a mechanism.
Ten thousand single requests waste round trips and rate-limit budget, while one request with ten thousand items risks a long transaction and an all-or-nothing failure. A bounded batch with per-item results is the usual middle ground.
Concretely, accept a bounded batch such as 500 items, return a status per item, make the batch idempotent by key, and document whether it is atomic or best effort.
The reason for that specificity is a failure I have seen: A batch endpoint accepted 10,000 items in one transaction, and a single invalid row at item 9,000 rolled back the entire batch after 4 minutes of work.
Ten thousand records.
| Design | Round trips | One bad item |
|---|---|---|
| one request per record | 10,000 | that record fails |
| single 10,000-item transaction | 1 | all fail |
| 20 batches of 500 | 20 | that item fails |
I would not consider it settled without evidence: Submit a batch containing one invalid item and confirm the documented behaviour with per-item results.
Bound the batch and report per item.
Curated: · Written: · Reviewed:
QA-19Which responsibilities belong in the gateway and which in the service?(show answer)
I would size what belongs in an API gateway against real traffic and a real data volume, not a local fixture.
A gateway is the right place for concerns identical across every service, such as TLS termination, coarse rate limiting, and routing. Business rules and authorization on specific resources belong in the service that owns the data.
Concretely, terminate TLS, authenticate, apply coarse limits and request ids at the gateway, and keep resource-level authorization, validation, and domain rules inside the service.
The reason for that specificity is a failure I have seen: Tenant authorization was implemented only in gateway rules, and an internal service added later was reachable without them, which exposed 3 tenants' data to each other.
Placement of four concerns.
| Concern | Gateway | Service |
|---|---|---|
| TLS and routing | yes | no |
| coarse rate limit | yes | optional |
| resource authorization | no | yes |
| domain validation | no | yes |
I would not consider it settled without evidence: Bypass the gateway in a test and confirm every service still enforces authentication and authorization.
The gateway is a front door, not the lock on every room.
Curated: · Written: · Reviewed:
QA-20When is it right to split a service into two?(show answer)
The part of deciding to split a service interviewers probe is the partial failure, not the happy path.
A split is justified when two parts have genuinely different scaling, availability, or deployment needs, or when ownership requires it. Splitting for tidiness converts fast in-process calls into slow network calls with new failure modes.
Concretely, measure the coupling, name the different requirement that justifies the boundary, keep data ownership with one service, and budget for the retries, timeouts, and tracing the new hop needs.
The reason for that specificity is a failure I have seen: A service was split into 7 pieces for tidiness, and one page load then needed 11 network calls with a p95 of 1.9 seconds against 240 milliseconds before.
One journey, two shapes.
| Shape | Calls per page load | p95 |
|---|---|---|
| one service | 1 | 240 ms |
| seven services | 11 | 1,900 ms |
I would not consider it settled without evidence: Compare p95 latency and failure modes for the same user journey before and after the split.
Split for a requirement, not for neatness.
Curated: · Written: · Reviewed:
QA-21Two services write to the same table. What is wrong with that?(show answer)
I would anchor the shared-database problem in a number from a load test or a trace rather than a preference.
A shared table makes the schema a public contract with no owner, so either service can break the other with a migration and neither can hold an invariant. One writer per table, with others reading through an API, a replica, or events, restores ownership.
Concretely, give each table one owning service, expose an API or events for the others, use a replica or a projection for read-heavy consumers, and migrate by moving writes before removing access.
The reason for that specificity is a failure I have seen: Two services wrote to the same orders table, and a column added by one broke the other's insert at 02:00, which took 40 minutes to diagnose because neither team owned the schema.
Writers before and after.
| Table | Writers before | Writers after |
|---|---|---|
| orders | 2 services | 1 |
| inventory | 3 services | 1 |
I would not consider it settled without evidence: List every writer per table and confirm each table has exactly one.
One writer per table.
Curated: · Written: · Reviewed:
QA-22How would you make a read-heavy endpoint cacheable?(show answer)
What separates a strong answer on making a read endpoint cacheable is knowing how the naive version fails under load.
HTTP caching works when the response states how long it stays fresh and how to validate it, through cache-control and an entity tag. Without those headers every client and proxy refetches, and with the wrong ones stale data outlives its usefulness.
Concretely, set cache-control with a maximum age suited to the data, return an entity tag and honour conditional requests, vary on the headers that change the body, and never cache personalized responses publicly.
The reason for that specificity is a failure I have seen: A public product endpoint sent no cache headers, so a proxy refetched it 40,000 times an hour, and a 60 second maximum age later cut origin traffic by 94 percent.
One hour of traffic.
| Headers | Origin requests | 304 responses |
|---|---|---|
| none | 40,000 | 0 |
| maximum age 60 with an entity tag | 2,400 | 31% |
I would not consider it settled without evidence: Measure origin request volume and the share of 304 responses before and after adding the headers.
State freshness and validation, or nothing can cache.
Curated: · Written: · Reviewed:
QA-23A user reports a failure at 14:02. How do you find what happened?(show answer)
I would treat request identifiers and structured logs as something that must survive retries, restarts, and a slow dependency.
A request id generated at the edge and propagated through every service and log line makes one request reconstructable. Free-text logs without that id force a search by timestamp, which fails under concurrency.
Concretely, generate a request id at the gateway, propagate it through headers and into asynchronous work, log structured fields including the id and tenant, and return the id to the client for support.
The reason for that specificity is a failure I have seen: Three services logged free text with no shared identifier, so reconstructing one failed checkout took 3 hours of manual correlation across 40,000 lines.
Reconstructing one request.
| Logging | Time to reconstruct |
|---|---|
| free text, no id | about 3 h |
| structured with a request id | under 1 min |
I would not consider it settled without evidence: Take one request id and retrieve every log line and span for that request across services.
One id, every log line.
Curated: · Written: · Reviewed:
QA-24Your service must handle ten times its traffic next quarter. Where do you start?(show answer)
Before changing anything for preparing for ten times the traffic, I would write down how the result will be measured.
Scaling starts with measurement, because the bottleneck is usually one resource such as a connection pool, a lock, or a single-threaded step. Adding instances multiplies a bottleneck rather than removing it.
Concretely, load test to find the first limit, remove or shard it, keep the service stateless so instances scale, add caching for read amplification, and re-test until the limit is the one you intended.
The reason for that specificity is a failure I have seen: A team tripled its application instances and throughput did not move, because every instance shared a 100 connection database pool that was already saturated.
Where the constraint sits.
| Change | Throughput | Binding constraint |
|---|---|---|
| baseline | 900 per second | database pool |
| 3 times the instances | 920 per second | database pool |
| larger pool and a read replica | 3,400 per second | application processor |
I would not consider it settled without evidence: Load test to find the binding constraint and show that it moved after the change.
Find the limit before adding instances.
Curated: · Written: · Reviewed:
QA-25Design the backend for a URL shortener that serves fifty thousand redirects a second.(show answer)
I would start designing a URL shortener from the failure a caller would see, not from the framework feature.
The write path is tiny and the read path is enormous, so the design is dominated by cheap key generation and cache-friendly redirects. Redirects must be served from cache rather than from the primary database.
Concretely, generate short keys from a counter or a random key with a uniqueness check, store the mapping in a key-value store, serve redirects from cache with a durable fallback, and record clicks asynchronously.
The reason for that specificity is a failure I have seen: Click counting wrote to the primary database on every redirect, and at 12,000 redirects a second that write load pushed p99 redirect latency to 800 milliseconds.
Three paths, three stores.
| Path | Rate | Store |
|---|---|---|
| create a link | 50 per second | primary database |
| redirect | 50,000 per second | cache at a 99% hit rate |
| click events | 50,000 per second | queue with batched writes |
I would not consider it settled without evidence: Load test the redirect path at the target rate and report p99 latency and the cache hit rate.
Optimize the read path and make the writes asynchronous.
Curated: · Written: · Reviewed:
QA-26A query filters on tenant and status and sorts by created time. What index would you add?(show answer)
The first question I ask about index design and column order is what happens on the second attempt.
A composite index is usable left to right, so the column order decides which queries it serves. Equality columns come first, then the range or sort column, and a query that skips a leading column cannot use the index efficiently.
Concretely, put equality predicates first in the order of selectivity, add the sort column last so the sort is satisfied by the index, and confirm the plan uses it rather than assuming.
The reason for that specificity is a failure I have seen: An index created on created time alone left a tenant query scanning 4.1 million rows, and the correct three-column order brought the same query to 1,200 rows.
Same query, three indexes.
| Index | Rows examined | Sort |
|---|---|---|
| created_at | 4,100,000 | index |
| tenant_id | 26,000 | in memory |
| tenant_id, status, created_at | 1,200 | index |
I would not consider it settled without evidence: Read the plan before and after and compare rows examined rather than only the runtime.
Column order is the index design.
Curated: · Written: · Reviewed:
QA-27A query is slow in production but fast locally. How do you diagnose it?(show answer)
With reading a query plan, a passing staging run is where I start checking rather than stop.
The plan depends on data volume and statistics, so a plan chosen against 500 local rows tells you nothing about 50 million. Reading the production plan shows whether the engine scans, seeks, or sorts, and where the estimate diverges from reality.
Concretely, capture the plan with actual row counts on production-like data, compare estimated with actual rows to spot stale statistics, and look for scans on large tables and sorts that spill to disk.
The reason for that specificity is a failure I have seen: A query planned for 900 estimated rows actually matched 2.4 million, and the resulting nested loop join took 46 seconds under load while the local run took 30 milliseconds.
Estimate against reality.
| Step | Estimated rows | Actual rows |
|---|---|---|
| index scan on status | 900 | 2,400,000 |
| nested loop join | 900 | 2,400,000 |
| runtime | none | 46 s |
I would not consider it settled without evidence: Compare estimated and actual row counts in the production plan and refresh statistics before changing the query.
Read the plan against real data volumes.
Curated: · Written: · Reviewed:
QA-28A list endpoint issues one query per row. How do you find and fix it?(show answer)
I would answer the N plus one query by following one request through the service and its dependencies.
Loading a collection and then lazily loading each row's relation produces one query per item, which turns a 20 millisecond page into hundreds of round trips. The fix is to load the relation in one query and join or map in memory.
Concretely, log query counts per request, eager-load or batch the relation, use an in-query join or a second query keyed by the parent ids, and assert a query-count ceiling in tests.
The reason for that specificity is a failure I have seen: A list of 200 orders issued 201 queries and took 1.8 seconds, and batching the line items into one query brought it to 40 milliseconds.
Two hundred orders.
| Implementation | Queries | Duration |
|---|---|---|
| lazy per row | 201 | 1,800 ms |
| one join | 1 | 40 ms |
| two queries and a map | 2 | 45 ms |
I would not consider it settled without evidence: Assert the query count per request in a test so the pattern cannot return unnoticed.
Count queries per request, not just their duration.
Curated: · Written: · Reviewed:
QA-29What anomalies does each isolation level allow, and which do you use by default?(show answer)
My approach to transaction isolation levels separates what the database guarantees from what the code assumes.
Read committed prevents dirty reads but allows a value to change between two reads in one transaction, repeatable read prevents that, and serializable prevents write skew at the cost of more conflicts. The default is usually read committed, which is why invariants across rows need explicit locking or a higher level.
Concretely, name the invariant, choose the level that protects it, take an explicit row lock where read committed is not enough, and handle serialization failures with a retry.
The reason for that specificity is a failure I have seen: Two concurrent transfers read the same balance of 100 under read committed and each withdrew 80, leaving the account at minus 60 because the invariant spanned two rows.
Concurrent withdrawal of 80 from 100.
| Level | Outcome |
|---|---|
| read committed | both succeed, balance -60 |
| read committed with a row lock | second waits, then fails the check |
| serializable | one commits, one retries |
I would not consider it settled without evidence: Run the concurrent case in a test at the chosen level and confirm the invariant holds or the transaction fails.
Name the invariant before choosing the level.
Curated: · Written: · Reviewed:
QA-30Two transactions deadlock in production. What causes it and how do you fix it?(show answer)
For deadlocks, I would name the invariant that must hold before choosing a mechanism.
A deadlock happens when two transactions hold locks the other needs, which usually means they touch the same rows in a different order. The database resolves it by killing one, so the application must expect and retry that failure.
Concretely, take locks in a consistent order everywhere, keep transactions short and narrow, avoid interleaving user or network waits inside them, and retry the loser with backoff.
The reason for that specificity is a failure I have seen: A batch job updated accounts ascending while the API updated them descending, which produced 240 deadlocks a day and failed 0.4 percent of payments.
One day of the same workload.
| Change | Deadlocks per day | Failed payments |
|---|---|---|
| mixed lock order | 240 | 0.4% |
| consistent order | 3 | 0% |
| consistent order plus retry | 3 | 0% |
I would not consider it settled without evidence: Read the deadlock log for the two lock orders and confirm every writer now takes them in the same order.
Order your locks and expect to retry.
Curated: · Written: · Reviewed:
QA-31Two users edit the same record. How do you stop one overwriting the other?(show answer)
I would size optimistic and pessimistic locking against real traffic and a real data volume, not a local fixture.
Optimistic locking carries a version with the row and rejects a write whose version has moved, which suits low-contention edits. Pessimistic locking takes the row lock up front, which suits high contention but blocks and risks deadlock.
Concretely, add a version column and include it in the update predicate, return a conflict to the caller when zero rows change, and reserve explicit row locks for genuinely contended counters and balances.
The reason for that specificity is a failure I have seen: Two support agents edited one ticket without version checks, and the second save silently discarded the first agent's 3 field changes.
Concurrent edits of one row.
| Strategy | Second write | Waiting |
|---|---|---|
| last write wins | overwrites silently | none |
| version check | 409 conflict | none |
| row lock | applied after the first | blocks |
I would not consider it settled without evidence: Run the concurrent edit in a test and confirm the second write is rejected rather than silently winning.
Detect the conflict rather than letting the last write win.
Curated: · Written: · Reviewed:
QA-32Your service opens a database connection per request. What goes wrong at scale?(show answer)
The part of connection pooling interviewers probe is the partial failure, not the happy path.
Connections are expensive server-side resources, so a pool bounds them and queues the rest. The pool size should be tuned to what the database can serve, because more connections than it can handle makes everything slower rather than faster.
Concretely, size the pool from the database's capacity and the service's concurrency, set an acquisition timeout so waiting fails fast, keep transactions short so connections return quickly, and add a proxy when many instances share one database.
The reason for that specificity is a failure I have seen: Forty instances each opened a 50 connection pool against a database with a 500 connection limit, so 1,500 connections were refused and the service returned 503 for 8 minutes.
Forty instances against a 500 connection limit.
| Pool per instance | Total connections | Result |
|---|---|---|
| 50 | 2,000 | refusals, 503s |
| 10 | 400 | healthy |
| 10 with a proxy | 400 multiplexed | headroom |
I would not consider it settled without evidence: Load test and confirm total connections across instances stay within the database's limit with the acquisition timeout enforced.
Size the pool for the database, not for the service.
Curated: · Written: · Reviewed:
QA-33How do you rename a column that a running service depends on?(show answer)
I would anchor migrations without downtime in a number from a load test or a trace rather than a preference.
A running service and a migration cannot change at the same instant, so a rename must be done in expand-and-contract steps. Add the new column, write both, backfill, move reads, then drop the old one, with each step safe on its own.
Concretely, add nullable columns rather than rewriting tables, avoid long-held locks, backfill in batches, deploy code that tolerates both shapes, and drop the old column only after no reader remains.
The reason for that specificity is a failure I have seen: A direct rename in one migration took an exclusive lock on a 40 million row table for 90 seconds, during which every write timed out and the deploy was rolled back.
Five safe steps.
| Step | Lock held | Safe to stop after |
|---|---|---|
| add nullable column | milliseconds | yes |
| write both columns | none | yes |
| backfill in batches of 5,000 | none | yes |
| move reads | none | yes |
| drop the old column | milliseconds | yes |
I would not consider it settled without evidence: Run each step against a production-sized copy and confirm no step holds a lock longer than the agreed budget.
Expand, migrate, then contract.
Curated: · Written: · Reviewed:
QA-34Should deletes be soft or hard, and what does that cost?(show answer)
What separates a strong answer on soft deletes and audit columns is knowing how the naive version fails under load.
A soft delete keeps the row with a marker, which preserves history and makes recovery easy, but every query must then filter it out and unique constraints must account for it. A hard delete is simpler and is sometimes legally required.
Concretely, decide per table, filter deleted rows in a view or a repository layer rather than in every query, make unique indexes partial so a deleted row does not block reuse, and document the retention and purge path.
The reason for that specificity is a failure I have seen: A soft-deleted email address kept its unique constraint, so a returning customer could not re-register, and support handled 60 such cases before the index was made partial.
Reusing a deleted value.
| Design | Re-registration | History |
|---|---|---|
| hard delete | works | lost |
| soft delete, full unique index | blocked | kept |
| soft delete, partial unique index | works | kept |
I would not consider it settled without evidence: Re-create a soft-deleted record in a test and confirm the constraint and the queries behave as documented.
A soft delete is a filter obligation on every query.
Curated: · Written: · Reviewed:
QA-35When would you denormalize a schema deliberately?(show answer)
I would treat normalization and denormalization as something that must survive retries, restarts, and a slow dependency.
Normalization keeps one authoritative copy of each fact, which prevents contradictory data, and denormalization copies data to avoid a join or an aggregate at read time. A copy is only worth it when the read cost is measured and the write path can keep it consistent.
Concretely, normalize by default, denormalize a specific measured hot read, keep the derived column updated in the same transaction or by an event, and add a reconciliation job that detects drift.
The reason for that specificity is a failure I have seen: An order total was denormalized onto the order row but updated only in the API path, so 1,400 orders changed by a batch job kept a stale total for 3 weeks.
Derived order total.
| Design | Read cost | Drift risk |
|---|---|---|
| sum line items per read | 14 ms | none |
| denormalized, updated in one path | 0.4 ms | 1,400 rows stale |
| denormalized, updated in the transaction plus nightly check | 0.4 ms | detected |
I would not consider it settled without evidence: Run a reconciliation query that compares the derived value with its source and confirm it reports zero drift.
A copy needs an owner and a reconciliation.
Curated: · Written: · Reviewed:
QA-36Would you use an auto-incrementing integer or a UUID as a primary key?(show answer)
Before changing anything for choosing a primary key, I would write down how the result will be measured.
Sequential integers are compact and index well but leak volume and order, while random identifiers are safe to expose and generate client-side at the cost of index locality and size. Time-ordered identifiers recover most of the locality.
Concretely, use a sequential or time-ordered key for the clustered primary key, expose a separate opaque public identifier when enumeration matters, and avoid fully random keys as clustered keys on write-heavy tables.
The reason for that specificity is a failure I have seen: Random identifiers as the clustered key on a 60 million row table caused page splits that raised insert latency from 3 to 22 milliseconds and grew the index by 40 percent.
Sixty million rows.
| Key | Insert latency | Index size | Enumerable |
|---|---|---|---|
| auto-increment integer | 3 ms | 1.0x | yes |
| random UUID | 22 ms | 1.4x | no |
| time-ordered UUID | 4 ms | 1.2x | no |
I would not consider it settled without evidence: Measure insert latency and index size for each key type at production row counts.
Separate the storage key from the public identifier.
Curated: · Written: · Reviewed:
QA-37When is a JSON column the right choice, and when is it a trap?(show answer)
I would start JSON columns from the failure a caller would see, not from the framework feature.
A JSON column suits genuinely variable attributes and third-party payloads you must keep verbatim. It becomes a trap when core fields live inside it, because constraints, types, and index support are weaker than a real column.
Concretely, keep queried and constrained fields as columns, keep variable extras in JSON, add expression indexes for the few JSON paths you filter on, and validate the document shape at the application boundary.
The reason for that specificity is a failure I have seen: Order status was stored inside a JSON document, and a query filtering on it scanned 12 million rows because no expression index existed, taking 9 seconds.
Filtering on status.
| Storage | Query time | Constraint possible |
|---|---|---|
| inside JSON, no index | 9 s | no |
| JSON with an expression index | 40 ms | limited |
| real column | 8 ms | yes |
I would not consider it settled without evidence: List the fields queried or constrained and confirm each is a column or has an expression index.
Promote anything you filter on to a column.
Curated: · Written: · Reviewed:
QA-38A table has grown to two billion rows. What are your options?(show answer)
The first question I ask about partitioning and sharding is what happens on the second attempt.
Partitioning splits one table inside one database so queries and maintenance touch a slice, while sharding splits data across databases and adds cross-shard queries and rebalancing. Partitioning is the cheaper first step when the queries carry the partition key.
Concretely, partition by the column most queries filter on, usually time or tenant, keep the partition key in every query, automate partition creation and retention, and treat sharding as a later step with a chosen key and a rebalancing plan.
The reason for that specificity is a failure I have seen: A two billion row events table was queried without the time filter after partitioning, so every query touched all 48 partitions and ran slower than before the change.
Forty-eight monthly partitions.
| Query | Partitions touched | Duration |
|---|---|---|
| with a month filter | 1 | 60 ms |
| without a month filter | 48 | 7 s |
| unpartitioned baseline | 1 table | 4 s |
I would not consider it settled without evidence: Confirm from the plan that queries prune to the expected partitions rather than scanning all of them.
Partitioning only helps queries that carry the key.
Curated: · Written: · Reviewed:
QA-39You move reads to a replica and users start seeing stale data. What happened?(show answer)
With read replicas and replication lag, a passing staging run is where I start checking rather than stop.
A replica applies changes after the primary commits, so a read immediately after a write can miss it. Read-your-writes consistency needs the write's session to read the primary, or a token that waits for the replica to catch up.
Concretely, route reporting and cacheable reads to replicas, route reads after a write to the primary for that session, monitor lag with an alert, and expose staleness in the interface where it matters.
The reason for that specificity is a failure I have seen: A profile save redirected to a replica read with 900 milliseconds of lag, so 6 percent of users saw their old details and saved them again, creating duplicate change records.
Read immediately after a write.
| Routing | Sees the new value | Primary load |
|---|---|---|
| always the replica | no, 900 ms lag | lowest |
| primary after a write in that session | yes | moderate |
| always the primary | yes | highest |
I would not consider it settled without evidence: Write and immediately read in a test against a lagging replica and confirm the session sees its own write.
Read your own writes from the primary.
Curated: · Written: · Reviewed:
QA-40How would you add a cache in front of a hot read?(show answer)
I would answer cache-aside and write-through by following one request through the service and its dependencies.
Cache-aside loads on a miss and leaves the write path untouched, which is simple but serves stale data until the key expires or is invalidated. Write-through updates the cache with the write, which keeps it fresh at the cost of coupling the write path to the cache.
Concretely, start with cache-aside and a short time to live, key by the exact query inputs including tenant, invalidate on write where correctness needs it, and measure the hit rate before expanding.
The reason for that specificity is a failure I have seen: A cache keyed without the tenant identifier served one tenant's pricing to another for 90 seconds until the key expired, and the incident required customer notification.
Two strategies on the same read.
| Strategy | Hit rate | Staleness |
|---|---|---|
| cache-aside, 60 s | 92% | up to 60 s |
| write-through | 92% | none, if every writer updates |
| no cache | none | none |
I would not consider it settled without evidence: Read the key format and confirm every input that changes the response is part of it, then verify with a multi-tenant test.
The cache key must contain everything that changes the answer.
Curated: · Written: · Reviewed:
QA-41A popular cache key expires and the database is hammered. What is happening?(show answer)
My approach to cache invalidation and stampedes separates what the database guarantees from what the code assumes.
When a hot key expires, every concurrent request misses and recomputes the same value, which multiplies load at the worst moment. A single-flight lock, a staggered expiry, or serving stale while refreshing prevents that.
Concretely, let one request recompute while others wait or serve the stale value, add jitter to expiry times so keys do not expire together, and refresh hot keys ahead of expiry.
The reason for that specificity is a failure I have seen: A dashboard key with 4,000 requests a second expired at a round minute, and 4,000 simultaneous misses drove the database to 100 percent processor use for 40 seconds.
Expiry under 4,000 requests a second.
| Design | Database queries at expiry |
|---|---|
| plain expiry | 4,000 |
| single-flight lock | 1 |
| stale while revalidating | 1, no user-visible wait |
I would not consider it settled without evidence: Expire the hot key under load in a test and confirm only one recomputation reaches the database.
One miss should cause one recomputation.
Curated: · Written: · Reviewed:
QA-42What does at-least-once delivery mean for your consumer code?(show answer)
For delivery guarantees in a queue, I would name the invariant that must hold before choosing a mechanism.
Most queues guarantee at-least-once delivery, so a message can arrive twice after a timeout, a rebalance, or a redelivery. The consumer must therefore be idempotent, because exactly-once delivery across a network is not something the broker can give you.
Concretely, make every handler idempotent by key, record processed message ids, keep the side effect and the record in one transaction where possible, and design for reordering as well as duplication.
The reason for that specificity is a failure I have seen: A consumer rebalance redelivered 1,200 messages, and because the handler was not idempotent, 1,200 duplicate refunds were issued before the consumer was stopped.
Twelve hundred redelivered messages.
| Consumer | Duplicate side effects |
|---|---|
| not idempotent | 1,200 |
| deduplicating on message id | 0 |
| idempotent by business key | 0 |
I would not consider it settled without evidence: Deliver the same message twice in a test and confirm exactly one side effect.
Assume the message will arrive twice.
Curated: · Written: · Reviewed:
QA-43How do you make a message handler idempotent in practice?(show answer)
I would size writing an idempotent consumer against real traffic and a real data volume, not a local fixture.
Idempotence means the second application of a message changes nothing, which requires either a natural key the handler can check or a record of processed ids in the same store as the effect. Checking a separate store risks a crash between the check and the write.
Concretely, derive a stable key from the message, insert the effect and the processed key in one transaction, use an upsert or a unique constraint to reject repeats, and keep the processed table pruned.
The reason for that specificity is a failure I have seen: A handler recorded processed ids in a cache and wrote the effect to the database, and a crash between the two steps caused 80 messages to be applied twice after restart.
Crash between the two writes.
| Design | After restart |
|---|---|
| cache for ids, database for the effect | 80 duplicates |
| unique constraint on the business key | 0 |
| one transaction for both | 0 |
I would not consider it settled without evidence: Kill the consumer between the effect and the bookkeeping in a test and confirm no message is applied twice after restart.
Put the effect and the bookkeeping in one transaction.
Curated: · Written: · Reviewed:
QA-44You must update the database and publish an event. How do you keep them consistent?(show answer)
The part of the dual-write problem interviewers probe is the partial failure, not the happy path.
Writing to two systems cannot be atomic, so a crash between them leaves the database updated with no event or an event with no row. The outbox pattern writes the event into the same transaction and a relay publishes it afterwards.
Concretely, insert the event into an outbox table in the business transaction, publish from the outbox with a relay that marks rows sent, make consumers idempotent, and monitor outbox age.
The reason for that specificity is a failure I have seen: A service committed an order then crashed before publishing, and 300 orders existed with no downstream fulfilment event until an engineer replayed them by hand.
Crash after the commit.
| Design | Missing events | Recovery |
|---|---|---|
| write then publish | 300 | manual replay |
| outbox plus relay | 0 | automatic |
I would not consider it settled without evidence: Kill the process between the commit and the publish in a test and confirm the event is still delivered after restart.
One transaction, then relay the event.
Curated: · Written: · Reviewed:
QA-45Events for one customer arrive out of order. Why, and how do you fix it?(show answer)
I would anchor event ordering and partition keys in a number from a load test or a trace rather than a preference.
Most brokers guarantee order only within a partition, so events for one entity must share a partition key to stay ordered. A random or round-robin key spreads one entity's events across partitions, where concurrent consumers process them in any order.
Concretely, partition by the entity identifier, keep one consumer per partition for ordered work, include a version or sequence number so a handler can reject stale updates, and accept that ordering across entities is not guaranteed.
The reason for that specificity is a failure I have seen: Address-change events were partitioned randomly, so an older change overwrote a newer one for 40 customers and shipments went to the previous address.
Two events for one customer.
| Partition key | Order preserved | Stale overwrite |
|---|---|---|
| random | no | 40 customers |
| customer id | yes | none |
| customer id plus version check | yes | rejected explicitly |
I would not consider it settled without evidence: Publish two ordered events for one entity and confirm the final state reflects the later one.
Order comes from the partition key.
Curated: · Written: · Reviewed:
QA-46A message fails every time it is processed. What should happen to it?(show answer)
What separates a strong answer on dead-letter queues is knowing how the naive version fails under load.
A message that cannot succeed will block or spin forever if it is retried without limit, so after a bounded number of attempts it belongs in a dead-letter queue with its error. That keeps the main flow healthy and makes the failure visible rather than silent.
Concretely, cap retries with backoff, move the message to a dead-letter queue with the error and attempt count, alert on dead-letter depth, and provide a replay path after the bug is fixed.
The reason for that specificity is a failure I have seen: One malformed message was retried without a cap and blocked its partition for 6 hours, delaying 48,000 downstream events behind it.
One poison message.
| Handling | Partition delay | Visibility |
|---|---|---|
| unlimited retries | 6 h | none |
| 5 attempts then dead letter | seconds | alert on depth |
I would not consider it settled without evidence: Send a permanently failing message and confirm it reaches the dead-letter queue and the partition keeps moving.
Failing messages must leave the main path.
Curated: · Written: · Reviewed:
QA-47How do you run a nightly job safely across several instances?(show answer)
I would treat scheduled and batch jobs as something that must survive retries, restarts, and a slow dependency.
Every instance running the same cron entry means the job runs several times, so scheduled work needs a lock or a scheduler that elects one runner. It also needs to be idempotent, because a retry after a partial run is normal.
Concretely, take a distributed lock or use a single-runner scheduler, make the job resumable with a checkpoint, process in bounded batches, and alert when a run is missed rather than only when it fails.
The reason for that specificity is a failure I have seen: A nightly invoice job ran on all 6 instances, which generated 6 copies of 4,200 invoices before anyone noticed the next morning.
Six instances, one schedule.
| Design | Runs per night | Duplicate invoices |
|---|---|---|
| cron on every instance | 6 | 21,000 |
| distributed lock | 1 | 0 |
| single-runner scheduler | 1 | 0 |
I would not consider it settled without evidence: Start the job on several instances in a test and confirm exactly one run proceeds.
One schedule, one runner, resumable work.
Curated: · Written: · Reviewed:
QA-48A customer asks for their data to be deleted. What does your backend need to do?(show answer)
Before changing anything for retention and deletion requests, I would write down how the result will be measured.
Data spreads into replicas, backups, caches, logs, analytics stores, and third-party systems, so deletion is a process across systems rather than one statement. Retention rules should be decided per data class before such a request arrives.
Concretely, maintain an inventory of where each data class lives, delete or anonymize in each store with a recorded completion, define what backups retain and for how long, propagate deletion to processors, and log the request and its outcome.
The reason for that specificity is a failure I have seen: A deletion removed the primary row but left the customer's address in the analytics warehouse and in 90 days of logs, which the next audit recorded as a compliance finding.
Where one record lives.
| Store | Action | Deadline |
|---|---|---|
| primary database | delete | immediate |
| analytics warehouse | delete or anonymize | 7 days |
| logs | expire | 30 days |
| backups | expire on schedule | 90 days |
I would not consider it settled without evidence: Run a deletion in a test tenant and search every store, including logs and analytics, for remaining identifiers.
Deletion is a list of systems, not one query.
Curated: · Written: · Reviewed:
QA-49Writes to a large table slow down over months even though the query plan is unchanged. What would you check?(show answer)
I would start maintaining a large table from the failure a caller would see, not from the framework feature.
Heavy update and delete traffic leaves dead rows and bloated indexes, which grows the data the engine must read and can stall writes. Autovacuum or the equivalent maintenance process is what keeps that bounded.
Concretely, monitor table and index bloat, tune autovacuum for the hottest tables, rebuild indexes when bloat is severe, archive old rows into partitions, and watch transaction age so maintenance is not blocked by a long-running transaction.
The reason for that specificity is a failure I have seen: A long-running analytics transaction held the snapshot open for 9 hours, autovacuum could not clean up, and dead rows grew the hot table by 38 percent with insert latency doubling.
Hot table over three months.
| Month | Dead rows | Insert latency |
|---|---|---|
| 1 | 2% | 3 ms |
| 2 | 19% | 5 ms |
| 3 | 38% | 6 ms |
I would not consider it settled without evidence: Track dead-row counts and index bloat weekly and confirm maintenance keeps them within the agreed bound.
Storage needs maintenance, not just indexes.
Curated: · Written: · Reviewed:
QA-50Analysts want to run heavy queries against your production database. What do you propose?(show answer)
The first question I ask about reporting queries on a transactional database is what happens on the second attempt.
Analytical queries scan far more data than transactional ones, so they compete for the same buffers and connections and can hold snapshots open. Serving them from a replica, a warehouse, or a materialized view protects the transactional path.
Concretely, route analysts to a read replica or a warehouse fed by change data capture, materialize the common aggregates on a schedule, cap analyst connections and statement timeouts, and keep the transactional database for the product.
The reason for that specificity is a failure I have seen: One unbounded analyst query scanned 90 million rows during peak hours, filled the buffer cache with cold data, and pushed checkout p99 from 180 milliseconds to 2.4 seconds.
During the analyst query.
| Analyst path | Checkout p99 | Analyst query time |
|---|---|---|
| primary database | 2,400 ms | 40 s |
| read replica | 180 ms | 45 s |
| warehouse with daily load | 180 ms | 6 s |
I would not consider it settled without evidence: Run the analyst workload against the replacement path and confirm transactional p99 latency is unchanged.
Keep analysis off the transactional path.
Curated: · Written: · Reviewed:
QA-51What does CAP actually force you to choose in a real system?(show answer)
With the CAP trade-off in practice, a passing staging run is where I start checking rather than stop.
During a network partition a system can keep answering with possibly stale data or refuse to answer until it is sure, and that is the only choice CAP forces. In normal operation the real trade is between latency and consistency rather than a binary.
Concretely, decide per operation rather than per system, keep money and inventory strongly consistent, allow stale reads where the cost of staleness is low, and state the choice in the design so nobody assumes otherwise.
The reason for that specificity is a failure I have seen: A stock count served from an eventually consistent replica during a partition oversold 240 units of a limited item, and every one had to be refunded and apologized for.
Three operations.
| Operation | Choice under partition | Cost |
|---|---|---|
| reserve stock | refuse | failed checkouts |
| product description | serve stale | minor |
| order history | serve stale | minor |
I would not consider it settled without evidence: Name each operation's choice in the design and test the behaviour with the dependency partitioned.
Choose per operation, not per system.
Curated: · Written: · Reviewed:
QA-52Which consistency guarantees do clients actually feel, and how do you provide them?(show answer)
I would answer consistency models a client notices by following one request through the service and its dependencies.
Users notice three things: seeing their own writes, never seeing time go backwards, and related updates appearing together. Those map to read-your-writes, monotonic reads, and causal consistency, and each can be provided without full linearizability.
Concretely, pin a session to one replica or the primary after a write, carry a version token so a read waits for at least that version, and group causally related writes so they become visible together.
The reason for that specificity is a failure I have seen: A feed read from two replicas with different lag showed a reply above the comment it answered for 8 percent of loads, which users reported as missing comments.
What each guarantee prevents.
| Guarantee | Prevents | Cost |
|---|---|---|
| read your writes | missing your own update | session pinning |
| monotonic reads | state moving backwards | version token |
| causal | reply before its parent | grouping writes |
I would not consider it settled without evidence: Write, then read from every replica in a test and confirm the session never sees older state than it has already seen.
Give users their own writes and a forward-moving clock.
Curated: · Written: · Reviewed:
QA-53An order must reserve stock, charge a card, and book delivery across three services. How do you keep that consistent?(show answer)
My approach to sagas instead of distributed transactions separates what the database guarantees from what the code assumes.
A transaction cannot span independent services, so the work becomes a sequence of local transactions with compensating actions for the steps already done. The system is eventually consistent, and every step must be idempotent and retryable.
Concretely, model the flow as explicit steps with a compensation for each, persist the state of the saga, retry forward where possible and compensate where not, and make every participant idempotent.
The reason for that specificity is a failure I have seen: A two-phase attempt across three services left 90 orders with the card charged and no stock reserved, and reconciliation took two engineers a full day.
Failure at the third step.
| Step | Forward action | Compensation |
|---|---|---|
| 1 | reserve stock | release stock |
| 2 | charge card | refund |
| 3 | book delivery | fails, triggers 1 and 2 |
I would not consider it settled without evidence: Fail each step in turn in a test and confirm the saga either completes or compensates fully.
Compensate rather than pretending one transaction spans services.
Curated: · Written: · Reviewed:
QA-54Only one instance should run a task. How do you guarantee that?(show answer)
For coordination and leader election, I would name the invariant that must hold before choosing a mechanism.
Agreement across instances needs a store with a strong guarantee, such as a lease with an expiry in a consensus system or a conditional write in a database. A lock held in memory or in a cache without fencing can be held by two instances after a pause.
Concretely, take a lease with a time to live, renew it while working, include a fencing token so a resumed old leader is rejected, and design the task so a lost lease stops the work safely.
The reason for that specificity is a failure I have seen: A leader paused for 12 seconds by garbage collection kept writing after its lock expired, and a second leader produced 2 conflicting settlement files for the same day.
Pause longer than the lease.
| Mechanism | Two writers possible | Detected |
|---|---|---|
| in-memory flag | yes | no |
| cache lock, no fencing | yes | no |
| lease plus fencing token | no | stale writer rejected |
I would not consider it settled without evidence: Pause the leader beyond the lease in a test and confirm its writes are rejected by the fencing token.
A lock without fencing is not exclusion.
Curated: · Written: · Reviewed:
QA-55Can you order events across servers by their timestamps?(show answer)
I would size clocks and ordering across machines against real traffic and a real data volume, not a local fixture.
Wall clocks on different machines drift and can jump backwards when corrected, so a timestamp comparison is not a reliable order. Logical counters, sequence numbers, or a single assigning authority give order that survives clock correction.
Concretely, use a monotonic clock for durations, a sequence or version per entity for ordering, and a single writer or a consensus store when a global order is genuinely needed, and never compare timestamps from two machines for correctness.
The reason for that specificity is a failure I have seen: A last-write-wins merge compared timestamps from two servers 400 milliseconds apart in clock, so the earlier edit won 1 in 9 conflicts and support saw reverted changes.
Two nodes, 400 ms of skew.
| Ordering basis | Correct order | Reverted edits |
|---|---|---|
| wall-clock timestamps | no | 1 in 9 conflicts |
| per-entity version | yes | 0 |
I would not consider it settled without evidence: Skew two nodes' clocks in a test and confirm ordering still follows the sequence number.
Order by sequence, measure with a monotonic clock.
Curated: · Written: · Reviewed:
QA-56What has to be true before you can add instances of a service freely?(show answer)
The part of statelessness and horizontal scaling interviewers probe is the partial failure, not the happy path.
An instance may be created or destroyed at any time, so nothing a request needs may live only in that process. Sessions, caches of authoritative data, uploaded files, and scheduled work all have to move to shared stores or be made instance-independent.
Concretely, keep request state in the request or a shared store, treat local memory as a cache only, keep uploads in object storage, elect a single runner for scheduled work, and confirm any instance can serve any request.
The reason for that specificity is a failure I have seen: Sessions were held in process memory behind a round-robin balancer, so 4 of every 5 requests appeared logged out once a second instance was added.
Two instances behind a balancer.
| State location | Works with 1 instance | Works with 2 |
|---|---|---|
| process memory sessions | yes | no, 80% logged out |
| shared session store | yes | yes |
| signed token in the request | yes | yes |
I would not consider it settled without evidence: Send consecutive requests to different instances in a test and confirm behaviour is identical.
Any instance must serve any request.
Curated: · Written: · Reviewed:
QA-57What happens to in-flight requests when you deploy?(show answer)
I would anchor graceful shutdown in a number from a load test or a trace rather than a preference.
A process that exits on the stop signal drops whatever it was doing, so a deploy becomes a burst of client errors. Graceful shutdown stops accepting new work, finishes in-flight requests within a grace period, and leaves the load balancer rotation first.
Concretely, handle the termination signal, fail readiness immediately so traffic drains, finish or reject in-flight work within the grace period, close pools and consumers cleanly, and set the platform's grace period above the longest normal request.
The reason for that specificity is a failure I have seen: A rolling deploy killed pods immediately on the signal, so each of 12 restarts dropped about 40 in-flight requests, and the release produced 480 client errors.
Twelve pod restarts under load.
| Shutdown | Dropped requests | Client errors |
|---|---|---|
| immediate exit | about 480 | 480 |
| drain then exit within 30 s | 0 | 0 |
I would not consider it settled without evidence: Deploy under load in staging and confirm zero dropped requests through the rollout.
Drain before exiting.
Curated: · Written: · Reviewed:
QA-58What is the difference between a liveness and a readiness probe, and what should each check?(show answer)
What separates a strong answer on liveness and readiness checks is knowing how the naive version fails under load.
Readiness says whether this instance should receive traffic now, and liveness says whether the process is beyond recovery and should be restarted. Checking dependencies in liveness turns a dependency outage into a restart loop that makes it worse.
Concretely, keep liveness local and cheap, put dependency checks in readiness so traffic drains instead of restarting, avoid expensive queries in either, and set thresholds that tolerate a brief blip.
The reason for that specificity is a failure I have seen: A liveness probe queried the database, so a 90 second database failover restarted all 20 pods repeatedly and the outage lasted 14 minutes rather than 90 seconds.
During a 90 second database failover.
| Probe design | Pod restarts | Outage length |
|---|---|---|
| dependency in liveness | 20 plus | 14 min |
| dependency in readiness only | 0 | 90 s |
I would not consider it settled without evidence: Fail the dependency in a test and confirm instances leave the rotation without restarting.
Liveness is about the process, readiness is about traffic.
Curated: · Written: · Reviewed:
QA-59How do you find out what your service can actually handle?(show answer)
I would treat load testing and capacity as something that must survive retries, restarts, and a slow dependency.
Capacity is the point where latency rises or errors begin, not the point where processor use looks high, and it must be measured with production-like data and a realistic mix of requests. A test against an empty database measures the framework rather than the service.
Concretely, replay a realistic request mix at production data volume, ramp until latency or errors break the target, record the binding resource, and re-test after each change.
The reason for that specificity is a failure I have seen: A load test against a 5,000 row database reported 4,000 requests a second, and the same service handled 260 against the real 40 million rows.
Same service, two datasets.
| Dataset | Sustained rate | p99 |
|---|---|---|
| 5,000 rows | 4,000 per second | 40 ms |
| 40 million rows | 260 per second | 600 ms |
| 40 million with the new index | 2,100 per second | 90 ms |
I would not consider it settled without evidence: Run the test at production data volume and report the request rate at which the latency target is missed.
Load test the data, not just the code.
Curated: · Written: · Reviewed:
QA-60Your average latency is 40 milliseconds but users complain. What are you missing?(show answer)
Before changing anything for tail latency, I would write down how the result will be measured.
Averages hide the tail, and a user request that fans out to several services experiences the slowest of them. With ten parallel calls, a p99 of one second means roughly one in ten user requests is slow.
Concretely, track p50, p95, and p99 per endpoint and per dependency, reduce fan-out or hedge slow calls, set per-call deadlines, and treat the tail as the target rather than the mean.
The reason for that specificity is a failure I have seen: A page fanned out to 10 services each with a p99 of 900 milliseconds, so 1 request in 10 took nearly a second while the reported average stayed at 40 milliseconds.
Ten parallel calls.
| Measure | Per call | User request |
|---|---|---|
| p50 | 12 ms | 40 ms |
| p99 | 900 ms | about 1 in 10 slow |
| after hedging | 900 ms | about 1 in 100 slow |
I would not consider it settled without evidence: Report the percentile distribution per dependency and the composed probability of a slow user request.
The user feels the tail, not the mean.
Curated: · Written: · Reviewed:
QA-61How would you set an availability target for your service?(show answer)
I would start service level objectives and error budgets from the failure a caller would see, not from the framework feature.
An objective is a measurable promise about what users experience, such as the share of requests served successfully within a latency bound over a window. The remaining budget is what licenses risky changes, so a target of 100 percent forbids all change and is never real.
Concretely, pick indicators from the user journey, set a target the business actually needs, measure from the client's perspective, and define what happens when the budget is spent, such as freezing risky releases.
The reason for that specificity is a failure I have seen: A team promised 99.99 percent with no measurement or policy, missed it twice in a quarter, and had no agreed response, so the number changed nothing except the tone of meetings.
One quarter.
| Item | Value |
|---|---|
| objective | 99.9% of requests under 300 ms |
| attained | 99.94% |
| budget consumed | 60% |
| policy | freeze risky releases above 100% |
I would not consider it settled without evidence: Report the objective, its measured attainment, and the budget consumed for the last window.
An objective without a policy is a slogan.
Curated: · Written: · Reviewed:
QA-62Which metrics would you add first to a new service?(show answer)
The first question I ask about metrics worth having is what happens on the second attempt.
Request rate, error rate, and duration describe what callers experience, and saturation of the constrained resource explains why. Counting internal events without those four leaves you unable to say whether the service is healthy.
Concretely, emit request rate, errors by class, and latency histograms per endpoint, add saturation for the pool or queue that binds first, label by tenant where cardinality allows, and keep dashboards to what a responder needs at 3 am.
The reason for that specificity is a failure I have seen: A service exported 240 business counters but no latency histogram, so during an incident nobody could say whether requests were slow or failing, and diagnosis took 50 minutes.
The first four.
| Metric | Answers |
|---|---|
| request rate | is traffic normal |
| error rate by class | is it failing |
| latency histogram | is it slow |
| pool saturation | why |
I would not consider it settled without evidence: Ask whether the dashboard answers is it up, is it slow, and is it failing, and confirm each has a chart.
Rate, errors, duration, saturation, in that order.
Curated: · Written: · Reviewed:
QA-63A request crosses six services and is slow. How do you find where the time goes?(show answer)
With distributed tracing, a passing staging run is where I start checking rather than stop.
A trace ties the spans of one request together across services, which shows the sequence, the parallelism, and the slow span. Without propagated context each service can only report its own view, and nobody can tell which call caused the delay.
Concretely, propagate trace context through every call and into asynchronous work, create spans for dependency calls with useful attributes, sample enough to catch the tail, and always sample errors and slow requests.
The reason for that specificity is a failure I have seen: A slow checkout was blamed on the payment service for 3 days until a trace showed 700 milliseconds of the 900 was a serial loop of 14 inventory calls in the checkout service itself.
One slow checkout trace.
| Span | Duration | Share |
|---|---|---|
| inventory, 14 serial calls | 700 ms | 78% |
| payment | 120 ms | 13% |
| everything else | 80 ms | 9% |
I would not consider it settled without evidence: Open one slow trace and name the span that owns most of the duration.
Propagate context or you are guessing.
Curated: · Written: · Reviewed:
QA-64What should page an on-call engineer at 3 am?(show answer)
I would answer alerting on symptoms by following one request through the service and its dependencies.
An alert should fire when users are affected or will be soon, because cause-based alerts on internal conditions produce noise that trains people to ignore pages. Objective burn and queue growth are symptoms, while high processor use alone usually is not.
Concretely, page on error-budget burn, user-visible error rate, latency breach, and stuck work, send everything else to a dashboard or ticket, require a runbook per alert, and delete alerts nobody acts on.
The reason for that specificity is a failure I have seen: A team received 340 pages in a month, 40 of which were real, and the on-call engineer missed a genuine outage because it looked like the usual processor-use noise.
One month of pages.
| Alert | Pages | Actioned |
|---|---|---|
| processor over 80% | 210 | 2 |
| disk over 70% | 90 | 1 |
| error budget burn | 12 | 12 |
| queue age over 30 min | 28 | 25 |
I would not consider it settled without evidence: Review a month of pages and report how many led to action, then delete the rest.
Page on what users feel.
Curated: · Written: · Reviewed:
QA-65Your service is failing in production. What do you do in the first ten minutes?(show answer)
My approach to responding to an incident separates what the database guarantees from what the code assumes.
Restoring service comes before understanding it, so the first moves are to declare the incident, take the fastest safe mitigation such as rollback or a flag, and communicate. Debugging while users are down extends the outage.
Concretely, declare and name a coordinator, check what changed recently, roll back or disable the suspect change, communicate at fixed intervals, and keep a timeline for the review.
The reason for that specificity is a failure I have seen: An engineer spent 40 minutes reading logs before rolling back a deploy that had started the errors, and the outage cost 40 minutes instead of 4.
Two responses to the same regression.
| Approach | Time to mitigate | Total outage |
|---|---|---|
| debug first | 40 min | 52 min |
| roll back first | 4 min | 4 min plus later analysis |
I would not consider it settled without evidence: Record time to mitigate and time to resolve separately for each incident.
Mitigate first, understand second.
Curated: · Written: · Reviewed:
QA-66What makes a postmortem useful rather than a formality?(show answer)
For postmortems that change something, I would name the invariant that must hold before choosing a mechanism.
A useful postmortem explains how the system allowed the failure, not who typed the command, and it leaves owned actions with dates. Blame suppresses the information the next incident needs.
Concretely, write a timeline with detection and mitigation times, identify contributing factors including missing guardrails, list actions with owners and dates, and review whether previous actions were completed.
The reason for that specificity is a failure I have seen: A postmortem named an engineer as the cause and created no action, and the same missing validation caused a second outage 5 weeks later.
Actions from one review.
| Action | Owner | Due | Done |
|---|---|---|---|
| add the missing validation | service owner | 1 week | yes |
| alert on queue age | on-call lead | 2 weeks | yes |
| rollback rehearsal | team | 1 month | yes |
I would not consider it settled without evidence: Check that each action has an owner and a date, and report the completion rate of actions from previous incidents.
Ask what allowed it, not who did it.
Curated: · Written: · Reviewed:
QA-67How do you release a backend change without risking every user?(show answer)
I would size canary and blue-green releases against real traffic and a real data volume, not a local fixture.
A canary exposes a small share of traffic to the new version and compares error rate and latency before expanding, while blue-green switches all traffic at once with a fast switch back. Both need the two versions to be compatible with the same data.
Concretely, keep migrations backwards compatible so both versions run, route a small share to the canary, compare error rate and latency against the control, expand in steps, and keep the previous version ready to take traffic.
The reason for that specificity is a failure I have seen: A release changed the schema and the code at once, so rollback was impossible after 12 minutes of writes, and the team fixed forward under pressure for 3 hours.
Canary stages.
| Stage | Traffic | Error rate | Decision |
|---|---|---|---|
| canary | 2% | 0.4% | expand |
| stage 2 | 25% | 0.4% | expand |
| full | 100% | 0.4% | complete |
I would not consider it settled without evidence: Confirm the previous version still serves correctly against the migrated schema before routing any traffic to the new one.
Keep both versions able to run.
Curated: · Written: · Reviewed:
QA-68How do you use feature flags in a backend service without creating a mess?(show answer)
The part of server-side feature flags interviewers probe is the partial failure, not the happy path.
A flag separates deploying from enabling, which allows dark launches, gradual exposure, and an instant switch back. Each flag is also a branch that must be tested and eventually removed, so an old flag is a latent bug.
Concretely, keep flags short-lived with an owner and a removal date, evaluate them in one place, default to the safe value when the flag service is unreachable, and log which variant served a request.
The reason for that specificity is a failure I have seen: A flag left in place for 14 months defaulted to the wrong branch when the flag service timed out, and 9 percent of requests used an unmaintained code path for 20 minutes.
Flag service unreachable.
| Flag default | Behaviour | Requests affected |
|---|---|---|
| last cached value | stale branch | 9% |
| documented safe default | known path | 0 |
I would not consider it settled without evidence: Make the flag service unreachable in a test and confirm every flag falls back to its documented safe default.
Every flag needs an owner and an expiry.
Curated: · Written: · Reviewed:
QA-69How should a service receive its configuration and its secrets?(show answer)
I would anchor configuration and secrets in a number from a load test or a trace rather than a preference.
Configuration belongs outside the artifact so the same build runs in every environment, and secrets belong in a manager that supports rotation and audit rather than in the repository or the image. A secret in version control must be treated as already leaked.
Concretely, read configuration from the environment or a config service, validate every value at start-up and refuse to boot if one is missing, fetch secrets from a manager with short-lived credentials, and rotate on a schedule and after any exposure.
The reason for that specificity is a failure I have seen: A database password committed in a configuration file stayed valid for 8 months after the repository was shared with a contractor, and rotation required a coordinated restart of 6 services.
Where each value lives.
| Value | Source | Rotated |
|---|---|---|
| feature toggles | config service | on change |
| database password | secret manager | every 90 days |
| signing key | secret manager | every 30 days |
I would not consider it settled without evidence: Scan the repository and images for secrets and confirm every credential comes from the manager with a rotation date.
A committed secret is a leaked secret.
Curated: · Written: · Reviewed:
QA-70The business asks for a second region. What does that actually cost?(show answer)
What separates a strong answer on multi-region trade-offs is knowing how the naive version fails under load.
A second region buys availability and lower latency for distant users, and it costs consistency, data-residency complexity, and a much harder failure model. Active-passive with asynchronous replication is far simpler than active-active writes.
Concretely, decide whether the second region serves reads only or takes writes, measure the replication lag you can tolerate, keep a single writer per data partition if possible, rehearse failover including the data path, and price the duplication.
The reason for that specificity is a failure I have seen: An active-active deployment without conflict handling produced 700 divergent rows in one week, and reconciling customer balances took a week of manual work.
Three shapes.
| Shape | Write conflicts | Failover time | Cost |
|---|---|---|---|
| single region | none | none available | 1.0x |
| active-passive | none | 15 min | 1.6x |
| active-active | needs resolution | seconds | 2.1x |
I would not consider it settled without evidence: Rehearse a full region failover and report the data loss window and the recovery time.
Two regions double the operations and complicate the truth.
Curated: · Written: · Reviewed:
QA-71How do you know your backups are good?(show answer)
I would treat backups, recovery point and recovery time as something that must survive retries, restarts, and a slow dependency.
A backup is only proven by a restore, and the two numbers that matter are how much data you can lose and how long recovery takes. Untested backups routinely fail on encryption keys, permissions, or missing dependencies.
Concretely, define the recovery point and recovery time objectives, take backups with point-in-time recovery where required, restore into an isolated environment on a schedule, verify row counts and application start-up, and store keys separately.
The reason for that specificity is a failure I have seen: Nightly backups ran for 14 months and the first restore attempt failed because the encryption key had been rotated and the old key was not retained, losing 9 hours of data.
Objective against tested reality.
| Measure | Objective | Last test |
|---|---|---|
| recovery point | 5 min | 4 min |
| recovery time | 1 h | 48 min |
| last successful restore | monthly | 9 days ago |
I would not consider it settled without evidence: Restore into an isolated environment on a schedule and record the achieved recovery point and recovery time.
A backup you have not restored is a hope.
Curated: · Written: · Reviewed:
QA-72How do you find out whether your failure handling works?(show answer)
Before changing anything for injecting faults on purpose, I would write down how the result will be measured.
Timeouts, retries, breakers, and fallbacks are only real if they have been exercised, because the first time they run in production is usually during an incident. Injecting a fault in a controlled window is how you learn which handler is wrong.
Concretely, start in staging with one dependency slowed or failed, state the expected behaviour first, run in production during business hours with an abort switch, and fix what the experiment reveals before the next one.
The reason for that specificity is a failure I have seen: A retry configuration that had never been exercised turned a 30 second dependency blip into a 9 minute outage, because every client retried 5 times without jitter at the same moment.
One experiment.
| Injected fault | Expected | Observed |
|---|---|---|
| dependency 3 s slower | breaker opens in 10 s | opened in 9 s |
| dependency returns 500 | fallback served | 5 retries, amplified load |
I would not consider it settled without evidence: Run the experiment and compare the observed behaviour with the behaviour you predicted in writing.
Untested failure handling is decoration.
Curated: · Written: · Reviewed:
QA-73Your cloud bill doubled after a release. How do you find out why?(show answer)
I would start cost awareness in backend design from the failure a caller would see, not from the framework feature.
Cost follows a few measurable drivers, usually data transfer, storage growth, per-request compute, and managed-service tiers. Without per-service tagging the bill is a single number nobody can act on.
Concretely, tag resources by service and environment, break the bill down per driver, measure cost per thousand requests, set an alert on an unexpected daily rise, and treat a large cost change as a design review.
The reason for that specificity is a failure I have seen: A logging change raised log volume from 40 to 900 gigabytes a day, which added 14,000 dollars a month and was noticed only when finance asked at the end of the quarter.
After one release.
| Driver | Before | After |
|---|---|---|
| log storage per day | 40 GB | 900 GB |
| monthly cost | $1,100 | $15,100 |
| cost per 1,000 requests | $0.004 | $0.052 |
I would not consider it settled without evidence: Report cost per thousand requests before and after each significant release.
Cost is a design property, measured per request.
Curated: · Written: · Reviewed:
QA-74Design a rate limiter that works across fifty service instances.(show answer)
The first question I ask about designing a distributed rate limiter is what happens on the second attempt.
A limit enforced per instance is not the limit the client sees, so the counter must be shared, and the shared store then becomes the hot path. The design balances accuracy against the latency and availability cost of that store.
Concretely, keep counters in a fast shared store with atomic operations, use a sliding window or token bucket, allow a small local allowance to reduce round trips, decide whether to fail open or closed when the store is unavailable, and return a retry-after header.
The reason for that specificity is a failure I have seen: A per-instance limiter of 100 requests a minute across 50 instances let one client send 5,000 a minute, which was 50 times the intended quota.
Intended quota of 100 per minute.
| Design | Admitted rate | Store calls per request |
|---|---|---|
| per instance | 5,000 per minute | 0 |
| shared counter | 100 per minute | 1 |
| shared plus local allowance of 5 | 100 to 105 | 0.2 |
I would not consider it settled without evidence: Drive a single client above the quota across instances and confirm the admitted rate matches the intended limit.
A shared limit needs a shared counter.
Curated: · Written: · Reviewed:
QA-75How does one service find the instances of another, and how is traffic spread across them?(show answer)
With service discovery and load balancing, a passing staging run is where I start checking rather than stop.
Instances come and go, so callers resolve a logical name through a registry or the platform's DNS rather than a fixed address. The balancing algorithm then decides how unevenly load lands, and least-outstanding-requests handles uneven instance speed far better than round robin.
Concretely, resolve through the platform's service name, respect short record lifetimes so replaced instances are not cached, remove failing instances from rotation through readiness, and prefer least-outstanding-requests when instance latency varies.
The reason for that specificity is a failure I have seen: A client cached resolved addresses for 30 minutes, so after a deploy 40 percent of calls went to instances that no longer existed and failed until the cache expired.
During a rolling replacement.
| Client behaviour | Failed calls | Recovery |
|---|---|---|
| addresses cached 30 min | 40% | after expiry |
| resolves per call, 5 s lifetime | under 1% | seconds |
| least outstanding requests | under 1% | balanced by speed |
I would not consider it settled without evidence: Replace every instance during a test and confirm callers resolve the new ones within the record lifetime.
Resolve the name on every call path, not once at start-up.
Curated: · Written: · Reviewed:
QA-76How does SQL injection still happen, and what actually prevents it?(show answer)
I would answer SQL injection by following one request through the service and its dependencies.
Injection happens whenever a query is assembled from text that includes untrusted input, so the fix is to send the query and the values separately as parameters. Escaping by hand and allowlisting characters both fail on edge cases that parameter binding handles by construction.
Concretely, use parameterized statements everywhere, keep identifiers such as table and column names out of user input or map them through an allowlist, review any dynamic query construction, and grant the application the narrowest database role it needs.
The reason for that specificity is a failure I have seen: A sort parameter was concatenated into an order-by clause, and a crafted value dumped 40,000 customer rows before the endpoint was disabled.
Three parameter kinds.
| Input | Safe approach | Unsafe approach |
|---|---|---|
| filter value | bound parameter | string concatenation |
| sort column | allowlist of 6 columns | interpolated name |
| page size | validated integer | interpolated text |
I would not consider it settled without evidence: Send injection payloads through every parameter, including sort and filter names, and confirm none changes the query's structure.
Send values as parameters, never as text.
Curated: · Written: · Reviewed:
QA-77What is the most common serious API vulnerability you would look for in a code review?(show answer)
My approach to broken object-level authorization separates what the database guarantees from what the code assumes.
The usual flaw is an endpoint that authenticates the caller but never checks that the requested object belongs to them, so changing an identifier in the URL returns someone else's data. Authentication answers who, and this check answers whether this object is theirs.
Concretely, scope every query by the caller's tenant or owner rather than filtering after the fetch, put the check in a shared layer so new endpoints inherit it, and test with two accounts on every endpoint that takes an identifier.
The reason for that specificity is a failure I have seen: An invoice endpoint fetched by id and checked only that the caller was logged in, so incrementing the id exposed 9,000 invoices across tenants.
Two accounts against one endpoint.
| Query | Own invoice | Another tenant's invoice |
|---|---|---|
| fetch by id only | 200 | 200, leak |
| fetch by id and tenant | 200 | 404 |
I would not consider it settled without evidence: Call every identifier-taking endpoint with another tenant's id and confirm each returns 404 or 403.
Authenticate the caller, then authorize the object.
Curated: · Written: · Reviewed:
QA-78How do you keep one tenant's data from reaching another in a shared database?(show answer)
For tenant isolation, I would name the invariant that must hold before choosing a mechanism.
Isolation has to be enforced where the data is read, because a filter that lives in application code is one forgotten query away from a leak. Row-level security, a mandatory tenant predicate in a single data layer, or separate schemas each move the check closer to the data.
Concretely, set the tenant on the connection or session and enforce it in the database where available, route every read through one layer that requires the tenant, and run a test suite that queries each table as two tenants.
The reason for that specificity is a failure I have seen: One reporting query written outside the data layer omitted the tenant predicate and returned 12 tenants' revenue in a single customer-facing export.
Where the predicate lives.
| Enforcement | Forgotten-query risk | Cost |
|---|---|---|
| in each query | high, 1 of 340 queries | none |
| single data layer | low | small refactor |
| row-level security | very low | policy maintenance |
I would not consider it settled without evidence: Run the full test suite as two tenants and confirm no query returns a row belonging to the other.
Enforce the tenant where the rows are read.
Curated: · Written: · Reviewed:
QA-79What does encrypting data at rest actually protect against?(show answer)
I would size encryption in transit and at rest against real traffic and a real data volume, not a local fixture.
Encryption at rest protects against stolen disks, misplaced backups, and snapshot leaks, and it does nothing against an application that queries the data or a compromised credential. Encryption in transit protects against network interception, including inside the cluster.
Concretely, enable transport security for every hop including service to database, use managed keys with rotation, encrypt fields such as tokens and identifiers at the application layer where the threat is a compromised query path, and keep keys out of the same store as the data.
The reason for that specificity is a failure I have seen: A team relied on disk encryption while an exposed internal endpoint allowed unauthenticated queries, so 2.3 million records were read straight from the application over plain text.
Which threat each control covers.
| Threat | At rest | In transit | Field level |
|---|---|---|---|
| stolen backup | yes | no | yes |
| network capture | no | yes | yes |
| compromised query path | no | no | partly |
I would not consider it settled without evidence: Name the threat each layer of encryption addresses and confirm the application path has its own control.
Disk encryption does not protect a live query path.
Curated: · Written: · Reviewed:
QA-80Which events should a backend service record for audit, and how?(show answer)
The part of audit logging for a service interviewers probe is the partial failure, not the happy path.
An audit log answers who did what to which record and when, which is a different job from debugging logs and needs to be tamper-evident and retained. Recording everything makes it unusable, and recording only errors makes it useless.
Concretely, record authentication events, authorization denials, privileged actions, and changes to sensitive records with actor, target, and outcome, write to append-only storage with the agreed retention, and keep personal data out of the message body.
The reason for that specificity is a failure I have seen: A support tool logged only failures, so when 40 accounts were altered by a compromised staff credential there was no record of what had been changed.
One audit record.
| Field | Example |
|---|---|
| actor | staff user 4471 through support tool |
| action | changed the email of customer 90210 |
| outcome | success |
| retained | 7 years, append only |
I would not consider it settled without evidence: Reconstruct one privileged change entirely from the audit log, including who performed it.
Audit the successful privileged actions, not only the failures.
Curated: · Written: · Reviewed:
QA-81A dependency you use has a critical advisory. What is your process?(show answer)
I would anchor dependency and supply-chain risk in a number from a load test or a trace rather than a preference.
Application code is a small share of what ships, so a dependency's vulnerability is your vulnerability, and reachability decides urgency. A pinned lockfile and a known inventory are what make a fast response possible.
Concretely, keep lockfiles committed and builds reproducible, scan dependencies continuously, judge whether the vulnerable path is reachable from your code, patch and deploy on a stated timeline by severity, and keep a bill of materials per release.
The reason for that specificity is a failure I have seen: An advisory sat unactioned for 6 weeks because nobody could say which of 9 services used the library, and the eventual audit finding required a written remediation plan.
Response by severity.
| Severity | Reachable | Deploy within |
|---|---|---|
| critical | yes | 24 hours |
| critical | no | 7 days |
| moderate | yes | 30 days |
I would not consider it settled without evidence: Answer which services and which releases contain a given package version within minutes, from the bill of materials.
You cannot patch what you cannot locate.
Curated: · Written: · Reviewed:
QA-82How would you divide tests for a backend service?(show answer)
What separates a strong answer on the test pyramid for a service is knowing how the naive version fails under load.
Fast tests on pure logic catch most mistakes cheaply, integration tests against a real database and broker catch the wiring that mocks hide, and a few end-to-end tests prove the whole path. Inverting that shape produces a suite that is slow, flaky, and eventually ignored.
Concretely, unit-test domain logic, run integration tests against real dependencies in containers, keep a handful of end-to-end journeys, assert query counts and contracts where they matter, and delete tests that fail for unrelated reasons.
The reason for that specificity is a failure I have seen: A suite of 900 tests mocked the database entirely, passed on a change that violated a unique constraint, and the failure appeared in production 20 minutes after deploy.
One suite rebalanced.
| Layer | Before | After | Runtime |
|---|---|---|---|
| unit | 900 | 700 | 20 s |
| integration with containers | 0 | 180 | 4 min |
| end to end | 12 | 8 | 5 min |
I would not consider it settled without evidence: Confirm the integration layer exercises the real database and broker rather than a mock.
Mocks cannot verify what the database does.
Curated: · Written: · Reviewed:
QA-83When should a test use a fake, and when a real dependency in a container?(show answer)
I would treat test doubles against real dependencies as something that must survive retries, restarts, and a slow dependency.
A fake keeps a test fast and deterministic but encodes your belief about the dependency, which is exactly what is wrong when the belief is mistaken. A real dependency in a container costs seconds and catches constraint, transaction, and serialization behaviour.
Concretely, use fakes for third-party services you cannot run and for error paths that are hard to trigger, use real containers for your database, cache, and broker, and verify fakes against a recorded contract so they do not drift.
The reason for that specificity is a failure I have seen: An in-memory fake for the message broker preserved ordering the real broker did not guarantee, and the ordering bug only appeared under production partitioning 3 weeks after release.
What each catches.
| Dependency | Test double | Real container |
|---|---|---|
| own database | misses constraints | catches them |
| own broker | misses ordering | catches it |
| third-party payments | correct choice | not available |
I would not consider it settled without evidence: Run the same integration suite against the fake and the real dependency and confirm both pass.
A fake tests your assumptions, a container tests the dependency.
Curated: · Written: · Reviewed:
QA-84How do you keep integration tests reliable as the schema changes?(show answer)
Before changing anything for managing test data, I would write down how the result will be measured.
Tests that depend on a shared mutable dataset interfere with each other and rot as the schema moves, so each test should create the data it needs through the same code path production uses. Hand-maintained dumps drift from the migrations.
Concretely, build data with factories that call real creation paths, run each test in a transaction that rolls back or a fresh schema, apply migrations before the suite, and keep one small realistic dataset for performance tests only.
The reason for that specificity is a failure I have seen: A 4 gigabyte shared test dump fell 11 migrations behind, and 60 of 300 integration tests failed for reasons unrelated to any change, so the team stopped trusting the suite.
Two approaches.
| Approach | Unrelated failures | Setup time |
|---|---|---|
| shared dump, 11 migrations behind | 60 of 300 | 90 s |
| factories plus rollback per test | 0 | 6 s |
I would not consider it settled without evidence: Run the suite twice in any order and confirm identical results without a manual reset.
Build test data through the code that creates it in production.
Curated: · Written: · Reviewed:
QA-85What does a good deployment pipeline for a backend service look like?(show answer)
I would start continuous delivery for a service from the failure a caller would see, not from the framework feature.
The pipeline should make the safe path the easy one: fast checks first, a single artifact promoted through environments, and a rollback that works without a rebuild. Long pipelines encourage batching, and batching makes every release riskier.
Concretely, run static checks, then unit tests, then integration tests, build one immutable artifact, promote it through environments, run smoke tests after deploy, and keep migrations backwards compatible so rollback is real.
The reason for that specificity is a failure I have seen: Releases were batched fortnightly with 60 changes each, and when one broke the team needed 3 hours to identify which change caused it.
Before and after.
| Measure | Fortnightly batches | Continuous |
|---|---|---|
| changes per release | 60 | 2 |
| time to identify a bad change | 3 h | 6 min |
| rollback | rebuild needed | promote the previous artifact |
I would not consider it settled without evidence: Measure deployment frequency, change failure rate, and time to restore for the last twenty releases.
Small releases make the cause obvious.
Curated: · Written: · Reviewed:
QA-86What do you look for when reviewing a pull request that touches a live service?(show answer)
The first question I ask about reviewing a backend change is what happens on the second attempt.
The risk in backend code is usually in the paths nobody demonstrated: the retry, the partial failure, the migration under load, and the query at production volume. A review should ask what happens on the second attempt and at ten times the data.
Concretely, check idempotence of anything that writes, look for missing timeouts and unbounded queries, confirm migrations are backwards compatible, check authorization on every new endpoint, and ask for the plan of any new query.
The reason for that specificity is a failure I have seen: A review approved a migration that added a non-null column with a default to a 30 million row table, and the deploy locked writes for 4 minutes during business hours.
A review checklist.
| Question | Why |
|---|---|
| what happens on retry | duplicate writes |
| is the query bounded | timeouts at volume |
| does the migration lock | deploy-time outage |
| who may call this | authorization gaps |
I would not consider it settled without evidence: Ask for the query plan and the migration's lock behaviour on production-sized data before approving.
Review the retry and the migration, not just the logic.
Curated: · Written: · Reviewed:
QA-87Errors are spiking in production. How do you work out what is happening?(show answer)
With debugging a live incident, a passing staging run is where I start checking rather than stop.
Diagnosis is fastest when you start from what changed and what the signals say, rather than from a hypothesis about the code. Deploys, configuration changes, traffic shifts, and dependency health cover most causes.
Concretely, check recent deploys and flag changes, read the error classes rather than the count, follow one failing trace end to end, compare dependency latency and saturation, and form one hypothesis at a time with a test for each.
The reason for that specificity is a failure I have seen: An engineer restarted every instance on a hunch, which cleared the queue backlog and destroyed the evidence, and the same incident recurred 2 days later with no more information.
First four checks.
| Check | Time | Outcome in this incident |
|---|---|---|
| recent deploys | 1 min | one 12 min ago |
| error classes | 2 min | all timeouts on one dependency |
| one trace | 3 min | 9 s in that call |
| dependency health | 1 min | saturated pool |
I would not consider it settled without evidence: Record the hypothesis, the signal that supported it, and the test that confirmed or refuted it.
Start from what changed and what the signals say.
Curated: · Written: · Reviewed:
QA-88A service is slow but no single query looks bad. How do you find the cost?(show answer)
I would answer profiling a slow service by following one request through the service and its dependencies.
Latency accumulates across serialization, allocation, lock contention, and repeated small calls, none of which shows up as one slow query. A profile of a real workload attributes the time rather than inviting a guess.
Concretely, profile under representative load, look at where wall-clock and processor time differ to separate waiting from computing, check allocation and lock contention, and re-profile after each change.
The reason for that specificity is a failure I have seen: A team spent 2 weeks tuning queries while 62 percent of request time was spent serializing a large response to JSON, which a 20 minute profile would have shown.
Where the time went.
| Component | Share of request time |
|---|---|
| JSON serialization | 62% |
| database queries | 21% |
| authorization checks | 9% |
| everything else | 8% |
I would not consider it settled without evidence: Attribute request time by component in a profile and show the largest contributor before and after the change.
Profile the request, not just the database.
Curated: · Written: · Reviewed:
QA-89A service's memory climbs until it is restarted. How do you investigate?(show answer)
My approach to memory and garbage collection problems separates what the database guarantees from what the code assumes.
A managed runtime reclaims what is unreachable, so growth means something is still referenced, typically an unbounded cache, a growing collection, or a registered listener. Pause times are a separate symptom, driven by heap size and allocation rate.
Concretely, take heap snapshots over time and compare retained sets, bound every cache, watch allocation rate as well as heap size, and tune the collector only after the retention bug is fixed.
The reason for that specificity is a failure I have seen: An unbounded per-tenant cache grew to 6 gigabytes over 4 days, which pushed pause times to 900 milliseconds and made the service restart every night.
Four days of uptime.
| Day | Heap | Longest pause |
|---|---|---|
| 1 | 1.2 GB | 40 ms |
| 3 | 4.1 GB | 380 ms |
| 4 | 6.0 GB | 900 ms |
I would not consider it settled without evidence: Compare heap snapshots hours apart and name the object graph that grows.
Growth is retention, not the collector.
Curated: · Written: · Reviewed:
QA-90How do you find and prevent a race condition in a service?(show answer)
For concurrency bugs, I would name the invariant that must hold before choosing a mechanism.
A race exists whenever two operations read and write shared state without ordering, and it appears under load rather than in tests. The fix is to remove the sharing or make the operation atomic.
Concretely, prefer immutable and per-request state, make the check and the write atomic in the database rather than in application code, and test with concurrent load rather than sequentially.
The reason for that specificity is a failure I have seen: A check for a remaining seat and the booking write were two statements, and under load 14 double bookings were created in one evening.
Fifty concurrent bookings for one seat.
| Implementation | Double bookings |
|---|---|
| read then write | 14 |
| conditional update in one statement | 0 |
| row lock then write | 0 |
I would not consider it settled without evidence: Run the operation from 50 concurrent workers in a test and confirm the invariant holds.
Make the check and the write one operation.
Curated: · Written: · Reviewed:
QA-91Would you use threads or an asynchronous model for a service that mostly waits on the network?(show answer)
I would size threads or asynchronous handling against real traffic and a real data volume, not a local fixture.
A thread per request is simple but each thread costs memory and context switching, so a workload dominated by waiting scales further with an asynchronous model on fewer threads. The asynchronous model breaks down if any handler blocks the loop with processor-bound work.
Concretely, choose asynchronous handling for input and output heavy services, keep processor-bound work on a separate pool, never call a blocking library from the loop, and measure concurrency against memory rather than assuming.
The reason for that specificity is a failure I have seen: One synchronous library call inside an asynchronous handler blocked the event loop for 300 milliseconds per call, and throughput collapsed from 4,000 to 60 requests a second.
Input-output bound workload.
| Model | Concurrent requests | Memory |
|---|---|---|
| thread per request | 400 | 1.6 GB |
| asynchronous, 4 threads | 8,000 | 300 MB |
| asynchronous with a blocking call | 60 | 300 MB |
I would not consider it settled without evidence: Load test both models on the real workload and report requests per second and memory per concurrent request.
Asynchronous only pays if nothing blocks the loop.
Curated: · Written: · Reviewed:
QA-92How do you estimate a backend feature?(show answer)
The part of estimating backend work interviewers probe is the partial failure, not the happy path.
The visible endpoint is a fraction of the work, while migrations, backfills, idempotence, authorization, observability, and rollout usually dominate. Breaking the feature into those parts produces an estimate that survives contact with production.
Concretely, list the schema change and backfill, the failure paths, the authorization rules, the metrics and alerts, the tests, and the rollout, estimate each, and give a range naming the largest unknown.
The reason for that specificity is a failure I have seen: An endpoint estimated at 3 days took 11, because a backfill over 40 million rows and a new authorization rule were never in the estimate.
Where eleven days went.
| Part | Estimated | Actual |
|---|---|---|
| endpoint and logic | 3 d | 3 d |
| migration and backfill | 0 | 3 d |
| authorization rules | 0 | 1 d |
| metrics, alerts, tests | 0 | 2 d |
| staged rollout | 0 | 2 d |
I would not consider it settled without evidence: Compare the estimate with actual time afterwards and record which category was missed.
Estimate the migration and the rollout, not the handler.
Curated: · Written: · Reviewed:
QA-93What should exist so that someone else can operate your service at 3 am?(show answer)
I would anchor runbooks and service documentation in a number from a load test or a trace rather than a preference.
An on-call engineer needs to know what the service does, what it depends on, what its alerts mean, and how to perform the three or four safe interventions. Architecture prose without those interventions does not help at 3 am.
Concretely, write one page per service with dependencies, dashboards, alert meanings, and the safe actions such as rollback, scale, and drain, keep a runbook per alert, and test both by having someone else use them during a drill.
The reason for that specificity is a failure I have seen: A page fired for a queue backlog with no runbook, and the responder spent 50 minutes finding the consumer scaling command while 120,000 messages accumulated.
What the runbook needs.
| Item | Present |
|---|---|
| what the alert means | yes |
| the 3 safe actions | yes |
| dashboards and traces links | yes |
| escalation and owner | yes |
I would not consider it settled without evidence: Have an engineer outside the team resolve a simulated alert using only the runbook.
Document the interventions, not the architecture.
Curated: · Written: · Reviewed:
QA-94A frontend team needs a field you think belongs elsewhere. How do you handle it?(show answer)
What separates a strong answer on working with frontend and product teams is knowing how the naive version fails under load.
API disagreements are usually about where a responsibility belongs, and the useful move is to find the underlying need before arguing about shape. Shipping whatever is asked creates a contract you maintain forever.
Concretely, ask which screen and decision the field serves, offer the alternative that keeps the responsibility in the right place, agree a contract in writing with an example payload, and version it if the answer changes.
The reason for that specificity is a failure I have seen: A backend added 9 presentation-shaped fields on request, and each change to the design then required a backend deploy, which added 3 days to every frontend iteration.
Three requests, three resolutions.
| Request | Underlying need | Resolution |
|---|---|---|
| formatted price string | locale display | client formats from amount and currency |
| combined name field | list rendering | client concatenates |
| total with tax | checkout accuracy | backend owns it, added |
I would not consider it settled without evidence: Write the agreed contract with an example payload and confirm both teams sign off before implementation.
Find the need before agreeing the shape.
Curated: · Written: · Reviewed:
QA-95Tell me about a production incident you owned.(show answer)
I would treat owning a production incident as something that must survive retries, restarts, and a slow dependency.
Interviewers want the detection, the decision under pressure, the communication, and what changed afterwards, with numbers. A story without a time to mitigate and a follow-up action sounds like a story about someone else's outage.
Concretely, state the impact and how it was detected, describe the mitigation and why you chose it, give the time to mitigate and to resolve, and name the change that prevents a repeat.
The reason for that specificity is a failure I have seen: A candidate described an outage entirely in terms of what a colleague did wrong, gave no timings, and the panel could not tell what the candidate had contributed.
The shape of the answer.
| Part | Example |
|---|---|
| impact | 6% of checkouts failing for 18 min |
| detection | error-budget alert |
| mitigation | rolled back the deploy, 4 min |
| follow-up | contract test added, shipped that week |
I would not consider it settled without evidence: Rehearse the story with the impact, the timings, and the follow-up action that shipped.
Own the decision and the follow-up.
Curated: · Written: · Reviewed:
QA-96Tell me about a technical decision you disagreed with.(show answer)
Before changing anything for a technical disagreement with a colleague, I would write down how the result will be measured.
The answer should show that you argued with evidence, understood the other position, and then committed to the decision the owner made. Interviewers are listening for whether you would keep relitigating.
Concretely, state the decision and the goal behind it, describe the evidence you brought such as a benchmark or a prototype, say who decided, and describe what you did afterwards.
The reason for that specificity is a failure I have seen: A candidate described refusing to implement a chosen approach and escalating twice, and the panel read it as someone who would not accept a settled decision.
A disagreement that worked.
| Part | Example |
|---|---|
| decision | adopt a message broker for one workflow |
| my concern | operational cost for 40 messages a day |
| evidence | prototype with a database queue, 1 day |
| outcome | database queue chosen, broker deferred |
I would not consider it settled without evidence: Check that the story names the evidence you brought and the decision you then supported.
Disagree with data, then commit.
Curated: · Written: · Reviewed:
QA-97Why do you want to work on backend systems rather than product surfaces?(show answer)
I would start why backend engineering from the failure a caller would see, not from the framework feature.
A convincing answer connects your experience to what the job actually rewards, such as reasoning about data correctness, failure, and scale, while showing you understand the costs, including on-call and slower visible feedback.
Concretely, give one or two specific examples of backend problems you enjoyed and shipped, say what you learned from operating them, and name the trade-offs you have accepted.
The reason for that specificity is a failure I have seen: A candidate said they preferred backend work because they did not enjoy talking to users, and the interviewer noted that the role required weekly work with product teams.
Motivation backed by evidence.
| Motivation | Example evidence |
|---|---|
| correctness under failure | made a payment path idempotent, duplicates to 0 |
| scale reasoning | cut p99 from 900 ms to 120 ms at 3 times traffic |
| ownership | on call for 18 months, 2 incidents I led |
I would not consider it settled without evidence: Prepare one example with the problem, your design decision, and the operational result.
Show you want the correctness and the pager, not just the abstraction.
Curated: · Written: · Reviewed:
QA-98Design the backend for a chat product with one million daily users.(show answer)
The first question I ask about designing a chat backend is what happens on the second attempt.
Chat is a fan-out and delivery problem: messages must be persisted, routed to connected recipients, and made available to those who were offline. Connection state, ordering per conversation, and read receipts drive most of the design.
Concretely, persist each message with a per-conversation sequence, hold connections on a gateway tier with a registry of which node holds which user, fan out through a broker, store an inbox per user for offline delivery, and page history by sequence.
The reason for that specificity is a failure I have seen: A first version fanned out by querying every group member's connection on each message, and a 400-member group produced 400 lookups per message, which saturated the registry at 900 messages a second.
Four decisions.
| Concern | Decision |
|---|---|
| ordering | per-conversation sequence number |
| delivery | broker topic per conversation |
| offline users | inbox rows, pulled on connect |
| history | keyset paging by sequence |
I would not consider it settled without evidence: Load test a large group at the target message rate and report delivery latency and registry load.
Route by conversation, not by scanning members.
Curated: · Written: · Reviewed:
QA-99Design the backend for a timeline that shows posts from accounts a user follows.(show answer)
With designing a feed backend, a passing staging run is where I start checking rather than stop.
The choice is where the work happens: fan-out on write precomputes each follower's timeline for fast reads, while fan-out on read queries followed accounts at request time. Accounts with millions of followers break the write-time approach and need a hybrid.
Concretely, precompute timelines for ordinary accounts, query at read time for very large accounts and merge, cap the stored timeline length, key pagination on a stable cursor, and process fan-out asynchronously with retries.
The reason for that specificity is a failure I have seen: Fan-out on write for an account with 4 million followers produced 4 million timeline inserts per post, which backed the queue up by 40 minutes and delayed every other user's feed.
One post, by account size.
| Followers | Strategy | Writes per post | Read cost |
|---|---|---|---|
| 500 | fan-out on write | 500 | 1 query |
| 4,000,000 | fan-out on read | 1 | merge at read |
| mixed feed | hybrid | bounded | 2 queries merged |
I would not consider it settled without evidence: Measure post-to-visible latency and queue depth for both an ordinary and a very large account.
Hybrid fan-out, because the tail of followers breaks one model.
Curated: · Written: · Reviewed:
QA-100How do you stop sensitive data ending up in your logs?(show answer)
I would answer keeping personal data out of logs by following one request through the service and its dependencies.
Logs are copied into aggregation systems, retained for weeks, and readable by more people than the database, so a logged token or identifier is a wider exposure than the record itself. Redaction has to happen where the log line is produced rather than in the pipeline afterwards.
Concretely, log identifiers rather than values, redact known sensitive fields in the logging layer, never log whole request or response bodies at information level, scan logs for patterns such as card numbers and tokens, and keep retention short.
The reason for that specificity is a failure I have seen: A debug statement logged the full request body of a signup endpoint, which wrote 240,000 plain-text passwords into an aggregation system retained for 30 days.
What each field becomes.
| Field | Logged as | Reason |
|---|---|---|
| customer email | customer id 90210 | identifier is enough |
| card number | last 4 digits only | support needs a hint |
| auth token | not logged | credential |
| request body | not logged at information level | contains both |
I would not consider it settled without evidence: Search a day of logs for token, password, and card patterns and confirm zero matches before release.
A log line is a copy you do not control.
Curated: · Written: · Reviewed:
