Top 100 Full-Stack Developer Interview Questions and Answers
The questions most likely to actually come up in your Full-Stack Developer interview, ranked by likelihood — with detailed, senior-level answers covering what an interviewer is really listening for.
Curated: · Written: · Reviewed:
QA-1After a cart write, which store is a refresh allowed to trust, and how do you keep the page and the database from disagreeing?(show answer)
I would start where client state lives versus the server from the screen the user sees and the row the database actually wrote.
Durable facts belong on the server and the client may hold only a working copy that a reload is allowed to discard. Mixing those jobs leaves the page looking correct while the row the API will persist is already stale.
Concretely, persist cart lines, prices, and ownership on the server as the only source a refresh reads. Keep UI chrome, draft keystrokes, and scroll position on the client. Reconcile the working copy from the API after every successful write so both halves name the same totals.
The reason for that specificity is a failure I have seen: A grocery app kept discounted prices only in client memory for 41 minutes, and 187 checkouts charged 17 percent less than the server catalog until the next deploy.
What a reload is allowed to forget.
| Fact | Client store | Server store |
|---|---|---|
| line price | working copy, 41 min stale | catalog row, billed |
| scroll position | kept | not stored |
| ownership of the cart | guessed from the tab | user id on the row |
I would not consider it settled without evidence: Reload the cart page after a write and confirm the totals match the database row, not the last in-memory snapshot.
Refresh is the test of which half actually owns the fact.
Curated: · Written: · Reviewed:
QA-2When can the page apply a click before the API answers, and what must both halves do if that guess is wrong?(show answer)
The first question I ask about optimistic UI versus server as source of truth is which side owns the truth after a refresh.
The client may guess the next state so a click feels instant, but the server response is the only row that survives a refresh. Optimism without a rollback path teaches the user a lie the API never stored.
Concretely, apply the expected change on the client with a pending marker and send the write with an idempotency key. Replace the working copy with the server body on success. Roll the screen back and show a retry when the API rejects the write.
The reason for that specificity is a failure I have seen: A news feed incremented a like count in the client and ignored status 409, so 2,640 screens kept a like the server had rejected, and 93 of those users shared a permalink that then returned not found.
Three writes, one source of truth.
| Action | Client first paint | Server row after refresh |
|---|---|---|
| like, status 200 | filled heart | filled heart |
| like, status 409 | filled, then empty | empty |
| pay, no optimism | spinner until body | charged once |
I would not consider it settled without evidence: Force the write to fail and confirm the client rolls back to the last server body rather than keeping the guessed count.
The click may guess but the row may not.
Curated: · Written: · Reviewed:
QA-3Should the browser call each microservice itself, or should a product API sit in front and return a page-shaped payload?(show answer)
With browser calling microservices versus a BFF, a green local UI is where I start checking rather than stop.
A browser that fans out to many microservices inherits every origin, auth, and version mismatch those services expose. A BFF owned by the product team collapses that into one contract the page can keep.
Concretely, route browser traffic through one product API that aggregates the downstream calls. Keep service credentials and internal paths on the server. Return a page-shaped payload so the client does not stitch six schemas together.
The reason for that specificity is a failure I have seen: A dashboard called 16 internal services from the browser, peak latency landed at 735 milliseconds, and 9 of the services advertised different CORS rules so one widget failed while the others painted.
Same dashboard, two call graphs.
| Caller | Requests on first paint | p95 |
|---|---|---|
| browser to 16 services | 16 | 735 ms |
| browser to one BFF | 1 | 210 ms |
| BFF to internals | 16, server side | hidden from CORS |
I would not consider it settled without evidence: Count browser-origin requests on the dashboard load and confirm they hit one product API rather than the internal fleet.
The browser should have one door, not a floor plan of the cluster.
Curated: · Written: · Reviewed:
QA-4For a browser SPA talking to your API, when does a cookie session beat a bearer token the page must store, and what does logout have to clear on both sides?(show answer)
I would answer cookie session versus bearer token in an SPA by following one click through the browser, the API, and the store.
A cookie session is sent by the browser on every matching request and is revoked on the server in one write. A bearer token in an SPA is copied by JavaScript and lives until it expires unless the client and the server both forget it.
Concretely, issue the session as a secure httpOnly same-site cookie from the API. Keep access tokens out of local storage. On logout, delete the cookie server-side and drop any in-memory token on the client so both halves stop accepting the session.
The reason for that specificity is a failure I have seen: An SPA stored a bearer token for 8 hours in local storage, a compromised script harvested 23 tokens, and those sessions stayed valid for 4.7 days because the API had no revocation list.
Where the credential lives.
| Arrangement | Readable by page script | Revoked on logout |
|---|---|---|
| cookie, httpOnly, server session | no | yes, one row |
| bearer in local storage | yes | no, until expiry |
| bearer in memory only | yes while the tab lives | yes, that tab only |
I would not consider it settled without evidence: From the page console confirm no long-lived token is readable, then logout and confirm the next API call from that tab and from a second tab both return status 401.
Logout is a server write plus a client forget, not a redirect.
Curated: · Written: · Reviewed:
QA-5If the session is a cookie, how do the page and the API stop another site from firing a write the browser will happily credential?(show answer)
My approach to CSRF on cookie-authenticated writes separates what the client may believe from what the server will persist.
Cookie-authenticated writes travel with the cookie whether the user meant to send them or not. The server must reject a state change that lacks a secret the attacking page cannot read, and the client must attach that secret on every write.
Concretely, require a CSRF token or a custom header on every cookie-authenticated POST, PATCH, and DELETE. Echo the token from the server into the page and attach it on the client. Set same-site on the session cookie as a second line, not the only line.
The reason for that specificity is a failure I have seen: A settings form posted a password change with cookies and no token, a foreign page triggered 156 successful writes in 11 minutes, and those users could not sign in until support reset the accounts.
What each half must do on a write.
| Control | Client | Server |
|---|---|---|
| CSRF token | attach on POST | reject if missing |
| same-site cookie | nothing | set on session |
| GET that changes state | do not use | do not accept |
I would not consider it settled without evidence: Submit the state-changing request from a different origin with the real cookie and confirm the API rejects it, then confirm the first-party page still succeeds with the token attached.
Automatic cookies make CSRF a two-sided problem.
Curated: · Written: · Reviewed:
QA-6A fetch works in a terminal and fails in the browser. What should the API advertise, and what must the page stop doing as a workaround?(show answer)
For CORS as a contract not a workaround, I would name the contract both halves must keep before choosing a library.
CORS is the server telling the browser which origins may read a response, not a client bug to proxy away. Echoing an unvalidated Origin header is an open door. A star with credentials is blocked by the browser and is not the usual leak.
Concretely, return an exact allowlist origin that you compared to the request Origin. Handle the preflight on the server. Never copy an unknown Origin into the allow-origin header, and never use a star when the client sends cookies.
The reason for that specificity is a failure I have seen: An API echoed any Origin it was given, 312 unknown sites could read account JSON with the user cookie, and a later proxy workaround added 47 milliseconds to every browser call without closing the hole.
What the browser actually checks.
| Server header | Cookie request from product origin | Cookie request from unknown origin |
|---|---|---|
| star, credentials on | browser blocks | browser blocks |
| exact origin, credentials on | allowed | blocked |
| exact origin, credentials off | cookies omitted | blocked |
I would not consider it settled without evidence: Inspect the preflight and the GET from the real origin and from a foreign origin, and confirm only the product origin receives the body.
Name the origin on the API rather than hiding the browser rule behind a proxy.
Curated: · Written: · Reviewed:
QA-7The bundle and the API ship on different days. How do you change a response so open tabs and a new deploy both keep working?(show answer)
I would size API versioning the UI can survive against a real page load and a real write, not a mocked fetch.
The UI bundle and the API ship on different clocks, so a field removal is a break for every stale tab. Additive changes plus a documented sunset let both halves move without a lockstep deploy.
Concretely, add new fields beside old ones and keep the old shape until the last active bundle stops reading it. Version the path or the header when a type must change. Teach the client to ignore unknown fields and to survive a missing optional one.
The reason for that specificity is a failure I have seen: A required field was renamed in place at noon, 1,420 open tabs still ran the previous bundle, and checkout returned status 422 for 27 hours until users hard-refreshed.
One field, three ship states.
| Stage | API body | Old bundle | New bundle |
|---|---|---|---|
| additive | old and new keys | reads old | reads new |
| rename in place | new key only | status 422 | works |
| sunset after zero old traffic | new key only | gone | works |
I would not consider it settled without evidence: Keep a canary of the previous bundle against the new API and confirm it still parses, then measure how many live sessions still send the old field before you remove it.
Stale tabs are callers so version for them.
Curated: · Written: · Reviewed:
QA-8A user double-clicks checkout and a timeout retries. How do the form and the API make that still one order?(show answer)
The part of idempotent form submit across the stack interviewers probe is the mismatch after deploy, not the happy render.
A double click and a retry after a timeout are the same write arriving twice. The client must send a key it generated once, and the server must return the original result instead of inserting a second row.
Concretely, stamp every unsafe form submit with a client-generated idempotency key and disable the button until the response lands. Persist that key with the first response on the server. Return the stored body for a repeat inside a 24 hour window.
The reason for that specificity is a failure I have seen: A checkout button stayed enabled during a 1.8 second wait, 1,655 orders were inserted twice, and finance spent the weekend matching 3,310 payment captures to 1,655 shipments.
Same click, two keys, two outcomes.
| Attempt | Key from the form | Rows the API inserted |
|---|---|---|
| first submit | k-a | 1 |
| timeout retry | k-a | 0, original body |
| second click, new key | k-b | 1 |
I would not consider it settled without evidence: Submit the same form twice with one key and confirm one row exists and both responses match, then submit a second key and confirm a second row.
The button and the table have to agree that twice is once.
Curated: · Written: · Reviewed:
QA-9The page wants an id before the create returns. Which identifier may the UI mint, and which identifier must later updates use?(show answer)
I would anchor client-generated ids versus server ids in a trace that starts at the click and ends at the query.
A client generated id is a legitimate idempotency key when the server authorizes the caller and enforces uniqueness. It is not a substitute for checking that this user may create that row, and hiding the server id from the client still makes later updates unaddressable.
Concretely, accept a client id on create as an idempotency key after you authorize the caller. Persist it with a unique constraint and return the server id as the canonical identifier. Address every later update with the server id. Reject a second user who presents the same client id.
The reason for that specificity is a failure I have seen: The API treated any client id as the primary key with no owner check, a second account reused an intercepted id, and 63 household lists were rewritten by the wrong user before anyone noticed.
Who mints, who owns.
| Identifier | Minted by | Used for |
|---|---|---|
| client id | UI, once per intent | create idempotency |
| server id | API | later GET and PATCH |
| both stored | database unique on client id | retry without a second row |
I would not consider it settled without evidence: Create twice with one client id and confirm one row, then PATCH with the server id and confirm the client no longer sends the minted id as the path.
Mint for retry and address with what the server stored.
Curated: · Written: · Reviewed:
QA-10A list changes while the user scrolls or walks back. How do the page and the API resume without skipping or repeating rows?(show answer)
What separates a strong answer on cursor pagination the UI can resume is knowing which half can lie to the other.
Offset pages shift when the server inserts rows, so a client that resumes by page number skips or repeats. A keyset cursor only stays stable when the sort is immutable and uniquely tie-broken.
Concretely, return an opaque cursor built from the sort columns plus a unique id. Persist that cursor on the client across back navigation. Reject a page-number resume on lists that change under the user, and reject a cursor whose sort no longer matches.
The reason for that specificity is a failure I have seen: An inbox used offset paging while mail arrived, 4,150 messages were skipped on resume, and 880 appeared twice so users marked the wrong threads read.
Resume after three inserts.
| Resume token | Rows seen twice | Rows skipped |
|---|---|---|
| page 7 of 20 | 880 | 4,150 |
| keyset created_at plus id | 0 | 0 |
| no token, refetch from top | 0 | user loses place |
I would not consider it settled without evidence: Insert rows during a paged scroll, walk back, and confirm the client plus API window contains each id once.
Store the cursor the API gave you, not the page index you counted.
Curated: · Written: · Reviewed:
QA-11The API returns 401, 404, 409, and 503. What should the screen do for each, and what must it never paint?(show answer)
I would treat mapping HTTP errors onto UI states as something that must survive a stale bundle talking to a new API.
Status codes are part of the contract the page must map, not strings to swallow into a generic toast. Treating every failure as an empty list hides outages, and treating every 404 as a crash hides missing records.
Concretely, map 401 to sign-in, 403 to a permission state, 404 to not found, 409 to a conflict the user can resolve, and 5xx to a retry. Keep that map in the client next to the API error body. Never paint an empty list for a server failure.
The reason for that specificity is a failure I have seen: A billing page treated status 503 as zero invoices, 673 customers saw an empty account for 37 minutes, and 112 of those customers submitted duplicate payments.
One endpoint, four paintings.
| Status | Wrong UI | Contract UI |
|---|---|---|
| 401 | empty table | sign-in |
| 404 | crash overlay | not found |
| 503 | empty table | retry |
I would not consider it settled without evidence: Stub each status from the API and confirm the client shows the matching surface rather than an empty table.
Empty and down are different screens chosen by the status line.
Curated: · Written: · Reviewed:
QA-12When the query returns no rows, when it is still in flight, and when it failed, how must the API and the page disagree on purpose?(show answer)
Before changing either half for loading and empty states as contract, I would write down how both sides will be checked.
Loading, empty, success, and error are four different contracts, and the server must distinguish no rows from a failed read. A spinner that never yields, or an empty table drawn on a timeout, trains users to click twice.
Concretely, return an explicit empty collection with status 200 when the query succeeds and finds nothing. Hold a loading state on the client only while the request is in flight. Render a distinct error surface when the API times out or fails.
The reason for that specificity is a failure I have seen: Search showed a spinner for 5.8 seconds then an empty table on a gateway timeout, 428 users reran the query, and the API recorded 1,284 extra reads against a replica that was already sick.
Three answers the page can tell apart.
| API outcome | Client surface | User action |
|---|---|---|
| 200, items empty | empty state | none |
| in flight | spinner | wait |
| timeout | error with retry | one more fetch |
I would not consider it settled without evidence: Drive the client with a slow 200 empty body, a 200 with rows, and a timeout, and confirm three different surfaces.
Zero rows is a successful 200 and a timeout is not zero rows.
Curated: · Written: · Reviewed:
QA-13A new checkout tender is flagged. Who evaluates the flag for the button, who evaluates it for the write, and what happens when they disagree?(show answer)
I would start feature flags that must match both halves from the screen the user sees and the row the database actually wrote.
A flag that is true in the bundle and false on the API, or the reverse, ships a screen that cannot complete its own writes. Both halves must evaluate the same key, the same user, and the same default.
Concretely, evaluate flags on the server for authorization and payload shape. Send the evaluated set to the client on bootstrap so the UI hides what the API will reject. Refuse a write whose flag is off even if the client still shows the control.
The reason for that specificity is a failure I have seen: Checkout for a new tender was on in 29 percent of browsers and off on the API, 1,100 taps produced not found, and support closed them as frontend bugs for 46 hours.
Same key, three evaluations.
| Client flag | Server flag | Result |
|---|---|---|
| on | on | button and POST both work |
| on | off | button lies, POST 404 |
| off | on | hidden, write still gated |
I would not consider it settled without evidence: Bootstrap a flagged-off user and a flagged-on user, confirm the button and the write agree, then flip the flag on the server only and confirm the stale bundle is rejected on POST.
The button is not the flag, the write is.
Curated: · Written: · Reviewed:
QA-14The UI compiles a TypeScript type and the API validates a schema. How do you keep those from drifting, and who fails first when they do?(show answer)
The first question I ask about shared types as the loop contract is which side owns the truth after a refresh.
The type the UI compiles and the schema the API serves are one contract, not two guesses. Generating both from one source is how a rename fails at build time instead of at runtime.
Concretely, publish the request and response types from the same schema the server validates. Import those types in the client instead of hand-writing mirrors. Fail the build when a required field is removed or retyped.
The reason for that specificity is a failure I have seen: Amount was changed from integer cents to a decimal string on the API only, 6 screens still sent integers, and 940 payouts stored 0 until a 117 minute rollback.
Where a rename dies.
| Source of types | Client still on old field | First failure |
|---|---|---|
| handwritten mirrors | yes | runtime, 940 zeros |
| generated from API schema | no, CI red | build |
| OpenAPI ignored | yes | production |
I would not consider it settled without evidence: Change a required field type in the schema and confirm the client package fails CI before a browser can send the old shape.
One schema, two compilers, no handwritten twins.
Curated: · Written: · Reviewed:
QA-15A user picks a large file. Should the bytes travel through your API, or should the API only mint permission while the browser talks to the store?(show answer)
With file upload through the API versus direct-to-store, a green local UI is where I start checking rather than stop.
Pushing a large file through the product API couples browser memory, request timeouts, and server disks. A signed direct-to-store upload keeps the API as the issuer of permission and the recorder of metadata, not the pipe.
Concretely, issue a short-lived upload URL from the API after checking the user and the content type. PUT the bytes from the client to the store. Confirm the object on the server and persist the metadata before the UI marks the upload done.
The reason for that specificity is a failure I have seen: Profile photos of 95 megabytes posted through the API exhausted a 512 megabyte dyno, 48 uploads failed with status 413, and the retry button sent the same body through the app server again.
Who carries 95 megabytes.
| Path | App server memory | When the UI may show done |
|---|---|---|
| POST file to API | whole body | response 200 |
| signed PUT to store, no confirm | none | too early, row missing |
| signed PUT, then API confirm | metadata only | after the row exists |
I would not consider it settled without evidence: Upload a file at the product size cap and confirm the app server never sees the bytes, then confirm the client does not show done until the API has the metadata row.
The API grants the write, the store takes the bytes, and the page waits for both.
Curated: · Written: · Reviewed:
QA-16A board should feel live. How do you choose poll versus a socket, and how do both halves stay authenticated when the transport changes?(show answer)
I would answer websocket versus polling for a live screen by following one click through the browser, the API, and the store.
A live screen needs a freshness budget, not a transport fashion. Polling is simple and cacheable, while a websocket is worth the sticky connection only when the server can push sooner than the poll would have waited.
Concretely, bound the poll interval on the client to the staleness the product can tolerate. Open a websocket only for rooms the server will actually push, and fall back to poll when the socket drops. Authenticate the socket with the same session the HTTP API uses.
The reason for that specificity is a failure I have seen: A board polled every 1.5 seconds from 18,000 open tabs, the API spent 410 milliseconds per tick, and switching the active room to a websocket cut origin CPU by 61 percent.
Same board, two freshness paths.
| Transport | Time to fresh pixel | Origin cost at 18,000 tabs |
|---|---|---|
| poll every 1.5 s | up to 1.5 s | 410 ms per tick |
| websocket push | 80 ms | push to the room |
| socket drop, poll fallback | 1.5 s again | poll resumes |
I would not consider it settled without evidence: Measure origin CPU and time-to-fresh-pixel for poll and for a socket on the real room size, and confirm a dropped socket resumes via poll without a second login.
Pick the transport that meets the freshness budget without a second auth story.
Curated: · Written: · Reviewed:
QA-17The server HTML and the first client render disagree. What did each half read, and how do you make the first paint one tree?(show answer)
My approach to SSR hydration mismatch as a contract bug separates what the client may believe from what the server will persist.
The HTML the server sent and the first client render must name the same tree, because a mismatch is the contract breaking before the user clicks. Reading time, randomness, or cookies only on one half is the usual cause.
Concretely, render the first paint from props the server already computed. Defer browser-only values until after hydration on the client. Keep cookies that affect markup readable during the server pass so both halves see them.
The reason for that specificity is a failure I have seen: A greeting branched on a timezone the server lacked, 1,240 sessions logged a hydration warning, and the header flashed empty for 220 milliseconds on every load.
What each half may read on first paint.
| Value | Server HTML | First client render |
|---|---|---|
| timezone from cookie | missing | present, mismatch |
| greeting from props | same string | same string |
| window width | unavailable | deferred after hydrate |
I would not consider it settled without evidence: Render on the server and hydrate in a test with a fixed clock and a cookie the markup depends on, and confirm no mismatch is reported.
First paint is a two-sided snapshot, not a client correction.
Curated: · Written: · Reviewed:
QA-18A profile screen is slow. How do you tell overfetch from underfetch at the client-server boundary, and which half should shrink the payload?(show answer)
For GraphQL overfetch versus REST underfetch at the boundary, I would name the contract both halves must keep before choosing a library.
Overfetch and underfetch are both boundary costs. GraphQL lets the client ask for a shape, REST returns a resource, and the wrong one at the page boundary either bloates the payload or multiplies the round trips.
Concretely, measure the bytes and the request count for the actual screen. Collapse underfetching REST with a BFF field set rather than N client calls. Cap GraphQL query cost on the server so one client cannot pull the whole graph.
The reason for that specificity is a failure I have seen: A profile page ran 5 REST calls then a team switched to one GraphQL query that pulled 180 kilobytes of unused edges, and TTI moved from 1.1 seconds to 2.8 seconds on mid-range phones.
One profile, three boundaries.
| Boundary | Requests | Payload |
|---|---|---|
| 5 REST resources | 5 | 28 kB, underfetch waterfalls |
| uncapped GraphQL | 1 | 180 kB unused edges |
| BFF page slice | 1 | 31 kB used fields |
I would not consider it settled without evidence: Compare request count, payload bytes, and TTI for the same screen on REST, on a BFF slice, and on a cost-capped GraphQL query.
Count bytes and round trips for the page, not for the fashion.
Curated: · Written: · Reviewed:
QA-19The API caps a search at 120 per minute. What must it return, and what must the typeahead do instead of retrying in a loop?(show answer)
I would size rate limits visible in the UI against a real page load and a real write, not a mocked fetch.
A limit the API enforces and the UI never surfaces becomes a silent retry storm. The client must back off on status 429, and the server must say when to retry.
Concretely, return status 429 with a retry-after value from the API. Honor that delay on the client and disable the repeating action. Show a quota state rather than spinning until the next click.
The reason for that specificity is a failure I have seen: A typeahead ignored 429 and retried immediately, 6,200 extra searches hit a 120 per minute cap, and the origin spent 13 minutes serving only those retries.
What the box does at the cap.
| Client behaviour | Extra requests | User sees |
|---|---|---|
| retry at once on 429 | 6,200 | spinner |
| honor retry-after | 0 | wait |
| quota surface | 0 | typed message |
I would not consider it settled without evidence: Drive the search box past the cap and confirm the client waits the advertised delay and paints a quota state instead of issuing another fetch.
Status 429 is a message to the page, not a hint to hammer harder.
Curated: · Written: · Reviewed:
QA-20The user logs out in one tab while two others stay open. What must the API destroy, and how do the other tabs find out?(show answer)
The part of multi-tab session and logout interviewers probe is the mismatch after deploy, not the happy render.
A cookie-authenticated tab that logs out already clears that cookie for the origin, so later writes from another tab should 401. What stays wrong is the other tab still painting a signed-in shell until you broadcast UI state.
Concretely, destroy the server session on logout. Broadcast a storage event so other tabs drop in-memory user state and show signed out. Recheck the session when a tab becomes visible, and never treat a painted name as proof the cookie still exists.
The reason for that specificity is a failure I have seen: Users logged out in one of 4 tabs, 361 other tabs still showed the previous name for 26 minutes, and 21 of those users tried to post, hit 401, and thought the site was down.
Three tabs, one session row.
| After logout in tab A | Server session | Tab B write |
|---|
| cookie cleared only | still valid | no cookie, status 401, tab still paints the name |
| server row deleted, no broadcast | gone | status 401, tab still paints the name |
| row deleted plus storage event | gone | status 401, signed-out UI |
I would not consider it settled without evidence: Log out in tab A, then from tab B POST a write and confirm status 401, and confirm tab B has dropped its in-memory user.
Every open tab is a client and logout has to reach them all.
Curated: · Written: · Reviewed:
QA-21Someone shares a project URL. What must be in the path, and what must the first paint fetch even if the original tab still has the entity in memory?(show answer)
I would anchor deep links that require server data in a trace that starts at the click and ends at the query.
A URL the user can share must be enough for the server to load the record and for the client to render it. Client-only routers that hold the entity in memory make the shared link an empty shell.
Concretely, encode the server id in the path. Fetch that id on first paint even when the client has a cache. Return status 404 from the API when the record is gone so the page can show not found instead of a spinner.
The reason for that specificity is a failure I have seen: Shared project links depended on a client cache, 2,275 cold loads showed a blank canvas, and 310 of those users created duplicate projects with the same name.
Same URL, three first paints.
| First paint | Client cache | API |
|---|---|---|
| warm tab | hit | skipped, looks fine |
| cold shared link, no fetch | miss | never called, blank |
| cold shared link, GET by id | miss | 200 or 404 |
I would not consider it settled without evidence: Open the shared URL in a logged-in browser with an empty cache and confirm the client fetches the id and paints the record or a 404 from the API.
The link is an API read, not a souvenir of another tab.
Curated: · Written: · Reviewed:
QA-22Search has debounce, filters, sort, and a URL. What does the client send, what does the API echo, and how do the box and the address bar stay one query?(show answer)
What separates a strong answer on search as a joint client-server protocol is knowing which half can lie to the other.
Search is a protocol with debounce, query text, filters, sort, and a cursor, not a text box that hopes the API guesses. Overlapping uncancelled fetches, not extra server sleep, are what duplicate the query.
Concretely, debounce on the client, abort the previous in-flight search, then send the full query document in one request. Rank and page on the server against that document. Echo the normalized query back so the URL and the box stay in sync.
The reason for that specificity is a failure I have seen: Each keystroke launched a fetch without aborting the last one, 1,900 overlapping queries landed per peak hour, and result tokens in the URL did not match the box after back navigation.
Who waits, who ranks.
| Step | Client | Server |
|---|
| debounce 250 ms, previous fetch aborted | one live fetch | ranks immediately |
| debounce, no abort | overlapping fetches | duplicates |
| echo normalized query | write URL and box | returns the document |
I would not consider it settled without evidence: Type a query, change a filter, walk back, and confirm the echoed document, the URL, and the box are the same object the API ranked.
The box, the URL, and the index must speak one query document.
Curated: · Written: · Reviewed:
QA-23Two tabs edit the same record. How does the page send what it last read, and how does the API refuse a stale save without last-write-wins?(show answer)
I would treat ETag optimistic lock in the UI as something that must survive a stale bundle talking to a new API.
Last write wins is a UI lie when two tabs edit the same row. The client must send the ETag it last read, and the server must reject a stale write so the page can reload.
Concretely, return an ETag with every GET of a mutable resource. Require If-Match on PUT and PATCH. On status 412, refresh the client from the server and show the conflict instead of overwriting.
The reason for that specificity is a failure I have seen: Two agents edited a ticket without If-Match, 88 notes were overwritten, and the UI still showed the first agent draft as saved for 39 minutes.
Two tabs, one row.
| Write | If-Match sent | Result |
|---|---|---|
| tab A, fresh tag | matches | 200, new ETag |
| tab B, stale tag | mismatch | 412, no overwrite |
| tab B after refetch | new tag | 200 |
I would not consider it settled without evidence: GET a record in two tabs, PUT from both, and confirm the second receives status 412 and then refetches rather than clobbering the first write.
The tag you read is the lock you send.
Curated: · Written: · Reviewed:
QA-24A payment provider posts a webhook. How does that become a badge on an open tab, and what must be stored before either the email or the toast fires?(show answer)
Before changing either half for webhook to in-app notification path, I would write down how both sides will be checked.
A webhook lands on the server, not on the open tab. The path from provider to badge is persist, then push or poll, and skipping the persist step makes the toast a fact the database never stored.
Concretely, accept the webhook on the API, persist the notification row, then push to subscribed clients. Fall the open page back to a poll of that row. Never toast from the webhook handler without a stored id the UI can fetch.
The reason for that specificity is a failure I have seen: Payment webhooks triggered emails only, 3,400 in-app badges stayed at zero for 22 minutes, and 205 users paid again because the screen still said unpaid.
Provider to badge.
| Step | Server | Open client |
|---|---|---|
| webhook only emails | no row | badge stays 0 |
| persist then push | row id exists | badge increments |
| persist, socket down, poll | row id exists | badge on next poll |
I would not consider it settled without evidence: Post a signed webhook, confirm the notification row exists, then confirm an open tab receives the badge by push or by poll of that id.
The badge is a row and the webhook is only how the row got there.
Curated: · Written: · Reviewed:
QA-25Web uses cookies and a bundle you can ship daily. Mobile uses a bearer header and a binary you cannot. What must the shared API accept so neither client papers over the other?(show answer)
I would start sharing one API across web and mobile from the screen the user sees and the row the database actually wrote.
One API for web and mobile means the contract cannot assume a cookie-only browser or a header-only app. Auth, errors, pagination, and versioning have to work for both callers or one half will paper over the other.
Concretely, accept cookie sessions and bearer tokens through the same authorization layer. Keep response shapes free of HTML and free of mobile-only envelopes. Version breaking changes so a store binary and a web bundle can lag independently.
The reason for that specificity is a failure I have seen: The API required a cookie the mobile client never stored, 2,160 logins failed with status 401, and the web team called it a mobile bug for 53 hours while both halves disagreed about where the session lived.
Same write, two callers.
| Caller | Credential | Shared API result |
|---|---|---|
| web bundle | cookie session | 200, JSON page slice |
| mobile binary | bearer header | 200, same JSON |
| mobile against cookie-only API | none sent | 401, blamed on the app |
I would not consider it settled without evidence: Run the same write from a cookie browser session and from a bearer mobile session and confirm status, body shape, and pagination tokens match.
Web and mobile are two clients of one contract, not two APIs that happen to share a hostname.
Curated: · Written: · Reviewed:
QA-26A product page is still slow after you added a cache. Which layer holds what, and what must the UI do after a write?(show answer)
The first question I ask about cache at CDN versus app versus database is which side owns the truth after a refresh.
The CDN holds bytes shared across users, the application cache holds session-keyed payloads, and the database cache holds query pages, so a miss at one layer is not a miss at the others. Putting personalized HTML at the CDN makes the next shopper see the previous shopper cart.
Concretely, put public product HTML and images on the CDN with a short shared lifetime. Keep session-keyed JSON in the application cache and leave query pages to the database buffer pool. After a write, purge or version the CDN object and delete the application key on the same mutation path.
The reason for that specificity is a failure I have seen: A homepage fragment with a 30 minute CDN lifetime still showed the previous price of 49 after a write to 39, and 18,000 shoppers bought at the stale price before the object was purged at minute 27.
Three layers, one product page.
| Layer | Stores | Lifetime | Personalized |
|---|---|---|---|
| CDN | public HTML and images | 120 s | no |
| application Redis | JSON by user id | 20 s | yes |
| database | query pages | engine | no |
I would not consider it settled without evidence: After a price write, fetch the page as an anonymous client and as a signed-in client and confirm both see the new price within the stated purge bound.
Cache the bytes that do not change per user, and purge on write.
Curated: · Written: · Reviewed:
QA-27A catalog paints 48 cards and each card then loads its seller. How do you fix that loop on the client and the API?(show answer)
With N plus one from a page of cards, a green local UI is where I start checking rather than stop.
A page of cards that then loads a relation per card turns one paint into dozens of browser fetches and the same number of server queries. The full-stack fix is one list payload that already contains the fields those cards render.
Concretely, count fetches in the network panel and queries in the request log on the same load. Change the list endpoint to join or batch the related rows. Stop the client from issuing a follow-up fetch per card.
The reason for that specificity is a failure I have seen: A grid of 48 product cards issued 49 HTTP calls and 49 queries, the page became interactive at 3.4 seconds, and a single batched payload of 48 rows dropped it to 2 queries and 280 milliseconds.
Forty-eight cards, two designs.
| Design | HTTP calls | Queries | Time to interactive |
|---|---|---|---|
| fetch per card | 49 | 49 | 3,400 ms |
| batched list payload | 1 | 2 | 280 ms |
I would not consider it settled without evidence: Load the catalog with an empty cache and assert at most 2 queries and at most 2 list fetches for a full first page.
One page of cards is one payload, not one fetch per card.
Curated: · Written: · Reviewed:
QA-28A dashboard paints 9 tiles from live joins and is slow. What changes on the write path and on the page?(show answer)
I would answer denormalized read model for a dashboard by following one click through the browser, the API, and the store.
A dashboard that joins live fact tables on every paint pays that join cost on every refresh the user triggers. A denormalized read model owned by the write path lets the UI read one row per tile while the same transaction or an event keeps that row honest.
Concretely, measure the live join cost for the tiles the first paint needs. Write a summary row in the same transaction as the source change or from the event the UI already waits on. Point the dashboard fetch at the summary and show a generated-at time so the user can see how fresh the tiles are.
The reason for that specificity is a failure I have seen: Nine tiles each joined 4 fact tables on every refresh, the first paint took 6.8 seconds, and a summary row of 9 counters brought the same paint to 90 milliseconds while 220 of 14,000 days drifted until a nightly check was added.
Nine tiles, live joins versus a summary.
| Read path | First paint | Drift over 14,000 days |
|---|---|---|
| 9 live joins of 4 tables | 6,800 ms | none, but slow |
| summary only in the API path | 90 ms | 220 days stale |
| summary in the transaction plus nightly check | 90 ms | 0 days stale |
I would not consider it settled without evidence: Refresh the dashboard after a source write and confirm the tiles match a hand sum of the facts and that the generated-at time moved.
Dashboard tiles read a summary the write path keeps honest.
Curated: · Written: · Reviewed:
QA-29How do you take payment and decrement stock so the confirmation screen and the warehouse agree?(show answer)
My approach to checkout write spanning UI and inventory separates what the client may believe from what the server will persist.
A card charge and an inventory row cannot share one database transaction with the processor. Checkout is a state machine of reservation, payment intent, capture, and compensation, and the confirmation screen must follow that durable order, not a wallet callback.
Concretely, reserve stock with a time to live, create an idempotent payment intent on the server, and write both through an outbox. Poll the order id from the confirmation page. Compensate stock if capture fails, and never treat a client wallet callback as the stock write.
The reason for that specificity is a failure I have seen: The client showed paid after a wallet callback, the stock call then timed out at 8 seconds, and 410 of 12,600 checkouts that day sold units the warehouse did not have.
One day of checkouts.
| Owner of success | Paid without stock | Warehouse shorts |
|---|
| wallet callback then a later stock POST | 410 | 410 |
| reserve, payment intent, outbox, UI polls | 0 | 0 |
I would not consider it settled without evidence: Timeout capture after a reservation and confirm the UI stays on pending, stock is released, and no paid confirmation is shown.
Confirm checkout from the durable order, not from the wallet callback.
Curated: · Written: · Reviewed:
QA-30A write succeeds but a refresh still shows the old value. What should the page do?(show answer)
For eventual consistency shown in the UI, I would name the contract both halves must keep before choosing a library.
After a write, replicas and caches can still serve the previous row, so a refresh that hits a lagging reader looks like a lost save. The UI must read from the writer, carry the optimistic value until a matching version returns, or tell the user the save is still propagating.
Concretely, return a version from the mutation. Keep that version in client state. Retry the read against the primary or wait until the returned version is at least the written one before painting the saved value as confirmed.
The reason for that specificity is a failure I have seen: A profile save returned status 200, a refresh 400 milliseconds later hit a replica 1.6 seconds behind, and 7 percent of editors re-submitted and created duplicate bios.
Save, then refresh.
| Next read | Version seen | Screen |
|---|---|---|
| replica 1.6 s behind | 11, written was 12 | old bio, user saves again |
| primary | 12 | confirmed |
| replica after waiting for 12 | 12 | confirmed |
I would not consider it settled without evidence: Save, then read from a replica you delay, and confirm the page keeps the optimistic fields until the versions match.
Paint a save as confirmed only when the read version has caught up.
Curated: · Written: · Reviewed:
QA-31The homepage is personalized and you also want a high CDN hit ratio. How do you split the page?(show answer)
I would size personalization versus cacheability against a real page load and a real write, not a mocked fetch.
Bytes that vary per user cannot sit in a shared CDN object without leaking one shopper to another. Split the shell that everyone can share from the fragment that must be fetched with a cookie, and cache each at the layer that matches its variance.
Concretely, cache the public shell and images at the CDN. Fetch the personalized rail with a private request after first paint. Never store session cookies or authorization headers on a shared edge key.
The reason for that specificity is a failure I have seen: A shared CDN key cached the first shopper HTML including a recently viewed rail of 6 SKUs, and 92,000 later visitors saw that shopper products for 15 minutes.
Homepage split.
| Object | Cache | Hit ratio | Leak |
|---|---|---|---|
| whole HTML including rail | shared CDN | 94% | 92,000 visitors |
| public shell | shared CDN | 93% | none |
| personalized rail | private, 20 s | n/a | none |
I would not consider it settled without evidence: Fetch the homepage twice with two sessions and confirm the shared object has no user fields while the private fragment differs.
Share the public shell and fetch the personal fragment after paint.
Curated: · Written: · Reviewed:
QA-32Search results look wrong even though the documents are in the index. What do you check across the stack?(show answer)
The part of full-text search ranking the user sees interviewers probe is the mismatch after deploy, not the happy render.
The ranking the user sees is the query the client sent, the analyzer the index used, and the fields the API chose to return. A mismatch in tokenization, filters, or the highlight payload makes a relevant row look missing even when it is stored.
Concretely, log the exact query string the typeahead submitted. Replay it against the index with the same analyzer. Return the score and the matched field to the UI so a support tool can show why a hit ranked where it did.
The reason for that specificity is a failure I have seen: Client-side hyphen stripping turned sku A-19 into A19 while the index tokenized on hyphens, so 1,240 product searches in a week returned zero hits for in-stock SKUs.
Query A-19.
| Analyzer pair | Query sent | Top hit | Weekly zero-result |
|---|---|---|---|
| client strips hyphens, index splits on them | A19 | none | 1,240 |
| both keep the hyphen | A-19 | sku A-19 | 40 |
I would not consider it settled without evidence: Capture one live search from the network tab and confirm the same tokens, filters, and top 10 ids from the index API.
Rank what the user typed with the same analyzer the index uses.
Curated: · Written: · Reviewed:
QA-33The growth dashboard and the billing table disagree. Which number does the product UI show?(show answer)
I would anchor analytics events versus source of truth in a trace that starts at the click and ends at the query.
Analytics events are a sampled, delayed, and lossy trail of what the client fired. Money, inventory, and entitlement come from the transactional store, so a screen that treats an analytics count as the billed quantity will disagree with finance the next morning.
Concretely, fire analytics from the client for funnels only. Read billed totals from the API that reads the ledger. If a marketing tile needs both, label the analytics figure as approximate and link the ledger figure as the one that can be invoiced.
The reason for that specificity is a failure I have seen: A usage meter painted 1.8 million events from the warehouse while invoices used 1.41 million billable rows, and 64 customers disputed the gap of 390,000 shown on the in-app usage page.
One tenant, three counters.
| Source | Count | Shown as payable |
|---|---|---|
| client events in the warehouse | 1,800,000 | wrongly, yes |
| billing ledger | 1,410,000 | yes |
| warehouse minus bot traffic | 1,520,000 | no |
I would not consider it settled without evidence: For one tenant, compare the in-app total with the ledger sum and the warehouse count and show only the ledger sum as payable.
Show money from the ledger, not from the event pipe.
Curated: · Written: · Reviewed:
QA-34A support screen needs to show a customer. What must never be in the bundle, the logs, or local storage?(show answer)
What separates a strong answer on PII that must not sit in the browser is knowing which half can lie to the other.
Government ids, full payment numbers, and authentication secrets are server-side facts. Putting them in a SPA store, a client log, or a CDN-cached payload copies them onto every laptop that opened the tab.
Concretely, keep raw PII in the database and in access-controlled APIs. Return masked values to the browser. Never write those fields to localStorage, to client error beacons, or to HTML that a CDN can cache.
The reason for that specificity is a failure I have seen: A support page put full national ids into client memory and a session replay script captured 11,400 of them, and the ids remained in browser disk caches for 7 days after the tab closed.
What the support tab held.
| Location | Raw national ids | Masked form |
|---|---|---|
| API response | 11,400 | no |
| session replay | 11,400 | no |
| API after the change | 0 | yes, last 4 |
I would not consider it settled without evidence: Search the shipped bundle, local storage, and a recorded session for the raw identifier and confirm only a masked form is present.
The browser gets a mask and the server keeps the identifier.
Curated: · Written: · Reviewed:
QA-35A user asks to be forgotten. What do you delete besides the database row?(show answer)
I would treat right-to-erasure across client caches as something that must survive a stale bundle talking to a new API.
Erasure is not done when the row is gone if the SPA, the CDN, search, backups, and analytics still hold the same person. The client must drop persisted state and the server must expire every copy the UI could still fetch.
Concretely, delete or anonymize the source row. Purge CDN and search documents that contain the person. Return not-found to every identifier the app still holds, and on the next load clear localStorage, IndexedDB, and service worker caches for that origin.
The reason for that specificity is a failure I have seen: Account deletion removed the users row, but a service worker still served a cached profile payload for 21 days, and 3 former customers remained visible in a teammates picker from a 14 day search index lag.
Copies after delete.
| Store | Days until gone | Visible to teammates |
|---|---|---|
| users table | 0 | no |
| service worker /me | 21 | yes, to the old session |
| search index | 14 | yes, 3 people |
I would not consider it settled without evidence: After erasure, load the app logged out and logged in as a teammate and confirm the person is absent from API, search, and every client store.
Forget the person on every cache the UI can still read.
Curated: · Written: · Reviewed:
QA-36A virtualized inbox paints 20 rows but the API returns 500. How do you size both halves?(show answer)
Before changing either half for virtualized lists versus server page size, I would write down how both sides will be checked.
Virtualization only skips DOM work. If the API still ships 500 rows the browser still parses and holds them, so page size on the server and window size on the client have to match the rows the user can actually see plus a small prefetch.
Concretely, return a cursor page of 30 rows near the viewport height. Have the list fetch the next cursor when the user is a few rows from the end. Do not download the whole mailbox to feed a window of 20.
The reason for that specificity is a failure I have seen: The inbox requested 500 messages of 8 KB each, first paint downloaded 4.0 MB, and a 30 row page of 240 KB with a cursor dropped time to interactive from 5.1 seconds to 700 milliseconds.
Inbox first paint.
| Page size | Bytes | Time to interactive | DOM rows |
|---|---|---|---|
| 500 | 4.0 MB | 5,100 ms | 20 virtualized |
| 30 with cursor | 240 KB | 700 ms | 20 virtualized |
I would not consider it settled without evidence: Throttle the network and confirm the first list response is near one viewport and that scrolling fetches the next cursor rather than a second giant page.
Page on the server to the window the client actually paints.
Curated: · Written: · Reviewed:
QA-37Product images are huge on mobile. What does the client request and what does the pipeline produce?(show answer)
I would start image pipeline through the CDN from the screen the user sees and the row the database actually wrote.
The browser should request a sized, formatted variant, and the origin or CDN should produce that variant once and cache it. Shipping a 4,000 pixel original and shrinking it in CSS still costs the download.
Concretely, encode variants at write time or on first miss. Put width, format, and quality in the URL the image tag requests. Serve WebP or AVIF when the client accepts those types and cache each variant at the CDN.
The reason for that specificity is a failure I have seen: A gallery requested 12 originals of 3.2 MB each on a phone, the CSS scaled them to 180 pixels, and switching to 360 pixel WebP variants of 42 KB cut gallery bytes from 38.4 MB to 504 KB.
Twelve product images.
| Request | Each | Gallery total | Displayed width |
|---|---|---|---|
| 4,000 px original | 3.2 MB | 38.4 MB | 180 px |
| 360 px WebP | 42 KB | 504 KB | 180 px |
I would not consider it settled without evidence: On a mobile viewport, record the image URLs and confirm each response width matches the displayed size and that a second view is served from the CDN cache.
Request the variant you display and cache that variant at the edge.
Curated: · Written: · Reviewed:
QA-38A user clicks Export and the tab spins. How should that click reach the database?(show answer)
The first question I ask about CSV export from the UI hitting the database is which side owns the truth after a refresh.
A CSV of hundreds of thousands of rows cannot ride the same request as the page that launched it. The click should create a job, the UI should poll or subscribe, and the exporter should read in cursors so the database is not locked by one browser request.
Concretely, post an export request that enqueues work and returns a job id. Stream or page the query on a worker. Let the UI download from object storage when the job is ready, and cap live exports so they cannot scan the primary during peak.
The reason for that specificity is a failure I have seen: Export ran inside the click handler against 620,000 order rows, held a read for 94 seconds, and blocked checkout queries so the 95th percentile of checkout rose from 180 milliseconds to 2.4 seconds for 11 minutes.
Export of 620,000 orders.
| Path | Click returns | Checkout p95 during export |
|---|---|---|
| query in the request | 94 s | 2,400 ms for 11 min |
| job then object download | 180 ms | 180 ms |
I would not consider it settled without evidence: Start an export of production size and confirm the original request returns in under 500 milliseconds and that checkout latency stays within its budget while the worker runs.
The click starts a job and the file is a later download.
Curated: · Written: · Reviewed:
QA-39The user starts a video upload that must be transcoded. How does the page wait without lying?(show answer)
With background job the UI waits on, a green local UI is where I start checking rather than stop.
Work that takes longer than a request timeout cannot be hidden behind a spinner on the original POST. The UI needs a job identifier, a status it can poll or push, and a terminal state that matches the row the worker wrote.
Concretely, create the job row first and return its id. Show queued, running, failed, or ready from that row. Subscribe with a socket or poll every 3 seconds, and do not paint complete until the worker has written the output object.
The reason for that specificity is a failure I have seen: The upload POST waited 58 seconds for transcode, the load balancer cut it at 60, the client showed an error while the worker later succeeded, and 1,900 users uploaded the same video twice.
Transcode wait.
| Wait strategy | Client result | Duplicate uploads |
|---|---|---|
| block the POST for 58 s | error at 60 s, worker succeeds | 1,900 |
| poll job id every 3 s | ready when the object exists | 0 |
I would not consider it settled without evidence: Kill the tab after enqueue and reopen the page, then confirm the job status matches the worker and the object store.
Wait on the job row, not on the original request.
Curated: · Written: · Reviewed:
QA-40A comment is posted and 40 people should see it. What does the writer do and what do the open pages do?(show answer)
I would answer fan-out notification from a write by following one click through the browser, the API, and the store.
The write stores the comment once. Fan-out tells every interested session without turning the writer into extra requests from the authoring browser, and open pages receive an event or poll a feed rather than each querying the whole recipient list.
Concretely, commit the comment. Publish one event with the thread id. Let a notifier expand recipients and push, and on each open client append the comment when the event arrives or catch up from a feed fetch if the socket is down.
The reason for that specificity is a failure I have seen: The authoring client looped 40 recipient POSTs after submit, 6 timed out, those 6 never saw the comment, and the write path 95th percentile rose from 90 milliseconds to 1.7 seconds.
Forty recipients.
| Fan-out | Writes | Missed inboxes | Submit p95 |
|---|---|---|---|
| 40 POSTs from the browser | 40 | 6 | 1,700 ms |
| one event, notifier pushes | 1 | 0 | 90 ms |
I would not consider it settled without evidence: Post one comment with 40 recipients and confirm one write, one event, and 40 clients that either got the push or caught up from the feed.
One write, one event, many listeners.
Curated: · Written: · Reviewed:
QA-41A user saves a form and the next page is empty. How do you handle replica lag in the UI?(show answer)
My approach to read replica lag visible on the page separates what the client may believe from what the server will persist.
A read-your-writes miss after a save is usually a replica that has not applied the commit. The page that must show the new row has to read from the primary or wait on a position the replica has caught up to.
Concretely, send the mutation to the primary. For the next read in that session, use the primary or pass a last-write token the replica must satisfy. If the token is not met, show the optimistic form data and a refreshing state rather than an empty page.
The reason for that specificity is a failure I have seen: Settings saved on the primary, the following GET hit a replica 2.8 seconds behind, the page rendered empty defaults, and 310 users clicked save again and overwrote later edits from another tab.
Save then open the next view.
| Next GET | Lag | Screen | Double saves |
|---|---|---|---|
| nearest replica | 2.8 s | empty defaults | 310 |
| primary | 0 | saved fields | 0 |
| replica after token | 0 after wait | saved fields | 0 |
I would not consider it settled without evidence: Force a replica delay in staging, save, and confirm the next view still shows the saved fields.
Read your writes from a caught-up source, or keep the optimistic copy on screen.
Curated: · Written: · Reviewed:
QA-42Two tenants use the same JavaScript app. How does a row from tenant A become visible to tenant B?(show answer)
For multi-tenant leak through a shared SPA, I would name the contract both halves must keep before choosing a library.
A shared SPA reuses in-memory stores, HTTP caches, and service workers across route changes. If the client keeps tenant A data after a switch, or if the API keys only on id without tenant, tenant B will paint tenant A rows.
Concretely, scope every API query by the session tenant on the server. Key client caches by tenant id. On tenant switch, drop memory, IndexedDB, and the service worker cache before fetching, and never trust a tenant id the client sent if it disagrees with the session.
The reason for that specificity is a failure I have seen: Switching from tenant A to tenant B reused a query cache of 86 invoices, 12 of those invoices belonged to tenant A, and a CDN cache keyed only on path leaked an invoice list of 200 rows to 4 other tenants for 8 minutes.
Switch A to B in one tab.
| Cache key | Tenant A invoices on B | Other tenants hit |
|---|---|---|
| path only, memory reused | 12 of 86 | 4 tenants for 8 min |
| session tenant plus cache bust | 0 | 0 |
I would not consider it settled without evidence: Switch tenants in one tab and confirm the network requests carry the new session and that no prior tenant rows remain in memory or on screen.
Tenant is a server check and a client cache key, not a URL segment.
Curated: · Written: · Reviewed:
QA-43The app works on a train. How do you queue writes and merge them when the network returns?(show answer)
I would size offline queue then sync against a real page load and a real write, not a mocked fetch.
Offline UI can record intent, but the server still assigns the real ids, rejects conflicts, and is the source after sync. A queue that later posts without idempotency keys will double-create when the first attempt actually succeeded.
Concretely, persist each intended write with a client key. Replay in order when online. Let the server upsert by that key, paint pending locally, then replace with the server row, and surface conflicts instead of last-write-wins on the same record.
The reason for that specificity is a failure I have seen: A field app queued 73 inspections, the first flush retried after a 30 second timeout, and 19 inspections were stored twice because the client key was not sent.
Seventy-three offline inspections.
| Replay | Server rows | Duplicates |
|---|---|---|
| POST without a client key, retry at 30 s | 92 | 19 |
| upsert by client key | 73 | 0 |
I would not consider it settled without evidence: Go offline, create a record, kill the tab, come back online, and confirm one server row and a client view that matches it.
Queue intent with a key and let the server dedupe the replay.
Curated: · Written: · Reviewed:
QA-44A cart shows a total that does not match the charge. Who computes the money?(show answer)
The part of derived totals client versus server interviewers probe is the mismatch after deploy, not the happy render.
Display math in the browser is a preview. The charge, tax, and discount the customer pays are computed on the server from the same rules that write the ledger, and if the UI total is authoritative a patched client can pay the wrong amount.
Concretely, compute a preview in the client for snappy updates. On checkout, ignore that preview and persist the server total. Re-render the confirmation from the server figures, and block pay if the preview drifted, while still charging the server amount.
The reason for that specificity is a failure I have seen: The client applied a 15 percent coupon twice, showed 42.50, the server charged 85.00, and 2,200 orders that week generated support tickets because the confirmation then switched numbers.
Coupon applied twice in the browser.
| Total owner | Screen | Charged | Tickets in a week |
|---|---|---|---|
| client preview | 42.50 | 85.00 | 2,200 |
| server ledger, UI re-renders | 85.00 | 85.00 | 0 |
I would not consider it settled without evidence: Mutate the client total in a test and confirm the charged amount still matches the server ledger.
Preview in the client and charge what the server computed.
Curated: · Written: · Reviewed:
QA-45A meeting at nine in the morning in Paris looks like nine in the morning in New York after a refresh. What is stored and what is formatted?(show answer)
I would anchor timezone display versus stored UTC in a trace that starts at the click and ends at the query.
Instants are stored in UTC. The browser formats them into the viewer zone, and storing a local wall time without a zone, or formatting UTC as if it were local, shifts every event when the page reloads in another zone.
Concretely, persist UTC instants from the API. Send an explicit IANA zone when the event has a local wall meaning. Format with the viewer zone for display, and never parse a zoneless string as the browser local zone on one half and UTC on the other.
The reason for that specificity is a failure I have seen: A nine in the morning Paris kickoff was stored as a zoneless wall clock of 9, New York clients rendered nine in the morning Eastern, and 480 attendees joined 6 hours late.
Paris kickoff opened in New York.
| Stored value | New York screen | Joined late |
|---|---|---|
| zoneless 9 | 9 Eastern | 480 people, 6 h |
| 08.00 UTC plus Europe/Paris | 3 Eastern | 0 |
I would not consider it settled without evidence: Create an event in one zone, open it in another, and confirm the instant is identical and the displayed clock time matches the viewer zone.
Store UTC and format for the person looking at the page.
Curated: · Written: · Reviewed:
QA-46Tax of 8.875 percent on 3 line items disagrees between the UI and the invoice. How do you round?(show answer)
What separates a strong answer on money rounding on both halves is knowing which half can lie to the other.
Rounding mid-stream in the client and again on the server produces pennies that do not match. Pick a currency minor unit, a rounding mode, and a single place that rounds, then have the UI display that already rounded amount.
Concretely, keep integer minor units in the API. Apply tax and discounts in a documented order on the server. Return the rounded line and grand totals, display those integers, and do not round each line in JavaScript and then again in SQL.
The reason for that specificity is a failure I have seen: The client rounded each of 3 lines to 2 decimals then summed, the server summed then rounded, and a 0.01 gap on 16,800 invoices that month failed a card settlement match.
Three 12.99 lines, tax 8.875 percent, half-up to cents on the grand total.
| Who rounds | Line cents | Tax cents | Grand cents |
|---|
| each line tax half-up then sum | 1299, 1299, 1299 | 345 | 4242 |
| sum 3897 then tax half-up on the server | 1299, 1299, 1299 | 346 | 4243 |
| server grand 4243, UI displays 4243 | 1299, 1299, 1299 | 346 | 4243 |
I would not consider it settled without evidence: Price a cart with repeating tax decimals and confirm the UI, the API response, and the ledger row share one integer total.
Round once, on the server, in minor units.
Curated: · Written: · Reviewed:
QA-47The product page says In stock. The user pays and it is not. What reservation does the click create?(show answer)
I would treat inventory reservation from the product page as something that must survive a stale bundle talking to a new API.
A stock badge is a cached read. A sale needs a reservation tied to the checkout session with a time to live, because without that hold two browsers can both paint In stock and both pay for the last unit.
Concretely, when the user starts checkout, create a reservation row for the quantity with a 10 minute expiry. Decrement available for others immediately. Convert the reservation to a sale on payment, or release it on expiry, and have the product page read available minus holds.
The reason for that specificity is a failure I have seen: Two clients loaded a remaining count of 1 from a 20 second cache, both paid, and the second order of 640 that week was later cancelled for missing stock.
Last unit, two browsers.
| Signal | Both see In stock | Both charged | Later cancelled |
|---|---|---|---|
| badge from 20 s cache | yes | yes | 640 |
| reservation at checkout start | first only | first only | 0 |
I would not consider it settled without evidence: Open two checkouts on the last unit and confirm only one reservation succeeds and the other page shows unavailable before pay.
Hold stock at checkout start, not at the badge fetch.
Curated: · Written: · Reviewed:
QA-48Faceted search is slow and some filters return empty while data exists. What must the query and the index share?(show answer)
Before changing either half for search filters that must match indexes, I would write down how both sides will be checked.
Every filter the UI offers has to be a field the index can apply without a scan, and the same field the documents were ingested with. A client-side filter on a page of 20 hits hides rows that the index would have returned if it had been told.
Concretely, map each facet control to an indexed field. Send all selected filters in one search query. Do not fetch a broad page and filter in the browser, and when a new facet ships, add the field to the index before enabling the control.
The reason for that specificity is a failure I have seen: Color was filtered in the client on 20 hits from a 40,000 document index, 92 percent of red items never appeared, and adding a color term to the query plus a keyword index cut empty-result rate from 31 percent to 2 percent.
Color facet on 40,000 products.
| Where the filter runs | Red items shown | Empty-result rate |
|---|---|---|
| client on 20 hits | 8% of red | 31% |
| index term on color | 100% of red | 2% |
I would not consider it settled without evidence: Select two facets, capture the query, and confirm the index plan uses those fields and the hit count matches a database count on the same predicates.
Do not offer a filter the index cannot apply.
Curated: · Written: · Reviewed:
QA-49Typeahead hammers the API on every keystroke. How do you bound cost on both halves?(show answer)
I would start autocomplete cost across the stack from the screen the user sees and the row the database actually wrote.
Each keystroke can become a search against the primary if the client does not debounce and the API does not cache or prefix-index. At traffic, autocomplete is often the most expensive read path the UI owns.
Concretely, debounce input by 150 milliseconds and require 2 characters. Hit a prefix index or an edge cache keyed on the normalized query. Cap the result list at 8, and cancel in-flight fetches when the next key arrives so the server is not scoring stale prefixes.
The reason for that specificity is a failure I have seen: Un-debounced typeahead issued 11 requests per second per user, a 2,400 user peak produced 26,400 queries per second against the primary, and adding debounce, a prefix index, and a 30 second cache dropped it to 900 queries per second.
Peak of 2,400 typists.
| Client and index | QPS at peak | Suggestions |
|---|---|---|
| every key, primary scan | 26,400 | stale in-flight |
| 150 ms debounce, 2 chars, prefix cache 30 s | 900 | latest prefix |
I would not consider it settled without evidence: Type a 10 letter query on a throttled CPU and confirm request count, server query rate, and that the painted suggestions match the latest prefix.
Debounce the keys, index the prefixes, and cache the short lists.
Curated: · Written: · Reviewed:
QA-50A manager opens Insights and the request times out. What does the page load instead?(show answer)
The first question I ask about a report too heavy for a page load is which side owns the truth after a refresh.
An aggregate over months of facts cannot be the homepage fetch. The page should load a precomputed report or start a job, and it should never run the heavy query on the primary in the request that paints the layout.
Concretely, precompute the report on a schedule into a summary table or file. Let the Insights page read that snapshot and show its as-of time. If the user asks for an ad hoc range, enqueue a job and keep the page interactive until the result is ready.
The reason for that specificity is a failure I have seen: Insights issued a 14 table join over 18 months of 90 million rows on page load, the request hit the 30 second gateway timeout, and moving to a nightly summary of 12,000 rows made first paint 400 milliseconds.
Insights first open.
| Query | Rows touched | First paint |
|---|---|---|
| 14 table join, 18 months | 90,000,000 | timeout at 30 s |
| nightly summary | 12,000 | 400 ms |
| ad hoc range as a job | 90,000,000 off the page | page stays interactive |
I would not consider it settled without evidence: Open Insights during peak writes and confirm the page query hits the summary, not the fact tables, and that an ad hoc range does not block the tab.
Paint a snapshot and compute the heavy range in a job.
Curated: · Written: · Reviewed:
QA-51How do you ship a frontend change that depends on a new API field?(show answer)
With deploying frontend and API together, a green local UI is where I start checking rather than stop.
Open tabs keep an old bundle after you ship, so the API must expand first and stay compatible until those tabs drain. Lockstep flipping both artifacts does not drain already-open clients.
Concretely, add the new field on the API while the old field still works. Ship the bundle that reads the new field. Measure remaining old sessions, then remove the old field. Keep a contract check that the currently live bundle still parses the candidate API.
The reason for that specificity is a failure I have seen: An API shipped 11 minutes before the matching bundle, and 18 percent of sessions still on the old client received 400 responses on checkout until the web deploy finished.
Staggered ship versus a paired ship.
| Stage | API version | Bundle version | Checkout success |
|---|---|---|---|
| staggered by 11 min | 47 | 46 | 82% |
| expand API, then bundle, then contract | 47 compatible with 46 | 46 then 47 | 99.7% |
I would not consider it settled without evidence: Run the live production bundle against the candidate API in the pipeline and block the release if any contract check fails.
Ship the pair, not two calendars.
Curated: · Written: · Reviewed:
QA-52Which environment variables may go into the browser bundle, and which must stay on the server?(show answer)
I would answer env vars in the browser versus the server by following one click through the browser, the API, and the store.
Anything inlined into client JavaScript is public because anyone can download that file. Server secrets stay on the process that handles requests and must never be copied into that public map.
Concretely, split configuration into a public client set and a private server set at build time. Inject only the public set into the bundle. Read private values from the runtime environment on the API and reject a build that prints a secret into client assets.
The reason for that specificity is a failure I have seen: A database password named with a public prefix was inlined into 4.2 MB of JavaScript, and scanners found it 6 hours after the release on 3 mirrored CDNs.
Where each value may live.
| Variable | Safe in bundle | Safe on API process |
|---|---|---|
| public marketing site url | yes | yes |
| feature flag CDN host | yes | yes |
| database password | no | yes |
| payment private key | no | yes |
I would not consider it settled without evidence: Scan the built client assets for private key names and fail the pipeline if any match.
If the browser can read it, treat it as public.
Curated: · Written: · Reviewed:
QA-53A pull request preview is up. What must it be forbidden to call?(show answer)
My approach to preview deploys hitting production APIs separates what the client may believe from what the server will persist.
A preview site that talks to production data is a write path from unreviewed code. Previews must bind to preview backends and preview stores so a pull request cannot mutate live users.
Concretely, give every preview its own API origin and its own database clone. Resolve that origin from the preview host rather than from a production default. Block outbound production hostnames from preview runtimes.
The reason for that specificity is a failure I have seen: A preview checkout form posted to the live payments API and created 27 real charges against 19 production customers before anyone noticed.
What one preview actually called.
| Destination | Intended | Observed | Writes |
|---|---|---|---|
| preview API | yes | no | 0 |
| production payments | no | yes | 27 charges |
| preview database | yes | no | 0 |
I would not consider it settled without evidence: Fetch the preview page, capture every API host it calls, and confirm none of those hosts is a production origin.
A preview that can write production is already production.
Curated: · Written: · Reviewed:
QA-54How do you change a response field without breaking users still on the previous bundle?(show answer)
For migrations that break an old bundle, I would name the contract both halves must keep before choosing a library.
A schema or payload change that the current live bundle cannot parse will fail for everyone still holding that bundle. Expand the API and the schema first, wait until old clients drain, then contract.
Concretely, add new columns and fields as optional. Keep the old field populated until the oldest allowed bundle no longer reads it. Only then drop the old shape.
The reason for that specificity is a failure I have seen: A required renamed JSON field shipped at 16 percent canary, and 84 percent of sessions on the previous bundle could not load the cart for 22 minutes.
Cart loads during a rename.
| Phase | Old field present | New field present | Old bundle cart loads |
|---|---|---|---|
| expand | yes | yes | 99.8% |
| canary rename only | no | yes | 16% of traffic, rest fail |
| after drain and contract | no | yes | 99.8% |
I would not consider it settled without evidence: Replay traffic captured from the previous bundle against the migrated API and require a clean parse.
Expand, drain, then contract.
Curated: · Written: · Reviewed:
QA-55You cut API traffic to green. Why can the site still break on static files?(show answer)
I would size blue-green with sticky static assets against a real page load and a real write, not a mocked fetch.
Blue-green at the API edge does not move hashed files already cached on a CDN. Users can sit on green HTML that still names blue hashes, or the reverse, unless asset names and cache keys move with the color.
Concretely, publish hashed assets to both colors under a shared immutable prefix. Point HTML only at hashes that exist on both. Drain the old color only after HTML and assets agree.
The reason for that specificity is a failure I have seen: Green HTML referenced a chunk that existed only on blue storage, and 31 percent of first loads after the cutover returned 404 for the main script for 9 minutes.
First load after the color flip.
| Object | Color that holds it | Status after flip |
|---|---|---|
| index.html | green | 200 |
| main.8f2a.js named by HTML | blue only | 404 |
| share of first loads | n/a | 31% blank |
I would not consider it settled without evidence: After the color flip, request the new HTML and every script it names, and require all of those objects to return 200 from the live edge.
The HTML and the hashes have to change together.
Curated: · Written: · Reviewed:
QA-56A refund flow is hurting users. How do you turn it off without waiting on a deploy?(show answer)
The part of feature-flag kill switch on both sides interviewers probe is the mismatch after deploy, not the happy render.
A flag that exists only in the client cannot stop a damaging server path, and a flag that exists only on the server cannot hide a broken screen. The same named switch must gate the UI entry and the API handler.
Concretely, resolve the flag on the server for every mutating request. Hide the matching control in the client from that same flag. Keep a default-off remote kill that does not require a deploy.
The reason for that specificity is a failure I have seen: The team hid a refund button in the client while the refund endpoint stayed live, and a leftover admin script issued 387 refunds in 8 minutes.
Kill on one half versus both.
| Switch location | Button visible | Endpoint accepts POST | Refunds in 8 min |
|---|---|---|---|
| client only | no | yes | 387 |
| server only | yes | no | 0, UI errors |
| both | no | no | 0 |
I would not consider it settled without evidence: Flip the kill switch in staging and confirm the control disappears and the endpoint returns a documented disabled status.
One named switch must cover the click and the write.
Curated: · Written: · Reviewed:
QA-57The new API is bad. Can you roll it back while leaving the new SPA in place?(show answer)
I would anchor rolling back SPA and API independently in a trace that starts at the click and ends at the query.
Rolling back only the API can leave a new bundle calling methods that no longer exist, and rolling back only the bundle can leave an old screen against a contracted schema. Independent rollback is safe only while both versions still speak the overlap.
Concretely, keep the previous API version serving until the previous bundle is the live HTML again. Restore artifacts as a pair when the overlap is gone. Record which bundle and which API were live together.
The reason for that specificity is a failure I have seen: Ops rolled the API back 2 versions and left the new SPA in place, which called a removed price endpoint and showed zero for 14,600 sessions.
Price shown after a one-sided rollback.
| API | Bundle | Price endpoint | Sessions seeing zero |
|---|---|---|---|
| v14 live | v14 | present | 0 |
| v12 restored | v14 | removed | 14,600 |
| v12 restored | v12 restored | present | 0 |
I would not consider it settled without evidence: After any rollback, load the live HTML and exercise the main journeys against the live API until they succeed.
Rollback is a pair when the overlap is gone.
Curated: · Written: · Reviewed:
QA-58A user signs in, closes the tab, and comes back the next day still signed in, but a token copied off that device must not work anywhere else. How does the full stack hold both of those?(show answer)
What separates a strong answer on session persistence versus stolen tokens is knowing which half can lie to the other.
Persistence and theft resistance pull against each other: the credential that survives a tab close is exactly the one an attacker wants to copy. Put the long life in a server-side session or a rotating refresh credential the page cannot read, and let the browser hold only a short-lived access token that expires before a leak becomes an incident.
Concretely, keep the session in a server store — a Redis key or a sessions table — with the cookie carrying only an opaque id marked HttpOnly, Secure and SameSite, and hand the page a 15-minute access token while the refresh credential rotates on every use. Treat a replayed refresh token as theft: revoke the whole token family and force both devices back to sign-in.
The reason for that specificity is a failure I have seen: A team issued a 90-day JWT into localStorage so users stayed signed in; a script injected through a third-party chat widget read it, and the same token drove 610 admin calls from another country over 12 days until the signing key was rotated and every session in the product died at once.
One family, one replay, both devices signed out.
| Time | Event | Server state |
|---|---|---|
| 09:00 | sign-in on laptop | family f1 created, refresh R1, access A1 (15 min) |
| 09:20 | tab reopened | A1 expired, R1 consumed, R2 issued |
| 09:22 | attacker replays stolen R1 | reuse detected on f1, family revoked |
| 09:22 | laptop calls API with A1 | 401, forced to sign in again |
I would not consider it settled without evidence: Sign in, close the tab, reopen after a day and confirm the page can read no long-lived secret, then replay yesterday's refresh token and confirm both the laptop and the attacker's device land on 401.
Long life belongs to the server store; the browser only borrows fifteen minutes at a time.
Curated: · Written: · Reviewed:
QA-59The load balancer says the backend for frontend is healthy, but the first page is blank. What should the check have done?(show answer)
I would treat health checks for the BFF as something that must survive a stale bundle talking to a new API.
A backend for frontend can return 200 while a required downstream is dead, which makes the web look healthy and the product unusable. The check that gates traffic must exercise the dependencies the first screen needs.
Concretely, probe the BFF with a request that talks to the session store and the primary API. Fail the check if either dependency misses its budget. Keep a shallow liveness probe separate so the process can be restarted without waiting on those dependencies.
The reason for that specificity is a failure I have seen: A load balancer used a process ping that stayed green while the session store was down, and 72 percent of first-page renders failed for 13 minutes with no pool marked unhealthy.
Ping versus a ready probe.
| Check | Session store down | Balancer state | First-page success |
|---|---|---|---|
| process ping | still 200 | in rotation | 28% |
| ready probe through store | 503 | drained | n/a, traffic shifted |
| after drain to healthy pool | store up | in rotation | 99.4% |
I would not consider it settled without evidence: Kill the session store in a drill and confirm the readiness check fails and the balancer stops sending users.
Ready means the first screen can still be built.
Curated: · Written: · Reviewed:
QA-60Checkout is slow. How do you prove whether the wait is in the browser, the API, or the query?(show answer)
Before changing either half for a trace from click to SQL, I would write down how both sides will be checked.
A fullstack incident is a path, not two dashboards. One trace id must follow the click through the browser, the BFF, the API, and the query so the slow span is not guessed.
Concretely, stamp a trace id on the click and send it on every fetch. Continue that id through the BFF into the API and into the database span. Join client timing and server spans in one view.
The reason for that specificity is a failure I have seen: Checkout p99 sat at 4.8 seconds, web blamed the API, API blamed the database, and 3 teams spent 90 minutes before a joined trace showed 3.1 seconds in a client retry loop.
Where 4.8 seconds actually went.
| Span | Time | Owner |
|---|---|---|
| click to first fetch | 40 ms | client |
| client retry loop | 3,100 ms | client |
| BFF plus API | 420 ms | server |
| SQL | 190 ms | database |
| remaining paint | 1,050 ms | client |
I would not consider it settled without evidence: Complete one purchase in staging and open a single trace that includes the click, the HTTP spans, and the SQL span.
If you cannot join the click to the query, you are guessing.
Curated: · Written: · Reviewed:
QA-61API graphs look fine. Users still say the page is slow. What else do you measure?(show answer)
I would start RUM plus backend APM from the screen the user sees and the row the database actually wrote.
Real user monitoring shows what devices actually felt, while backend APM shows what the process spent. Neither half explains a slow page by itself when the wait is on the other side.
Concretely, correlate RUM page views to backend traces by shared ids. Compare p75 and p95 on both sides for the same journey. Alert when the user-side budget is burned even if the API still looks cheap.
The reason for that specificity is a failure I have seen: APM p95 of 180 ms looked fine while RUM p95 of 6.4 seconds on mid-range phones came from a 2.1 MB script the API graphs never showed.
Same route, two instruments.
| Signal | p75 | p95 |
|---|---|---|
| API APM | 95 ms | 180 ms |
| RUM on desktop | 1.1 s | 1.8 s |
| RUM on mid-range phone | 3.9 s | 6.4 s |
I would not consider it settled without evidence: For the same route, place RUM p95 next to API p95 and explain any gap larger than 500 ms.
Feel lives in RUM, cost lives in APM, and you need both.
Curated: · Written: · Reviewed:
QA-62API availability is above the SLO. Signups still fail. What should the budget have counted?(show answer)
The first question I ask about error budgets for a user journey is which side owns the truth after a refresh.
An error budget on CPU or on one service misses a journey that fails when either half fails. The budget should be the share of complete journeys that succeed, counted from the click to the persisted result.
Concretely, define the journey as a named sequence of client events and server outcomes. Count success only when the last server write and the last client confirmation both happen. Burn the budget when either half drops the sequence.
The reason for that specificity is a failure I have seen: API availability sat at 99.95 percent while 8.6 percent of signups never reached the confirm screen because a client timeout fired at 2 seconds, and no budget page showed that loss.
Signup week of the timeout.
| Metric | Value |
|---|---|
| API availability | 99.95% |
| signup journeys attempted | 50,000 |
| confirm screen reached | 45,700 |
| journey success | 91.4% |
I would not consider it settled without evidence: Plot weekly successful journeys over attempted journeys and page the owners when the remaining budget hits 25 percent.
Budget the journey, not the process.
Curated: · Written: · Reviewed:
QA-63A sale is coming. Do you scale HTML rendering the same way you scale JSON?(show answer)
With capacity of SSR versus the API, a green local UI is where I start checking rather than stop.
Server rendering spends CPU and downstream calls on every HTML request, which is a different pool from JSON handlers. Sizing only the API leaves the render fleet as the first thing to fall over on a traffic spike.
Concretely, load-test HTML routes and JSON routes on separate pools. Cap concurrent renders per instance. Shed to a static shell when render queue delay exceeds the budget.
The reason for that specificity is a failure I have seen: A sale sent 12,000 HTML requests per minute through the same 8 API boxes, render queue delay hit 7 seconds, and JSON checkout also stalled because the process threads were busy drawing pages.
Mixed load on one shared pool.
| Pool | HTML rpm | JSON rpm | Render queue delay | Checkout p95 |
|---|---|---|---|---|
| 8 shared boxes | 12,000 | 4,000 | 7 s | 6.2 s |
| 6 render plus 8 JSON | 12,000 | 4,000 | 180 ms | 240 ms |
I would not consider it settled without evidence: Run a mixed load of HTML and JSON and record saturation of each pool before raising either limit.
HTML and JSON do not share a capacity story.
Curated: · Written: · Reviewed:
QA-64You want to reject anonymous HTML at the edge. What must the rule not block?(show answer)
I would answer edge middleware auth by following one click through the browser, the API, and the store.
Auth at the edge can reject anonymous HTML cheaply, but a wrong allow list will either leak private pages to the CDN or bounce every asset. The edge rule must match the app session cookie and must not run on public hashed files.
Concretely, gate only HTML routes that require a session. Skip immutable static paths. Forward the verified identity to the origin with a signed header the API also checks.
The reason for that specificity is a failure I have seen: Middleware required a cookie on every path including hashed scripts, and 64 percent of logged-in users saw a blank app because the script returned 401 after a cookie domain mismatch.
What the edge returned.
| Path | Cookie present | Status | Result |
|---|---|---|---|
| /app | no | 302 to login | correct |
| /app | yes, wrong domain | 200 HTML, 401 script | blank app for 64% |
| /static/main.hash.js | no | 200 | required |
I would not consider it settled without evidence: Request a private HTML route without a session and a hashed script without a session, and confirm the HTML is blocked while the script returns 200.
Guard the document, not the hashes.
Curated: · Written: · Reviewed:
QA-65Fetches fail in the browser after a CSP deploy, and the API logs are empty. What did the policy miss?(show answer)
My approach to CSP and API hosts separates what the client may believe from what the server will persist.
A content security policy that omits the API host will block fetches even when CORS is correct, while a policy that allows any host will let a compromised script call anywhere. The allow list is the set of origins the app actually calls.
Concretely, enumerate every fetch origin in the client. Put those origins in the connect-src list. Deploy the policy in report-only first, then enforce when reports are clean.
The reason for that specificity is a failure I have seen: connect-src listed only the web origin after the API moved to a sibling host, and 96 percent of writes failed in the browser while the API access logs stayed empty.
Writes under two policies.
| connect-src | Browser writes | API log lines | CSP reports |
|---|---|---|---|
| web origin only | 4% | 0 | 12,400 blocked connects |
| web plus API origin | 99.1% | 11,800 | 0 |
I would not consider it settled without evidence: Load the app with the enforcing policy and confirm zero CSP connect violations and a successful write in the network log.
CORS is not CSP, and both have to name the API.
Curated: · Written: · Reviewed:
QA-66A third-party private key showed up in the main JavaScript chunk. What do you do, and how do you stop the next one?(show answer)
For secrets in the browser bundle, I would name the contract both halves must keep before choosing a library.
A secret compiled into client code is a published secret. Rotation after the fact does not unsay the copies already downloaded and cached.
Concretely, keep third-party private keys on the server and mint short-lived public tokens for the browser. Fail CI if a secret pattern appears in client assets. Rotate any key that ever shipped in a bundle.
The reason for that specificity is a failure I have seen: A payment-provider secret key sat in the main chunk for 11 weeks, was scraped from a cached file, and generated 28,000 dollars of test charges that posted as live in 4 days.
Cost after the scrape.
| Week | Key location | Extra live charges |
|---|---|---|
| 1 to 11 | main chunk on CDN | 0 extra charges, key still secret |
| 12, days 1 to 4 | scraped copies | 28,000 dollars of live test charges |
| 12, after rotate | server mint only | extra charges stop |
I would not consider it settled without evidence: Grep the production JavaScript for key-shaped strings and for names used by secret stores, and require a clean result.
The bundle is a public document.
Curated: · Written: · Reviewed:
QA-67How should infrastructure as code describe the static host and the API so they cannot drift?(show answer)
I would size IaC for static and API together against a real page load and a real write, not a mocked fetch.
Two separate stacks for the bucket and the service will drift on origins, certificates, and env maps. The web origin, the API origin, and the shared secrets should be one module so a preview and production share the same wiring.
Concretely, declare the static host, the API service, the DNS records, and the env map in one workspace. Apply them in one plan. Fail the plan if the web origin would point at an API from another environment.
The reason for that specificity is a failure I have seen: Terraform for the API moved the hostname while the static stack still pointed at the old origin, and 100 percent of writes from the live site hit a decommissioned cluster for 26 minutes.
DNS after a split apply.
| Record | Intended | After split apply | Writes |
|---|---|---|---|
| www | new static | new static | pages 200 |
| api | new cluster | new cluster | none from site |
| site API host env | new cluster | old cluster | 100% to dead fleet |
I would not consider it settled without evidence: After apply, resolve the web origin and the API origin from the same state file and confirm they match the live DNS.
One plan should own both origins.
Curated: · Written: · Reviewed:
QA-68A pull request changes a query and a screen. What database should that preview use?(show answer)
The part of preview database for a fullstack PR interviewers probe is the mismatch after deploy, not the happy render.
A pull request that changes queries and screens needs data that looks like production without being production. Sharing the live database with previews mixes unreviewed migrations with real rows.
Concretely, fork a sanitized copy per pull request. Run migrations on that copy. Point the preview API at it and destroy the copy when the request closes.
The reason for that specificity is a failure I have seen: A preview migration dropped a column on the shared staging database, production deploys used the same staging for a dry run, and the dry run deleted 1.1 million search rows.
Where the drop actually ran.
| Target | Column dropped | Search rows left |
|---|---|---|
| dedicated preview copy | yes | copy only |
| shared staging | yes, by mistake | 0 of 1.1 million |
| production | no | unchanged |
I would not consider it settled without evidence: Apply a destructive migration on a preview and confirm production and shared staging schemas are unchanged.
Unreviewed SQL belongs on a disposable copy.
Curated: · Written: · Reviewed:
QA-69How do you send a known user to the new fullstack pair without rolling the whole site?(show answer)
I would anchor canary by cookie in a trace that starts at the click and ends at the query.
A named canary cookie lets you send volunteer sessions to the new pair while everyone else stays on the old pair. A sticky percent canary can shrink without a full rollback, but it still puts unpaid users on a broken client.
Concretely, set a signed canary cookie on volunteer accounts. Route HTML and API for that cookie to the candidate pair. Promote when those sessions meet the journey budget, then widen with a sticky percent you can shrink.
The reason for that specificity is a failure I have seen: A 5 percent random canary put 40,000 shoppers on a new checkout bundle and conversion dropped 19 percent. Shrinking the percent recovered most sessions, but 2,200 sticky clients stayed on the broken pair until the cookie expired.
Random percent versus cookie steer.
| Method | Users on candidate | Conversion vs stable | Pull-back |
|---|
| 5% random sticky | 40,000 | -19% | shrink percent, 2,200 linger |
| cookie on 80 staff | 80 | -2% | clear cookie |
| cookie then 1% | 8,000 | -0.4% | clear cookie plus shrink |
I would not consider it settled without evidence: Toggle the cookie on one account, prove HTML and API both land on the candidate, then prove a normal account still lands on the stable pair.
Steer named sessions first, then widen.
Curated: · Written: · Reviewed:
QA-70The site is now on HTTPS. Checkout still fails in the console. What is left to fix?(show answer)
What separates a strong answer on mixed content after an HTTPS cutover is knowing which half can lie to the other.
After TLS is required on the web origin, any remaining http asset or API call is blocked by the browser. The cutover is incomplete until every client request is https, including those inside old cached HTML.
Concretely, rewrite every origin the client calls to https. Purge HTML that still embeds http. Add a report policy for mixed content and fix reports before dropping http on the API.
The reason for that specificity is a failure I have seen: The site moved to https while 7 API SDKs still used http from an old vendor snippet, and 22 percent of checkouts were blocked as mixed content for 3 hours after the certificate cutover.
Blocked calls after the certificate.
| Request | Scheme | Browser action | Checkout share |
|---|---|---|---|
| document | https | allowed | n/a |
| 7 vendor SDK calls | http | blocked | 22% failed |
| same SDKs rewritten | https | allowed | 0.3% failed |
I would not consider it settled without evidence: Load the homepage and the checkout page over https and confirm zero mixed content warnings and zero http requests.
HTTPS is done when the last http call is gone.
Curated: · Written: · Reviewed:
QA-71Users look logged in on www and logged out on data. What is wrong with the session cookie?(show answer)
I would treat cookie domain split across api and www as something that must survive a stale bundle talking to a new API.
A session cookie scoped only to www never travels to a sibling API host, so the browser looks logged in while every API call is anonymous. A parent-domain cookie fixes that and also widens theft surface. A same-origin BFF or a Host-prefix cookie avoids the Domain attribute.
Concretely, prefer a same-origin BFF so the session cookie never needs a Domain attribute. If you must split hosts, set Secure and HttpOnly on the smallest parent that both need, and never use a Host-prefix cookie with Domain. Confirm the cookie is sent on document loads and on API fetches.
The reason for that specificity is a failure I have seen: The session cookie was bound to www only, API calls to a sibling host went out without it, and 100 percent of mobile Safari users appeared logged in on the page and logged out on data, which generated 2,400 support tickets in a day.
Where the cookie was sent.
| Cookie domain | Sent to www | Sent to api sibling | Tickets in 24 h |
|---|
| www host only | yes | no | 2,400 |
| same-origin BFF, Host-prefix cookie | yes | same host | 0 |
| parent domain | yes | yes | 12, wider theft surface |
I would not consider it settled without evidence: Inspect a real navigation and a real fetch and confirm the session cookie is present on both.
Prefer a same-origin BFF. Widen the cookie only if the hosts must split.
Curated: · Written: · Reviewed:
QA-72Which pages should a search engine be allowed to index in a logged-in product?(show answer)
Before changing either half for indexing an authenticated app, I would write down how both sides will be checked.
Search engines should index marketing pages and must not index private app routes. robots.txt is not access control, and disallowing a path can stop a crawler from even seeing noindex.
Concretely, require a session for app HTML and return 401 or 404 to anonymous crawlers. Send noindex on any HTML that requires a session. Keep public marketing routes crawlable and linked, and do not rely on robots disallow as the only gate.
The reason for that specificity is a failure I have seen: The app shell was reachable without a session and was indexed for 18,000 private dashboard URLs, which leaked customer first names in titles, and the cleanup took 5 weeks to drop from results.
What was in the index.
| Route class | Anonymous fetch | Indexed URLs |
|---|---|---|
| marketing | 200, indexable | 42 |
| app dashboard, no auth | 200 public HTML | 18,000 |
| app dashboard, 401 plus noindex | 401 | 0 after 5 weeks |
I would not consider it settled without evidence: Fetch an app URL without cookies and confirm 401 plus noindex, then fetch marketing HTML and confirm it is indexable.
Index the brochure, not the account.
Curated: · Written: · Reviewed:
QA-73Your client error service stores the full page URL and headers. Why is that a secret incident?(show answer)
I would start redacting tokens in browser error reports from the screen the user sees and the row the database actually wrote.
Browser error services capture URLs, headers, and storage by default, which often includes access tokens. Those reports leave the browser and become a secret store unless the payload is stripped first.
Concretely, strip Authorization headers, query tokens, and local storage keys before send. Deny list cookie values. Sample a real report in staging and read it as an attacker would.
The reason for that specificity is a failure I have seen: A client tracer shipped 9,400 error payloads containing bearer tokens, and 62 of those tokens were still valid when a contractor exported the project.
What the project export held.
| Field | Before redact | After redact | Still valid tokens |
|---|---|---|---|
| Authorization | bearer present | stripped | n/a |
| query access_token | present | stripped | n/a |
| payloads exported | 9,400 | 9,400 | 62 before, 0 after |
I would not consider it settled without evidence: Trigger a handled error in staging and confirm the stored report has no token, no session cookie, and no query secret.
An error report is an exfil path until you strip it.
Curated: · Written: · Reviewed:
QA-74A product HTML route feels cheap per request. Why is it dominating the API bill?(show answer)
The first question I ask about cost of SSR fan-out is which side owns the truth after a refresh.
One HTML request that fans out to many API calls multiplies both latency and origin spend. A render that looks cheap per page can dominate the bill when each page issues a dozen downstream reads.
Concretely, count downstream calls per HTML route. Batch or cache the reads the first screen needs. Set a hard cap on fan-out and fail the build when a route exceeds it.
The reason for that specificity is a failure I have seen: A product HTML route made 22 API calls, p95 render sat at 3.4 seconds, and origin cost for that route alone was 41 percent of the monthly API bill at 2.8 million views.
Calls per HTML view.
| Route | Downstream calls | p95 render | Share of API bill |
|---|---|---|---|
| product, unbatched | 22 | 3.4 s | 41% |
| product, batched | 4 | 420 ms | 9% |
| cap in CI | 6 | n/a | n/a |
I would not consider it settled without evidence: Trace one HTML request, count origin calls, and record added milliseconds and added cost per thousand views.
Fan-out is a bill and a latency budget.
Curated: · Written: · Reviewed:
QA-75Web dashboards are green and API dashboards are green, yet checkout fails. How do you run that incident?(show answer)
With an incident that needs both halves, a green local UI is where I start checking rather than stop.
Some outages only make sense when the client and the API are read together, because each half looks healthy in isolation. The response has to change both sides or the symptom will bounce between teams.
Concretely, declare a joint commander when a journey fails and both dashboards are green. Capture one failing session with HAR, logs, and a trace. Ship a paired fix or a paired rollback, then write the timeline as one story.
The reason for that specificity is a failure I have seen: Web and API each closed their tickets after green dashboards, the checkout still failed for 11 percent of users with an old service worker, and the incident lasted 4 hours and 20 minutes instead of 25 minutes.
Two closed tickets, one still-broken journey.
| View | Status at +25 min | Checkout success |
|---|---|---|
| web dashboard | green | hidden |
| API dashboard | green | hidden |
| joined session with old worker | stale API shape | 89% |
| after paired worker plus API fix | both updated | 99.6% |
I would not consider it settled without evidence: In the drill, produce one timeline that includes a client artefact and a server artefact before the incident is marked resolved.
If both halves look green and the user is still stuck, it is one incident.
Curated: · Written: · Reviewed:
QA-76How do you split unit tests, contract tests, and end-to-end tests for a full-stack feature?(show answer)
I would answer e2e versus contract versus unit tests by following one click through the browser, the API, and the store.
A passing unit test on one half does not prove the other half honors the same payload, and a wall of browser journeys is too slow to catch that mismatch on every change. Contract tests sit between those layers so the UI and the API agree on status, fields, and errors without paying for a full stack on each commit.
Concretely, unit-test pure rendering and query logic on each half in isolation. Add consumer-driven contract checks for every payload the page depends on. Keep a small set of end-to-end journeys that write a row and then read it back through the browser.
The reason for that specificity is a failure I have seen: A suite of 410 unit tests and 6 end-to-end journeys stayed green while the list endpoint renamed amount to total, and checkout showed blank totals for 14 percent of sessions for 2 hours.
One suite rebalanced across both halves.
| Layer | Before | After | What it proves |
|---|---|---|---|
| unit on UI and handlers | 410 | 360 | local logic |
| contract on payload | 0 | 48 | field names and status |
| end to end through browser and store | 6 | 5 | one paid write |
I would not consider it settled without evidence: Break one response field on a stub API and confirm the contract suite fails while the unit suite still passes, then restore it and confirm one end-to-end journey still writes and reads the row.
Prove each half, then prove they still agree, then prove one real write.
Curated: · Written: · Reviewed:
QA-77When should a full-stack team trust MSW, and when must they hit a real staging stack?(show answer)
My approach to MSW versus a real staging stack separates what the client may believe from what the server will persist.
MSW proves the page behaves under the responses you invented, and a staging stack proves the page against the schema, auth, and data the server actually ships. Treating a green MSW run as staging coverage hides every mismatch that only a real database and a real session produce.
Concretely, use MSW for component and page tests so the UI can cover loading, empty, and error states quickly. Run a pre-merge path against staging with the same journeys using real auth cookies and real rows. Fail the release if staging disagrees with the fixtures on status or required fields.
The reason for that specificity is a failure I have seen: MSW fixtures returned 200 with a nested customer object for 22 routes while staging returned 409 on 7 of them after a unique constraint landed, and the team spent 6 days chasing a frontend bug that was a real 409.
Twenty-two routes compared.
| Check | MSW | Staging |
|---|---|---|
| routes returning 200 | 22 | 15 |
| routes returning 409 | 0 | 7 |
| nested customer present | 22 | 15 |
| days lost on the wrong half | 6 | 0 after the diff |
I would not consider it settled without evidence: Diff fixture status codes against staging for the same journeys and require zero mismatches on required fields before merge.
Green MSW is a fixture. Green staging is the product.
Curated: · Written: · Reviewed:
QA-78A pixel diff is green and a contract test is red. What do you trust for a full-stack change?(show answer)
For visual regression versus API contract, I would name the contract both halves must keep before choosing a library.
Visual regression catches layout drift the API never sees, and an API contract catches field and status drift a screenshot will still paint as an empty box. A full-stack team needs both because a pretty page over a broken payload and a correct payload behind a clipped button fail different users.
Concretely, snapshot stable components under pinned fonts and dates rather than whole pages. Keep contract tests on every field the screenshot depends on. Review visual diffs and contract failures as two queues, never as substitutes.
The reason for that specificity is a failure I have seen: Eighty-four visual diffs were accepted as font noise while the card contract dropped imageUrl, and 31 percent of mobile checkouts rendered a blank hero for 2 days.
Two failure modes, two checks.
| Change | Visual suite | Contract suite | User impact |
|---|---|---|---|
| font hinting drift | 84 diffs | green | none |
| imageUrl dropped | green at 0.1% tolerance | red | 31% blank hero |
| primary button clipped 12 px | 1 real diff | green | missed taps |
I would not consider it settled without evidence: Remove one required field from the staging response and confirm the contract fails even when the screenshot still matches within tolerance.
Screenshots cannot see a renamed field, and contracts cannot see a clipped button.
Curated: · Written: · Reviewed:
QA-79The API returns a useful error body. How do you make that error accessible in the page?(show answer)
I would size accessibility of errors returned by the API against a real page load and a real write, not a mocked fetch.
An error the API returns is not accessible until the page names it, moves focus, and keeps it in the accessibility tree after the next render. A JSON message that never reaches a live region leaves keyboard and screen-reader users with a silent failure even when the network tab looks complete.
Concretely, map each API status and code to a visible message with a role of alert. Move focus to that message on failure and keep the field that caused it described by the same text. Confirm the raw body is never the only copy, because codes like 422 are not words a user can hear.
The reason for that specificity is a failure I have seen: Validation 422s returned 3 field errors in JSON while the form only flashed a red border, and 1400 support tickets over 4 weeks came from users who never heard why submit failed.
Same 422, two surfaces.
| Surface | What the user gets | Tickets in 4 weeks |
|---|---|---|
| network tab JSON | 3 field codes | 1400 |
| red border only | no words | 1400 |
| alert plus labelled fields | 3 spoken reasons | 40 |
I would not consider it settled without evidence: Submit an invalid payload with a screen reader running and confirm each API field error is spoken once and focus lands on the first invalid control.
An API error is only accessible once the page speaks it.
Curated: · Written: · Reviewed:
QA-80The API returns a snippet of HTML for a comment. How do you stop that becoming script in the page?(show answer)
The part of XSS from HTML the API returned interviewers probe is the mismatch after deploy, not the happy render.
HTML that arrives from your own API is still untrusted in the page because any stored field can carry markup from another user or from a compromised write path. Escaping on the server does not save a client that assigns the string to inner HTML, and a sanitizer on the client does not save a server that stores raw markup for every other consumer.
Concretely, store and return comments as text, not as markup. Render that text as text nodes on every client including admin and mobile. If a rich snippet is truly required, sanitize on the write with a vetted library and enforce a content security policy that blocks inline script.
The reason for that specificity is a failure I have seen: A comment endpoint returned stored HTML and the web client assigned it to inner HTML, so a payload executed in 2800 sessions before the field was escaped on both the write and the render.
Three consumers of one field.
| Consumer | Render path | Sessions exposed |
|---|---|---|
| web comment list | inner HTML | 2800 |
| admin moderation | text node | 0 |
| mobile after the fix | text node | 0 |
I would not consider it settled without evidence: Submit a script payload through the comment API and confirm every consumer renders it as text rather than executing it.
Treat API HTML as hostile as any other string.
Curated: · Written: · Reviewed:
QA-81A product search box talks to a SQL backend. Where does injection still hide in a full-stack app?(show answer)
I would anchor SQL injection through a search box in a trace that starts at the click and ends at the query.
Injection still starts in a search box that concatenates the query string, and it finishes in a handler that interpolates that string into SQL. Escaping in the UI does not bind a parameter, and an ORM on most paths does not save the one sort or filter that is still built as text.
Concretely, send the search value as a bound parameter from the handler. Allowlist sort columns and operators in one shared module both the page and the API import. Grant the search role only SELECT on the intended tables and test payloads through the box, not only through a private client.
The reason for that specificity is a failure I have seen: The search box passed a q string that was concatenated into an ORDER BY clause, and a crafted value dumped 68000 customer rows in 9 minutes before the endpoint was disabled.
Three inputs from one box.
| Input | Safe approach | Unsafe approach | Rows leaked |
|---|---|---|---|
| search text | bound parameter | concatenated WHERE | 0 vs 68000 |
| sort column | allowlist of 5 names | interpolated ORDER BY | 0 vs 68000 |
| page size | validated integer | interpolated LIMIT | 0 |
I would not consider it settled without evidence: Submit injection payloads through the visible search box including sort and filter controls and confirm none changes the query structure in the database log.
Bind the search value. Never splice it into SQL.
Curated: · Written: · Reviewed:
QA-82How do clickjacking and cookie-based auth interact on a full-stack site?(show answer)
What separates a strong answer on clickjacking and cookie auth is knowing which half can lie to the other.
If a session cookie is sent on framed navigations, another site can overlay your HTML and capture clicks that perform authenticated writes. Frame denial belongs on HTML via CSP frame-ancestors or X-Frame-Options. Client frame-busting is not the control.
Concretely, set SameSite and Secure on session cookies. Send CSP frame-ancestors none or an equivalent frame-denying header on every HTML response that can start a session. Do not treat a JSON API header or a client top-window check as the substitute.
The reason for that specificity is a failure I have seen: A settings page loaded inside a hidden iframe on a phishing site, the session cookie was sent because SameSite was unset, and 53 password resets completed before the frame-denying header shipped 40 minutes later.
One framed settings page.
| Control | Cookie sent in iframe | Resets completed |
|---|
| no SameSite, no frame header | yes | 53 |
| SameSite strict only | no on cross-site | 0 |
| CSP frame-ancestors none on HTML | cookie may be sent | frame blocked, 0 |
I would not consider it settled without evidence: Load the signed-in settings page inside a foreign iframe and confirm the cookie is not sent and the UI refuses to render.
Session cookies must refuse foreign frames.
Curated: · Written: · Reviewed:
QA-83A full-stack app has a JavaScript manifest and a server manifest. How do you handle supply-chain risk in both?(show answer)
I would treat supply chain in both package manifests as something that must survive a stale bundle talking to a new API.
A full-stack app ships two dependency graphs, and a compromised package in either one can read tokens from the browser or from the server. Lockfiles and scanners on only the JavaScript side leave the API runtime uninventoried.
Concretely, commit lockfiles for both manifests and fail CI if either drifts. Scan both graphs on every merge and judge reachability from the page and from the handler. Pin versions, verify integrity hashes where the ecosystem allows, and keep a bill of materials per release for both halves.
The reason for that specificity is a failure I have seen: A postinstall script in one npm package and a wheel on the API side from the same publisher ran for 5 days across 14 services, and the audit found the JavaScript scanner had been green the whole time.
One publisher, two manifests.
| Manifest | Scanner coverage | Services touched | Days until found |
|---|---|---|---|
| npm lockfile only | web only | 8 | still open |
| server lockfile only | API only | 6 | still open |
| both lockfiles | both halves | 14 | 5 |
I would not consider it settled without evidence: Answer which release of the web bundle and which release of the API contain a given package version within minutes from both bills of materials.
Scan both lockfiles or you scanned half the attack surface.
Curated: · Written: · Reviewed:
QA-84The network panel shows a fast API call and a slow TTFB. How do you read that waterfall as a full-stack engineer?(show answer)
Before changing either half for waterfall TTFB versus API time, I would write down how both sides will be checked.
Time to first byte is the document wait before any script can call the API, so a fast JSON endpoint cannot rescue a slow HTML or gateway hop. Interviewers expect you to split the waterfall into document, parse, and fetch rather than quoting one round-trip as the page cost.
Concretely, record TTFB, DOM content loaded, and the first API call on the same trace. Attribute document wait to origin, gateway, and server render separately from XHR wait. Fix the hop that owns the bytes before tuning a handler that is already fast.
The reason for that specificity is a failure I have seen: The detail handler sat at 90 ms while document TTFB was 1.1 s and HTML download of a 1.9 MB body added 2.7 s after first byte, and the team spent 3 sprints indexing a query that was never on the critical path.
Landing page waterfall.
| Hop | Time | On the first paint path |
|---|
| document TTFB | 1.1 s | yes |
| HTML body 1.9 MB download after first byte | 2.7 s | yes |
| first API GET | 90 ms | after paint | | indexed query later | 12 ms | never |
I would not consider it settled without evidence: Export a waterfall for the landing page and show TTFB, HTML size, and first API duration as three numbers before proposing a database change.
Split the waterfall before you blame the handler.
Curated: · Written: · Reviewed:
QA-85Product blames the API for poor Core Web Vitals. How do you decide which half is actually guilty?(show answer)
I would start Core Web Vitals blamed on the API from the screen the user sees and the row the database actually wrote.
Core Web Vitals mix server wait, asset discovery, and main-thread work, so blaming the API for a poor field score is usually a measurement error. Largest contentful paint can include TTFB, but layout shift and interaction delay almost never come from query time.
Concretely, break LCP into TTFB, resource delay, and render delay on field data. Map CLS to specific layout shifts in the page and INP to event handlers, not to average API latency. Change the half that owns the vital you named, and re-read the field 75th percentile after the change.
The reason for that specificity is a failure I have seen: LCP sat at 4.1 s and INP at 340 ms while the API p75 was 70 ms, yet the team scaled the database for 2 weeks while a 1.6 MB hero and a 280 ms click handler were the actual budget.
Field 75th percentile versus API p75.
| Metric | Field 75th | API share | Actual owner |
|---|---|---|---|
| LCP | 4.1 s | 70 ms | 1.6 MB hero |
| CLS | 0.28 | 0 ms | late font swap |
| INP | 340 ms | 70 ms | 280 ms click handler |
I would not consider it settled without evidence: For each vital, name the owning hop with a trace and show the API share of that hop before any capacity ticket is opened.
Name which vital moved before you open the query planner.
Curated: · Written: · Reviewed:
QA-86Users see a 500 only after the page retries. How do you debug that as a full-stack failure?(show answer)
The first question I ask about a 500 that only happens after a client retry is which side owns the truth after a refresh.
A client retry of a write that is not idempotent will look like a network flake in the browser and like a duplicate or a 500 in the API. The bug lives in the pair of a retrying fetch and a handler that is not safe to run twice.
Concretely, send an idempotency key from the client on every state-changing request and persist it on the server. Disable automatic retries on POST until that key is honored. Log the key on both the page report and the handler so the second attempt is visible as a retry, not as a new user.
The reason for that specificity is a failure I have seen: A checkout fetch retried a POST after a 250 ms abort, the handler created a second charge without a key, and 1160 duplicate charges plus 500s on the unique constraint landed in 27 minutes.
One abort, two charges.
| Attempt | Client | Server | Outcome |
|---|---|---|---|
| first POST | aborted at 250 ms | charge 1 written | pending UI |
| retry without key | new POST | unique 500 or charge 2 | 1160 dupes |
| retry with key | same key | original row returned | 0 dupes |
I would not consider it settled without evidence: Abort the first write in the client, allow one retry, and confirm the server returns the original result rather than a 500 or a second row.
Retries need idempotency keys on both halves.
Curated: · Written: · Reviewed:
QA-87A bug only happens for one customer. How do you reproduce it with a HAR file and SQL?(show answer)
With reproducing a bug with HAR and SQL, a green local UI is where I start checking rather than stop.
A full-stack bug is reproduced when the HAR shows the request the browser sent and the query log shows the statement the database ran for that request. Either artifact alone lets each half blame the other.
Concretely, capture a HAR from the failing session and extract method, path, status, timing, and body. Pull the query log for the same request id and compare bind values to the HAR body. Replay the request against staging and the SQL against a scrubbed copy of the row.
The reason for that specificity is a failure I have seen: A list page took 1.8 s for one tenant because the HAR showed 1 fetch while the SQL log showed 9 extra queries per row, and each half spent 3 days insisting the other tool was clean.
One tenant, one request id.
| Artifact | What it showed | Count |
|---|---|---|
| HAR | 1 GET /orders | 1.8 s |
| query log | 1 list plus 9 per row | 541 statements |
| after join on request id | N plus 1 on line items | 9 extras per row |
I would not consider it settled without evidence: Join HAR entries to query log lines on request id and show one table with status, duration, and statement count before changing either half.
Keep the HAR and the query log for the same request id.
Curated: · Written: · Reviewed:
QA-88A ticket needs a UI change and an API change. How do you pair so the work actually lands together?(show answer)
I would answer pairing on a ticket that spans both halves by following one click through the browser, the API, and the store.
A ticket that spans both halves fails when the pair splits by file type and merges on different days, because the contract then exists only in chat. Pairing means one branch, one payload example, and one journey that both people can run before either half is called done.
Concretely, start on the contract with an example request and response. Implement the handler and the page on one branch, switching seats at the write and the render. Do not merge until the journey creates a row and the page reads it back on a shared preview.
The reason for that specificity is a failure I have seen: Two specialists split a refund ticket by folder, merged 3 days apart, and spent 11 days reconciling a status enum the UI never sent, against a 3 day pairing estimate.
Refund ticket, two ways of working.
| Approach | Calendar days | Handoffs | Enum mismatches |
|---|---|---|---|
| split by folder | 11 | 4 | 3 |
| one branch pairing | 3 | 0 | 0 |
| contract first then split anyway | 8 | 2 | 1 |
I would not consider it settled without evidence: Show a preview where one click writes the new status and the same session reads it back before either half merges alone.
Pair on the write path until both halves ship together.
Curated: · Written: · Reviewed:
QA-89You are on call for a service that is both a web app and an API. How do you page and mitigate?(show answer)
My approach to on-call for a fullstack service separates what the client may believe from what the server will persist.
A full-stack pager that only watches API error rate will miss a white screen, and a pager that only watches frontend error reports will miss a silent 200 with empty rows. One person has to own the click through to the row, with dashboards and runbooks that name both halves.
Concretely, page on user-facing success, not only on 5xx. Keep a runbook that starts with a synthetic click, then the API trace, then the query. Practice rollback of the bundle and the API artifact independently and together.
The reason for that specificity is a failure I have seen: Checkout error reports sat at 18 percent for 74 minutes while API 5xx stayed at 0.2 percent because a new bundle called a removed field, and two teams each waited for the other to own the page.
Seventy-four minutes of a white checkout.
| Signal | Value during incident | Who watched it |
|---|---|---|
| frontend error rate | 18% | nobody paged |
| API 5xx | 0.2% | API pager, stayed quiet |
| synthetic checkout click | fail at 2 min | added after |
| time to mitigate after joint runbook | 8 min | one on-call |
I would not consider it settled without evidence: Run a game day where the bundle and the API break separately and confirm one on-call person can mitigate each with the same runbook.
One pager owns the click through to the row.
Curated: · Written: · Reviewed:
QA-90Tell me about a conflict over whether the frontend or the backend owned a piece of behaviour.(show answer)
For STAR conflict over frontend versus backend ownership, I would name the contract both halves must keep before choosing a library.
Interviewers grade a frontend versus backend ownership fight as a STAR story, so the answer needs a situation, a task, an action backed by a measurement, and a result that changed who owns the field. A story that only names the argument without a number reads as politics rather than engineering.
Concretely, state the screen and the write that were in dispute. Bring a trace or a prototype that shows where the invariant actually breaks. Let the contract owner decide, then ship the chosen half and delete the duplicate logic from the other.
The reason for that specificity is a failure I have seen: Tax rounding lived in the client and in the API for 6 weeks of disagreement, invoices differed by 41 cents on 1900 orders, and finance paused payouts worth 41000 dollars until one owner was named.
STAR for tax rounding ownership.
| Part | What happened |
|---|---|
| situation | web and API each rounded tax, 41 cent gap on 1900 invoices |
| task | pick one owner before payouts resumed |
| action | traced 50 orders, prototype moved rounding into the API |
| result | gap to 0 cents, 41000 dollars unblocked in 4 days |
I would not consider it settled without evidence: Rehearse the story with the four STAR parts and the 41 cent gap, then show the API as the only rounding owner and the UI displaying those integers.
Settle ownership with a measured contract, then commit.
Curated: · Written: · Reviewed:
QA-91Tell me about a time you missed an SLA that was defined on the whole user journey, not on one service.(show answer)
I would size STAR missed an end-to-end SLA against a real page load and a real write, not a mocked fetch.
Missing an end-to-end SLA is a STAR story about a number that neither half owned. Interviewers want the situation where the page target was missed, the task to find which hop burned the budget, the action that instrumented both halves, and the result that restored the number.
Concretely, name the journey SLA and the two half SLAs that were still green. Instrument success as a single click-to-row metric. Change the hop that actually burned the budget and show the end-to-end number recovering.
The reason for that specificity is a failure I have seen: Checkout availability sat at 99.2 percent for 11 hours because the API stayed at 99.95 while a client retry storm on 502s exhausted the pool, and nobody paged because each half dashboard was still green.
STAR for a missed checkout SLA.
| Part | What happened |
|---|---|
| situation | journey SLA 99.9%, observed 99.2% for 11 hours |
| task | find the hop while both half dashboards stayed green |
| action | joined client retries to API 502s, capped retries at 1 |
| result | journey back to 99.94%, pool wait from 800 ms to 40 ms |
I would not consider it settled without evidence: Join client attempts to server outcomes on a request id and page when successful journeys divided by attempted journeys diverge from either half by more than 0.2 points.
An end-to-end SLA is successful journeys over attempted journeys, not two half rates added together.
Curated: · Written: · Reviewed:
QA-92Tell me about a time you refused to ship a frontend-only fix because the store would still be wrong.(show answer)
The part of STAR refused a client-only shortcut interviewers probe is the mismatch after deploy, not the happy render.
Refusing a client-only shortcut is a STAR story about a write that would still fail after a pretty save. Interviewers want the situation, the cheaper patch you declined, the action that fixed the handler, and the result in lost or prevented bad rows.
Concretely, show the screen state and the row that would disagree after a refresh. Estimate the shortcut in days and the correct path in days. Ship the handler change with a contract test and drop the fake success from the client.
The reason for that specificity is a failure I have seen: A 3 day client cache of the profile hid a 409 from a unique email constraint, and 4200 writes never reached the store while the page showed saved, against a 9 day fix that bound the constraint to the form.
STAR for a refused profile shortcut.
| Part | What happened |
|---|---|
| situation | email unique 409, product asked for a client cache in 3 days |
| task | stop the page from showing saved on a failed write |
| action | refused the cache, returned 409 to labelled fields, 9 days |
| result | 4200 false saves avoided after launch, refresh matched the row |
I would not consider it settled without evidence: Refresh after save on the shortcut and on the real fix, and show the shortcut lying while the row is unchanged.
Refuse the client-only patch when the store would still be wrong.
Curated: · Written: · Reviewed:
QA-93How do you mentor a strong frontend or backend specialist so they can work across the full-stack boundary?(show answer)
I would anchor mentoring a specialist across the boundary in a trace that starts at the click and ends at the query.
Mentoring across the boundary is teaching the other half's failure mode, not assigning a tutorial repo. A specialist becomes useful on both sides when they can follow one click to a row and change the half that is actually wrong.
Concretely, pick one production journey and sit together on the trace from click to query. Give the specialist the unfamiliar half of the next ticket with a review on the contract, not on style. Measure success as incidents they can diagnose without handing off.
The reason for that specificity is a failure I have seen: A backend specialist shipped 8 weeks of API work a frontend partner had to restitch, and 5 incidents in that window were white screens on a 200 with a renamed field the mentor never made them render.
Eight weeks with and without a shared trace.
| Mentoring style | Incidents in 8 weeks | Handoffs per ticket |
|---|---|---|
| reviews on their home half only | 5 white screens | 2 |
| shared click-to-row traces | 0 | 0 |
| tutorial repo, no production | 4 | 2 |
I would not consider it settled without evidence: Have the specialist reproduce one incident from a HAR and a query log without the other half in the room, then ship the fix on both sides.
Mentor the missing half with a shared trace, not a lecture.
Curated: · Written: · Reviewed:
QA-94How do you estimate a feature that needs a schema change and a new screen?(show answer)
What separates a strong answer on estimating a feature that is both halves is knowing which half can lie to the other.
A full-stack estimate that prices only the mockup or only the endpoint will miss states, migrations, contract tests, and rollout of two artifacts. Breaking the work into both halves plus the contract produces a range that survives review.
Concretely, list schema and backfill, API failure paths, every screen state, accessibility, contract tests, and two-artifact rollout. Estimate each part and name the largest unknown. Give a range, not a single day count.
The reason for that specificity is a failure I have seen: A billing screen estimated at 5 days from the mockup took 17 because 7 states, 3 migrations, and a dual rollback were never in the estimate.
Where seventeen days went.
| Part | Estimated | Actual |
|---|---|---|
| happy-path screen | 5 d | 5 d |
| 7 states including 409 and empty | 0 | 4 d |
| 3 migrations and backfill | 0 | 4 d |
| contract tests and dual rollback | 0 | 4 d |
I would not consider it settled without evidence: Compare the estimate with actual time afterwards and record which half and which contract item were missed.
Estimate both halves and the contract, not the screen.
Curated: · Written: · Reviewed:
QA-95What belongs in a design doc for a vertical slice that cuts through UI, API, and store?(show answer)
I would treat a design doc for a vertical slice as something that must survive a stale bundle talking to a new API.
A vertical slice doc is one request, one row, and one check the user can perform, not a tour of every layer in isolation. Docs that specify only the schema or only the screens force a rewrite when the two meet.
Concretely, write the user journey, the example payload, the tables touched, the failure statuses, and the rollback of both artifacts. Include the one metric that says the slice shipped. Keep diagrams that follow the click, not the org chart.
The reason for that specificity is a failure I have seen: A 12 page API design and a 9 page UI spec were approved separately, then a 4 week rewrite started when the first slice could not load a row the handler had stored under a different id.
Two docs versus one slice.
| Artifact | Pages | Weeks to a working slice | Id mismatches |
|---|---|---|---|
| separate UI and API specs | 21 | 4 after approval | 1 silent |
| single slice doc with payload | 4 | 1 | 0 |
| diagrams by team, no payload | 8 | 3 | 2 |
I would not consider it settled without evidence: Walk a reviewer through one example payload that creates a row and renders it without adding a second identifier.
Write the vertical slice as one request and one row.
Curated: · Written: · Reviewed:
QA-96What goes wrong when a full-stack review only comments on the half the reviewer knows?(show answer)
Before changing either half for a code review that only looks at one side, I would write down how both sides will be checked.
A review that only reads the React files or only reads the handler will miss the payload the other half trusts. Full-stack risk lives in the fields, statuses, and retries that cross the boundary, not in the local style of one folder.
Concretely, require the example request and response in the pull request. Check authorization, idempotency, and empty and error renders even if you did not author those files. Refuse approval when only one half of the journey is exercised.
The reason for that specificity is a failure I have seen: A review approved a new export button without reading the query, and the handler streamed 900 rows from the wrong tenant because the UI passed orgId as a display string the server never re-checked.
Export button review.
| Review scope | Finding | Rows leaked |
|---|---|---|
| UI folder only | button styled | 900 |
| handler only | query looks indexed | 900 |
| payload plus both halves | missing tenant check | 0 after the fix |
I would not consider it settled without evidence: List the consumers of the payload and confirm the reviewer exercised the highest-traffic screen and the query plan before approving.
Review the payload and the render, not only the file you know.
Curated: · Written: · Reviewed:
QA-97The same POST succeeds in curl and fails in the browser with a CORS error while the API returns 200. Which request did the browser make first, what must that response carry, and what changes once the fetch also sends cookies?(show answer)
I would start CORS preflight and credentialed requests from the screen the user sees and the row the database actually wrote.
A CORS error is the browser refusing to hand your page a response it is not allowed to read; the API accepted the request and logged a normal 200. curl, Node and Python have no same-origin policy, so they never send a preflight and never see the failure — and any call with a JSON body, a custom header, or a verb other than GET or POST is preceded by an OPTIONS preflight the API must answer correctly.
Concretely, trace the OPTIONS round trip in the network panel before the real request, and answer it with the allowlisted Origin echoed back, Access-Control-Allow-Methods and Access-Control-Allow-Headers covering exactly what the page sends, and an Access-Control-Max-Age so the check is not repeated per call. Then repeat the same origin on the real response with Vary: Origin, and when the fetch uses credentials, send Access-Control-Allow-Credentials: true and a concrete origin instead of a wildcard.
The reason for that specificity is a failure I have seen: A team silenced a preflight failure by returning Access-Control-Allow-Origin: * together with Access-Control-Allow-Credentials: true; the browser still blocked the read because a wildcard is never valid with credentials, so the next patch echoed whatever Origin arrived. Every curl check and every server test stayed green for three weeks while any site could read authenticated /orders responses out of a logged-in browser session — the fix for a blocked read became a cross-origin data leak.
What each request must carry.
| Page request | Preflight? | Response headers required |
|---|---|---|
| GET /orders, no custom header | none | ACAO: https://app.example.com, Vary: Origin |
| POST /orders, Content-Type: application/json | OPTIONS first | + Allow-Methods: POST, OPTIONS; Allow-Headers: content-type, authorization |
| same call with credentials: 'include' | OPTIONS first | + Allow-Credentials: true, exact origin, never * |
Hypothetical, 80 ms RTT: the preflight adds 80 ms before the POST lands. Access-Control-Max-Age: 600 turns that into 80 ms per 10 minutes per browser instead of 80 ms per call.
I would not consider it settled without evidence: Show the OPTIONS response and the real response side by side and confirm both hit the same allowlist middleware, then replay the request with Origin: https://evil.example.com and confirm no allow header comes back for it.
CORS is the browser's permission slip to read the response; curl never needed one.
Curated: · Written: · Reviewed:
QA-98You are handing a full-stack feature to another team. What contract documentation do they actually need?(show answer)
The first question I ask about documenting the contract for the next team is which side owns the truth after a refresh.
The next team needs a versioned example payload, status meanings, and compatibility rules, not a prose tour of components and tables. Without that artifact, each half will extend the other in incompatible ways within a sprint.
Concretely, publish the request, response, and error examples next to the code. State which fields are required, which may be added, and which cannot be renamed without a version bump. Include one consumer test the next team can run on day one.
The reason for that specificity is a failure I have seen: Handoff notes listed screens and tables with 0 example payloads, and 6 teams drifted the same enum in 3 weeks, breaking mobile while web still shipped.
Handoff with and without a payload.
| Handoff artifact | Teams consuming | Weeks to first break | Enum variants in prod |
|---|---|---|---|
| prose screens and tables | 6 | 3 | 4 |
| versioned payload plus consumer test | 6 | none in 12 | 1 |
| OpenAPI without examples | 6 | 2 | 3 |
I would not consider it settled without evidence: Have the receiving team implement one new field using only the contract doc and confirm web, mobile, and API agree on a shared fixture.
Leave the next team a versioned example payload.
Curated: · Written: · Reviewed:
QA-99When does a backend-for-frontend belong on the team tech radar, and when is it extra hop theatre?(show answer)
With choosing a BFF on the tech radar, a green local UI is where I start checking rather than stop.
A BFF earns a radar ring when the web origin must hide domain credentials, adapt a session cookie, or aggregate several backends for one screen. Client count is a hint, not the rule. An extra hop that copies the same payload is theatre.
Concretely, list the boundary jobs the page needs. Put a BFF in assess or trial when aggregation, cookie session, or secret hiding is real. Measure the added hop against a direct call before adopting it as the default.
The reason for that specificity is a failure I have seen: A BFF was adopted for a single web app that already talked to one domain API, added 180 ms p75 on every click, and still copied the same payload so 4 later consumers bypassed it.
Radar choice for a BFF.
| Setup | Clients | Extra hop p75 | Duplicate payloads |
|---|---|---|---|
| BFF in front of 1 web app | 1 | 180 ms | 1 |
| BFF aggregating web and mobile | 2 | 40 ms | 0 |
| direct domain API | 4 | 0 ms | 0, but 3 shapes in clients |
I would not consider it settled without evidence: Show a table of clients, fields used, and p75 with and without the extra hop before the radar vote.
Put a BFF on the radar when the boundary job is real, not when you simply have two clients.
Curated: · Written: · Reviewed:
QA-100When should a full-stack team split a monorepo, and when should they keep UI and API together?(show answer)
I would answer when to split a fullstack repo by following one click through the browser, the API, and the store.
Split a full-stack repo when deploys, tests, and ownership already move on different clocks, not when the folder feels large. A premature split desynchronizes the contract, and a late split makes every CI minute pay for both halves on unrelated changes.
Concretely, measure CI time, deploy frequency per half, and how often a merge touches both. Keep one repo while contract tests still need a single version. Split when you can version the payload independently and each half has its own rollback.
The reason for that specificity is a failure I have seen: One repo CI hit 40 minutes, a split then shipped UI 9 minutes later than API for 3 deploys in a week, and a required field vanished from the handler while the bundle still sent it.
Keep versus split.
| Shape | CI time | Deploys desynced in a week | Contract breaks |
|---|---|---|---|
| one repo, no path filters | 40 min | 0 | 0 |
| split too early | 9 min each | 3 | 1 missing field |
| one repo with path filters | 11 min | 0 | 0 |
I would not consider it settled without evidence: Track CI minutes, co-change rate, and contract breaks for a month before and after any split proposal.
Split the repo when deploys and tests already split, not when the folder feels large.
Curated: · Written: · Reviewed:
