Skip to content
Tech Interview Prep home

Top 100 Technical Product Manager Interview Questions and Answers

The questions most likely to actually come up in your Technical Product Manager interview, ranked by likelihood — with detailed, senior-level answers covering what an interviewer is really listening for.

Curated: · Written: · Reviewed:

Reviewed 45Review pending 55
QA-1Sales wrote "as a user I want a Salesforce sync so that I can see accounts". What is wrong with that story, and what do you write instead?(show answer)

The first thing I would pin down about problem versus solution in a story is which user outcome or explicit non-commitment it actually changes.

A story names an actor, a job, and a benefit; naming the implementation in the want-clause freezes a solution before the need is testable.

Concretely, rewrite as the job ("reconcile the account the seller is looking at with the system of record") plus evidence, frequency, and a success measure. Keep Salesforce as one option in the design, not as the requirement.

The reason for that specificity is a failure I have seen: The team shipped a two-way sync in 11 weeks; 64% of sellers still copied fields by hand because the job was "trust the number on the call", not "sync every object". Support took 180 tickets/week on duplicate accounts.

Same request, two stories.

DraftActorWantBenefit
salessellerSalesforce syncsee accounts
rewriteselleropen the canonical account in <10squote the right entity
test12 sellers9/12 succeed3 still paste from email

I would not consider it settled without evidence: Five observed seller sessions and a completion metric (correct account opened in <10s) before any CRM vendor is in the story.

If the want-clause names a vendor, you wrote a procurement ticket, not a user story.

Curated: · Written: · Reviewed:

QA-2The story is "as a buyer I want two-factor authentication so that the account is secure". Engineering asks whether SMS or a hardware key is in scope. How does the benefit clause decide?(show answer)

I would start benefit clause as a trade-off guide from the decision it informs, not from the first feature request that arrived.

The benefit states why the outcome matters, which is what lets the team drop a solution that does not serve that why.

Concretely, define the protected action, attacker, recovery risk, and applicable policy. SMS may reduce risk but remains exposed to SIM-swap and phishing; TOTP avoids SIM-swap but remains phishable; WebAuthn security keys provide phishing resistance. Compliance is a constraint to verify, not a reason to choose a knowingly weak control.

The reason for that specificity is a failure I have seen: SMS 2FA shipped to 2.1×10^6 accounts; SIM-swap took over 340 wallets in 6 weeks (median loss $1,180). The benefit had said "secure" with no asset or threat.

Benefit that can choose a control.

BenefitStored assetControl that fits
"be secure"unspecifiedanything labelled 2FA
stop takeover of stored cardsPAN + walletapp TOTP or hardware
satisfy a named controlpolicy-defined scopecontrol verified against that policy

I would not consider it settled without evidence: A threat row: stored value, likely attacker, acceptable residual risk, named risk owner — before the story is Ready.

A benefit that only says "secure" cannot choose between SMS and a key.

Curated: · Written: · Reviewed:

QA-3The criterion is "the checkout feels fast". What do you replace it with so QA and the browser can disagree with the PRD in the same language?(show answer)

This is a place where a green launch dashboard and a correct handling of observable acceptance criteria are not the same event.

Acceptance criteria are observable outcomes: population, action, percentile, measurement point, and failure behaviour — not adjectives.

Concretely, write: eligible buyer, click Pay, p95 < 800ms to a committed order-id at the client, timeout at 8s with a retry that cannot double-charge. Include empty cart, 3DS, and declined card as named cases.

The reason for that specificity is a failure I have seen: QA passed "feels fast" on office Wi-Fi; p95 on 4G was 4.6s and 9% of pays double-submitted. Chargebacks were $47k in the first 14 days.

Adjective versus observable.

CriterionPasses onFails on
feels fastoffice Wi-Fi4G p95 4.6s
p95 Pay→order-id <800mssame4G
timeout 8s + idempotent retrylabdouble-charge if missing

I would not consider it settled without evidence: A RUM panel segmented by network class, plus a replay test showing that concurrent requests with the same intent key produce one capture and the same response, while reuse with a different payload is rejected.

If a stopwatch cannot fail the criterion, it is not acceptance.

Curated: · Written: · Reviewed:

QA-4Two items score the same RICE except confidence: one is 100% because the VP is sure, the other is 50% from a 40-person diary study. Which number is wrong, and what do you put in the confidence cell?(show answer)

My answer to RICE confidence as evidence quality begins at the constraint: if I cannot name the opportunity cost and the kill criterion, I do not have a product decision.

RICE confidence discounts every input for the quality of its evidence; enthusiasm and seniority are not evidence.

Concretely, assess confidence across reach, impact, and effort evidence, record the sources and dates, and apply a documented overall confidence score. A weak critical assumption should materially lower the score, but confidence is not mechanically the minimum of three percentages.

The reason for that specificity is a failure I have seen: The VP-sure item shipped at confidence 100; reach was 8 enterprise logos, not the 2×10^5 users in the cell. It occupied 14 engineer-weeks and moved North Star 0.0%.

Same formula, different confidence.

ItemReach sourceConfidenceTrue reach
VP-sureopinion100 (wrong)8 logos
diaryn=40 matching502×10^5 eligible
after rescoresame20 vs 50order reversed

I would not consider it settled without evidence: A shared confidence rubric tied to evidence quality—validated target-population data, limited matching research, or unsupported judgment—with assumptions and dates beside the score.

Confidence is a property of the evidence, not of the sponsor.

Curated: · Written: · Reviewed:

QA-5Engineering estimated 5 days of coding for a new export. Support already spends 12 hours/week on the CSV the export would replace. What belongs in the effort cell?(show answer)

I would treat RICE effort as lifecycle cost as a choice under capacity rather than as a slide that lists everything as Must.

RICE effort estimates the cross-functional person-time required to deliver and support the item over the chosen horizon; coding time alone is incomplete, but effort must not become a negative denominator.

Concretely, include product, design, engineering, data, security, rollout, documentation, and expected maintenance in person-months. Represent support hours retired as impact or in a separate cost-benefit calculation rather than subtracting them below zero.

The reason for that specificity is a failure I have seen: The 5-day export shipped; the format was wrong for 3 ERPs and support rose from 12h to 31h/week. Lifecycle effort was closer to 40 days plus a standing queue.

Effort with and without lifecycle.

ViewBuildSupport/week90-day total
coding only5dignored5d
with support savingsdelivery effort stays positive12h/week saved counted as benefitevaluate separately
as shipped5d+19h5d + 19h×12 incremental

I would not consider it settled without evidence: A cost table: build days, weekly support hours before/after a 4-week canary, and who eats the hours.

Coding days without support and rollout is a fiction that always wins RICE.

Curated: · Written: · Reviewed:

QA-6The board wants all three of: new onboarding, EU data residency, and a partner API, each a quarter of capacity. How do you make the yes visible as a no?(show answer)

The useful question for opportunity cost of a yes is what still happens when the loudest customer, the empty state, or the excluded cohort shows up.

Every yes consumes scarce capacity and is an implicit not-now to the next-best outcome; the roadmap must name the displaced work.

Concretely, show one capacity envelope, rank the three against the current goal, pick one or a sequenced pair, and write the Won't for this horizon with a revisit trigger. Do not stretch all three to 33% and call it a plan.

The reason for that specificity is a failure I have seen: All three started at 33%; none finished in 90 days. Onboarding completion fell 4 points, the residency date slipped past the contract, and the partner sent a 30-day cure letter.

Three items, one envelope.

MixOnboardingResidencyPartnerFinished in 90d
33/33/33startedstartedstarted0
100/0/0doneWon'tWon't1
50/50/0slicesliceWon't2 partial

I would not consider it settled without evidence: A single-page capacity map: people, calendar, and the one displaced item signed by the decision owner.

A yes that does not name a no is not a priority; it is a wish list.

Curated: · Written: · Reviewed:

QA-7Legal, sales, and the CEO each labelled their item Must Have for the same release. What test do you apply before the word Must is allowed?(show answer)

I would settle MoSCoW Must test against a counter-example first, so the roadmap date has to survive it.

A Must is indispensable to a viable, safe, or compliant outcome in a named horizon; preference, revenue, and seniority do not mint Musts.

Concretely, for each item ask: if it is absent, does the release fail a law, a safety bar, or the defined viable outcome? Cap Musts so contingency remains. Recategorise the rest as Should/Could/Won't with a named owner for the consequence.

The reason for that specificity is a failure I have seen: 14 Musts went into a 6-week release; 9 slipped. The one true Must (consent receipt) shipped on day 41 as a hotfix after a regulator query.

Must versus loudly wanted.

ItemSponsorAbsent meansCategory
consent receiptlegalunlawful processingMust
logo on invoicesalesawkwardCould
CEO colour tweakCEOnothingWon't this horizon

I would not consider it settled without evidence: A Must list of ≤3 with the failure mode each prevents, plus a Should list that can slip without killing the release.

If everything is Must, MoSCoW has revealed no choice.

Curated: · Written: · Reviewed:

QA-8A strategic customer asks for on-prem in the same quarter as the cloud GA. You cannot do both. What does a usable Won't look like?(show answer)

The judgement in explicit Won't Have is which metric or contract is a commitment versus a forecast, not which status colour looks calm.

Won't Have is an explicit exclusion for this horizon with rationale and a revisit trigger, not a polite never or a secret yes.

Concretely, state the need you heard, the governing outcome, the opportunity cost, the horizon, and what evidence would reopen it. Offer a workaround if one exists. Do not put a fake date on Later to end the call.

The reason for that specificity is a failure I have seen: The AE said "next half". The customer staffed a 9-person on-prem project; 4 months later you still had no binary. The account paused $1.2M ARR.

Three ways to say no.

ReplyCustomer hearsRevisit
next halfdatenone
neverinsultnone
Won't this GA; reopen at 20 EU cloud logosexclusionnamed

I would not consider it settled without evidence: A written Won't in the canonical roadmap, copied to the account plan, with the revisit trigger (e.g. 20 paying cloud customers in regulated EU).

A vague later is a commitment you will be quoted on.

Curated: · Written: · Reviewed:

QA-9A 2-week checkout bug loses $40k/week. A 10-week personalisation bet might add $15k/week later. Which do you sequence first, and why is that WSJF rather than "the bigger vision"?(show answer)

Where candidates lose the interview on cost of delay versus job size is usually a solution they named before they could name the need.

WSJF sequences by Cost of Delay ÷ job size so a short, expensive leak outranks a long bet with later value.

Concretely, cost of Delay here is $40k/week versus a speculative $15k/week. Job size is 2 versus 10 weeks. 40/2 = 20; 15/10 = 1.5. Fix the leak first unless a hard date (law, contract) overrides.

The reason for that specificity is a failure I have seen: The team protected the 10-week bet; the bug ran 8 more weeks ($320k) and the bet slipped anyway because checkout was on fire.

WSJF on two jobs.

JobCoD / weekSizeWSJF
checkout leak$40k measured2w20
personalisation$15k hoped10w1.5
sequenceleak first—then bet

I would not consider it settled without evidence: A two-row WSJF table with the leak's measured weekly loss, not a vision slide.

Vision does not earn interest while a leak is compounding.

Curated: · Written: · Reviewed:

QA-10The CEO wants June 12 on a Later theme. What does Now/Next/Later actually communicate, and what do you put on the slide instead of a date?(show answer)

I would answer Now Next Later as confidence by separating what the dashboard proved from what the user still could not complete.

Now/Next/Later encodes confidence: Now is being built with the strongest evidence, Next is ordered but still being refined, Later is directional and expected to move.

Concretely, keep Later free of calendar dates. If a real external deadline exists, mark it as a commitment with owner, conditions, and change control — not as Later with a date taped on.

The reason for that specificity is a failure I have seen: Later said "Q3 marketplace". Sales sold June 12. Engineering had not started discovery. The date moved three times; two logos churned citing the slide.

Same theme, three labels.

LabelEvidenceDate allowed?
Nowin build, measuredsprint range
Nextordered, thinningno fixed day
Laterdirectionalno; or a named commitment elsewhere

I would not consider it settled without evidence: One canonical roadmap where Later items have outcomes and uncertainty, and a separate commitments list with names.

A date on Later is a forecast wearing a commitment costume.

Curated: · Written: · Reviewed:

QA-11A customer attaches last quarter's roadmap PDF to a contract as Appendix B. What was the roadmap for, and what do you put in the contract instead?(show answer)

The product content of roadmap is not a delivery contract is the trade-off and the revisit trigger, not the template that scored it.

A roadmap communicates direction and uncertainty; a contract needs scoped commitments with change control.

Concretely, keep the roadmap audience-internal or clearly labelled non-binding. Put only named outcomes, SLOs, and change process into the contract. Train sales not to paste slides into MSAs.

The reason for that specificity is a failure I have seen: Appendix B listed 11 Later items as deliverables. Legal spent 6 weeks unwinding it; one item was already Won't. Credits of $180k went out.

Slide versus exhibit.

ArtifactItemsBinding?
Q2 roadmap PDF11 Lateraccidentally yes
MSA exhibit3 outcomesyes, by design
labelled non-binding roadmapthemesno

I would not consider it settled without evidence: A contract exhibit with ≤3 commitments, each with acceptance tests, and a clause that roadmaps are not exhibits.

If a lawyer can sue the slide, it was never a roadmap.

Curated: · Written: · Reviewed:

QA-12The board pack says "we will hit 18% activation by 30 September". Engineering's current evidence says 12–15%. Which word is missing, and who is allowed to turn a forecast into a commitment?(show answer)

Before writing a story I would write what a correct result for commitment versus forecast versus target looks like for n = 1 user and for the 10^6 who will hit the same path.

A forecast is the current evidence-based range; a target coordinates ambition; a commitment has authority, capacity, conditions, and change control.

Concretely, label the 12–15% as the forecast, 18% as a target if the board wants stretch, and only commit what capacity and dependencies support. The product owner (or named authority) is the one who can commit.

The reason for that specificity is a failure I have seen: 18% went into the board minutes as a commitment. September closed at 13.4%. Two directors treated it as a broken promise rather than a missed target.

Same number, three statuses.

Word18% meansOwner
forecastnot this, 12–15% isevidence
targetstretchboard
commitmentcapacity existsproduct owner

I would not consider it settled without evidence: A three-column table in the pack: forecast range, target, commitments — signed.

Will is a commitment verb; might is a forecast verb; mix them and you mint fake promises.

Curated: · Written: · Reviewed:

QA-13The dashboard's headline is total signups, up 44% after a contest. Weekly active collaborators is flat. Which is the North Star candidate, and what test disqualifies signups?(show answer)

The first thing I would pin down about North Star versus vanity count is which user outcome or explicit non-commitment it actually changes.

A North Star measures recurring value delivered to users; a count that can rise through spam, coercion, or contests without that value is vanity.

Concretely, ask whether a large increase, holding quality and guardrails equal, would mean the product is working. Signups fail that test. Collaborators/week with a quality threshold is closer for a docs product.

The reason for that specificity is a failure I have seen: The contest produced 220k signups, 91% never created a doc, and infra cost rose $28k/month. The board celebrated the 44%.

Headline after a contest.

MetricBeforeAfter contestValue?
signups500k720kno
weekly collaborators41k41.2kyes
docs with ≥2 editors18k18.1kyes

I would not consider it settled without evidence: A one-page NSM definition: event, eligible actor, window, quality bar, and a contest/spam guardrail.

If a giveaway can move the headline without moving value, it is not a North Star.

Curated: · Written: · Reviewed:

QA-14Activation is 62% or 19% depending on who you ask. What did each person put in the denominator, and which one should the team use?(show answer)

I would start denominator as the product decision from the decision it informs, not from the first feature request that arrived.

A rate is a numerator over a defined eligible population; changing the denominator in silence is a different product claim.

Concretely, define the cohort before observing activation: for example, completed signup, non-bot, non-employee, within the last 28 days. Do not condition eligibility on clicking Start or reaching the value screen unless the KPI explicitly measures conversion from that step. Both 62% (of people who clicked Start) and 19% (of all signups) can be true; only one is the activation KPI.

The reason for that specificity is a failure I have seen: Growth reported 62% using clicked-Start; finance used all signups at 19%. The board funded a channel that looked 3× better than it was.

Two activation rates.

DenominatornActivatedRate
clicked Start10k6.2k62%
all signups32k6.2k19%
KPI contract28d eligiblesame 6.2kthe one you version

I would not consider it settled without evidence: A metric contract: numerator, denominator, exclusions, window, owner, and the date the definition last changed.

The numerator and denominator together define the product claim.

Curated: · Written: · Reviewed:

QA-15Time-on-site is the goal and it is up 22%. Refunds and accessibility complaints are also up. What did the goal miss, and how do you instrument the next quarter?(show answer)

This is a place where a green launch dashboard and a correct handling of guardrail metrics beside the goal are not the same event.

A goal without guardrails can be won by harming users; pair the outcome with safety, quality, cost, and inclusion limits that can veto a "win".

Concretely, publish the goal plus hard guardrails (refund rate, WCAG defects, p95 latency, support hours, privacy incidents). A move that breaches a guardrail is a failed experiment even if the goal moved.

The reason for that specificity is a failure I have seen: An infinite-scroll bet added 22% time-on-site and a 1.8× refund rate; 140 accessibility tickets cited missing skip links. The experiment was declared a win on day 10.

Goal versus guardrails.

SignalWeek 0Week 2Veto?
time-on-site4.1m5.0mgoal
refund rate1.1%2.0%yes
a11y tickets12140yes

I would not consider it settled without evidence: A launch scorecard with goal, guardrails, and a pre-committed abort if any guardrail breaches for 7 days.

If harm cannot veto the metric, you are not measuring success.

Curated: · Written: · Reviewed:

QA-16Revenue is down 6% this quarter. Feature adoption of the new editor is up 18%. Which number can you still act on this month, and what must you not claim?(show answer)

My answer to leading versus lagging indicators begins at the constraint: if I cannot name the opportunity cost and the kill criterion, I do not have a product decision.

Leading indicators move earlier in a hypothesized chain; lagging indicators confirm later results. Neither label proves causality.

Concretely, use adoption as a leading signal only with a written hypothesis (adoption → documents finished → retained paid seats) and a test. Do not tell the board that 18% adoption has already saved next quarter's revenue.

The reason for that specificity is a failure I have seen: The 18% was celebrated as a turnaround. Finished documents were flat; revenue fell another 4% the next quarter. The hypothesis had never been tested.

Same month, two clocks.

MetricTypeThis monthNext quarter
editor adoptionleading?+18%unknown
finished docshypothesized link0%still 0%
revenuelagging−6%−4% more

I would not consider it settled without evidence: A metric tree with one tested link (cohort retained 8 weeks after first finished doc) before adoption is treated as leading.

Leading means earlier, not proven.

Curated: · Written: · Reviewed:

QA-17A tax-filing product team is told to maximise Engagement from HEART. Why might that be the wrong category, and what do you put in its place?(show answer)

I would treat HEART as a menu as a choice under capacity rather than as a slide that lists everything as Must.

HEART is a menu of Happiness, Engagement, Adoption, Retention, and Task Success; maximising engagement is wrong when the job is to finish quickly and leave.

Concretely, for a filing product, Task Success (complete accurate return, time-to-complete, error rate) and Happiness beat session length. Engagement can even be a negative signal.

The reason for that specificity is a failure I have seen: A "stay in product" badge increased sessions 30% and completion fell 7 points because people wandered. Refunds of filing fees hit $62k.

HEART on a filing product.

CategorySignalWanted direction
Engagementminutes/sessiondown
Task Successcomplete accurate returnup
HappinessCES after submitup

I would not consider it settled without evidence: A goals-signals-metrics row that names Task Success as the primary category and treats extra minutes as a cost.

If the user's win is leaving, engagement is the wrong trophy.

Curated: · Written: · Reviewed:

QA-18Research wants a 6-week study "to learn about onboarding". What do you require before that study occupies a discovery slot?(show answer)

The useful question for experiment as a decision, not theatre is what still happens when the loudest customer, the empty state, or the excluded cohort shows up.

Discovery belongs on the roadmap when it resolves a named uncertainty that would change a decision; otherwise it is theatre.

Concretely, write the decision it enables, the hypothesis, the evidence threshold, the budget, and the stop/pivot rule. Separate permission to learn from permission to scale.

The reason for that specificity is a failure I have seen: The 6-week study produced a 42-page deck and no backlog change. Meanwhile a 5-day prototype would have killed the proposed flow: 0/8 users found the CTA.

Study versus decision.

WorkWeeksDecision enabledOutcome
open study6nonedeck
prototype1ship/kill CTA0/8 → kill
A/B after2scale?only if prototype lives

I would not consider it settled without evidence: A one-page experiment brief with the decision, n, threshold, and a calendar kill date.

If no decision will move, do not spend the slot.

Curated: · Written: · Reviewed:

QA-19A partner integration is in Now. What kill criterion do you write on day 0 so it cannot become a zombie for three quarters?(show answer)

I would settle kill criteria before launch against a counter-example first, so the roadmap date has to survive it.

A Now item needs a pre-committed stop: metric, window, and owner who will actually stop it.

Concretely, example: if 8 weeks after GA fewer than 200 accounts complete a successful sync, or support hours exceed 15/week, the owner sunsets the integration. Write the sunset steps (flags, contracts, comms) now.

The reason for that specificity is a failure I have seen: No kill line existed. The integration sat at 37 accounts and 22 support hours/week for 11 months, occupying 1.5 engineers.

Kill line on a partner.

Week 8AccountsSupport h/wAction
bar≥200≤15keep
actual3722sunset
with no bar3722zombie 11 months

I would not consider it settled without evidence: A Now card with the stop metric, the date, and a named sunset owner — reviewed in the same forum that approved the build.

Without a kill line, Now means forever.

Curated: · Written: · Reviewed:

QA-20The team is at 95% utilisation on committed delivery. Leadership wants "more discovery". Where does the time come from?(show answer)

The judgement in discovery versus delivery capacity is which metric or contract is a commitment versus a forecast, not which status colour looks calm.

Discovery is a capacity allocation, not a slogan; it displaces delivery unless you raise capacity.

Concretely, set an explicit split (e.g. 70/30 delivery/discovery), protect it in the sprint, and show the displaced delivery. Do not ask people to discover after 40 hours of committed work.

The reason for that specificity is a failure I have seen: Discovery was "evenings". Two spikes were skipped; a $400k bet shipped on an untested assumption and was rolled back in 9 days.

95% delivery, then a slogan.

PolicyDeliveryDiscoveryWhat shipped
95/evenings95%0 realuntested bet
70/3070%30%spike then bet
100/0100%0same as row 1

I would not consider it settled without evidence: A velocity view that shows discovery hours as a first-class slice, not leftover.

Unscheduled discovery is unpaid overtime, not a strategy.

Curated: · Written: · Reviewed:

QA-21The launch invite has 47 names. How do you rebuild it so a decision can actually be made?(show answer)

Where candidates lose the interview on stakeholder map versus mailing list is usually a solution they named before they could name the need.

Map stakeholders by decision role, impact, expertise, and dependency — not by who might be interested.

Concretely, name one accountable owner, contributors, risk approvers, implementers, and notify-only. Confirm who can decide. A RACI that lists everyone as C is a mailing list.

The reason for that specificity is a failure I have seen: 47-person "alignment" took 3 weeks; nobody could cut scope. The launch missed the event window by 11 days.

47 names versus a decision.

RoleCount in inviteCount that can decide
interested470
accountable owner11
risk approvers22 bounded

I would not consider it settled without evidence: A one-page map with a single D, ≤5 C's, and an escalation path.

If everyone is consulted, no one is accountable.

Curated: · Written: · Reviewed:

QA-22A regional VP wants a custom workflow that would delay the shared GA by 6 weeks. How do you say no without promising a fake quarter?(show answer)

I would answer saying no with a revisit trigger by separating what the dashboard proved from what the user still could not complete.

A usable no restates the need, names the governing outcome and opportunity cost, states Not now for the horizon, and gives a revisit trigger.

Concretely, offer a configuration that does not fork the core if one exists. Put the custom workflow on Later with a trigger such as "after GA and 3 regions on the shared flow".

The reason for that specificity is a failure I have seen: PM said "Q4 maybe". The VP staffed a shadow team of 4. The fork landed in production under a flag nobody owned; incidents doubled.

No that can be quoted.

ReplyDelay to GAShadow team?
Q4 maybeunknownyes, 4 people
Not now; revisit after GA + 3 regions0no
silent ignore0 then explosionyes

I would not consider it settled without evidence: A written Not now in the decision log, copied to the VP, with the trigger and the owner of the next review.

A maybe is a yes with plausible deniability.

Curated: · Written: · Reviewed:

QA-23Three people left the same meeting with three different "decisions". What do you write before anyone stands up?(show answer)

The product content of decision record over meeting memory is the trade-off and the revisit trigger, not the template that scored it.

A decision record captures choice, options, evidence, owner, dissent, consequences, and revisit trigger — not a transcript.

Concretely, before close: read the decision aloud, record objections, name the owner and date, link the artifact. Version superseding decisions rather than editing history.

The reason for that specificity is a failure I have seen: Engineering built option B, sales sold option A, legal thought it was deferred. Unwinding took 19 days and a public status page.

Same meeting, three memories.

PersonThought we choseCost of the mismatch
engBbuilt B
salesAsold A
legaldefer19 days to unwind

I would not consider it settled without evidence: A dated record in the canonical log with options A/B/C and the selected one highlighted.

If it is not written, it was not decided.

Curated: · Written: · Reviewed:

QA-24The CEO pings "can we just add a chatbot on the homepage this sprint". What is the first translation, and what do you refuse to do?(show answer)

Before writing a story I would write what a correct result for executive request intake looks like for n = 1 user and for the 10^6 who will hit the same path.

Translate an executive request into a need, evidence, urgency, and displaced work; do not silently inject it into the sprint.

Concretely, clarify the job (deflect support? convert?). Compare with the current Sprint Goal. If it preempts, show the displaced item and get the CEO to own that trade-off. If it does not, it waits in intake.

The reason for that specificity is a failure I have seen: The chatbot landed in 4 days, deflected 0 tickets, and the sprint goal (checkout tax) slipped. Tax miscalc cost $91k in refunds.

Ping versus intake.

PathSprint goalChatbotTax refunds
silent injectslippedshipped$91k
intake + trade-offheldLater$0
CEO owns slipslipped on purposeshippedstill $91k if tax waits

I would not consider it settled without evidence: An intake note with the underlying job, the displaced item, and a yes/no from the same person who asked.

Seniority is a source of evidence, not a bypass of the backlog.

Curated: · Written: · Reviewed:

QA-25Sales has a customer roadmap, product has a Now/Next/Later, and the board has a Gantt. They disagree on whether marketplace is this year. What do you do?(show answer)

The first thing I would pin down about one canonical roadmap is which user outcome or explicit non-commitment it actually changes.

One canonical roadmap holds product intent; audiences get tailored views of the same facts, not contradictory decks.

Concretely, pick the canonical artifact, map every other view to it, and mark marketplace's true status (Later, not committed). Kill or watermark the Gantt that implies a year.

The reason for that specificity is a failure I have seen: A customer was shown the Gantt with marketplace in 2026 H1. Canonical Later said "directional". Legal bought a $250k concession.

Three decks, one status.

DeckMarketplaceSafe?
board GanttH1 2026 barno
sales PDFthis yearno
canonicalLater, directionalyes

I would not consider it settled without evidence: A single source with status labels, plus a diff of the sales deck against it before any customer meeting.

Two roadmaps are a bug; three are a lawsuit.

Curated: · Written: · Reviewed:

QA-26The quarter recap lists 47 stories shipped. Activation is unchanged. What do you put on the recap instead?(show answer)

I would start outcome versus output counting from the decision it informs, not from the first feature request that arrived.

Output counts activity; an outcome is a change in a user or business measure the work was for.

Concretely, report the hypothesized outcome, the result, and only then the stories as evidence of what was tried. A 47-story quarter with a flat activation is a failed bet, not a productive team.

The reason for that specificity is a failure I have seen: The recap led with 47. Finance cut the team 20% the next quarter because activation had been the board goal and had not moved.

Two recaps.

LeadStoriesActivation
47 shipped470 pts
activation +0470 pts, 3 bets killed
activation +412the 12 that mattered

I would not consider it settled without evidence: A recap with one outcome chart, confidence, and a table of bets killed versus scaled.

Stories are inventory; outcomes are the product.

Curated: · Written: · Reviewed:

QA-27The epic split is "frontend, backend, data model". Why is that a bad split for a first release, and what is a vertical slice here?(show answer)

This is a place where a green launch dashboard and a correct handling of vertical slice versus layer stories are not the same event.

Split by a coherent user outcome that can be learned from, not by architectural layer that delivers no value alone.

Concretely, for onboarding: one segment, one happy path, with auth, data, and UI together, plus a manual fallback. Leave other segments and the pretty empty states for later.

The reason for that specificity is a failure I have seen: Three layer stories closed independently. Nothing was usable for 7 weeks. The first combined demo failed identity for 40% of the pilot.

Layer split versus slice.

SplitWeek 4User can
FE/BE/data3 tickets donenothing
one-segment slice1 pathcreate org
all segmentsstill buildingnothing

I would not consider it settled without evidence: A slice that one user can complete end-to-end in staging, with a named segment.

A closed frontend story with no backend is inventory, not an increment.

Curated: · Written: · Reviewed:

QA-28A story sits in Ready with 14 unchecked template boxes but the team cannot explain the empty-state. Is it Ready?(show answer)

My answer to ready is shared understanding begins at the constraint: if I cannot name the opportunity cost and the kill criterion, I do not have a product decision.

Ready means the team shares enough understanding to take a safe next step, not that a checklist is full.

Concretely, if empty, error, authorization, and measurement are unagreed, it is not Ready. A short conversation that resolves those beats a decorated ticket.

The reason for that specificity is a failure I have seen: The template was 14/14. Empty-state was "TBD". Production showed a blank page; 22% of new orgs never created a record. Recovery took 5 days.

Template versus understanding.

SignalBoxesEmpty stateReady?
template14/14TBDno
walkthrough6/14agreedyes
production—blank22% stall

I would not consider it settled without evidence: A walkthrough where two engineers and one designer can describe the empty, error, and success states without opening the ticket.

A full checklist with a TBD in the behaviour is not Ready.

Curated: · Written: · Reviewed:

QA-29Day 8 of a 10-day sprint, research shows the chosen flow excludes keyboard users. Freeze or change? How do you assess it?(show answer)

I would treat mid-delivery requirement change as a choice under capacity rather than as a slide that lists everything as Must.

A late change is judged by new evidence versus current scope: assess impact, decide now/later/not, and do not defend sunk cost as a requirement.

Concretely, if the exclusion is a legal or Definition-of-Done miss, change now and cut elsewhere. If it is a later enhancement, record it. Update criteria, tests, and comms in the same decision.

The reason for that specificity is a failure I have seen: The team froze "to protect the date". A WCAG complaint landed 12 days after GA; the patch took longer than the original cut would have.

Day 8 choice.

ChoiceDateKeyboardLater cost
freezeheldexcludedcomplaint + patch
change now+2dincludednone
later ticketheldexcludedsame as freeze

I would not consider it settled without evidence: An impact note: users, tests, date, residual risk, named accepter.

Sunk tickets are not a reason to ship an inaccessible flow.

Curated: · Written: · Reviewed:

QA-30The story's acceptance tests are green in a mock. Production observability and a rollback path are missing. May it be Done?(show answer)

The useful question for Definition of Done versus story acceptance is what still happens when the loudest customer, the empty state, or the excluded cohort shows up.

Story acceptance is item-specific; Definition of Done is the shared quality of a releasable Increment, including ops and risk controls.

Concretely, doD should require reviewed code, relevant tests, security/privacy/accessibility, docs, observability, and rollback evidence. A mock-green story is not an Increment.

The reason for that specificity is a failure I have seen: It merged as Done. The first 5xx had no trace id; the rollback was "redeploy yesterday" and took 47 minutes. 8k users saw a partial write.

Acceptance versus Done.

GateMockProduction
story testsgreenunproven
trace idn/amissing
rollback drilln/a47 min

I would not consider it settled without evidence: A DoD checklist that staging-with-mocks cannot tick, plus a drill of the rollback.

Green in a mock is not Done in a product.

Curated: · Written: · Reviewed:

QA-31The spec lints clean. A partner still cannot complete "create order" because 200 is returned for a stock failure. What did lint not prove?(show answer)

I would settle OpenAPI validity versus product contract against a counter-example first, so the roadmap date has to survive it.

A valid OpenAPI document proves structural rules, not that the contract matches the product behaviour partners must branch on.

Concretely, map each domain outcome to a status and a stable problem type. Stock failure is 409 or 422 with a typed code, not 200 with a message. Contract-test the partner's client.

The reason for that specificity is a failure I have seen: Partners interpreted the 200 as a created order and advanced fulfilment even though inventory was unavailable. Other clients parsed the message and retried, exposing the separate absence of idempotency. Warehouse variance was $210k before the status was fixed.

Lint versus branch.

OutcomeLintPartner branch
created200success
out of stock as 200still lint-cleanretry → dupes
out of stock as 409lint-cleanstop

I would not consider it settled without evidence: Consumer contract tests for success, unavailable inventory, validation failure, and replay behavior, including the documented status, problem type, and retryability of each outcome.

Lint cannot see a lie in a 200.

Curated: · Written: · Reviewed:

QA-32Mobile drops after Pay; 8% of users tap again. The API is POST /charges with no key. What do you put in the PRD?(show answer)

The judgement in POST idempotency as a product requirement is which metric or contract is a commitment versus a forecast, not which status colour looks calm.

A create that users will retry needs a scoped idempotency key so two taps are one charge.

Concretely, the client sends one idempotency key per payment intent. The server atomically claims the scoped key, stores the request fingerprint and terminal response, returns that response on an identical replay, and rejects reuse with a different payload. Define TTL and behavior after expiry.

The reason for that specificity is a failure I have seen: Double-tap charged 8% twice. Chargebacks and goodwill were $63k in 21 days. Concurrent taps without an atomic key claim still create two captures.

Two taps, one intent.

DesignTapsCaptures
POST no key22
key per intent21
8% of 50k pays—4k extra

I would not consider it settled without evidence: Concurrent and sequential replay tests over 100,000 payment intents produce one capture per intent, identical replay responses, and explicit errors for key reuse with changed payloads.

If the user can tap twice, POST without a key is a product defect.

Curated: · Written: · Reviewed:

QA-33The export API returns 202 and a message "we're on it". Support asks when the file exists. What was missing from the product contract?(show answer)

Where candidates lose the interview on HTTP 202 is not completion is usually a solution they named before they could name the need.

202 Accepted means the request was accepted, not that the work finished; completion needs a status resource, event, or polling contract.

Concretely, return 202 with a job URL. Specify states (queued, running, done, failed), a terminal SLA (e.g. p95 < 60s for ≤10k rows), cancellation, and a signed download.

The reason for that specificity is a failure I have seen: Users refreshed for 15 minutes; 18% opened a second export. Disk filled. No job id existed to cancel.

Accepted versus done.

ResponseUser knowsSecond click
202 "on it"nothing18%
202 + job URLqueued/donerare
p95 complete60s bar—

I would not consider it settled without evidence: A state diagram and a p95 job-complete panel on the launch scorecard.

202 without a job is a shrug in HTTP clothing.

Curated: · Written: · Reviewed:

QA-34The activity feed uses OFFSET 10000 and sometimes skips a row the user just saw. What product requirement did the API miss?(show answer)

I would answer pagination as a product behaviour by separating what the dashboard proved from what the user still could not complete.

Deterministic pagination needs stable order, a tie-breaker, and a cursor whose behaviour under concurrent writes is defined.

Concretely, specify sort + id tie-break, cursor meaning, page-size bounds, and whether a new write can appear on a later page. Prefer cursors over large offsets.

The reason for that specificity is a failure I have seen: Support reproduced "missing invoice" 40 times/week at page 50+. OFFSET 10_000 on a 2k-write/min table was the cause.

Offset versus cursor under writes.

MethodAt 10kConcurrent insert
OFFSETslow + skiprow vanishes
cursor(ts,id)stabledefined
tickets/week400 after

I would not consider it settled without evidence: A contract test that inserts and deletes records during a page walk and verifies the documented snapshot semantics: no duplicate IDs and no missing IDs that were eligible at the walk's defined boundary.

Page numbers are a UX; skipping paid invoices is a product incident.

Curated: · Written: · Reviewed:

QA-35Every failure is HTTP 500 with "something went wrong". What do you require so users, clients, and support can act?(show answer)

The product content of error taxonomy for support and clients is the trade-off and the revisit trigger, not the template that scored it.

Errors need stable machine types that distinguish validation, auth, conflict, quota, dependency, and server failure, with safe human text.

Concretely, use problem+json or equivalent codes. Say whether to fix input, retry, wait, reauthenticate, or contact support. Never dump stack traces.

The reason for that specificity is a failure I have seen: Partners retried validation errors 30×. Token errors looked like outages. Support MTTR was 4.2h because the ticket had no type.

One 500 versus types.

EventOldNew typeClient action
bad SKU500400 validationfix
expired token500401reauth
warehouse down500503 dependencyretry

I would not consider it settled without evidence: A catalogue of ≤12 problem types covering 95% of volume, reviewed with support.

One 500 is not a contract; it is a black box.

Curated: · Written: · Reviewed:

QA-36v1 must die. Engineering can delete it Friday. What does a TPM put in front of that delete?(show answer)

Before writing a story I would write what a correct result for deprecation as a product programme looks like for n = 1 user and for the 10^6 who will hit the same path.

Deprecation is a consumer-migration programme: replacement, support window, adoption evidence, and retirement authority — not a URL bump.

Concretely, announce, ship v2, measure remaining v1 traffic by consumer, set a date with a kill-switch, help the last 5%, then delete. Inventory shadow clients.

The reason for that specificity is a failure I have seen: Friday delete took down a white-label that still called v1 (12% of GMV). Rollback took 6 hours. The announcement had been an internal Slack.

Delete Friday versus a programme.

Stepv1 trafficGMV at risk
Slack only12%12%
measured migrate0.8%last 5% helped
then delete~0kill-switch ready

I would not consider it settled without evidence: A traffic-by-consumer chart at <1% v1 before the delete is armed.

If you cannot name the last consumer, you are not ready to delete.

Curated: · Written: · Reviewed:

QA-37You add optional tax_region to the order JSON. One generated client throws. Why can an additive field be breaking, and what do you do?(show answer)

The first thing I would pin down about optional field that still breaks clients is which user outcome or explicit non-commitment it actually changes.

Strict deserializers, generated clients, signatures, and snapshots can make an optional field a breaking change even when OpenAPI calls it additive.

Concretely, declare response objects extensible and require supported clients to ignore unknown fields. Test representative generated clients before rollout. For already-deployed strict clients that cannot be updated, negotiate a migration, compatibility endpoint, or versioned response.

The reason for that specificity is a failure I have seen: A Java client with FAIL_ON_UNKNOWN_PROPERTIES died on 100% of new orders. That retailer was 9% of volume for 3 hours.

Additive on the wire.

ClientUnknown fieldResult
tolerant JSignoreok
Java FAIL_ON_UNKNOWNthrow9% down
new versionexpectedok

I would not consider it settled without evidence: A consumer matrix: tolerant vs strict, and a canary that includes the field against each.

An additive field is compatible only when the consumer contract permits unknown fields.

Curated: · Written: · Reviewed:

QA-38The API returns 429 with no headers. A startup on the free tier thinks they are banned. What belongs in the product contract and the plan page?(show answer)

I would start rate limits as packaging from the decision it informs, not from the first feature request that arrived.

A rate limit is a published quota with scope, response semantics, and retry guidance — and a packaging choice, not an undocumented wall.

Concretely, document requests/min per key, burst, 429 + Retry-After, and how to buy more. Make the free tier's number visible before signup.

The reason for that specificity is a failure I have seen: Free-tier 60 rpm was unpublished. 40% of new developers churned in week 1 citing "random bans". Support spent 25h/week.

Silent 429 versus a plan.

SurfaceNumber429 body
none60 hiddenempty
plan page60→600 paidRetry-After
churn week 140% → 11%after publish

I would not consider it settled without evidence: Plan page with numbers, a sandbox that 429s predictably, and a header on every 429.

An unpublished quota is experienced as a defect.

Curated: · Written: · Reviewed:

QA-39The flag is still in code 14 months after "100%". What product lifecycle did you skip, and what is the risk?(show answer)

This is a place where a green launch dashboard and a correct handling of feature flag as exposure control are not the same event.

A flag separates deploy from exposure and must have an owner, default, expiry, and a removal once the behaviour is the product.

Concretely, define cohorts, kill, and audit. After verified full exposure, delete the flag and the dead branch. A flag is not a permanent override.

The reason for that specificity is a failure I have seen: The stale flag defaulted off in a misconfig and hid billing for 2.4% of tenants for 11 hours. $180k of invoices were skipped.

Flag age.

Age at 100%BranchIncident
14 monthsstill forkedbilling hide
30 days then deleteone pathnone
no ownerunknown defaultsame class

I would not consider it settled without evidence: A flag inventory with age, owner, and a max-age (e.g. 30 days at 100%) before a removal ticket is automatic.

A flag that outlives the decision is a second, unofficial product.

Curated: · Written: · Reviewed:

QA-40The PM writes "your balance updates instantly everywhere". Replicas lag p95 2.4s. What do you change in the product, not in the database?(show answer)

My answer to eventual consistency as user copy begins at the constraint: if I cannot name the opportunity cost and the kill criterion, I do not have a product decision.

If the system is eventually consistent, the product must say so in the places users will act, or you must buy stronger consistency for those actions.

Concretely, show "updating…" on secondary read paths, or read-your-writes on the session that just submitted. Do not promise instant global reads if replicas lag.

The reason for that specificity is a failure I have seen: Users double-transferred $1.1M in a 40-minute incident because the app showed the old balance after a debit.

Promise versus lag.

Copyp95 lagUser action
instantly everywhere2.4sdouble send
updating…2.4swait
session read-your-writes0 on that nodenone

I would not consider it settled without evidence: A UX review of every post-write read, plus a lag SLO on the copy ("usually <3s").

Copy that denies lag is a defect when lag is real.

Curated: · Written: · Reviewed:

QA-41Marketing wants a 1-hour CDN cache for pricing. Legal wants the published price to match checkout. What number do you put in the PRD?(show answer)

I would treat cache staleness as a UX budget as a choice under capacity rather than as a slide that lists everything as Must.

Caching is a product budget for staleness, privacy, and invalidation — not a performance slogan.

Concretely, if checkout must match, cache pricing at a TTL the business will honour (e.g. 60s) or purge on publish. Do not 1-hour cache a legally binding price.

The reason for that specificity is a failure I have seen: A sale ended at :00; CDN served old 20% off until :47. 1,900 orders were repriced; 310 chargebacks.

TTL versus binding price.

PageTTLBinding?Incident
blog1hnonone
price CDN1hyes1,900 orders
price CDN60s + purgeyescontained

I would not consider it settled without evidence: A staleness table: page, TTL, purge path, and whether the number is binding.

A cache TTL is a promise about how wrong you may be.

Curated: · Written: · Reviewed:

QA-42Image resize moved behind a queue. The UI still says "ready". p95 drain is 90s. What is the product change?(show answer)

The useful question for queue as user-visible latency is what still happens when the loudest customer, the empty state, or the excluded cohort shows up.

Async processing is a latency trade-off the user must see: progress, failure, and a bound — not a silent still-processing ready.

Concretely, show queued/processing/failed. Bound p95 (e.g. 30s for <10MB). Offer notify-me. Do not render a broken thumbnail as success.

The reason for that specificity is a failure I have seen: Users downloaded originals for 90s of "ready" thumbs; 14% retried uploads and doubled the queue. Peak drain hit 11 minutes.

Ready versus queued.

UIp95 drainRetries
ready90s14%
processing + notify90s2%
SLO 30s28s2%

I would not consider it settled without evidence: A job-age SLO on the launch scorecard and a UI state for >30s.

Ready is a lie if the bytes are still in the queue.

Curated: · Written: · Reviewed:

QA-43During a partition, should the booking button fail closed or sell a seat twice? Who decides, and what do you write?(show answer)

I would settle CAP as a product choice against a counter-example first, so the roadmap date has to survive it.

During a network partition, a distributed operation cannot guarantee both linearizable consistency and an available response from every partition; the product must define behavior per operation.

Concretely, for seat allocation, choose a design that preserves the no-double-sell invariant—such as rejecting uncertain writes, partition ownership, or bounded inventory escrow. For like counts, accept stale reads and reconcile. Document guarantees rather than labelling the whole product CP or AP.

The reason for that specificity is a failure I have seen: The store stayed available and sold 220 double seats on a 14-minute partition. Compensation was $48k plus social posts.

Seat versus like.

FlowDuring partitionHarm if AP
seatfail closeddouble sell
like countstale oknone material
booking as AP220 doubles$48k

I would not consider it settled without evidence: A table of flows: CP vs AP, user copy, and who pays for inconsistency.

Availability during a partition is a refund policy in disguise.

Curated: · Written: · Reviewed:

QA-44Growth wants a live-video feature. The current error budget is 12% remaining this month. How does the SLO participate in the roadmap?(show answer)

The judgement in SLO as a product constraint is which metric or contract is a commitment versus a forecast, not which status colour looks calm.

An SLO is a product constraint: if the budget is gone, reliability work preempts features that would burn more.

Concretely, estimate the feature's error-budget cost. If remaining budget cannot absorb it, sequence reliability first or cut scope. Do not "borrow" next month in silence.

The reason for that specificity is a failure I have seen: Live video shipped into a thin budget; burn went 3.1×. The team froze features for 3 weeks after a 99-minute outage.

Budget versus live video.

StateBudget leftFeature
today12%proposed
after ship−3.1× burn
correct sequencerecover then shipdelayed

I would not consider it settled without evidence: A budget forecast on the Now card: current burn, feature's added burn, remaining days.

A feature that cannot fit the error budget is not Now.

Curated: · Written: · Reviewed:

QA-45Analytics wants every keystroke "for later ML". What do you require before a single extra field is collected?(show answer)

Where candidates lose the interview on privacy purpose limitation is usually a solution they named before they could name the need.

Collect only what a named decision needs; purpose, retention, access, and deletion are product requirements, not a later legal pass.

Concretely, write the decision, the fields, the legal basis, retention, and who may see them. Keystroke-level data usually fails necessity. Prefer aggregated events.

The reason for that specificity is a failure I have seen: Keystream landed in a warehouse with 400-day retention. A DSAR took 3 weeks to export. A journalist story followed a mis-set ACL.

Field versus purpose.

CollectPurposeRetention
every keystrokelater ML400d
submit + error codesimprove form30d
DSAR time3 weeksvs 2 hours

I would not consider it settled without evidence: A data-review record with purpose and a query that the decision can be made without the field.

Later ML is not a purpose; it is a hope.

Curated: · Written: · Reviewed:

QA-46Accessibility is a Should for "after GA". The form is the only way to apply. What is wrong, and what is the Must test?(show answer)

I would answer WCAG as acceptance, not a ticket by separating what the dashboard proved from what the user still could not complete.

If a journey is the product, applicable WCAG success criteria are acceptance of that journey, not a later ticket.

Concretely, include labels, keyboard, errors, contrast, and assistive tests in the story's criteria. Automated scans are necessary and not sufficient.

The reason for that specificity is a failure I have seen: GA without keyboard submit. 0 of 8 assistive-tech users could apply. A public complaint landed on day 4; the patch was 11 days.

Should versus the only journey.

StatusKeyboard applyAssistive n=8
Should laterno0/8
Must in criteriayes8/8
complaintday 411-day patch

I would not consider it settled without evidence: A manual keyboard + screen-reader pass on the apply journey as a launch gate.

A11y after GA is exclusion with a date.

Curated: · Written: · Reviewed:

QA-47The spec says "authenticated users can GET /documents/{id}". A user guesses another id and sees it. What requirement was missing?(show answer)

The product content of object-level authorization in the PRD is the trade-off and the revisit trigger, not the template that scored it.

Authentication is who; authorization is whether this principal may act on this object. Both belong in the product contract.

Concretely, write object- and tenant-level checks, 404-versus-403 policy, and tests for id-guessing. Do not treat a valid session as a universal read.

The reason for that specificity is a failure I have seen: An IDOR listed 18k documents. 2 contained KYC images. The incident ran 31 hours from first report.

Session versus object.

CheckGuessed idResult
auth only200IDOR
object ACL404safe
incident18k docs31h

I would not consider it settled without evidence: A forbidden-object test in CI and a PRD line: "unauthorised id → 404".

Logged in is not allowed to read that.

Curated: · Written: · Reviewed:

QA-48The pricing module has become the tax on every feature: estimates went from 3 days to 9, and last quarter's four rollbacks cost 26 engineer-hours. Engineering says "give us capacity to pay down debt". What do you make visible before they get a single sprint, and how much is defensible?(show answer)

Before writing a story I would write what a correct result for technical debt capacity looks like for n = 1 user and for the 10^6 who will hit the same path.

Technical debt is an interest payment you can measure: it shows up as extra estimate days, rollbacks and support volume on the work you already wanted. Capacity for it is bought against that measured cost, not granted as a percentage of goodwill.

Concretely, start by pricing the tax the debt charges — extra estimate days per feature touching the module, rollback hours, support tickets — then agree a fixed capacity slice with a named owner and an exit metric such as lead time for changes in that module. Review the slice at the same cadence as features, and let the exit metric decide when it stops.

The reason for that specificity is a failure I have seen: Debt sat in "20% when we can" for three quarters. Estimates on the pricing module went 3 to 9 days; 12 features a quarter each ate 6 extra days, which is 72 engineer-days, about 14 engineer-weeks of drag per quarter. Q3 committed 19 items and shipped 11, and two of the four rollbacks landed on paying customers.

Priced interest.

LineFigure
features touching pricing / quarter12
estimate per feature3 d → 9 d
quarterly drag12 × 6 d = 72 d ≈ 14 engineer-weeks
15% slice of 6 engineers × 13 weeks≈ 12 engineer-weeks
after 3 sprints of contract tests9 d → 4 d
rollbacks4/qtr (26 h) → 0

I would not consider it settled without evidence: Name one check: a debt line carrying a cost figure, an owner and an exit metric, plus the lead-time chart for changes touching that module before and after the slice.

Debt you can price is a line item; debt you can't is an argument.

Curated: · Written: · Reviewed:

QA-49You change the unit of weight from lb to kg in the same field. Is that a version or a silent fix?(show answer)

The first thing I would pin down about versioning versus a breaking change is which user outcome or explicit non-commitment it actually changes.

Changing meaning, units, or requiredness of an existing field is breaking; version or introduce a new field, do not silently reinterpret.

Concretely, add weight_kg, deprecate weight, or bump a version with a migration. Never switch units in place.

The reason for that specificity is a failure I have seen: Carriers treated 10 as kg after you meant 10 lb. Because 10 lb is about 4.54 kg, 1,100 parcels were billed at 2.2× their intended weight for 9 days.

Same field, new meaning.

ChangeFieldSafe?
lb→kg in placeweightno
add weight_kgnewyes
1,100 parcels—9 days

I would not consider it settled without evidence: A compatibility review that flags unit/meaning changes as breaking, with a consumer test.

The bytes stayed the same; the product lied.

Curated: · Written: · Reviewed:

QA-50A vendor is $40k/year. A build estimate is 6 engineer-weeks. What costs are still missing before you choose?(show answer)

I would start TCO on a build-versus-buy from the decision it informs, not from the first feature request that arrived.

TCO includes run, support, change, exit, and risk — not licence versus first build.

Concretely, add on-call, vendor lock-in, data export, compliance, and a 3-year change rate. Six weeks of build plus two people to run it can exceed $40k/year quickly.

The reason for that specificity is a failure I have seen: They built. Year-1 run was 1.2 FTE plus an incident week. Fully loaded cost was ~$240k versus the vendor's $40k.

Year-1 true cost.

OptionYear-1 cashRun
vendor$40ktheirs
build 6w~$40k labour1.2 FTE
build actual$240k loadedincidents

I would not consider it settled without evidence: A 3-year TCO table with run FTE, not a one-line build estimate.

The cheap build is often an expensive operations product.

Curated: · Written: · Reviewed:

QA-51The search vendor's contract forbids exporting relevance judgements. What do you require before GA depends on them?(show answer)

This is a place where a green launch dashboard and a correct handling of vendor lock-in as a product risk are not the same event.

Exit, data portability, and a documented replacement path are product requirements when a vendor sits on a user-facing journey.

Concretely, demand export of judgements, queries, and click logs. Time a trial restore onto an alternative. Price the switching cost into TCO.

The reason for that specificity is a failure I have seen: They raised 3× at renewal. No judgements could leave. A 5-month rewrite started under a deadline.

Renewal without an exit.

ExitJudgementsRenewal
forbiddentrapped3×
export drill100%negotiate
rewrite5 monthsunder fire

I would not consider it settled without evidence: A contract clause plus a passing export drill before the journey is Now.

If you cannot leave, you do not own the journey.

Curated: · Written: · Reviewed:

QA-52You add a required middle_name to signup. 18% of locales have no middle name. What broke, and what is the compatible move?(show answer)

My answer to schema evolution as a product change begins at the constraint: if I cannot name the opportunity cost and the kill criterion, I do not have a product decision.

Making a field required, or changing its meaning, is a product change for every client and locale, not a database tweak.

Concretely, add optional middle_name, or collect it later for locales that use it. Do not fail signup.

The reason for that specificity is a failure I have seen: Signup completion fell 18 points in 9 locales overnight. Ads kept spending. CAC looked 2.1× worse for a week.

Required middle name.

LocaleHas middle nameCompletion
USoften−2 pts
9 localesrare−18 pts
optional field—recovered

I would not consider it settled without evidence: A locale table and a canary on the new requiredness before 100%.

Required is a product word; schema is just how it is stored.

Curated: · Written: · Reviewed:

QA-53The feature has no trace, no business metric, and a log line "handled". How do you know it worked for the users you claimed?(show answer)

I would treat observability as a launch requirement as a choice under capacity rather than as a slide that lists everything as Must.

Observability of the user outcome is part of the feature: identifiers, a success metric, and a safe log — not a later platform ticket.

Concretely, define the golden event, its dimensions (segment, channel), and a dashboard on the launch scorecard. Redact PII.

The reason for that specificity is a failure I have seen: A 6% drop in completions hid for 11 days because no event existed. Ads spent through it ($44k).

Handled versus an outcome.

SignalExists at GA6% drop
log handledyesinvisible
completion eventno11 days
ads spendyes$44k

I would not consider it settled without evidence: An event in the warehouse matching the PRD's success definition, querying live before GA.

If you cannot see the outcome, you did not launch it; you deployed it.

Curated: · Written: · Reviewed:

QA-54Launch day, error rate 9×. Engineering can "forward-fix in a few hours". What does the product owner require instead?(show answer)

The useful question for rollback as a product promise is what still happens when the loudest customer, the empty state, or the excluded cohort shows up.

A launch needs a rehearsed rollback or kill that restores the prior user-visible behaviour inside a named bound.

Concretely, prefer flag-off in minutes. If migrate-forward is the only path, say so before launch and bound the data risk. Do not invent rollback during the incident.

The reason for that specificity is a failure I have seen: Forward-fix took 3h12m. $310k of bad orders needed manual repair. A flag would have been a 4-minute kill.

Forward-fix versus kill.

PathTime to safeBad orders
forward-fix3h12m$310k
flag off4 mincontained
no drillunknownsame class

I would not consider it settled without evidence: A drill timestamp: last successful flag-off in staging this week.

A few hours is not a rollback; it is a bet against the user.

Curated: · Written: · Reviewed:

QA-55Activation is 71% in the warehouse and 44% in product analytics. The SQL INNER JOINs to a properties table that 38% of users lack. What is the product mistake?(show answer)

I would settle INNER JOIN dropping users from a funnel against a counter-example first, so the roadmap date has to survive it.

INNER JOIN keeps matches only; using it as the funnel denominator silently drops people and invents a better rate.

Concretely, funnel denominators are LEFT JOIN or a base population table. Properties are attributes, not eligibility.

The reason for that specificity is a failure I have seen: The 71% was reported to the board. The missing 38% were mostly mobile. A mobile rewrite was delayed a quarter.

Inner join as a silent filter.

StepUsersRate
signups10k—
INNER properties6.2k71% fake
all signups activated4.4k44% true

I would not consider it settled without evidence: A count of distinct users before and after each join, committed as a test.

A join is a product filter whether you meant it or not.

Curated: · Written: · Reviewed:

QA-56The funnel is View → Click → Pay. 12% of Pays have no Click because of a cached button. How do you report conversion?(show answer)

The judgement in funnel step order versus event order is which metric or contract is a commitment versus a forecast, not which status colour looks calm.

A funnel is an eligibility rule, not a hope that instrumentation is complete; define whether a later event may skip a step.

Concretely, either repair the Click event or define Pay as convertible without Click (e.g. returning users). Do not drop 12% of revenue from the funnel in silence.

The reason for that specificity is a failure I have seen: The dashboard showed 2.1% click-to-pay; finance saw 8.4% of sessions paying. Growth killed a checkout experiment that was actually working.

Skip-step pays.

RulePays countedReported CR
must Click88%2.1%
Pay eligible100%8.4%
GMV gap12%killed a winner

I would not consider it settled without evidence: A reconciliation: paid GMV versus funnel-attributed GMV, with a <2% gap.

If money can skip a step, the funnel must say so.

Curated: · Written: · Reviewed:

QA-57Week-1 retention is 40% or 22% depending on whether "week 1" is 7 days from signup or calendar week. Which do you pick, and why must it be versioned?(show answer)

Where candidates lose the interview on retention cohort window is usually a solution they named before they could name the need.

Retention is a cohort plus a window; mixing calendar weeks with tenure weeks is a different product.

Concretely, pick tenure (day 7 / day 28) for product learning. Calendar weeks are for ops. Version the definition; do not retcon history.

The reason for that specificity is a failure I have seen: A "fix" switched to calendar weeks and retention "rose" 18 points. The board funded a channel that had not improved.

Two week-1s.

WindowRateUse
tenure d722%product
calendar week40%ops, not comparable
silent switch+18 ptsfake win

I would not consider it settled without evidence: A metric contract with tenure windows and a frozen backfill policy.

Retention moved because the clock moved.

Curated: · Written: · Reviewed:

QA-58Revenue by customer is $4.2M. Revenue by invoice rolled to customer is $5.1M. What grain bug is this, and which number is the KPI?(show answer)

I would answer GROUP BY grain by separating what the dashboard proved from what the user still could not complete.

An aggregation is only comparable at a declared grain; joining then summing at the wrong grain double-counts.

Concretely, decide the KPI grain (invoice vs customer-month). Pre-aggregate to that grain before joining attributes.

The reason for that specificity is a failure I have seen: A region leader was paid on the $5.1M number. Finance's $4.2M was the books. Clawback talks lasted 6 weeks.

Double-count at join.

GrainSumTies to books?
invoice$4.2Myes
join then sum$5.1Mno
bonuson $5.1Mclawback

I would not consider it settled without evidence: A grain diagram and a reconciliation to the ledger within $1.

If it does not tie to the ledger, it is not revenue.

Curated: · Written: · Reviewed:

QA-59Last-touch attribution uses MAX(ts) without a tie-breaker. Two events share a timestamp. What happens to 4% of conversions?(show answer)

The product content of window function for last-touch is the trade-off and the revisit trigger, not the template that scored it.

A maximum timestamp is not a unique last-touch record when events tie; joining it back can duplicate a conversion or force a nondeterministic choice.

Concretely, rOW_NUMBER() OVER (PARTITION BY conv ORDER BY ts DESC, event_id DESC). Document the tie-break.

The reason for that specificity is a failure I have seen: The MAX timestamp joined to two channel rows for 4% of conversions; downstream code selected an arbitrary row, moving $90k between runs. Bids followed the noise.

Tie without a breaker.

OrderStable?$ moved next day
MAX(ts)no$90k
ts, event_idyes$0
bid changefollowed noiseyes

I would not consider it settled without evidence: A replay of yesterday's table that hashes to the same attribution file.

If you cannot replay last-touch, you are not measuring it.

Curated: · Written: · Reviewed:

QA-60The KPI SQL is SELECT COUNT(DISTINCT user_id) FROM events WHERE type IN (...12 types). Why is this a product smell?(show answer)

Before writing a story I would write what a correct result for DISTINCT as a smell in a KPI looks like for n = 1 user and for the 10^6 who will hit the same path.

COUNT(DISTINCT user_id) is appropriate for a unique-user metric; it does not define which events make a user eligible and can conceal mixed event semantics or identity problems.

Concretely, define the success event and eligibility contract first, then use DISTINCT if the KPI's grain is unique users. Audit the twelve included event types separately.

The reason for that specificity is a failure I have seen: The twelve event types included both CTA views and completed activations, so 80k users were counted although only 41k completed the activation event. Headcount was hired against the mixed count.

Retry as activation.

QueryWeek 1Truth
COUNT DISTINCT 12 types80kmixed
one success event41kKPI
with retries distinct'd80khired against

I would not consider it settled without evidence: A duplicate-rate monitor on the success event, target ≈0, plus a per-type audit of the IN-list.

DISTINCT can enforce user grain; it cannot define success.

Curated: · Written: · Reviewed:

QA-61Average order value is $54 or $41 depending on whether NULL discounts are treated as 0. Which is the product number?(show answer)

The first thing I would pin down about NULLs in an average is which user outcome or explicit non-commitment it actually changes.

NULL is not zero unless the product says missing means no discount; averages over mixed NULL policy lie.

Concretely, define missing discount = 0 if the field is optional, or exclude incomplete rows if missing means unknown. State it on the dashboard.

The reason for that specificity is a failure I have seen: A "win" on AOV +$13 was a NULL-policy change. Pricing experiments were scaled on a mirage.

NULL as zero.

PolicyAOVHonest?
skip NULLs$54if unknown
NULL=0$41if optional
silent switch+$13no

I would not consider it settled without evidence: A definition line on the chart and a test that inserts a NULL and expects the documented behaviour.

NULL policy is a pricing decision.

Curated: · Written: · Reviewed:

QA-62DAU in UTC drops 8% on Sunday evening in the US. Product thinks usage fell. What clock should DAU use?(show answer)

I would start time zone in daily active from the decision it informs, not from the first feature request that arrived.

DAU depends on an explicitly chosen reporting clock; UTC, business timezone, and user-local dates answer different questions.

Concretely, use a stable business timezone for globally additive reporting, and user-local dates for behavioral analyses where local day matters. Label both and never splice them into one series.

The reason for that specificity is a failure I have seen: A "Sunday dip" triggered a push campaign that hit users at 01:00 local. Uninstalls +0.6% that week.

UTC versus local Sunday.

ClockSunday evening USDip?
UTCsplit across days8% fake
local dateone daynone
push at 01:00—+0.6% uninstall

I would not consider it settled without evidence: A DAU definition with the clock, plus a plot that the Sunday dip vanishes in local time.

A UTC midnight is not a user midnight.

Curated: · Written: · Reviewed:

QA-63Paid signups are up 35%. Completed onboarding is flat. 28% of "users" fail a cheap bot heuristic. What do you do to CAC?(show answer)

This is a place where a green launch dashboard and a correct handling of bot filtering in acquisition are not the same event.

Acquisition metrics that include bots invent cheap users and lie about CAC; filter with a documented heuristic and a review path.

Concretely, exclude known bots from signup KPI, keep a raw count for fraud. Recompute CAC on eligible humans.

The reason for that specificity is a failure I have seen: CAC looked $18 versus $31 true. The channel was scaled 2×. True CAC went to $44 before anyone noticed.

Signups with 28% bots.

CountnCAC
raw+35%$18
humanflat$31
scaled 2×—$44

I would not consider it settled without evidence: A weekly bot share on the acquisition dashboard, and CAC that uses the filtered denominator.

A bot does not have a lifetime value.

Curated: · Written: · Reviewed:

QA-64The experiment randomises page views. The metric is conversion per user. Why is the test lying, and what unit do you randomise?(show answer)

My answer to A/B assignment unit begins at the constraint: if I cannot name the opportunity cost and the kill criterion, I do not have a product decision.

Assignment must prevent treatment crossover and the analysis must account for the randomization unit; the metric grain need not always equal the assignment grain.

Concretely, for a persistent experience measured by user conversion, assign users so one person cannot encounter both variants. If assignment is by account, session, or page view, use an estimand and variance calculation appropriate to that unit and its clustering.

The reason for that specificity is a failure I have seen: View-level assignment put power users in both arms. The "winner" was +4% and vanished in a user-level replay.

View versus user.

UnitPower usersReplay
page viewboth arms+4% fake
userone arm0%
SRMfailstop

I would not consider it settled without evidence: A unit table: assignment = user, metric = user conversion, plus a SRM check.

If one person can be in both arms, you do not have a persistent user-level experiment.

Curated: · Written: · Reviewed:

QA-65A nominal 50/50 test shows 53,000/47,000 exposures after 10 days. What must you calculate before reading the +2% lift?(show answer)

I would treat sample ratio mismatch as a choice under capacity rather than as a slide that lists everything as Must.

An observed imbalance is SRM only when it is statistically inconsistent with the expected allocation at the actual sample size.

Concretely, run the pre-specified SRM test, then investigate bucketing, eligibility, logging, redirects, and bots if it fails. Restart when exposure was corrupted; if only downstream logging was defective and recoverable, document and repair the analysis.

The reason for that specificity is a failure I have seen: They shipped the +2%. The 53/47 was a redirect that dropped control on mobile. Mobile conversion actually fell 6%.

53/47.

RatioLift shownTruth
53k/47k, p<0.001+2%invalid
mobile control drop—−6%
SRM gatefailno ship

I would not consider it settled without evidence: An automated SRM gate that fails the experiment at p < 0.001 on expected 50/50.

A broken coin is not a treatment effect.

Curated: · Written: · Reviewed:

QA-66Logged-out mobile and desktop are two users. After login they are one. How do you stop DAU from double-counting the morning commute?(show answer)

The useful question for counting unique users across devices is what still happens when the loudest customer, the empty state, or the excluded cohort shows up.

Identity resolution is a product rule: when anonymous ids merge, historical unique counts must follow a documented policy.

Concretely, pick last-click merge or keep anonymous until login. Apply the same rule in DAU and in experiments. Do not mix.

The reason for that specificity is a failure I have seen: DAU fell 9% the week login improved because more anonymous device pairs merged. Leadership called it contraction and cut acquisition spend even though the underlying people count was flat.

Two devices, one person.

RuleDAU weekReal users
never mergebaseline inflated 9%same
merge on login−9% series breaksame
version and annotateflat underlyingsame

I would not consider it settled without evidence: A merge-rate chart beside DAU so a login change cannot look like growth.

A login is not a new customer; it is a join.

Curated: · Written: · Reviewed:

QA-6720% of mobile events arrive 36 hours late. Yesterday's activation "fell" 6 points this morning. What do you show the board?(show answer)

I would settle late events and backfill against a counter-example first, so the roadmap date has to survive it.

Late events need a stated freeze window and a backfill policy; otherwise every morning is a fake regression.

Concretely, report preliminary vs frozen (e.g. T+48h). Do not compare a frozen day to a 6-hour-old day.

The reason for that specificity is a failure I have seen: A "6-point drop" triggered a rollback of a working onboarding. Two days later the points came back as late events.

T+6h versus T+48h.

As-ofActivationAction
T+6h−6 ptsrollback
T+48h0 ptsnone
labelledpreliminarywait

I would not consider it settled without evidence: A freshness label on every chart: as-of, freeze time, % late.

An incomplete day is not a worse product.

Curated: · Written: · Reviewed:

QA-68A self-serve cohort tool scans 14 TB/query. The bill is $18k this month from 3 PMs. What do you change besides yelling?(show answer)

The judgement in analytics query cost as a product constraint is which metric or contract is a commitment versus a forecast, not which status colour looks calm.

Self-serve analytics is a product with quotas, pre-aggregates, and a cost guardrail — not unlimited warehouse.

Concretely, ship a daily cohort table, cap interactive scan, show $ before run, and keep a sandbox.

The reason for that specificity is a failure I have seen: Finance froze the warehouse for a day. All experiments went dark, including checkout.

14 TB versus a daily table.

PathScanMonth $
raw 14 TB14 TB$18k / 3 people
daily cohort2 GB$400
freeze0experiments dark

I would not consider it settled without evidence: A $ cap per user per day and a 99th percentile scan < 200 GB.

An unbounded query is a budget incident with a SQL editor.

Curated: · Written: · Reviewed:

QA-69You compute "currently premium" from a snapshot captured Monday. A user who churned Tuesday is still premium in Friday's board pack. Why?(show answer)

Where candidates lose the interview on snapshot versus event tables is usually a solution they named before they could name the need.

A snapshot is valid only for its capture time; period-end status requires an event reconstruction or a snapshot captured as of that period end.

Concretely, use subscription events (start/cancel) or as-of snapshots. Label the as-of on the pack.

The reason for that specificity is a failure I have seen: Churn was understated 4 points. Retention looked healthy while revenue fell. The correction was a board surprise.

Tuesday churn, Friday pack.

SourceFriday statusTruth Tuesday
Monday snapshot reused Fridaypremiumstale after Tuesday churn
eventschurnedchurned
board+4 pts fakesurprise

I would not consider it settled without evidence: An as-of timestamp on the slide and a reconciliation to billed revenue.

Latest snapshot is not end-of-period.

Curated: · Written: · Reviewed:

QA-70A prospect's security reviewer asks "is our data encrypted?" Sales wants to paste "yes" into the RFP. What analogy do you use in the meeting, and what business consequence changes if you give the one-word answer?(show answer)

I would answer explaining encryption to a buyer by separating what the dashboard proved from what the user still could not complete.

Three different claims hide under the word encrypted — in transit, at rest, and who holds the key — and a one-word answer collapses them into a contractual promise the architecture cannot keep.

Concretely, walk one analogy across all three layers, say which of the three we actually implement, and name the business consequence: what the buyer may tell their auditor and what they must never write in the RFP.

The reason for that specificity is a failure I have seen: Sales answered "yes, fully encrypted" on a hospital group's RFP. The buyer assumed zero-knowledge; a court audit produced 62,000 message bodies in readable text from our side. The $1.9M renewal closed at $1.2M with a right-to-audit clause and an eight-week remediation plan.

The three claims behind one word.

ClaimAnalogyWhat we doWhat the buyer may say
in transitcash in an armoured van between buildingsTLS 1.3"unreadable on the wire"
at restlocked cabinet; we hold the keyAES-256, keys in KMS"useless if the disk is stolen"
end-to-enda safe only they can opennot implementednothing; never claim it

I would not consider it settled without evidence: Name the one sentence on the security page an auditor can verify — TLS 1.3 in transit, AES-256 at rest, keys in our KMS, staff can read content — and the one sentence the RFP must never contain.

Encrypted is not a yes or no; it is a claim about who holds the key.

Curated: · Written: · Reviewed:

QA-71A deck says "we improved conversion 10% to 11%, a 10% win". Another slide says "+1 percentage point". Which phrase do you standardise on?(show answer)

The product content of percent of a percent is the trade-off and the revisit trigger, not the template that scored it.

Relative percent and percentage points are different claims; mixing them inflates small wins.

Concretely, always show baseline, new value, pp change, and relative change. Prefer pp for conversion rates in decisions.

The reason for that specificity is a failure I have seen: "+10%" (1 pp) was funded as if it were 10 pp. The expected GMV was off by 10×.

10% versus 1 pp.

Phrase10% → 11%Funded as
10% winrelative10× GMV
+1 ppabsolutecorrect
both shownyesdecide

I would not consider it settled without evidence: A slide template that forbids a bare % change on a rate.

Ten percent of 10% is one point, not a miracle.

Curated: · Written: · Reviewed:

QA-72Week-over-week signups are +20% every early September. A new campaign claims the credit. How do you check?(show answer)

Before writing a story I would write what a correct result for seasonality mistaken for a trend looks like for n = 1 user and for the 10^6 who will hit the same path.

Compare to the same period last year and to a holdout; seasonal lifts are not campaign wins.

Concretely, yoY and pre-period slope. If September always pops, the campaign needs a geo or user holdout.

The reason for that specificity is a failure I have seen: The campaign took credit for a school-year spike. It was scaled in November and produced +2% at 8× spend.

Every September.

YearEarly SepCampaign?
2024+19%no
2025+20%yes, claimed
Nov scale+2%8× spend

I would not consider it settled without evidence: A 3-year September overlay on the same chart as the campaign.

September is not your funnel.

Curated: · Written: · Reviewed:

QA-73Overall conversion rose 1 pp but every region fell. What happened, and which number do you act on?(show answer)

The first thing I would pin down about Simpson's paradox in a segment rollup is which user outcome or explicit non-commitment it actually changes.

A mix shift can move the aggregate opposite to every segment; act on the segments and the mix, not the flattering total.

Concretely, show mix (share of traffic) and rate per region. If mix moved to a high-converting region, say so.

The reason for that specificity is a failure I have seen: The aggregate "win" hid a product regression in all four regions. Mix had shifted to US enterprise.

All regions down, total up.

RegionRate ΔTraffic share
each of 4−1 to −3 ppshifted
US entstill high20% → 45%
overall+1 ppmix, not product

I would not consider it settled without evidence: A mix-rate table on every rollup that could change a ship decision.

The total is not a fifth region.

Curated: · Written: · Reviewed:

QA-74Two quarters of roadmap are committed. Platform work (a migration off a deprecated dependency) and a revenue feature both need the same two engineers; the migration has no visible customer impact, and the feature has a named prospect waiting. How do you decide, what do you write down, and what do you tell the prospect?(show answer)

I would start deprecation migration versus revenue feature from the decision it informs, not from the first feature request that arrived.

A deprecation is a dated obligation with a cost curve, not invisible debt: price it against the deadline and blast radius it retires, date the prospect's revenue the same way, and break one date on purpose instead of both by accident.

Concretely, write the migration as a dated obligation — vendor EOL date, blast radius if it slips, engineer-weeks — and the feature as dated revenue with the commitment date and the cost of moving it. Put both on one timeline, then sequence or split, and record the deferral with an owner and a review date.

The reason for that specificity is a failure I have seen: We shipped a named $240k prospect's feature and pushed the migration off a deprecated auth SDK to the next quarter. The vendor's EOL landed in week 11; the pinned fork needed 3 engineer-weeks of emergency patching and auth p99 went from 180 ms to 900 ms for 6 days. The prospect's security review slipped and it closed two months later anyway — the revenue was never as dated as the EOL was.

Same two engineers, one quarter, both items dated.

ItemDate it must beatEngineer-weeksCost of missing it
SDK migrationvendor EOL, week 912unsupported fork, ~3 ewk emergency patching
Prospect featurecontract date, week 1310$240k deal slips a quarter
Capacity—16 (2 eng × 8 weeks after KTLO/on-call)—

12 + 10 > 16, so one date breaks this quarter. The record says which one, why, and when it is re-examined — and the prospect hears the dated alternative, not a hopeful one.

I would not consider it settled without evidence: Name the one dated constraint that decides — the vendor EOL or support cutoff — and show both items' engineer-weeks and deadlines on a single timeline, with the deferral's owner and review date written down.

Every deferral carries a price and a review date, or it is just a hope.

Curated: · Written: · Reviewed:

QA-75Support wants a global search across all tenants. Engineering sharded by tenant_id. What do you tell support the product can and cannot do?(show answer)

This is a place where a green launch dashboard and a correct handling of sharding as a product limitation are not the same event.

A shard key is a product boundary: cross-shard queries are slow or forbidden unless you budget a new access path.

Concretely, either offer per-tenant search as the product, or fund a search index that is not the OLTP shard. Do not promise global search on sharded OLTP.

The reason for that specificity is a failure I have seen: A "quick" scatter-gather locked 40 shards for 8 minutes; checkout p95 hit 12s. The search was used 30 times/day.

Scatter-gather versus an index.

Path40 shardsCheckout p95
scatter OLTP8 min lock12s
per-tenant onlynoneheld
search indexasyncheld

I would not consider it settled without evidence: A documented cannot: global OLTP search, with the alternative index on the roadmap if needed.

Shard key is UX, not just ops.

Curated: · Written: · Reviewed:

QA-76A user edits a doc and immediately opens it on their phone; the phone replica is 3s behind. What product behaviour do you specify?(show answer)

My answer to replication lag as a user-visible read begins at the constraint: if I cannot name the opportunity cost and the kill criterion, I do not have a product decision.

Cross-device read-your-writes requires a consistency token or account-level write watermark that the second device can receive; session stickiness on the first device is insufficient.

Concretely, after the write, propagate a version or consistency token through account sync and route reads until that version is visible, or read from an authoritative source. Otherwise show a syncing state.

The reason for that specificity is a failure I have seen: Users thought saves failed and duplicated sections. Support's "refresh" macro was 200 tickets/week.

Phone 3s behind.

ReadSees edit?Tickets/week
lagged replicano200
token/primaryyes~0
syncing bannerwaitlow

I would not consider it settled without evidence: A read-your-writes test across two devices in the launch suite.

A replica is a product if the user can see it.

Curated: · Written: · Reviewed:

QA-77A write-back cache promises snappy saves. The node dies before flush. What does the product owe the user?(show answer)

I would treat write-back cache as data-loss UX as a choice under capacity rather than as a slide that lists everything as Must.

Write-back is a durability trade: if you offer it, you must say what can vanish and offer a safer path for high-value writes.

Concretely, use write-through for billing and legal text. If write-back is used for drafts, show "saved to device, syncing" and recover from local.

The reason for that specificity is a failure I have seen: A cache node died with 6 minutes of unflushed edits for 900 docs. There was no local recovery. Trust NPS fell 11 points.

Flush versus vanish.

WritePolicyNode death
billingwrite-throughsafe
draft write-back6 min loss900 docs
local + syncrecoverok

I would not consider it settled without evidence: A durability table by write type, and a chaos test that kills a cache node.

Snappy is not the same as saved.

Curated: · Written: · Reviewed:

QA-78Welcome email is fired from a queue with at-least-once delivery. 7% of new users get two welcomes. Is that a platform bug or a product bug?(show answer)

The useful question for at-least-once emails is what still happens when the loudest customer, the empty state, or the excluded cohort shows up.

At-least-once delivery requires a stable idempotency key for the logical message and a send path that closes the post-send crash window; a time-window dedupe is not enough.

Concretely, create one welcome-message intent per signup in a transactional outbox and enforce uniqueness on that intent ID. Every retry uses the same provider idempotency key, and the worker reconciles the provider receipt before resending after an ambiguous timeout. If the provider offers no idempotency contract, do not claim exactly-once delivery: make the discount itself single-redemption and monitor duplicates. Do not turn off retries and lose the 2% who need them.

The reason for that specificity is a failure I have seen: Double welcomes included two discount codes; 7% redeemed both. Margin on that cohort went negative.

Two codes.

SendUsersTwo codes
no receipt7%yes
intent + provider key~0no
retries offlose 2%no mail

I would not consider it settled without evidence: A crash test after provider acceptance but before the local acknowledgement still produces one delivered welcome, plus a duplicate-send metric and a single-redemption constraint on the discount.

The queue is allowed to retry; the inbox is not a dumpster.

Curated: · Written: · Reviewed:

QA-79A chatbot "must stick to one box" for memory. Horizontal scaling is now sticky sessions. What product alternative do you offer?(show answer)

I would settle session affinity as a scale tax against a counter-example first, so the roadmap date has to survive it.

Sticky sessions are a product choice that taxes failover and scale; prefer externalised memory unless the constraint is real.

Concretely, put conversation state in a store keyed by session id. If stickiness remains, document the blast of a node death.

The reason for that specificity is a failure I have seen: A node death dropped 14 minutes of chats for 8k sessions. CSAT on that hour was 1.9/5.

Sticky versus store.

MemoryNode deathLost turns
in-process sticky14 min8k sessions
session store00
CSAT1.9recovered

I would not consider it settled without evidence: A design that survives killing one app node with 0 lost turns.

Memory in the box is a product outage waiting for a deploy.

Curated: · Written: · Reviewed:

QA-80The app retries 429s with no jitter and no message. Users mash the button. What is the product design for overload?(show answer)

The judgement in 429 as a user-facing state is which metric or contract is a commitment versus a forecast, not which status colour looks calm.

Overload is a UX: explain wait, backoff with jitter, and disable the button; silent retry storms make the outage worse.

Concretely, show "we're busy, try in Ns" from Retry-After. Cap client retries. Offer an offline queue for safe actions.

The reason for that specificity is a failure I have seen: A 12-minute 429 period became 40 minutes because clients amplified traffic 6×. Paid conversion hit 0 for the window.

Mash versus Retry-After.

ClientTraffic during 429Window
mash + retry6×40 min
Retry-After + disable~1×12 min
conversion0then recover

I would not consider it settled without evidence: A client policy in the PRD and a test that a 429 storm does not multiply QPS >1.2×.

A spinner is not a rate-limit policy.

Curated: · Written: · Reviewed:

QA-81Marketing says "your data stays in the EU". The analytics pipeline is in us-east-1. What do you change before the campaign?(show answer)

Where candidates lose the interview on multi-region as a promise is usually a solution they named before they could name the need.

A residency promise is a product and legal contract; every pipeline, log, and support tool is in scope.

Concretely, inventory flows. Either move analytics, tokenize, or change the copy. Do not run the campaign on a lie.

The reason for that specificity is a failure I have seen: A DPA audit found US logs of EU tickets. The campaign was pulled; two deals stalled ($2.4M).

Copy versus pipeline.

FlowRegionPromise holds?
appEUyes
analyticsus-east-1no
campaign—pulled

I would not consider it settled without evidence: A data-flow diagram signed by legal and the TPM before the sentence is used.

Stays in the EU means the warehouse too.

Curated: · Written: · Reviewed:

QA-82Relevance wants a 3-stage ranker (p95 900ms). The search SLO is 200ms. Who wins, and how?(show answer)

I would answer search relevance versus latency budget by separating what the dashboard proved from what the user still could not complete.

Latency is a product budget; a ranker that blows the SLO is a different product than the one on the page.

Concretely, keep a fast candidate stage inside 200ms; run expensive ranking on a subset or asynchronously. Do not silently miss the SLO.

The reason for that specificity is a failure I have seen: They shipped the 900ms ranker. Bounce on search rose 9 points; revenue from search fell $60k/week.

NDCG versus bounce.

Rankerp95Search revenue
3-stage900ms−$60k/w
200ms candidates180msheld
async polishextraoptional

I would not consider it settled without evidence: A p95 search panel on the same scorecard as NDCG, with a veto.

A better list that nobody waits for is not better.

Curated: · Written: · Reviewed:

QA-83The next-page request returns an expired-cursor problem after two hours. What should the product do?(show answer)

The product content of cursor expiry as UX is the trade-off and the revisit trigger, not the template that scored it.

Cursor lifetime is a product rule: expire, then recover by restarting the query with a message — do not dump a raw auth failure.

Concretely, return a documented invalid/expired-cursor response, commonly 400 or 410 with a stable problem type. Reserve 401 for authentication failures, and let the UI restart the query with a clear message.

The reason for that specificity is a failure I have seen: 401 JSON flashed in the app. 8% of users reported "hacked". Trust tickets spiked 3 days.

401 versus refresh.

Next page at 2hUser seesTickets
raw 401hacked?spike
refresh resultsnew page 1none
TTL 30 minmessagenone

I would not consider it settled without evidence: A UX state for expired cursor in the PRD, tested at TTL+1s.

An expired cursor is not an auth incident unless you display it like one.

Curated: · Written: · Reviewed:

QA-84A 10-minute video export has a spinner and no percent. 22% of users start a second export. What do you add?(show answer)

Before writing a story I would write what a correct result for async job progress UX looks like for n = 1 user and for the 10^6 who will hit the same path.

Long jobs need progress, a bound, cancel, and notify; a spinner invites duplicate work.

Concretely, show stages or measured progress, an ETA range, cancel, and completion notification. Deduplicate retries using a job-intent idempotency key and request fingerprint, not merely user plus source.

The reason for that specificity is a failure I have seen: Second exports filled the worker pool. p95 went from 10 min to 47 min. The original spinner never finished for many.

Spinner versus stages.

UISecond startp95
spinner22%47 min
stages + notify2%10 min
cancel—pool healthy

I would not consider it settled without evidence: A duplicate-start rate <2% after progress ships.

A spinner is not progress.

Curated: · Written: · Reviewed:

QA-85Enterprise wants "our data never touches another tenant's disk". You run a shared DB with row-level security. What do you sell?(show answer)

The first thing I would pin down about tenant isolation as a SKU is which user outcome or explicit non-commitment it actually changes.

Isolation is a product SKU: logical RLS is not the same as dedicated disk; sell what you can prove.

Concretely, define isolation dimensions separately: tenant authorization, database or schema, encryption keys, compute, storage volume, backups, logs, and support access. Sell only the specific dedicated resources the architecture and cloud contract prove. Do not let sales paste "never touches".

The reason for that specificity is a failure I have seen: A pentest found a support tool that queried across tenants. The "never touches" clause triggered a $1.5M renegotiation.

Shared versus dedicated.

SKUDiskSales sentence
standardshared + RLSlogical isolation
dedicated data planedocumented dedicated resourcesuse only verified isolation claims
mixed copy—$1.5M

I would not consider it settled without evidence: A one-pager of isolation levels with test evidence, used as the only sales language.

Row-level security is not a separate disk.

Curated: · Written: · Reviewed:

QA-86A noisy neighbour on shared compute made a customer's p95 8× for 2 hours. They want dedicated "for free because we are large". How do you package it?(show answer)

I would start shared versus dedicated as packaging from the decision it informs, not from the first feature request that arrived.

Noisy-neighbour risk is a packaging and SLO choice: shared has a published SLO; dedicated is paid isolation.

Concretely, offer dedicated at a price. Improve noisy-neighbour controls on shared. Do not gift dedicated to whoever shouts.

The reason for that specificity is a failure I have seen: Four "free dedicated" exceptions ate the margin on the segment. Shared customers still had the 8× incident.

Free dedicated exceptions.

PathMarginShared p95
4 free dedicatedgonestill 8×
paid dedicatedheldoptional
better shared limitsheldbetter

I would not consider it settled without evidence: A price, an SLO, and an exception log with expiry.

Large is not a SKU; dedicated is.

Curated: · Written: · Reviewed:

QA-87A customer wants 10× quota "for a launch week". The default is 100 rpm. What is the product process?(show answer)

This is a place where a green launch dashboard and a correct handling of API quota as a commercial lever are not the same event.

Quota changes are commercial and capacity decisions with an end date, not a Slack yes.

Concretely, check headroom, price the burst, set an expiry, and instrument. A launch week is a time-bound SKU.

The reason for that specificity is a failure I have seen: An open-ended 10× stayed on. That customer became 40% of API CPU. Other tenants' p95 doubled.

Launch week 10×.

GrantExpiryCPU share
Slack yesnone40%
7-day burstyesspike then down
others' p952×recovered

I would not consider it settled without evidence: A ticket with expiry and a graph of CPU share.

A burst without an end date is a new plan given away.

Curated: · Written: · Reviewed:

QA-88A new CSV importer saves sales demos but creates 40 tickets/week. How does that enter prioritisation?(show answer)

My answer to support load as cost of a feature begins at the constraint: if I cannot name the opportunity cost and the kill criterion, I do not have a product decision.

Support hours are product cost and belong in RICE effort and in the kill criterion.

Concretely, price 40×15 min = 10h/week. Either fix the empty/error states until tickets <5/week or kill the importer.

The reason for that specificity is a failure I have seen: The importer stayed. Ten hours per week is about 520 hours per year—roughly 0.25 FTE before escalation and management overhead—and exceeded the original build effort.

40 tickets/week.

StateTickets/wHours/w
launch4010
after errors51.2
ignored40×52 tickets~0.25 FTE at 15 min/ticket

I would not consider it settled without evidence: A ticket tag on the feature and a weekly hours chart on the Now card.

A feature that hires a support person is a staffing decision.

Curated: · Written: · Reviewed:

QA-89Checkout is down. Status is green. Twitter is not. What is the product owner's job in the first 15 minutes?(show answer)

I would treat incident communication as product as a choice under capacity rather than as a slide that lists everything as Must.

Incident communication is a product surface: knowns, unknowns, impact, next update time — early, not after a perfect RCA.

Concretely, if user-visible, post in 15 minutes. Do not wait for root cause. Match cadence to impact. Separate internal blame.

The reason for that specificity is a failure I have seen: Silence for 70 minutes. Two large customers posted. Trust recovery took a quarter.

70 minutes of green.

tPublicCustomers
0–70mgreenTwitter
15m policyknowns/unknownsheld
RCA waitsilencequarter to recover

I would not consider it settled without evidence: A clock: first public update ≤15 min for Sev-1 user-facing.

Green status during a dead checkout is a second incident.

Curated: · Written: · Reviewed:

QA-90Engineers want a quarter to rewrite the billing module. How do you put that on a product roadmap without making it a blank cheque?(show answer)

The useful question for technical debt as a roadmap item is what still happens when the loudest customer, the empty state, or the excluded cohort shows up.

Debt is sequenced by user/risk outcome it unlocks or harm it prevents, with a bound, not by age or disgust.

Concretely, name the incidents, the kill-the-rewrite criterion, and the user-visible result (correct invoices, faster refunds). Time-box. Keep a feature slice if the rewrite slips.

The reason for that specificity is a failure I have seen: A 14-week rewrite slipped to 26 with no user change. Two product bets were cancelled to feed it.

Rewrite versus a defect bar.

FrameWeeksUser metric
rewrite billing26none
invoice defects <0.3%62.1%→0.4%
unbounded26two bets died

I would not consider it settled without evidence: A Now card: "reduce invoice defects from 2.1% to <0.3% in 6 weeks" rather than "rewrite billing".

A rewrite with no user metric is a hobby with a budget.

Curated: · Written: · Reviewed:

QA-91A platform team wants all product squads on a new CI by Friday. Product has a launch Friday. Who has the right to say not now?(show answer)

I would settle platform versus product team against a counter-example first, so the roadmap date has to survive it.

Platform adoption is a product decision for each squad: the platform team proposes, the product owner of the launched journey can refuse a Friday collision.

Concretely, sequence migrations off launch weeks. Measure adoption as an outcome with a freeze window around Sev-1 journeys.

The reason for that specificity is a failure I have seen: CI migrated Friday; the launch pipeline was a different YAML. The launch missed a 2-hour retail window ($90k).

Friday collision.

OwnerFridayRetail window
platformmigratemissed $90k
productfreeze CIhit
safety exceptionmigrateonly if broken

I would not consider it settled without evidence: A freeze calendar the platform team cannot override without the journey's owner.

Platform urgency does not outrank a named user launch unless a safety issue does.

Curated: · Written: · Reviewed:

QA-92Sales wants to discount 40% instead of removing a feature. What is the difference, and which should you prefer when the cost is COGS?(show answer)

The judgement in pricing versus packaging is which metric or contract is a commitment versus a forecast, not which status colour looks calm.

Packaging defines what is included; pricing defines what is charged. Variable COGS must be covered by the combination of package limits, allowances, overages, and price.

Concretely, for SMS, compare a lower tier without SMS, a priced allowance with overages, and a higher all-inclusive price. Reject a 40% discount if expected usage makes contribution margin negative.

The reason for that specificity is a failure I have seen: 40% off kept SMS. That cohort's margin was −12%. The discount became the street price.

Discount versus remove SMS.

LeverSMS includedMargin
40% offyes−12%
lower packageno+18%
street pricefollowed discountstuck

I would not consider it settled without evidence: A contribution-margin table by package, not a discount list.

A discount on a costly feature is a gift of COGS.

Curated: · Written: · Reviewed:

QA-93A confirm-shaming modal lifts conversion 8%. Unsubscribe is two extra clicks. Do you ship it?(show answer)

Where candidates lose the interview on dark patterns versus the metric is usually a solution they named before they could name the need.

A metric that rises through coercion or hidden choice is not a product win; guardrails include complaint rate, unsubscribe, and regulation.

Concretely, ban confirm-shaming in the design system. Measure complaints. If legal risk (FTC dark patterns) exists, it is a Must-not.

The reason for that specificity is a failure I have seen: The modal shipped. Conversion +8%, chargebacks +2.4×, an attorney letter in 3 weeks.

Shame modal.

Ship?ConversionChargebacks
yes+8%+2.4×
no0held
letter—3 weeks

I would not consider it settled without evidence: A pre-commit: no ship if unsubscribe friction increases, regardless of conversion.

If the user has to be tricked, the metric is the trap.

Curated: · Written: · Reviewed:

QA-94RICE reach is 2×10^6. The feature is unusable with a screen reader (0.8% of users, high harm). How does reach change the priority?(show answer)

I would answer reach that excludes a minority by separating what the dashboard proved from what the user still could not complete.

Aggregate reach can hide severe harm to a small group; accessibility and safety are constraints, not a rounding error in reach.

Concretely, apply a floor: the journey must work for assistive tech or it is not Now. Do not average 0.8% to zero.

The reason for that specificity is a failure I have seen: They shipped to 2×10^6. 16k users were blocked. A regulator complaint named the RICE sheet.

2×10^6 minus 16k blocked.

ViewnShip?
aggregate reach2e6yes
screen reader16k blockedveto
RICE only—complaint

I would not consider it settled without evidence: A segmented harm table beside RICE, with a veto for exclusion of a legal class.

Reach is not permission to exclude.

Curated: · Written: · Reviewed:

QA-95A customer wants to turn off MFA "just for the pilot". Who can accept that, and what expires?(show answer)

The product content of security exception as a product record is the trade-off and the revisit trigger, not the template that scored it.

A control exception is a product and risk record: owner, scope, compensating control, expiry — not a silent config.

Concretely, named risk owner, expiry (e.g. 30 days), monitoring, no expansion via a label. Prefer a separate pilot tenant.

The reason for that specificity is a failure I have seen: The exception had no expiry. 14 months later 8k production users had MFA off. An incident used one of those accounts.

Pilot MFA off.

RecordExpiry14 months later
Slack yesnone8k off
register30dforced on
incidentused an off accountyes

I would not consider it settled without evidence: An exception register with age alerts at 14 days.

A pilot switch that never expires is the product.

Curated: · Written: · Reviewed:

QA-96A legacy export is used by 0.4% of accounts but 40% of support time. How do you retire it without a surprise?(show answer)

Before writing a story I would write what a correct result for retiring a feature looks like for n = 1 user and for the 10^6 who will hit the same path.

Retirement is a product change: users, dates, migration, comms, data retention, then delete — not a quiet code removal.

Concretely, announce, offer a path, a date, a usage chart, then disable, then delete. Preserve audit where required.

The reason for that specificity is a failure I have seen: A quiet delete broke a partner's nightly job. The 0.4% included a $2M account. Restore from backup took 2 days.

0.4% with a $2M account.

StepRemaining$2M account
quiet delete0broken
announce + date0.4%migrated
T-14 listnamedcontacted

I would not consider it settled without evidence: A named list of remaining users at T-14 days, contacted.

Low percentage can still be high revenue.

Curated: · Written: · Reviewed:

QA-97You dual-write to old and new billing for 6 weeks. The two systems disagree on 1.8% of invoices. Who is source of truth for the user?(show answer)

The first thing I would pin down about dual-write migration as a product is which user outcome or explicit non-commitment it actually changes.

During dual-write, the product must name one source of truth the user sees, plus a reconciliation SLO.

Concretely, read from one system. Reconcile daily. Bound the 1.8% with a user-visible correction policy. Do not show both numbers.

The reason for that specificity is a failure I have seen: The app showed new; PDF showed old. 1.8% of customers paid the wrong amount. Collections chaos for 3 weeks.

App versus PDF.

SurfaceSoT1.8%
bothnonewrong payments
app onlynewreconcile PDF
cutover bar<0.1%then kill old

I would not consider it settled without evidence: A SoT line in the PRD and a daily mismatch <0.1% before cutover.

Two truths is a billing incident.

Curated: · Written: · Reviewed:

QA-98The conference is Tuesday. The checklist is 60% amber. Do you go live on stage?(show answer)

I would start launch date versus launch checklist from the decision it informs, not from the first feature request that arrived.

A date is a constraint; the checklist is the evidence. If must-gates are amber, you demo a slice or you miss — you do not redefine Done.

Concretely, identify Must gates (security, billing, rollback). If any Must is amber, do not GA. A recorded demo is a product option.

The reason for that specificity is a failure I have seen: They GA'd for the stage. A billing Must was amber. 48 hours of free usage leaked ($70k COGS) and a status-page incident during the keynote.

Amber Must on Tuesday.

ChoiceStageBilling
GA anywaylive$70k leak
recorded demolivesafe
slip GAmisssafe

I would not consider it settled without evidence: A Must-only checklist with owners, reviewed Monday.

A conference is not a security control.

Curated: · Written: · Reviewed:

QA-99Two weeks after GA, nobody has looked at the scorecard. The feature is "done". What did the team skip, and what meeting do you add?(show answer)

This is a place where a green launch dashboard and a correct handling of post-launch learning loop are not the same event.

Launch is the start of a learning loop: compare outcomes and guardrails to the hypothesis, then scale, iterate, or kill.

Concretely, a scheduled review (e.g. day 14) with the scorecard, a decision, and a written follow-up. Done is not a status that ends measurement.

The reason for that specificity is a failure I have seen: Day 40, a guardrail (refunds) had been red since day 3. The team noticed when finance called. The kill was 5 weeks late.

Day 14 versus day 40.

ReviewRefundsKill
nonered since d3d40 finance
d14seend14
"done"ignoredzombie

I would not consider it settled without evidence: A calendar invite created when Now was approved, with the scorecard link.

A feature without a day-14 decision is a zombie with metrics.

Curated: · Written: · Reviewed:

QA-100Keep-the-lights-on is "whatever is left after features". On-call overtime is 18 hours/week. Where does KTLO sit on a TPM roadmap?(show answer)

My answer to KTLO on the same roadmap as features begins at the constraint: if I cannot name the opportunity cost and the kill criterion, I do not have a product decision.

Operational load is capacity, not leftover: reliability, security, and support work compete with features on the same envelope.

Concretely, put a KTLO slice on the roadmap (e.g. 20%) with a trigger to raise it when overtime or error budget burns. Features do not get to assume 100%.

The reason for that specificity is a failure I have seen: Features took 100% for a quarter. Overtime hit 18h/week; two engineers left. The next quarter's "velocity" collapsed 40%.

100% features.

MixOvertime h/wNext-quarter velocity
100% features18−40%
80/20 KTLO4held
no timesheetsunknownsurprise attrition

I would not consider it settled without evidence: A capacity pie that includes KTLO hours from last month's timesheets, reviewed with engineering.

Leftover is not a plan for keeping the lights on.

Curated: · Written: · Reviewed: