Top 100 Product Manager Interview Questions and Answers
The questions most likely to actually come up in your Product Manager interview, ranked by likelihood — with detailed, senior-level answers covering what an interviewer is really listening for.
Curated: · Written: · Reviewed:
QA-1A large customer asks for a specific feature by name. How do you decide whether to build it?(show answer)
I would start the customer problem behind a feature request from the customer problem it is supposed to change, not from the feature already on the roadmap.
A feature request is a customer's proposed solution, and the job is to recover the problem underneath it before pricing the work. Building the request as stated hands product design to whoever wrote the email.
Concretely, ask what the customer does today, how often, and what it costs them when it goes wrong. Restate the problem in one sentence and check that the customer agrees with the restatement. Only then compare the requested solution against two cheaper ones that would solve the same problem.
The reason for that specificity is a failure I have seen: A billing team built the CSV export three enterprise accounts asked for by name, and 11 weeks later 2 of the 3 still reconciled by hand because the real problem was a missing invoice identifier, not the file format.
One request, three readings of the problem.
| Reading | Work | Reconciles the invoice |
|---|---|---|
| build the named CSV export | 11 weeks | 1 of 3 accounts |
| add the invoice identifier | 2 weeks | 3 of 3 accounts |
| do nothing and document | 0 weeks | 0 of 3 accounts |
I would not consider it settled without evidence: Read the problem statement back to the requesting customer and confirm they accept it without adding a new requirement.
The request is the customer's hypothesis, not the specification.
Curated: · Written: · Reviewed:
QA-2How do you write a problem statement that an engineering team can argue with?(show answer)
The first thing I would settle about separating a problem from its solution is what evidence would make me drop the idea.
A problem statement names who is affected, what they are unable to do, how often it happens, and what it costs, with no verb describing a build. A statement that already contains the solution cannot be falsified by discovery.
Concretely, write the statement in the customer's vocabulary with a frequency and a cost attached. Remove every noun that names a screen, an integration, or a technology. Circulate it to two engineers and ask them to propose a solution you did not have in mind.
The reason for that specificity is a failure I have seen: A team opened a quarter with the statement that the product needed a notification centre, and after 6 weeks of building it discovered that 78 percent of the missed events were escalations that should never have needed a human to notice them.
The same finding stated two ways.
| Statement | Contains a solution | Solutions proposed in review |
|---|---|---|
| we need a notification centre | yes | 1 |
| admins miss 4 escalations a week and learn from the customer | no | 5 |
| escalations should auto-assign | yes | 1 |
I would not consider it settled without evidence: Hand the statement to an engineer who was not in discovery and check they can propose an approach you had not already chosen.
If nobody can propose a different solution, you wrote a plan and called it a problem.
Curated: · Written: · Reviewed:
QA-3You have time for eight customer interviews before a planning cycle. How do you pick who to talk to?(show answer)
With choosing who to interview in discovery, I would separate what we believe from what we have actually observed.
Discovery is a sampling problem, and the sample that reaches a PM by default is the one that already complains loudly. The interview list has to include people who solved the problem another way and people who left.
Concretely, split the list across current heavy users, current light users, customers who churned in the last two quarters, and prospects who evaluated and chose something else. Recruit from product usage data rather than from the account team's suggestions. Cap any single segment at a third of the interviews.
The reason for that specificity is a failure I have seen: A workflow product ran 12 interviews sourced entirely from its customer advisory board, shipped the roadmap those 12 endorsed, and the following quarter churn concentrated in the self-serve tier, which had 0 seats on that board and 61 percent of the accounts.
Two interview lists over the same eight slots.
| Source | Slots | Segments covered | Churned accounts heard |
|---|---|---|---|
| advisory board | 8 | 1 | 0 |
| usage-sampled | 8 | 4 | 2 |
| inbound complaints | 8 | 2 | 0 |
I would not consider it settled without evidence: Compare the interview list against the revenue and account mix and confirm no segment is represented more than a third of the time.
The customers easiest to reach are the ones whose problems you already know.
Curated: · Written: · Reviewed:
QA-4How do you ask about a problem without teaching the customer the answer you want?(show answer)
I would answer interview questions that do not lead the witness by naming the decision it feeds and who has to live with that decision.
A question about the past returns behaviour, and a question about the future returns a courtesy. Asking whether someone would use a feature reliably produces a yes that predicts nothing.
Concretely, ask what they did the last time the problem occurred, in order, with dates. Ask what they tried before that and why they stopped. Keep every question about an event that already happened, and save the concept test for the end where it cannot contaminate the account.
The reason for that specificity is a failure I have seen: A team asked 30 users whether they would use a scheduling assistant and 27 said yes, then launched it to a 4 percent adoption rate because only 3 of those users had ever rescheduled anything twice in a month.
Two question forms against the same population.
| Question | Yes rate | Adoption at 30 days |
|---|---|---|
| would you use a scheduling assistant | 27 of 30 | 4 percent |
| walk me through the last reschedule | n/a | identified 3 real users |
| how often did you reschedule last month | 3 of 30 twice or more | n/a |
I would not consider it settled without evidence: Count the questions in the guide that ask about a specific past event and confirm they outnumber the hypothetical ones.
People predict their behaviour badly and report it well.
Curated: · Written: · Reviewed:
QA-5When is a jobs-to-be-done framing useful, and when does it get in the way?(show answer)
My approach to jobs to be done as a framing device starts with the outcome the business is buying rather than the output the team ships.
The framing is useful when it widens the competitive set to include the non-software workaround the customer uses today. It is a liability when it becomes a vocabulary exercise that renames existing features without changing a decision.
Concretely, state the job as a situation, a motivation, and an expected outcome, then list what the customer hires today including spreadsheets, agencies, and doing nothing. Compare your product against those alternatives on the dimension the customer actually optimises. Drop the framing the moment it stops changing the roadmap.
The reason for that specificity is a failure I have seen: A reporting tool ran a job-story workshop that produced 40 restated feature names, and the quarter's roadmap was identical to the one written before the workshop, which cost 3 days of six people's time.
What the customer hires today.
| Alternative | Cost to customer | Switching friction |
|---|---|---|
| spreadsheet plus analyst | 6 hours a week | none |
| agency retainer | 4,000 a month | contract |
| our product | 89 a month | data migration |
I would not consider it settled without evidence: Check whether the job statement changed at least one roadmap decision, and drop the exercise if it did not.
A framing earns its keep by changing a decision, not by renaming a backlog.
Curated: · Written: · Reviewed:
QA-6An executive asks how big this opportunity is. How do you answer without inventing a number?(show answer)
For market sizing that survives scrutiny, I would write down the assumption that has to hold before I argue about the solution.
A defensible size is built bottom-up from countable units and a price, with every assumption exposed so a reviewer can attack one of them. A top-down percentage of an analyst's market figure is unfalsifiable and therefore useless in a decision.
Concretely, count the addressable accounts from a source you can name, multiply by the seats or transactions each one would plausibly buy, and multiply by a price you have actually charged. Show the three assumptions separately and give a range rather than a point. State which assumption the estimate is most sensitive to.
The reason for that specificity is a failure I have seen: A team justified a bet with 1 percent of a 12 billion analyst market, and the launch reached 340,000 in its first year because only 2,100 accounts in that market had the integration the product required.
Bottom-up against top-down for the same bet.
| Method | Estimate | Assumption that breaks it |
|---|---|---|
| 1 percent of 12B | 120M | none stated |
| 2,100 accounts x 8 seats x 1,200 | 20M | account count |
| first-year realistic | 0.3M to 2M | integration coverage |
I would not consider it settled without evidence: Show the estimate to a skeptical reviewer and confirm they can name which single assumption would halve it.
A number nobody can attack is a number nobody should trust.
Curated: · Written: · Reviewed:
QA-7How do you decide which segments your product should serve first?(show answer)
I would treat segmentation that changes a decision as a question about behaviour, so I would look for the behaviour before the opinion.
A useful segment differs in the behaviour the product must support, not only in company size or industry label. Segments that behave identically are reporting categories rather than product decisions.
Concretely, cut the user base by the workflow they run and the frequency they run it at, and check whether the cuts differ in activation, retention, and support load. Keep a cut only if it would change what you build next. Name the segment you are deliberately not serving this year.
The reason for that specificity is a failure I have seen: A product reported by industry vertical for 5 quarters, and a later cut by weekly active editors showed retention of 71 percent against 22 percent between two groups that the vertical view had averaged into one 46 percent line.
Two cuts of one cohort.
| Cut | Groups | 90-day retention spread |
|---|---|---|
| by industry | 6 | 44 to 49 percent |
| by weekly editors | 2 | 22 to 71 percent |
| by contract size | 3 | 41 to 52 percent |
I would not consider it settled without evidence: Cut the same cohort two ways and confirm the segmentation you propose separates retention curves that the other one merges.
A segment is real when the two halves need different products.
Curated: · Written: · Reviewed:
QA-8How do you test whether customers will pay for something before you build it?(show answer)
What separates a strong answer on willingness to pay before a pricing page is knowing which number would have to move and by how much.
Stated willingness to pay is unreliable and a commitment is not. The strongest pre-build evidence is a signature, a deposit, or a customer changing their own process in anticipation.
Concretely, put a price in front of the customer in a context where saying yes costs them something, such as a letter of intent or a paid pilot. Record how many convert at each price point rather than the average of what people say they would pay. Treat verbal enthusiasm at no cost as a zero.
The reason for that specificity is a failure I have seen: A team surveyed 140 accounts who reported an average acceptable price of 240 a month, then converted 3 of 60 at 99 a month when the paid pilot opened, which was a 5 percent take rate against a forecast of 45 percent.
Stated price against paid conversion.
| Signal | Population | Converted |
|---|---|---|
| survey at 240 | 140 | not asked to pay |
| pilot at 99 | 60 | 3 |
| pilot at 49 | 60 | 14 |
I would not consider it settled without evidence: Open a paid pilot at the proposed price and count signatures rather than survey responses.
Ask for the money before you write the code.
Curated: · Written: · Reviewed:
QA-9A competitor ships a feature you do not have. How do you respond?(show answer)
I would frame competitive analysis without feature checklists around the segment it affects, because an average across segments hides the decision.
A competitor's release is evidence about their strategy and only weak evidence about your customers' problems. The response depends on whether the feature moves the dimension your customers actually choose on.
Concretely, check whether your own win-loss records mention the capability before the competitor shipped it. Ask the last 10 losses what decided the deal and count how many name this dimension. Match only when the evidence comes from your pipeline rather than from the press release.
The reason for that specificity is a failure I have seen: A team rebuilt a competitor's dashboard builder in 9 weeks and won 1 additional deal, while the same quarter's loss reviews named onboarding time in 7 of 12 losses and nothing was scheduled against it.
What the last twelve losses actually named.
| Reason named | Losses | Work scheduled |
|---|---|---|
| onboarding time | 7 | none |
| dashboard builder | 2 | 9 weeks |
| price | 3 | none |
I would not consider it settled without evidence: Count how many of the last ten losses named the capability before deciding to match it.
Copying a roadmap means outsourcing your strategy to a competitor's guesses.
Curated: · Written: · Reviewed:
QA-10How do you position a product whose main competitor is a spreadsheet?(show answer)
Before committing to anything on positioning against the real alternative, I would state the smallest test that could change my mind.
Positioning is chosen against the alternative the customer would otherwise pick, which is often a manual process rather than a named vendor. Claiming superiority over a vendor the customer never considered wastes the only sentence they will read.
Concretely, identify what the customer does today and what it costs them in hours or errors. Write the claim against that baseline with a number. Test the claim on people who have never seen the product and confirm they can repeat it back.
The reason for that specificity is a failure I have seen: A tool positioned itself as the leading alternative to two enterprise vendors, and 64 percent of its trial accounts had never evaluated either one, arriving from a spreadsheet the messaging never mentioned.
What trial accounts used before signup.
| Previous tool | Share of trials | Named in positioning |
|---|---|---|
| spreadsheet | 64 percent | no |
| enterprise vendor A | 11 percent | yes |
| nothing at all | 18 percent | no |
I would not consider it settled without evidence: Ask trial signups what they used before the product and check the positioning names that thing.
Position against what the customer would do on Monday without you.
Curated: · Written: · Reviewed:
QA-11How do you decide what stays out of a first release?(show answer)
I would handle defining a minimum viable product by making the trade-off explicit rather than promising both sides of it.
A first release is an instrument for learning one thing, and everything that does not serve that question is optional. Scope decided by what feels incomplete grows without limit because nothing is ever complete.
Concretely, state the single question the release must answer and the result that would count as a yes. Cut every item that cannot change that answer, and write down what you cut so the omission is deliberate rather than forgotten. Set the date by the question, not by the feature list.
The reason for that specificity is a failure I have seen: A team spent 5 months adding role permissions, audit logs, and theming to a first release, then learned in week 2 of the beta that 8 of 10 pilot users abandoned at the import step none of those features touched.
Scope against the learning question.
| Item | Weeks | Could change the answer |
|---|---|---|
| import flow | 3 | yes |
| role permissions | 6 | no |
| audit logs | 4 | no |
| theming | 2 | no |
I would not consider it settled without evidence: Write the learning question on the release plan and confirm every included item could change its answer.
The first release exists to answer a question, not to look finished.
Curated: · Written: · Reviewed:
QA-12What evidence would convince you a product has found its market?(show answer)
I would ground reading product-market fit signals in what customers do after launch rather than in what they said in the room.
Fit shows up as behaviour that repeats without prompting: cohorts that flatten rather than decay, usage that survives the end of onboarding support, and organic accounts arriving from existing ones. A growth curve driven by paid acquisition can hide the absence of all three.
Concretely, plot retention by cohort and look for a flattening tail rather than a slower decline. Split acquisition into paid and organic and watch whether the organic share grows. Track how many accounts expand seats without a salesperson touching them.
The reason for that specificity is a failure I have seen: A company read 40 percent month-on-month growth as fit, and when paid spend paused for 6 weeks, new accounts fell 82 percent while paid month-6 retention was already 9 percent.
What the growth curve was made of.
| Source | Share of new accounts | Month-6 retention |
|---|---|---|
| paid | 78 percent | 9 percent |
| organic referral | 14 percent | 44 percent |
| partner | 8 percent | 31 percent |
I would not consider it settled without evidence: Pause or reduce paid acquisition for a defined window and check whether organic signups and cohort retention hold.
Fit is what remains when you stop pushing.
Curated: · Written: · Reviewed:
QA-13When is a survey the right instrument, and what can it not tell you?(show answer)
I would start surveys as a supporting instrument from the customer problem it is supposed to change, not from the feature already on the roadmap.
A survey measures the distribution of an opinion across a population you already understand, and it cannot discover a problem you did not think to ask about. Using it as the first instrument produces precise answers to the wrong questions.
Concretely, run interviews first to learn what to ask, then survey to size how common each finding is. Report the response rate and the non-response profile alongside every result. Never let a free-text box stand in for a conversation.
The reason for that specificity is a failure I have seen: A team surveyed 4,200 users with a 7 percent response rate, built the top-ranked request, and discovered afterwards that respondents skewed to accounts older than 2 years, who were 14 percent of the base and had the opposite need to the newest cohort.
Respondents against the base.
| Group | Share of base | Share of respondents |
|---|---|---|
| accounts over 2 years | 14 percent | 61 percent |
| accounts under 90 days | 38 percent | 8 percent |
| trial accounts | 22 percent | 3 percent |
I would not consider it settled without evidence: Compare the respondent profile against the whole user base and report where they diverge.
Surveys size a finding; they do not find it.
Curated: · Written: · Reviewed:
QA-14How do you use support data without letting it set the roadmap?(show answer)
The first thing I would settle about support tickets as a discovery source is what evidence would make me drop the idea.
Support volume measures who is already in the product and already frustrated enough to write. It is excellent at ranking known friction and blind to the customers who left silently or never started.
Concretely, cluster tickets by the workflow they interrupt rather than by the feature named. Weight each cluster by the accounts affected rather than by ticket count, since one power user can file 40. Pair the cluster with a churn or activation number before it earns a roadmap slot.
The reason for that specificity is a failure I have seen: A team prioritised the top ticket cluster for two quarters, which came from 6 accounts filing 380 tickets, while activation for new accounts sat at 31 percent and generated almost no tickets because those users never reached a screen worth complaining about.
The same quarter, two rankings.
| Cluster | Tickets | Distinct accounts |
|---|---|---|
| export formatting | 380 | 6 |
| first import fails | 94 | 210 |
| billing confusion | 120 | 88 |
I would not consider it settled without evidence: Recount the top ticket clusters by distinct accounts affected and compare the ranking against the count by tickets.
Tickets are loudest where customers still care enough to complain.
Curated: · Written: · Reviewed:
QA-15What do you write down before asking for a team's quarter?(show answer)
With opportunity assessment before a quarter, I would separate what we believe from what we have actually observed.
An opportunity assessment states the problem, the affected population, the expected behaviour change, the value if it works, and what would falsify it. Without the falsifier the document is advocacy rather than analysis.
Concretely, size the affected population from product data and name the behaviour that must change. State the value per unit of that behaviour and the total. Write the result that would tell you the bet failed, and the date you will check.
The reason for that specificity is a failure I have seen: A team funded a 3-team quarter on a memo with no failure condition, and 9 months later the initiative was still running with a redefined goal each quarter and 0 recorded checks against the original claim.
The assessment's own scoreboard.
| Field | Value |
|---|---|
| affected accounts | 3,400 |
| behaviour that must change | weekly import from 1 to 3 |
| value if it holds | 1.2M annual |
| failure condition | under 1.6 imports by day 90 |
I would not consider it settled without evidence: Confirm the assessment names a date and a number that would count as failure, and put the check in the calendar.
An investment without a falsifier cannot be stopped.
Curated: · Written: · Reviewed:
QA-16How do you test a solution idea without engineering time?(show answer)
I would answer concept testing a solution before building by naming the decision it feeds and who has to live with that decision.
A concept test measures whether the proposed solution is understood and preferred, which is cheap and informative, but it cannot measure whether people will do the work the product requires. Comprehension failures found on paper are the cheapest failures you will ever buy.
Concretely, put a static mock or a written description in front of eight target users and ask them to explain what it does and when they would use it. Count comprehension failures before preference. Follow with a commitment step so preference has a cost attached.
The reason for that specificity is a failure I have seen: A team ran 6 concept tests where users praised a proposed automation, then built it, and 4 of the 6 could not describe what it would do when shown the working version without a facilitator narrating it.
Comprehension against stated preference.
| Participant group | Described it correctly | Said they liked it |
|---|---|---|
| with facilitator narration | 6 of 6 | 6 of 6 |
| unaided | 2 of 6 | 5 of 6 |
| after 24 hours | 1 of 6 | 4 of 6 |
I would not consider it settled without evidence: Ask each participant to describe the concept back unaided and count how many get it right.
If they cannot explain it back, they did not evaluate it.
Curated: · Written: · Reviewed:
QA-17How much fidelity does a prototype need?(show answer)
My approach to prototypes and the fidelity question starts with the outcome the business is buying rather than the output the team ships.
Fidelity should match the question being asked: flow questions need clickable structure, desirability questions need visual design, and feasibility questions need real data. Higher fidelity than the question requires buys delay and invites feedback about colours.
Concretely, write the question first and pick the cheapest artefact that can answer it. Use grey-box wireframes for structure, styled mocks only when the decision is about appeal, and a working slice when the risk is whether the data supports the flow. Say at the start of the session which kind of feedback you are collecting.
The reason for that specificity is a failure I have seen: A team spent 4 weeks on a pixel-complete prototype to test a three-step flow, and 9 of 11 testers spent the session commenting on typography while the step where 5 of them got lost went undiscussed.
Artefact against the question it can answer.
| Artefact | Build time | Answers |
|---|---|---|
| paper flow | 2 hours | is the sequence understood |
| grey-box clickable | 2 days | where do people get lost |
| pixel-complete | 4 weeks | is it appealing |
I would not consider it settled without evidence: State the question the prototype must answer and confirm the artefact is the cheapest one that can answer it.
Fidelity is a cost you pay for a specific answer.
Curated: · Written: · Reviewed:
QA-18A usability test shows users struggling. How do you decide whether that changes the product?(show answer)
For usability findings versus product decisions, I would write down the assumption that has to hold before I argue about the solution.
A usability failure tells you a task is hard to perform, not that the task is worth performing. Fixing the interface on a feature nobody needs converts a discovery finding into wasted engineering.
Concretely, separate whether users could complete the task from whether they wanted to. Check usage data for how many accounts attempt the task at all. Fix the interface when demand exists and reconsider the feature when it does not.
The reason for that specificity is a failure I have seen: A team ran 5 rounds of usability fixes on a report builder over 7 weeks, raising task completion from 40 to 88 percent, while only 4 percent of accounts opened the builder in any month before or after.
Completion against demand.
| Round | Task completion | Accounts attempting monthly |
|---|---|---|
| baseline | 40 percent | 4 percent |
| after 5 fixes | 88 percent | 4 percent |
| after removal | n/a | 0 percent |
I would not consider it settled without evidence: Check how many accounts attempt the task in production before scheduling interface work on it.
Making an unwanted task easier does not make it wanted.
Curated: · Written: · Reviewed:
QA-19Describe a case where the right product decision was to build nothing.(show answer)
I would treat deciding not to build as a question about behaviour, so I would look for the behaviour before the opinion.
Not building is a decision with its own evidence and its own communication, and it frees the scarcest resource the company has. Treating it as a non-event leaves the request to return every quarter with no record of why it was declined.
Concretely, write down the reason, the evidence, and the condition under which you would revisit. Tell the requester directly rather than letting the item age in a backlog. Put the revisit condition somewhere it will be checked.
The reason for that specificity is a failure I have seen: A request was silently deprioritised 4 quarters in a row, cost roughly 6 hours of re-analysis each time, and was eventually built in a rush when an executive asked about it, with none of the earlier analysis found.
The cost of an unrecorded no.
| Quarter | Re-analysis hours | Decision recorded |
|---|---|---|
| Q1 | 6 | no |
| Q2 | 6 | no |
| Q3 | 6 | no |
| Q4 | 0, built in a rush | no |
I would not consider it settled without evidence: Confirm the declined request has a written reason and a revisit condition that someone will actually check.
A no you cannot find is a no you will pay for again.
Curated: · Written: · Reviewed:
QA-20How does discovery change when your customers are internal engineering teams?(show answer)
What separates a strong answer on discovery for a platform or internal product is knowing which number would have to move and by how much.
Internal customers cannot leave, which removes the strongest signal an external market provides. Adoption has to be measured against the workaround teams build instead, because that workaround is the real competitor.
Concretely, count the teams who route around the platform and read what they built instead. Measure time to first successful use for a new team rather than satisfaction. Treat a team choosing the workaround as churn even though the budget never moves.
The reason for that specificity is a failure I have seen: An internal platform reported 100 percent adoption because every service was registered, while 14 of 31 teams maintained their own deployment scripts and the platform's own onboarding took 11 days against 2 hours for the workaround.
Registered against actually used.
| Measure | Value |
|---|---|
| services registered | 31 of 31 |
| teams with parallel scripts | 14 |
| platform onboarding | 11 days |
| workaround onboarding | 2 hours |
I would not consider it settled without evidence: Count the teams maintaining a parallel path and measure their time to first success against the platform's.
A captive customer expresses churn as a workaround.
Curated: · Written: · Reviewed:
QA-21How do you tell whether a problem is severe enough to pay for?(show answer)
I would frame distinguishing a vitamin from a painkiller around the segment it affects, because an average across segments hides the decision.
Severity shows up in what the customer already spends on the problem in time, money, or headcount. A problem nobody has paid to work around is usually tolerable, whatever people say in an interview.
Concretely, ask what the customer currently spends on the workaround and who owns it. Look for a budget line, a contractor, or a recurring meeting. Rank problems by the existing spend rather than by the emotion in the interview.
The reason for that specificity is a failure I have seen: A team built for a problem that 22 of 25 interviewees called frustrating, and discovered at launch that the median account spent 0 hours a month on it because a default setting had quietly made it rare 18 months earlier.
Stated frustration against existing spend.
| Problem | Called frustrating | Median monthly spend |
|---|---|---|
| duplicate records | 22 of 25 | 0 hours |
| month-end close | 14 of 25 | 9 hours |
| access requests | 11 of 25 | 1 contractor |
I would not consider it settled without evidence: Ask each interviewee what they currently spend on the workaround in hours or money and record the median.
Follow the spending, not the sighing.
Curated: · Written: · Reviewed:
QA-22An interviewer asks how you would improve a product you do not work on. How do you structure the answer?(show answer)
Before committing to anything on product sense questions in an interview, I would state the smallest test that could change my mind.
The interviewer is scoring the path from user to problem to trade-off to measurable outcome, not the cleverness of the idea. An answer that opens with a feature skips every part being assessed.
Concretely, name the user segment and the goal you are optimising, state two or three problems that block it, pick one with a reason, propose two solutions with a trade-off between them, and close with the metric that would tell you it worked. Keep the whole structure visible as you speak.
The reason for that specificity is a failure I have seen: A candidate opened with three feature ideas for a maps product, spent 14 of 20 minutes on interface detail, and never named a user or a success metric, which the debrief recorded as no evidence of structured product thinking.
Where the twenty minutes went.
| Section | Weak answer | Strong answer |
|---|---|---|
| user and goal | 1 min | 4 min |
| problems and choice | 2 min | 6 min |
| solution and trade-off | 17 min | 6 min |
| metric | 0 min | 4 min |
I would not consider it settled without evidence: Check your own answer contains a named segment, a chosen problem with a reason, a trade-off, and a success metric before adding detail.
The structure is the signal; the feature is the illustration.
Curated: · Written: · Reviewed:
QA-23How do you forecast adoption for a product with no comparable predecessor?(show answer)
I would handle estimating demand for something that does not exist by making the trade-off explicit rather than promising both sides of it.
Forecasts for genuinely new products come from analogues and from a live signal you can buy cheaply, not from confidence. The honest output is a range with the mechanism that would move it to each end.
Concretely, find two or three analogues with a similar adoption barrier and record their first-year curves. Run a smoke test such as a landing page or a paid waitlist to get a live conversion rate. Publish a range, the analogue it came from, and the assumption each end depends on.
The reason for that specificity is a failure I have seen: A launch forecast of 20,000 first-year accounts came from a single optimistic analogue, actual was 2,900, and the company had already signed a 12-month support contract sized for the forecast at 41,000 a month.
Forecast, smoke test, and outcome.
| Input | First-year accounts |
|---|---|
| optimistic analogue | 20,000 |
| smoke-test extrapolation | 2,000 to 4,500 |
| actual | 2,900 |
I would not consider it settled without evidence: Run a paid smoke test and compare its conversion rate against the forecast before any spend is committed.
Buy a small real signal before betting on a large imagined one.
Curated: · Written: · Reviewed:
QA-24Usage data and customer interviews disagree. Which do you trust?(show answer)
I would ground when qualitative evidence should outrank quantitative in what customers do after launch rather than in what they said in the room.
Instrumentation tells you what happened and interviews tell you why, and a disagreement usually means one of them is measuring something other than what you assumed. The resolution is to find the mechanism that explains both, not to pick a favourite.
Concretely, check the instrumentation definition before trusting either side, since most conflicts come from an event firing in a case nobody intended. Watch a real session end to end. Form a mechanism that predicts both the number and the story, then test it.
The reason for that specificity is a failure I have seen: A team saw a 23 percent completion rate against interviews where everyone described finishing easily, and the event had been firing only for accounts created after a release 5 months earlier, which was 31 percent of the users in the report.
Where the conflict came from.
| Source | Completion | Population measured |
|---|---|---|
| dashboard | 23 percent | accounts after the release |
| interviews | 9 of 10 | long-tenured accounts |
| corrected event | 71 percent | all accounts |
I would not consider it settled without evidence: Read the event definition and replay one real session before deciding which source is wrong.
Disagreement between the number and the story is usually a definition problem.
Curated: · Written: · Reviewed:
QA-25What belongs on a single page that starts a project?(show answer)
I would start writing the one-page product brief from the customer problem it is supposed to change, not from the feature already on the roadmap.
The brief exists so that ten people build the same thing, which requires the problem, the customer, the measurable outcome, the scope boundary, and the open questions. Anything that does not reduce disagreement belongs somewhere else.
Concretely, write the customer and problem in two sentences, the outcome as a number with a date, the explicit non-goals, and the open questions with an owner each. Circulate before the kickoff and collect disagreement in writing. Update the same page rather than issuing a second document.
The reason for that specificity is a failure I have seen: A project ran from a 14-page specification that 3 of 9 team members had read, and two engineers built overlapping services for 5 weeks because the scope boundary appeared only on page 11.
Document length against shared understanding.
| Document | Read fully | Stated goal correctly |
|---|---|---|
| 14-page spec | 3 of 9 | 4 of 9 |
| 1-page brief | 9 of 9 | 8 of 9 |
| kickoff only | n/a | 3 of 9 |
I would not consider it settled without evidence: Ask two people who were not in the kickoff to state the goal and the non-goals from the page alone.
A brief is measured by the disagreements it surfaces before the work starts.
Curated: · Written: · Reviewed:
QA-26How do you use a scoring framework like RICE without letting it make the decision for you?(show answer)
The first thing I would settle about what a prioritization framework can and cannot decide is what evidence would make me drop the idea.
A scoring framework makes assumptions comparable and visible; it does not supply judgement, because every input is an estimate someone chose. A ranked list treated as an output rather than as a prompt for argument launders opinion into arithmetic.
Concretely, score the candidates, then read the top and bottom three and ask whether the order matches informed judgement. Where it does not, find the input that is wrong rather than overriding the list silently. Record the inputs so next quarter can check which estimates were accurate.
The reason for that specificity is a failure I have seen: A team shipped the top five RICE items for two quarters, and a review found reach had been estimated from total accounts rather than accounts who reach the screen, overstating one item by a factor of 9 and sinking a fix that affected every new signup.
Estimated against actual for three shipped items.
| Item | Estimated reach | Actual reach |
|---|---|---|
| bulk edit | 12,000 | 1,300 |
| import retry | 3,000 | 11,400 |
| saved views | 8,000 | 6,900 |
I would not consider it settled without evidence: Compare last quarter's estimated reach and impact against what actually happened for three shipped items.
The score is an argument written down, not a verdict.
Curated: · Written: · Reviewed:
QA-27Two initiatives both look worthwhile. How do you choose?(show answer)
With prioritizing against opportunity cost, I would separate what we believe from what we have actually observed.
The cost of an initiative is the next best thing the same team could have done, so the comparison has to be between alternatives rather than against zero. Anything justified only by being better than nothing will pass forever.
Concretely, force a ranked list rather than a set of approvals, and make each item name what it displaces. State the value per team-week for each candidate. Require the sponsor of the winning item to say out loud which item it pushed out.
The reason for that specificity is a failure I have seen: A quarter approved 7 initiatives on the basis that each had positive return, delivered 3 of them by the deadline, and the 4 unfinished ones consumed 19 team-weeks with no releasable result.
Approvals against capacity.
| Plan | Items approved | Team-weeks needed | Available |
|---|---|---|---|
| by return | 7 | 41 | 26 |
| by rank | 4 | 25 | 26 |
| actual delivery | 3 shipped, 4 unfinished | 7 shipped plus 19 unfinished | 26 |
I would not consider it settled without evidence: Ask each sponsor which item their initiative displaces and confirm the total fits the team-weeks available.
Everything is worth doing compared to nothing.
Curated: · Written: · Reviewed:
QA-28How do you order work when one team's output blocks another's?(show answer)
I would answer sequencing a roadmap around dependencies by naming the decision it feeds and who has to live with that decision.
A dependency turns two independent estimates into one joint risk, and the joint risk is worse than either alone because the blocked team idles on the other's slip. Sequencing has to make the blocking item earlier and smaller, not merely earlier.
Concretely, identify the interface between the teams and ship a stub or contract version of it first. Give the blocking team a date earlier than the consumer needs, with slack sized to their historical variance. Keep a fallback the consumer can build against if the date slips.
The reason for that specificity is a failure I have seen: A mobile release waited on an API that slipped 3 weeks, and because there was no contract stub, 4 mobile engineers spent 11 days on work that was discarded when the response shape changed.
Two sequencings of the same pair.
| Approach | Idle days | Rework days |
|---|---|---|
| wait for the real API | 11 | 11 |
| contract stub first | 0 | 2 |
| build both blind | 0 | 17 |
I would not consider it settled without evidence: Confirm the consuming team can build against a stub or contract before the producing team's real date.
A dependency you cannot stub is a date you do not control.
Curated: · Written: · Reviewed:
QA-29Engineering wants a quarter for refactoring. How do you decide?(show answer)
My approach to allocating capacity to technical debt starts with the outcome the business is buying rather than the output the team ships.
Technical debt competes with features on the same team-weeks, so it has to be argued in the same currency: what it will change about delivery speed, incidents, or the ability to ship a named future thing. A moral argument about cleanliness cannot be prioritised against a revenue number.
Concretely, ask which specific future work is slowed and by how much, and which incidents trace to the area. Fund the slice that unblocks named upcoming work rather than a general cleanup. Agree the measure that will show it worked before the work starts.
The reason for that specificity is a failure I have seen: A team took a full refactoring quarter on a service that had produced 2 incidents in a year, while the checkout path that caused 9 incidents and blocked three roadmap items was untouched because nobody asked which debt was in the way.
Debt candidates against evidence.
| Area | Incidents last year | Roadmap items blocked |
|---|---|---|
| reporting service | 2 | 0 |
| checkout path | 9 | 3 |
| admin tooling | 1 | 1 |
I would not consider it settled without evidence: Map the proposed debt work against the next two quarters of roadmap items and the last year of incidents.
Fund the debt that is standing in front of something you want.
Curated: · Written: · Reviewed:
QA-30An executive asks for a feature that is not on the roadmap. What do you do?(show answer)
For saying no to an executive request, I would write down the assumption that has to hold before I argue about the solution.
The answer is neither a refusal nor a silent insertion, but a cost stated in the work it displaces and a decision made by whoever owns that trade-off. Absorbing the request quietly makes the roadmap fiction and the slip someone else's surprise.
Concretely, estimate the request, name the two items it would push out, and put both in front of the executive. Ask them to choose rather than telling them no. Record the decision and the displaced items where the affected teams can see it.
The reason for that specificity is a failure I have seen: A PM absorbed 3 executive requests over a quarter without changing the published plan, and the team missed 3 committed dates, which the operating review recorded as a delivery problem rather than as a scope decision.
What the absorbed requests displaced.
| Request | Team-weeks | Displaced item | Announced |
|---|---|---|---|
| exec dashboard | 5 | import retry | no |
| partner export | 4 | onboarding fix | no |
| logo refresh | 2 | billing bug | no |
I would not consider it settled without evidence: Show the requester the two items their request displaces and record which they chose.
Say what it costs and let the owner of the trade-off decide.
Curated: · Written: · Reviewed:
QA-31How specific should a roadmap be six months out?(show answer)
I would treat roadmaps as commitment versus direction as a question about behaviour, so I would look for the behaviour before the opinion.
Confidence decays with distance, so a roadmap should carry decreasing specificity: dated commitments near, themes with intent in the middle, and problems without solutions at the far end. Publishing distant work with dates converts every normal learning event into a broken promise.
Concretely, split the roadmap into committed, planned, and exploring, and say what moves an item between them. Attach dates only to the committed band. Republish on a fixed cadence so movement is expected rather than announced as a failure.
The reason for that specificity is a failure I have seen: A team published 11 dated items covering 3 quarters, delivered 4 on their original dates, and the account team had already promised 7 of them to customers in writing, generating 23 escalations when the plan changed.
Three bands and what they may carry.
| Band | Horizon | Carries a date | Repeatable to customers |
|---|---|---|---|
| committed | 0 to 6 weeks | yes | yes |
| planned | 1 to 2 quarters | no | with caveat |
| exploring | beyond | no | no |
I would not consider it settled without evidence: Check what the sales team has repeated to customers and confirm it matches the committed band only.
Precision you cannot honour is a promise you did not mean to make.
Curated: · Written: · Reviewed:
QA-32How do you choose a single metric for a product team to steer by?(show answer)
What separates a strong answer on defining a north star metric is knowing which number would have to move and by how much.
A north star should express delivered customer value, move within a team's control, and be hard to inflate without producing the value it stands for. Revenue fails the second test and pageviews fail the third.
Concretely, write the customer's successful moment and count it directly, such as weekly reports actually delivered rather than reports created. Test the candidate by asking how a cynical team could raise it without helping anyone. Pair it with one or two guardrails that would catch that gaming.
The reason for that specificity is a failure I have seen: A team steered on accounts created for two quarters, grew it 34 percent through an unusable trial, and the share of accounts completing a first real task fell from 44 to 19 percent without the north star moving down.
Candidate metrics against the gaming test.
| Candidate | Gameable by | Guardrail |
|---|---|---|
| accounts created | frictionless empty signups | first task completed |
| pageviews | splitting pages | task time |
| weekly delivered reports | hard to game | error rate |
I would not consider it settled without evidence: Ask how the metric could be raised without helping a customer and confirm a guardrail would catch it.
Choose the number a cynic cannot move without doing the work.
Curated: · Written: · Reviewed:
QA-33Your north star moves slowly. What do you give the team to steer by week to week?(show answer)
I would frame input metrics a team can actually move around the segment it affects, because an average across segments hides the decision.
A lagging outcome cannot guide weekly decisions, so it has to be decomposed into inputs the team changes directly and whose relationship to the outcome has been checked. Inputs chosen without checking that relationship become busywork with a dashboard.
Concretely, decompose the outcome into a small number of inputs and test each one's historical correlation with the outcome. Set weekly targets on the inputs and review the link every quarter. Drop an input when the link stops holding rather than keeping it for continuity.
The reason for that specificity is a failure I have seen: A team drove an input of demos booked for 3 quarters, raised it 58 percent, and revenue per demo fell by half because the added demos came from a segment that converted at 4 percent against 19 percent for the original one.
The input moved and the outcome did not.
| Quarter | Demos booked | Revenue per demo | Revenue |
|---|---|---|---|
| baseline | 210 | 4,100 | 861k |
| plus two quarters | 332 | 2,050 | 681k |
I would not consider it settled without evidence: Recheck the correlation between each input and the outcome using the last two quarters of data.
An input metric is a hypothesis about the outcome, and hypotheses expire.
Curated: · Written: · Reviewed:
QA-34How do you set a quarterly target that is neither sandbagged nor fantasy?(show answer)
Before committing to anything on goal setting with a credible target, I would state the smallest test that could change my mind.
A credible target is derived from a baseline, a mechanism, and a size of effect, so that missing it teaches you which part was wrong. A number chosen for ambition alone teaches nothing when it is missed, because nobody knows which assumption failed.
Concretely, state the baseline, the specific change you will make, and the effect size you expect from it with a reason. Add the three together rather than picking a round number. When you miss, identify whether the baseline, the mechanism, or the size was wrong.
The reason for that specificity is a failure I have seen: A team committed to doubling activation because the number sounded ambitious, landed at 12 percent improvement, and the review could not say whether the shipped changes underperformed or the baseline had been mismeasured, so the next quarter repeated the same guess.
Target built from parts.
| Part | Value | Source |
|---|---|---|
| baseline activation | 31 percent | last 8 weeks |
| import fix effect | plus 6 points | pilot cohort |
| guided setup effect | plus 4 points | analogue |
| target | 41 percent | sum |
I would not consider it settled without evidence: Write the target as baseline plus mechanism plus expected effect and confirm each part can be checked separately.
A target you cannot debug is a wish with a deadline.
Curated: · Written: · Reviewed:
QA-35You can fund an initiative at half the requested size. Should you?(show answer)
I would handle the cost of partially funding an initiative by making the trade-off explicit rather than promising both sides of it.
Some work has a threshold below which it delivers nothing, and funding under that threshold buys the cost without the outcome. The question is whether half the team produces half the value or none of it.
Concretely, ask what the half-size version would actually deliver end to end and whether a customer could use it. Fund fully, cut scope to a smaller complete thing, or do not start. Refuse the version that leaves an unusable half in production.
The reason for that specificity is a failure I have seen: A migration funded at half its estimate ran 7 months, moved 38 percent of accounts, and the company paid for both systems for 11 months at 24,000 a month while neither team could delete the old path.
Full, half, and none for the same migration.
| Funding | Accounts migrated | Dual-run months |
|---|---|---|
| full | 100 percent | 2 |
| half | 38 percent | 11 |
| none | 0 percent | 0 |
I would not consider it settled without evidence: Describe what the half-funded version delivers to a customer and confirm someone could use it as it stands.
Half a bridge costs more than no bridge.
Curated: · Written: · Reviewed:
QA-36How do you decide whether to build a capability or buy it?(show answer)
I would ground build, buy, or partner in what customers do after launch rather than in what they said in the room.
Build when the capability is how you differentiate, buy when it is necessary but undifferentiated, and partner when you need the capability and the relationship more than the control. The mistake is counting only the licence fee against the first release estimate.
Concretely, compare total cost over three years including integration, ongoing maintenance, and the cost of switching later. Ask whether customers would choose you because of this capability. Check the vendor's data and exit terms before the price.
The reason for that specificity is a failure I have seen: A team built its own search to save a 60,000 annual licence, spent 14 engineer-months on the first version, and then 2 engineers permanently on relevance tuning, which at loaded cost was roughly 4 times the licence every year.
Three-year total cost of the same capability.
| Path | Year 1 | Ongoing per year |
|---|---|---|
| licence | 60k plus 1 month integration | 60k |
| build | 14 engineer-months | 2 engineers |
| partner | revenue share | joint roadmap risk |
I would not consider it settled without evidence: Build the three-year total cost for both paths including ongoing maintenance headcount.
Build what customers buy you for and rent the rest.
Curated: · Written: · Reviewed:
QA-37How do you retire a feature that a small number of customers still use?(show answer)
I would start sunsetting a feature from the customer problem it is supposed to change, not from the feature already on the roadmap.
Retirement is a migration project with a communication plan, not an announcement. The cost of keeping a feature is carried by every future change to the code and every support person who must learn it, which is why small usage does not mean small cost.
Concretely, count the accounts and the revenue attached, offer a named replacement path, and give notice proportional to the switching work. Migrate the largest accounts by hand, instrument the remaining usage, and set a date that does not move. Keep an exception process with an owner rather than a permanent extension.
The reason for that specificity is a failure I have seen: A legacy import was announced as retiring in 30 days, 40 accounts were still using it at the deadline including 2 of the top 10 by revenue, and the rollback cost 6 engineer-weeks plus a renegotiated contract.
Usage at each stage of the retirement.
| Stage | Accounts using | Top-10 accounts |
|---|---|---|
| announcement | 210 | 4 |
| after assisted migration | 40 | 2 |
| after replacement path | 0 | 0 |
I would not consider it settled without evidence: Instrument the feature and confirm remaining usage is zero for two consecutive weeks before the code is removed.
A sunset is a migration you own, not a notice you send.
Curated: · Written: · Reviewed:
QA-38How do you justify platform investment that no customer asked for?(show answer)
The first thing I would settle about platform work against feature work is what evidence would make me drop the idea.
Platform work is justified by the future features it makes cheaper or possible, so the argument has to name those features and the difference in cost. Without that list it competes as an act of faith against work with customers attached.
Concretely, list the next four to six roadmap items that depend on the platform change and estimate them with and without it. Fund the platform slice that serves the nearest two. Recheck after delivery whether those items actually got cheaper.
The reason for that specificity is a failure I have seen: A team built a configuration framework for 9 weeks on the promise of faster feature delivery, and the following two quarters shipped 1 feature that used it while 6 features bypassed it because the framework assumed a data shape the product no longer used.
Features estimated both ways.
| Upcoming feature | Without platform | With platform |
|---|---|---|
| tenant branding | 4 weeks | 1 week |
| regional pricing | 6 weeks | 2 weeks |
| partner fields | 3 weeks | 3 weeks |
I would not consider it settled without evidence: Estimate two upcoming features with and without the platform change and confirm the difference is real.
Platform work is a bet on named future features, so name them.
Curated: · Written: · Reviewed:
QA-39A competitor halves its price. How do you respond?(show answer)
With handling a competitor's price cut, I would separate what we believe from what we have actually observed.
Price is a positioning decision, and matching a cut without changing costs or packaging transfers the loss straight to margin. The first question is how many of your deals actually lose on price rather than how the market feels.
Concretely, read the last 20 losses and count those where price was decisive against those where it was the stated reason. Model the margin impact of matching at current volume and the volume increase required to break even. Change packaging or value communication before changing the number.
The reason for that specificity is a failure I have seen: A company matched a competitor's 50 percent cut across the board, needed 2.1 times the volume to hold gross profit, achieved 1.2 times, and later found price was decisive in 4 of 20 losses rather than the 14 that had named it.
What matching required against what happened.
| Measure | Value |
|---|---|
| price cut | 50 percent |
| volume needed to hold profit | 2.1x |
| volume achieved | 1.2x |
| losses where price was decisive | 4 of 20 |
I would not consider it settled without evidence: Recount recent losses for price being decisive rather than merely mentioned, and model the volume needed to break even.
Matching a price is a permanent decision made from a temporary panic.
Curated: · Written: · Reviewed:
QA-40How do you decide which features belong in which tier?(show answer)
I would answer packaging and tiering decisions by naming the decision it feeds and who has to live with that decision.
Tiering should follow the value metric that grows with the customer, so that customers who get more value pay more without renegotiating. Splitting on arbitrary features teaches customers to buy the cheap tier and complain.
Concretely, identify the dimension along which customer value scales, such as seats, volume, or governance needs. Put capabilities that only matter at scale in the higher tier and keep the core workflow complete in every tier. Check that no tier is a trap where a customer must upgrade for something they already believed they had.
The reason for that specificity is a failure I have seen: A product put its audit log behind the top tier, and 34 percent of mid-tier renewals raised it as a blocker because a compliance rule had made it mandatory for every customer in that segment, producing 19 one-off contract exceptions.
Exceptions as a signal of bad packaging.
| Tier boundary | Exceptions granted | Deals lost |
|---|---|---|
| audit log at top tier | 19 | 6 |
| seats at mid tier | 1 | 0 |
| API volume | 2 | 1 |
I would not consider it settled without evidence: Count how many contract exceptions the current packaging has generated in the last two quarters.
Tiers should follow value, not withhold hostages.
Curated: · Written: · Reviewed:
QA-41How do you decide between fixing a known bug and shipping the next feature?(show answer)
My approach to prioritizing a bug against a feature starts with the outcome the business is buying rather than the output the team ships.
A bug's priority comes from the frequency, the severity of the consequence, and whether a workaround exists, not from the fact that it is a defect. Some defects deserve to sit for a quarter and some features deserve to be dropped for a defect.
Concretely, estimate accounts affected per week, what the consequence is, and the workaround cost. Compare that with the expected value of the feature over the same period. Write the comparison down so the decision can be argued with rather than defended by seniority.
The reason for that specificity is a failure I have seen: A team deferred a rounding defect for 2 quarters because it was a small bug, and it had been producing invoices wrong by up to 40 for 1,100 accounts a month, ending in 61,000 of credits and a finance audit.
The deferred defect priced.
| Input | Value |
|---|---|
| accounts affected monthly | 1,100 |
| error per invoice | up to 40 |
| quarters deferred | 2 |
| credits issued | 61,000 |
I would not consider it settled without evidence: Calculate accounts affected per week times the consequence per occurrence and compare it against the feature's expected value.
Small is not a property of bugs; it is a property of their consequences.
Curated: · Written: · Reviewed:
QA-42How do you plan for a product whose usage is shrinking?(show answer)
For roadmap for a declining product, I would write down the assumption that has to hold before I argue about the solution.
A declining product needs an explicit strategy chosen from harvest, turnaround, or exit, because the default of continued small investment delivers none of them. The decision depends on whether the decline comes from the market, the product, or neglect.
Concretely, split the decline into segments and find whether any cohort is stable. Decide with the business which of the three strategies applies and fund accordingly. If harvesting, reduce investment openly and hold quality; if exiting, publish the timeline early enough for customers to plan.
The reason for that specificity is a failure I have seen: A product received 2 engineers a quarter for 7 quarters with no declared strategy, lost 41 percent of revenue, and the eventual shutdown gave customers 60 days' notice, costing 3 of the company's 10 largest relationships.
Three strategies and what they fund.
| Strategy | Investment | Customer message |
|---|---|---|
| harvest | maintenance only | stable, no new features |
| turnaround | full team | named bets and dates |
| exit | migration team | dated end of life |
I would not consider it settled without evidence: Name the chosen strategy in writing and confirm the funding level matches it rather than the previous default.
Drift is the one strategy nobody chooses and many follow.
Curated: · Written: · Reviewed:
QA-43What makes a product strategy useful rather than decorative?(show answer)
I would treat a strategy document that constrains choices as a question about behaviour, so I would look for the behaviour before the opinion.
A strategy is useful when it rules things out, so a reader can predict which requests will be declined. A document that no proposal contradicts is a description of ambition rather than a strategy.
Concretely, state the customer you are serving, the value you are betting on, the capabilities you will build, and explicitly the things you will not do. Test it by taking three real requests and predicting the answer. Revisit on a schedule rather than whenever it is inconvenient.
The reason for that specificity is a failure I have seen: A strategy deck of 22 slides was used to justify both a self-serve push and an enterprise push in the same quarter, and the two teams built conflicting onboarding flows that took 5 weeks to reconcile.
Predicting three decisions from the document.
| Request | Predicted from strategy | Actual decision |
|---|---|---|
| enterprise SSO | yes | yes |
| free tier expansion | no | yes |
| partner marketplace | unclear | deferred |
I would not consider it settled without evidence: Take three live requests and check the document predicts the same answer a reviewer would give.
If it cannot say no, it is not a strategy.
Curated: · Written: · Reviewed:
QA-44How do you plan a quarter when part of the team's time is unpredictable?(show answer)
What separates a strong answer on quarterly planning with uncertain capacity is knowing which number would have to move and by how much.
Support, incidents, and compliance work consume real capacity, and a plan that assumes full availability converts every normal interruption into a missed commitment. Planning at measured capacity is honest rather than pessimistic.
Concretely, measure what proportion of the last two quarters went to unplanned work and reserve that share. Commit only to what fits the remainder. If the reserve is consistently large, treat reducing it as a funded project rather than as a hope.
The reason for that specificity is a failure I have seen: A team planned at 100 percent capacity for 2 quarters while unplanned work consumed about 34 percent, met 7 of 10 commitments, and only the third quarter reserved the unplanned share and then met 4 of 4.
Planned against actual capacity.
| Quarter | Planned capacity | Unplanned work | Commitments met |
|---|---|---|---|
| Q1 | 100 percent | 31 percent | 4 of 5 |
| Q2 | 100 percent | 36 percent | 3 of 5 |
| Q3 | 66 percent | 34 percent | 4 of 4 |
I would not consider it settled without evidence: Calculate the unplanned share of the last two quarters and confirm the plan reserves that much.
Plan with the capacity you have, not the one on the org chart.
Curated: · Written: · Reviewed:
QA-45Two weeks into a quarter the company changes direction. What do you do with the plan?(show answer)
I would frame handling a mid-quarter priority change around the segment it affects, because an average across segments hides the decision.
A direction change is a scope decision that has to be executed rather than absorbed: something stops, something starts, and both are communicated. Adding the new work while keeping the old commitments is the failure mode that looks like flexibility.
Concretely, stop the displaced work explicitly and put it back in the backlog with its state recorded. Restate the quarter's commitments with new dates. Tell every downstream team and customer-facing function what changed on the same day.
The reason for that specificity is a failure I have seen: A team added a new priority in week 2 without removing anything, finished the quarter with 4 half-built initiatives, and 3 of them were abandoned the following quarter after 22 team-weeks with nothing released.
Adding without removing.
| Response | Initiatives in flight | Released by quarter end |
|---|---|---|
| add only | 5 | 1 |
| add and stop two | 4 | 3 |
| defer the change | 4 | 4 |
I would not consider it settled without evidence: Confirm the plan after the change lists what stopped and that downstream teams received the update.
A change of direction means something must stop, on purpose and in writing.
Curated: · Written: · Reviewed:
QA-46An executive wants a date before the design exists. How do you answer?(show answer)
Before committing to anything on estimating without a fixed scope, I would state the smallest test that could change my mind.
An estimate before scope is a range whose width is the honest expression of what is unknown, and narrowing it requires reducing uncertainty rather than restating the number confidently. A single date given early becomes the commitment everyone remembers.
Concretely, give a range with the assumptions that would move it to each end. Offer a date for a decision point rather than for delivery, such as when a prototype will exist. Update on a fixed cadence and record what narrowed the range.
The reason for that specificity is a failure I have seen: A PM gave a single date in a hallway conversation, it appeared in a board deck 9 days later, and the team delivered 7 weeks past it while the range they had internally was 6 to 14 weeks all along.
The range against what was repeated.
| Artefact | Figure given |
|---|---|
| internal range | 6 to 14 weeks |
| hallway answer | 8 weeks |
| board deck | 8 weeks, committed |
| actual | 15 weeks |
I would not consider it settled without evidence: State the range and the assumptions to whoever asked, and confirm what they will repeat is the range.
An early date is a rumour with your name on it.
Curated: · Written: · Reviewed:
QA-47How do you slot regulatory work into a roadmap that is already full?(show answer)
I would handle prioritizing compliance and regulatory work by making the trade-off explicit rather than promising both sides of it.
Regulatory work has a deadline set outside the company and a consequence that is not proportional to its usefulness, which makes it a constraint rather than a candidate. Treating it as one more scored item guarantees a late scramble.
Concretely, place the deadline on the plan first and work backwards with slack, since external deadlines do not move. Separate the mandatory core from the optional improvements often bundled with it. Get the interpretation confirmed by whoever is accountable rather than assuming the strictest reading.
The reason for that specificity is a failure I have seen: A team scored a data-residency requirement against features, started it 6 weeks before the deadline, and spent 340,000 on contractors plus a 9-day feature freeze to finish on time.
Cost of starting late.
| Start | Contractor cost | Feature freeze |
|---|---|---|
| 6 weeks before | 340k | 9 days |
| 5 months before | 0 | 0 |
| after the deadline | penalty exposure | n/a |
I would not consider it settled without evidence: Put the external deadline on the plan and confirm the work starts with slack measured against the team's historical variance.
A deadline you did not set is a constraint, not a priority.
Curated: · Written: · Reviewed:
QA-48How do you split investment between acquisition and retention?(show answer)
I would ground balancing new customers against existing ones in what customers do after launch rather than in what they said in the room.
The split follows where value leaks, and that is measured rather than argued: if retention is poor, acquisition spend fills a leaking bucket. The balance should change when the leak changes, not once a year at planning.
Concretely, measure retention by cohort and the payback period on acquisition. If the cohort curve has not flattened, weight investment to retention until it does. Recheck each quarter and move the split with the evidence.
The reason for that specificity is a failure I have seen: A company tripled acquisition spend for 2 quarters while month-6 retention sat at 18 percent, added 9,400 accounts, and held 1,700 of them, at an acquisition cost of 310 each against a lifetime value of 190.
Cohort economics under the current split.
| Measure | Value |
|---|---|
| accounts added | 9,400 |
| retained at month 6 | 1,700 |
| acquisition cost each | 310 |
| realised value each | 190 |
I would not consider it settled without evidence: Compare acquisition cost against the realised value of a cohort that has aged six months.
Filling a leaking bucket faster is still a leak.
Curated: · Written: · Reviewed:
QA-49How do you prioritise between supply and demand in a marketplace?(show answer)
I would start prioritizing for a two-sided marketplace from the customer problem it is supposed to change, not from the feature already on the roadmap.
A marketplace is constrained by one side at a time, and investing in the unconstrained side produces experience that cannot be fulfilled. The constraint moves, so the priority has to be re-derived rather than fixed by preference.
Concretely, measure match rate and time to match by segment and geography to find which side is short. Invest in the short side until the match rate recovers, then re-measure. Keep the measurement at the level the market actually clears, which is usually local rather than global.
The reason for that specificity is a failure I have seen: A marketplace ran a demand campaign that added 12,000 buyers in a city where match rate was already 41 percent, and match rate fell to 22 percent because supply barely moved; about 10,800 of the new buyers left without a match, of whom 6 percent ever returned.
Match rate before and after a demand push.
| Stage | Buyers | Match rate | Return rate of unmatched |
|---|---|---|---|
| before | 8,000 | 41 percent | n/a |
| after campaign | 20,000 | 22 percent | 6 percent |
| after supply push | 20,000 | 58 percent | n/a |
I would not consider it settled without evidence: Measure the match rate in the specific segment before funding either side.
Growth on the wrong side of a shortage manufactures disappointment.
Curated: · Written: · Reviewed:
QA-50How do you argue for accessibility work against revenue features?(show answer)
The first thing I would settle about prioritizing accessibility work is what evidence would make me drop the idea.
Accessibility is a correctness property with legal and market consequences, and it is cheapest when treated as part of the definition of done rather than as a later project. Retrofitting costs more than building it right and leaves a period of exposure in between.
Concretely, audit against a named standard and classify issues by whether they block a task entirely. Fix blocking issues as defects, not as a separate initiative. Add the standard to the definition of done so the backlog stops growing.
The reason for that specificity is a failure I have seen: A company deferred accessibility for 6 quarters, then an enterprise deal worth 480,000 required a conformance report, and the retrofit took 11 weeks across 3 teams plus an external audit at 40,000.
Cost of retrofit against cost of doing it inline.
| Approach | Engineering cost | External audit |
|---|---|---|
| inline in definition of done | a small percent of each feature | 0 |
| retrofit after 6 quarters | 11 weeks across 3 teams | 40,000 |
| no action | deal at risk | 0, with 480,000 deal exposure |
I would not consider it settled without evidence: Audit the primary workflows against the named standard and count the issues that block task completion outright.
Deferred accessibility is a debt with a legal interest rate.
Curated: · Written: · Reviewed:
QA-51How do you keep customer-facing teams informed without creating promises?(show answer)
With roadmap communication to sales and support, I would separate what we believe from what we have actually observed.
Whatever reaches a salesperson will reach a customer, so the internal roadmap has to be published with an explicit rule about what may be repeated. The alternative is not secrecy but a controlled version with the confidence level attached.
Concretely, publish a version marked by confidence band and state plainly what may be said externally. Hold a short recurring session where changes are explained rather than emailed. Track which commitments have been repeated to customers so a change has a known blast radius.
The reason for that specificity is a failure I have seen: An internal roadmap slide was pasted into 14 customer emails within a week, 4 of the items moved out of the quarter, and the support team handled 63 related tickets with no prepared answer.
What may leave the building.
| Band | May be repeated | Requires caveat |
|---|---|---|
| committed | yes | no |
| planned | theme only | yes |
| exploring | no | n/a |
I would not consider it settled without evidence: Ask the sales team to state which roadmap items they have repeated to customers in the last month.
Assume every internal date will be read aloud to a customer.
Curated: · Written: · Reviewed:
QA-52Two dashboards disagree about active users. How do you settle it?(show answer)
I would answer a metric definition two teams can share by naming the decision it feeds and who has to live with that decision.
A metric is a definition before it is a number: the population, the event, the window, and the exclusions all have to be written down. Two correct implementations of an unwritten definition will disagree and both will be defended.
Concretely, write the definition with population, qualifying event, time window, and exclusions, and store it next to the query. Reconcile the two figures by isolating which clause differs. Publish one owner for the definition and route changes through them.
The reason for that specificity is a failure I have seen: A weekly active figure differed by 19 percent between two teams for 4 months because one counted API-only accounts and the other excluded them, and a pricing decision had already been made from the higher number.
Where the nineteen percent lived.
| Clause | Team A | Team B |
|---|---|---|
| population | all accounts | excludes API-only |
| event | any request | a session |
| window | 7 days rolling | calendar week |
I would not consider it settled without evidence: Reconcile the two queries clause by clause and confirm the remaining difference is zero.
Agree on the sentence before arguing about the number.
Curated: · Written: · Reviewed:
QA-53How do you tell whether a metric is worth reporting?(show answer)
My approach to vanity metrics and the decision test starts with the outcome the business is buying rather than the output the team ships.
A metric earns its place by changing a decision at some value, so if no number would cause a different action the measure is decoration. Cumulative totals fail this test because they can only rise.
Concretely, for each reported metric, write the threshold at which you would act and what you would do. Remove the ones with no answer. Prefer rates and cohort views over cumulative totals, which cannot fall and therefore cannot warn.
The reason for that specificity is a failure I have seen: A monthly review carried 31 metrics for a year, 22 of which had never been mentioned in a decision, and the weekly active rate fell 14 percent over 3 months without anyone raising it because the cumulative signup chart dominated the page.
Reported metrics against decisions taken.
| Metric | Times cited in a decision | Can fall |
|---|---|---|
| cumulative signups | 0 | no |
| weekly active rate | 9 | yes |
| support contacts per account | 4 | yes |
I would not consider it settled without evidence: Ask for the threshold and the action for each reported metric, and delete the ones nobody can supply.
If no value of it would change what you do, stop reporting it.
Curated: · Written: · Reviewed:
QA-54What do you look for in a cohort retention chart?(show answer)
For reading a retention curve, I would write down the assumption that has to hold before I argue about the solution.
The shape matters more than the level: a curve that flattens shows a group that found durable value, and a curve that decays towards zero shows a product people finish with. Averaging cohorts hides both.
Concretely, plot each acquisition cohort separately and look for the asymptote rather than the first-week drop. Compare recent cohorts against older ones to see whether changes are working. Segment by the behaviour you believe drives retention and check the curves separate.
The reason for that specificity is a failure I have seen: A team read a stable 38 percent month-2 average as healthy for 3 quarters, and separating cohorts showed the number was held up by one channel while every cohort after a redesign flattened at 11 percent.
The average against its parts.
| Cohort | Month 2 | Month 6 |
|---|---|---|
| pre-redesign | 44 percent | 39 percent |
| post-redesign | 33 percent | 11 percent |
| blended | 38 percent | 25 percent |
I would not consider it settled without evidence: Plot the last six cohorts separately and confirm the recent ones flatten rather than continuing to decay.
The average of a healthy cohort and a dying one looks fine.
Curated: · Written: · Reviewed:
QA-55How do you define activation for a product?(show answer)
I would treat activation as a measurable moment as a question about behaviour, so I would look for the behaviour before the opinion.
Activation is the earliest behaviour that predicts durable retention, which makes it an empirical question rather than a matter of taste. A definition chosen for convenience produces a metric teams optimise without moving retention.
Concretely, test candidate behaviours against month-3 retention and pick the one with the sharpest separation and a reachable rate. Confirm the relationship holds across segments before adopting it. Recheck annually since product changes move the moment.
The reason for that specificity is a failure I have seen: A team defined activation as completing the profile, raised it from 51 to 78 percent over two quarters through prompts, and month-3 retention did not move because profile completion separated retention by only 3 points.
Candidate behaviours against retention separation.
| First-week behaviour | Retention if done | If not |
|---|---|---|
| completed profile | 34 percent | 31 percent |
| invited a teammate | 62 percent | 24 percent |
| ran two imports | 58 percent | 22 percent |
I would not consider it settled without evidence: Compare month-3 retention for accounts that did and did not perform the candidate behaviour in their first week.
Activation is discovered in the data, not chosen in a meeting.
Curated: · Written: · Reviewed:
QA-56Signups are flat but traffic is up. How do you find the problem?(show answer)
What separates a strong answer on funnels and where the drop actually is is knowing which number would have to move and by how much.
A funnel localises loss between defined steps, and the useful comparison is conversion between steps rather than the totals at each one. Rates by source matter because a traffic mix change moves the aggregate without any step changing.
Concretely, instrument each step with a single definition and compute step-to-step conversion. Segment by source, device, and geography before concluding anything about the product. Compare against the same period's mix to separate a product change from a traffic change.
The reason for that specificity is a failure I have seen: A team rebuilt a signup form over 5 weeks on a report of a falling conversion rate, and the drop was entirely a paid campaign that had shifted 40 percent of traffic to a source converting at 0.4 percent against 3.1 percent for the rest.
Conversion by source across the drop.
| Source | Share before | Share after | Conversion |
|---|---|---|---|
| organic | 62 percent | 34 percent | 3.1 percent |
| paid social | 8 percent | 48 percent | 0.4 percent |
| referral | 30 percent | 18 percent | 4.0 percent |
I would not consider it settled without evidence: Recompute conversion within each traffic source and confirm the drop survives the segmentation.
An aggregate rate moves when the mix moves, with nothing else changing.
Curated: · Written: · Reviewed:
QA-57What do you watch to be sure a win is not causing damage elsewhere?(show answer)
I would frame choosing a guardrail metric around the segment it affects, because an average across segments hides the decision.
Most product changes move their target metric by moving something else, so a guardrail is the measure of the cost you are least willing to pay. Guardrails must be chosen before the result arrives, or they become excuses selected after it.
Concretely, name two or three guardrails before launch, such as support contacts, latency, or downstream conversion, with thresholds that stop the rollout. Monitor them on the same schedule as the primary metric. Stop when a guardrail breaches even if the primary metric is winning.
The reason for that specificity is a failure I have seen: An aggressive upsell prompt raised trial-to-paid by 9 percent and was kept for a quarter, while support contacts rose 31 percent and 90-day churn for the converted cohort was 2.4 times the baseline, which nobody had listed as a guardrail.
The win and its unwatched costs.
| Measure | Change | Watched in advance |
|---|---|---|
| trial to paid | plus 9 percent | yes |
| support contacts | plus 31 percent | no |
| 90-day churn | 2.4x baseline | no |
I would not consider it settled without evidence: Write the guardrails and their stopping thresholds into the experiment plan before the first user sees the change.
Decide what would make you stop while you can still be honest about it.
Curated: · Written: · Reviewed:
QA-58When would you not run an experiment?(show answer)
Before committing to anything on when an A/B test is the wrong instrument, I would state the smallest test that could change my mind.
Experiments need enough traffic to detect the effect you care about, a change that can be randomised, and a decision that depends on the result. Running one without those conditions buys delay and a number nobody should act on.
Concretely, compute the sample size needed for the minimum effect worth acting on before designing the test, using a named baseline conversion rather than an unstated 50 percent. If the traffic would take longer than the decision can wait, choose a different method such as a staged rollout with a holdout or a qualitative study. Say in advance what you will do for each outcome.
The reason for that specificity is a failure I have seen: A team ran a test on a page with 400 weekly visitors for 6 weeks, called a 2 percent lift a win at a p-value of 0.31, and the rollout showed no change over the following quarter.
Power for the effect you care about, assuming a 50 percent baseline.
| Minimum effect | Sample per arm | Weeks at 400 per week |
|---|---|---|
| 2 percent | roughly 39,000 | 195 |
| 10 percent | roughly 1,600 | 8 |
| 25 percent | roughly 270 | 2 |
I would not consider it settled without evidence: Compute the required sample for the smallest effect worth acting on and compare it with available traffic.
An underpowered test is an opinion with error bars.
Curated: · Written: · Reviewed:
QA-59An experiment looks significant on day three. Do you ship it?(show answer)
I would handle stopping rules and peeking by making the trade-off explicit rather than promising both sides of it.
Repeatedly checking a running test inflates the false-positive rate far beyond the nominal level, because each look is another chance for noise to cross the line. The stopping rule has to be fixed before the test starts or the reported significance is not the real one.
Concretely, set the duration and sample size in advance and read the result once at the end. If you need to look early, use a sequential method designed for it rather than the same threshold. Record the plan before the test opens so the rule cannot move.
The reason for that specificity is a failure I have seen: A team stopped 6 of 9 experiments early on favourable days, shipped all 6, and a later holdout found 4 of them had no effect, having spent 14 engineer-weeks on features that did nothing.
Early stops against a later holdout.
| Experiments | Stopped early | Confirmed by holdout |
|---|---|---|
| 9 run | 6 | 2 of 6 |
| full duration | 3 | 3 of 3 |
I would not consider it settled without evidence: Write the stopping rule and duration before launch and check the reported result against it, not against a favourable day.
If you look often enough, noise will eventually agree with you.
Curated: · Written: · Reviewed:
QA-60A new design wins for a week and then the gap closes. What happened?(show answer)
I would ground novelty and primacy effects in what customers do after launch rather than in what they said in the room.
A visible change draws attention regardless of its quality, so early differences can measure novelty rather than value, and habitual users can also react badly to a change that is better. Both effects decay, which is why test duration has to cover more than the first reaction.
Concretely, run long enough to cover at least two usage cycles for the behaviour in question. Split new and existing users, since novelty affects them differently. Look at the trend within the test rather than only the pooled result.
The reason for that specificity is a failure I have seen: A redesign showed a 14 percent engagement lift in week 1, was rolled out on that basis, and by week 6 sat 3 percent below the original; existing users were negative from week 1 while new users supplied the early novelty lift.
The lift by week and segment.
| Week | New users | Existing users |
|---|---|---|
| 1 | plus 19 percent | minus 2 percent |
| 3 | plus 8 percent | minus 2 percent |
| 6 | plus 2 percent | minus 6 percent |
I would not consider it settled without evidence: Plot the effect by week within the test and confirm it is stable rather than decaying.
Measure the habit, not the reaction to the change.
Curated: · Written: · Reviewed:
QA-61An experiment shows no overall effect. Is it a dead end?(show answer)
I would start segment heterogeneity in a flat result from the customer problem it is supposed to change, not from the feature already on the roadmap.
A null average can hide two segments moving in opposite directions, and that finding is often more valuable than a small uniform win. The discipline is that segments must be specified in advance or the analysis becomes a search for any subgroup that looks good.
Concretely, pre-register the two or three segments you have a reason to believe differ. Report those cuts with the overall result. Treat any unplanned segment that looks interesting as a hypothesis for a new test rather than as a finding.
The reason for that specificity is a failure I have seen: A team dropped a feature after a null result, and a later pre-registered rerun found a 22 percent gain for accounts with more than 5 seats and a 9 percent loss for single-seat accounts, which together had cancelled out across 60,000 users.
The null result taken apart.
| Segment | Share of users | Effect |
|---|---|---|
| over 5 seats | 21 percent | plus 22 percent |
| 2 to 5 seats | 16 percent | plus 9 percent |
| single seat | 63 percent | minus 9 percent |
| pooled | 100 percent | plus 0.4 percent |
I would not consider it settled without evidence: Pre-register the segments in the test plan and report them alongside the pooled result.
Zero on average can mean two real effects pointing in opposite directions.
Curated: · Written: · Reviewed:
QA-62A team is optimising a proxy metric. What do you watch for?(show answer)
The first thing I would settle about proxy metrics and Goodhart's law is what evidence would make me drop the idea.
A proxy is a stand-in for something expensive to measure, and it holds only while the relationship between them does. Once the proxy becomes a target, effort flows to the cheapest way of moving it, which is usually not the path that moves the real outcome.
Concretely, re-measure the link between proxy and outcome on a schedule. Pair the proxy with the outcome in the same report so divergence is visible. Change the proxy when the relationship weakens rather than defending the historical choice.
The reason for that specificity is a failure I have seen: A support team was measured on first-response time, improved it from 4.1 hours to 22 minutes with an automated acknowledgement, and resolution time grew by 31 percent while satisfaction fell 12 points.
Proxy improved, outcome worsened.
| Measure | Before | After |
|---|---|---|
| first response | 4.1 hours | 22 minutes |
| resolution time | 19 hours | 25 hours |
| satisfaction | 74 | 62 |
I would not consider it settled without evidence: Report the proxy and the outcome side by side and check the relationship still holds this quarter.
Every proxy is a measure until it becomes a target.
Curated: · Written: · Reviewed:
QA-63How do you handle accounts, users, and devices in one metric?(show answer)
With the unit a metric counts, I would separate what we believe from what we have actually observed.
The unit of analysis must match the decision: a workspace-level product should count workspaces, and counting logged-in devices inflates every rate. Mixing units across reports produces comparisons that cannot be reconciled.
Concretely, declare the unit for each metric and derive every rate from that unit. Deduplicate across devices and sessions before aggregating. Keep a stable identifier so a user reappearing on a new device does not read as a new one.
The reason for that specificity is a failure I have seen: A product reported 2.1 million monthly users from device identifiers, and deduplication by account showed 780,000, after a pricing model had been designed against the larger figure for two quarters.
One product, three units.
| Unit | Monthly count |
|---|---|
| devices | 2.1M |
| user accounts | 780k |
| paying workspaces | 41k |
I would not consider it settled without evidence: Recompute the headline number at the account level and compare it against the device-level figure.
A rate is only meaningful once you know what is in the denominator.
Curated: · Written: · Reviewed:
QA-64Assume engineering's doubled estimate is right. What gives: scope, date, or quality?(show answer)
I would answer trading scope, date, or quality by naming the decision it feeds and who has to live with that decision.
Scope is the only lever that does not send a bill to the future: cutting quality is debt the same team repays at interest, and moving a date is honest only when a specific commitment is renegotiated in the open. Trade scope first, quality never, and the date only against something named.
Concretely, ask for the estimate broken into slices with a cost per slice before you argue with any of it. Cut to the thinnest set of slices that still moves the outcome you were building for, and renegotiate the date only by naming the commitment that breaks if it moves.
The reason for that specificity is a failure I have seen: A billing migration stayed on its 15 March date by skipping load testing and shipping a reconciliation job expected to run in 30 seconds; at launch volume it ran 26 minutes, double-charged 1,400 invoices, and the rollback plus manual re-billing cost six engineering weeks against the two weeks the team had saved.
One estimate, two capacities.
| Slice | Engineer-weeks | Share of target usage |
|---|---|---|
| create and edit core flow | 6 | 68% |
| import from spreadsheet | 4 | 17% |
| approval routing | 5 | 9% |
| audit log | 3 | 5% |
| bulk export | 2 | 3% |
Engineering's estimate is 20 engineer-weeks; the team has 10 before the date. The top two slices are 10 weeks and carry 85 percent of target usage. The bottom three are the other 10 weeks and carry 15 percent.
I would not consider it settled without evidence: Split the estimate into slices and name the slice you would drop first, with the usage or outcome figure that says the thinner version still works.
Once you know who the date is a promise to, the only honest trade is scope.
Curated: · Written: · Reviewed:
QA-65Users who do X retain better. Should you push everyone to do X?(show answer)
My approach to correlation offered as a product insight starts with the outcome the business is buying rather than the output the team ships.
A correlation between a behaviour and retention usually reflects the kind of user who does the behaviour, so forcing the behaviour on everyone often moves nothing. The claim becomes actionable only when an intervention tests it.
Concretely, propose the causal mechanism and check it against what you know about the segments. Run an experiment that induces the behaviour in a random group. Compare the induced group's retention against the control rather than against natural adopters.
The reason for that specificity is a failure I have seen: A team prompted every new account to invite a teammate because inviters retained at 62 percent against 24 percent, lifted invitations by 40 percent, and the induced group retained at 26 percent against a 25 percent control.
Natural adopters against induced ones.
| Group | Invited a teammate | 90-day retention |
|---|---|---|
| natural adopters | yes | 62 percent |
| never invited | no | 24 percent |
| induced by prompt | yes | 26 percent |
I would not consider it settled without evidence: Randomly induce the behaviour in a test group and compare retention against a control.
Inducing the behaviour is the test of whether the behaviour was the cause.
Curated: · Written: · Reviewed:
QA-66How do you know whether a year of feature work made any difference?(show answer)
For holdouts for long-running effects, I would write down the assumption that has to hold before I argue about the solution.
Individual tests measure short-run effects and cannot be added up, because interactions and decay are not additive. A long-run holdout is the only way to measure what a year of changes did in aggregate.
Concretely, keep a small randomly assigned group on the old experience for a defined period. Compare the primary outcome between holdout and treated at the end. Size the holdout so the cost is acceptable and the comparison still has power.
The reason for that specificity is a failure I have seen: A team summed 19 winning experiments to claim a 47 percent improvement, and the one holdout that existed showed 6 percent, with 3 of the 19 later found to have interacted so their gains double-counted the same users.
Summed lifts against the holdout.
| Source | Claimed improvement |
|---|---|
| sum of 19 tests | 47 percent |
| holdout comparison | 6 percent |
| double-counted tests | 3 |
I would not consider it settled without evidence: Maintain a holdout group and read the aggregate difference at the end of the period rather than summing individual lifts.
Wins do not add up; the holdout is what actually happened.
Curated: · Written: · Reviewed:
QA-67A result is statistically significant at 0.3 percent lift. What do you do?(show answer)
I would treat statistical significance against practical significance as a question about behaviour, so I would look for the behaviour before the opinion.
Significance says an effect is unlikely to be noise, not that the effect is worth having. With enough traffic, trivially small differences become significant, and shipping them costs maintenance for no meaningful return.
Concretely, decide the minimum effect worth shipping before the test and compare the result against it. Weigh the ongoing cost of the change against the measured gain. Report the confidence interval rather than the point estimate so the plausible range is visible.
The reason for that specificity is a failure I have seen: A company shipped 12 significant but sub-1-percent wins over a year, added 3 permanent code paths and 2 configuration flags, and the aggregate holdout showed no detectable difference against a maintenance cost of roughly 2 engineer-weeks a quarter.
Effect sizes against the shipping bar.
| Result | Interval | Worth shipping at 2 percent bar |
|---|---|---|
| plus 0.3 percent | 0.1 to 0.5 | no |
| plus 2.4 percent | 1.1 to 3.7 | yes |
| plus 5.0 percent | minus 1 to 11 | not yet |
I would not consider it settled without evidence: Compare the confidence interval against the pre-registered minimum effect worth shipping.
Significant and worthwhile are different questions.
Curated: · Written: · Reviewed:
QA-68When does randomising individual users give the wrong answer?(show answer)
What separates a strong answer on experiment interference between users is knowing which number would have to move and by how much.
Randomising individuals assumes one user's treatment does not affect another's outcome, which fails in marketplaces, social products, and anywhere users share a workspace. Under interference, the measured effect can have the wrong sign.
Concretely, randomise at the unit where the interaction is contained, such as workspace, city, or time window. Accept the lower power that comes with fewer units and plan the duration accordingly. Check for spillover by comparing untreated users near treated ones.
The reason for that specificity is a failure I have seen: A marketplace tested a seller ranking boost by randomising buyers, reported a 12 percent conversion gain, and the city-level rollout produced 1 percent because the gain had come from moving the same limited supply between arms.
Same change, two randomisation units.
| Unit | Measured lift | Held on rollout |
|---|---|---|
| individual buyers | 12 percent | no |
| city | 1 percent | yes |
| time window | 2 percent | yes |
I would not consider it settled without evidence: Randomise at the level where interaction is contained and confirm the effect survives that design.
If the arms share a resource, they are not independent.
Curated: · Written: · Reviewed:
QA-69An experiment produced an effect you cannot explain. What next?(show answer)
I would frame qualitative follow-up after a surprising result around the segment it affects, because an average across segments hides the decision.
An unexplained result is not yet a finding, because without a mechanism there is no way to know where it will generalise. Watching real users is the cheapest route to a mechanism that predicts the next case.
Concretely, watch session recordings or run five interviews with users from the winning arm. Form a mechanism that explains both the direction and the size. Test the mechanism's prediction somewhere else before building on it.
The reason for that specificity is a failure I have seen: A team shipped a checkout change that won by 8 percent without understanding why, reapplied the same pattern to two other flows, and both lost, because the original gain came from removing a shipping-cost surprise those flows did not have.
One win, two failed generalisations.
| Flow | Change | Result |
|---|---|---|
| checkout | remove cost surprise | plus 8 percent |
| signup | same pattern | minus 2 percent |
| upgrade | same pattern | minus 4 percent |
I would not consider it settled without evidence: State the mechanism and confirm it predicts the outcome of a second application before generalising.
Without a mechanism you cannot tell where the result stops being true.
Curated: · Written: · Reviewed:
QA-70Your outcome takes six months to measure. How do you steer before then?(show answer)
Before committing to anything on leading indicators for a slow outcome, I would state the smallest test that could change my mind.
A leading indicator is only useful if its relationship to the slow outcome has been established on historical data, and that relationship has to be rechecked as the product changes. Choosing one by intuition produces confident steering in an unknown direction.
Concretely, find behaviours in the first weeks that historically predicted the six-month outcome and quantify the strength. Steer on the strongest one while continuing to measure the outcome itself. Re-fit the relationship each time the product changes materially.
The reason for that specificity is a failure I have seen: A subscription product steered on first-week logins as a predictor of annual renewal, and after a mobile launch the relationship weakened from 0.61 to 0.14 while the team kept optimising logins for 3 quarters.
Predictive strength before and after a platform change.
| Indicator | Correlation before | After |
|---|---|---|
| first-week logins | 0.61 | 0.14 |
| second-week import | 0.55 | 0.58 |
| teammate invited | 0.49 | 0.52 |
I would not consider it settled without evidence: Re-fit the indicator against the realised outcome for the most recent cohort that has aged fully.
A leading indicator is a relationship, and relationships change.
Curated: · Written: · Reviewed:
QA-71How do you evaluate an optional feature's success?(show answer)
I would handle measuring a feature nobody is required to use by making the trade-off explicit rather than promising both sides of it.
Adoption alone says nothing without a comparison, because the people who choose an optional feature differ from those who do not. Success has to be expressed against what those same users would otherwise have done.
Concretely, compare adopters against a matched group or run a randomised availability test. Measure the outcome the feature was meant to change rather than the feature's own usage. Decide in advance what adoption level would make the feature not worth maintaining.
The reason for that specificity is a failure I have seen: A team celebrated 34 percent adoption of a saved-views feature, and a randomised availability test later showed the adopting group's task completion was identical to the control, while the feature carried 2 engineer-weeks a quarter of maintenance.
Adoption against outcome.
| Group | Adopted | Task completion |
|---|---|---|
| feature available | 34 percent | 71 percent |
| withheld | n/a | 70 percent |
| heavy adopters | 100 percent | 88 percent |
I would not consider it settled without evidence: Withhold the feature from a random group and compare the outcome, not the usage.
Usage is not value; the comparison group tells you which you have.
Curated: · Written: · Reviewed:
QA-72How do you avoid launching something you cannot measure?(show answer)
I would ground instrumentation as part of the definition of done in what customers do after launch rather than in what they said in the room.
Instrumentation added after launch produces a gap in the record exactly when the decision is made, and backfilling it is usually impossible. The events and their definitions belong in the specification with the feature.
Concretely, write the events, properties, and the question each answers into the spec before build. Verify the events in staging and in the first hours of production. Refuse to call the work done until a query returns the expected shape.
The reason for that specificity is a failure I have seen: A launch ran for 5 weeks with no funnel instrumentation, the review had to rely on a support-ticket count, and the backfill was impossible because the intermediate steps had never emitted an event.
What the review could answer.
| Question | With events | Without |
|---|---|---|
| where users drop | yes | no |
| how many started | yes | partial |
| which segment adopted | yes | no |
I would not consider it settled without evidence: Run the analysis query against real production data within the first day of launch and confirm it returns usable rows.
A feature without events ships blind and stays that way.
Curated: · Written: · Reviewed:
QA-73Your satisfaction score fell twelve points. What do you do first?(show answer)
I would start reading a satisfaction score from the customer problem it is supposed to change, not from the feature already on the roadmap.
A composite satisfaction score aggregates several populations and several causes, so the movement is a prompt to disaggregate rather than a finding. Response rate and mix changes explain a large share of apparent movements.
Concretely, check the response rate and respondent mix before anything else. Cut by segment, tenure, and recent support contact. Read the free text for the affected cut rather than the whole sample.
The reason for that specificity is a failure I have seen: A team ran a retention campaign after a 12-point drop, and the drop came from a survey prompt newly shown to trial users, who were 38 percent of responses that month against 6 percent before.
Same population, different respondents.
| Group | Share of responses before | After | Score |
|---|---|---|---|
| trial users | 6 percent | 38 percent | 41 |
| paying users | 94 percent | 62 percent | 78 |
| blended | 100 percent | 100 percent | 64 |
I would not consider it settled without evidence: Compare the respondent mix against the previous period before acting on the movement.
Score movements are usually mix movements until proven otherwise.
Curated: · Written: · Reviewed:
QA-74Finance wants a number for next year. How do you produce one you can defend?(show answer)
The first thing I would settle about forecasting a metric for a plan is what evidence would make me drop the idea.
A defensible forecast separates the momentum the business already has from the effect of work not yet done, so a miss can be attributed to one or the other. A single blended number cannot be debugged and will be treated as a commitment anyway.
Concretely, project the baseline from existing cohorts with no new work, then add each planned initiative's expected contribution separately with its confidence. Publish the parts, not just the total. Review monthly and correct the component that moved.
The reason for that specificity is a failure I have seen: A plan committed to 8.4 million on a single blended line, landed at 6.1 million, and the post-mortem could not say whether the baseline had decayed or the three new initiatives underdelivered, so the next plan repeated the structure.
The forecast taken apart.
| Component | Forecast | Actual |
|---|---|---|
| existing cohort baseline | 5.6M | 5.4M |
| pricing change | 1.4M | 0.5M |
| new segment | 1.4M | 0.2M |
I would not consider it settled without evidence: Show the forecast as baseline plus named contributions and confirm each part can be checked on its own next month.
Forecast in parts so a miss tells you something.
Curated: · Written: · Reviewed:
QA-75Nobody looks at the team dashboard. What is wrong with it?(show answer)
With dashboards people actually use, I would separate what we believe from what we have actually observed.
A dashboard is used when it answers a question somebody has on a schedule they already keep, which means it has an audience and a cadence rather than a completeness goal. Adding charts to an unused dashboard makes it less likely to be read.
Concretely, name the audience, the recurring decision, and the cadence, then keep only the charts that bear on that decision. Put the primary metric and its target at the top with the current value. Review quarterly and remove anything nobody has cited.
The reason for that specificity is a failure I have seen: A team maintained a 40-chart dashboard that took 3 hours a week to keep accurate, and an access log showed 11 views in a quarter, 9 of them by the person maintaining it.
Effort against readership.
| Dashboard | Charts | Weekly upkeep | Views per quarter |
|---|---|---|---|
| everything | 40 | 3 hours | 11 |
| weekly decision | 6 | 20 minutes | 180 |
I would not consider it settled without evidence: Check the dashboard access log and remove every panel nobody has opened or cited this quarter.
A chart with no decision behind it is maintenance with no reader.
Curated: · Written: · Reviewed:
QA-76Your flagship bet did not work. How do you report it?(show answer)
I would answer presenting a negative result by naming the decision it feeds and who has to live with that decision.
A negative result is information the company paid for, and its value is realised only if it is reported clearly enough to change the next decision. Softening it preserves the belief that produced the bet and guarantees a repeat.
Concretely, state the original hypothesis, the result against the pre-registered threshold, and the decision you are taking. Name what you would have to believe to continue. Record the finding where the next person planning something similar will find it.
The reason for that specificity is a failure I have seen: A team reported a failed 2-quarter bet as promising early signals, kept 4 engineers on it for another quarter, and the eventual stop came after 28 team-weeks with the same data that had been available at the first review.
What each review knew and did.
| Review | Result against threshold | Decision |
|---|---|---|
| quarter 1 | 31 percent of target | continue |
| quarter 2 | 29 percent of target | continue |
| quarter 3 | 30 percent of target | stop |
I would not consider it settled without evidence: Report the result against the threshold written before the bet and state the decision taken.
The expensive part of a failed bet is the quarter you spend not admitting it.
Curated: · Written: · Reviewed:
QA-77What does a good specification contain, and what should it leave to the team?(show answer)
My approach to writing a specification engineers can build from starts with the outcome the business is buying rather than the output the team ships.
A specification fixes the user outcome, the rules, and the edge cases, and leaves the implementation to the people accountable for it. Specifying the solution removes the expertise you hired without removing the responsibility.
Concretely, describe the behaviour the user should see, enumerate the rules including what happens when each input is missing or wrong, and state the acceptance criteria. Leave data structures, libraries, and sequencing to engineering. Review the edge-case list with the team before estimation.
The reason for that specificity is a failure I have seen: A spec that described screens but not the rule for partially failed imports produced 3 different behaviours across 2 clients and the API, and reconciling them after launch took 4 weeks and a data repair for 640 accounts.
Edge cases named in the spec against defects found after launch.
| Area | Edge cases in spec | Post-launch defects |
|---|---|---|
| import partial failure | 0 | 7 |
| duplicate rows | 4 | 1 |
| permission denied | 3 | 0 |
I would not consider it settled without evidence: Walk the edge-case list with an engineer and confirm every one has a defined behaviour before estimation.
Own the rules and the outcome; rent the implementation from people who do it better.
Curated: · Written: · Reviewed:
QA-78Two weeks before launch the work will not fit. What do you cut?(show answer)
For cutting scope against a fixed date, I would write down the assumption that has to hold before I argue about the solution.
Scope is cut along the seam of the user's complete journey, keeping one path whole rather than every path partial. Cutting evenly across features produces a release where nothing works end to end.
Concretely, identify the narrowest complete journey a real customer could run and protect it. Remove whole capabilities rather than trimming each one. Tell the affected teams and customers what is not in the release rather than letting them discover it.
The reason for that specificity is a failure I have seen: A team trimmed 20 percent from every feature to hold a date, and the launch shipped an import that could not map fields, an export that ignored filters, and a report with no date range, producing 214 support contacts in 9 days.
Two ways to remove the same work.
| Approach | Features shipped | Journeys completable |
|---|---|---|
| trim every feature 20 percent | 4 partial | 0 |
| cut two features whole | 2 complete | 1 |
| move the date 2 weeks | 4 complete | 3 |
I would not consider it settled without evidence: Walk the remaining journey end to end as a customer and confirm it completes without a workaround.
Ship one whole path rather than four unfinished ones.
Curated: · Written: · Reviewed:
QA-79Engineering's estimate is twice what you hoped. How do you respond?(show answer)
I would treat working with engineering on estimates as a question about behaviour, so I would look for the behaviour before the opinion.
An estimate that surprises you is information about scope or risk you did not have, so the productive move is to find the driver rather than to negotiate the number. Pressure applied to an estimate moves the reported figure and not the work.
Concretely, ask which part carries the uncertainty and what would reduce it, such as a spike or a decision you can make now. Offer scope reductions that remove the expensive part rather than asking for the same work sooner. Re-estimate after the uncertainty is resolved rather than at the moment of disagreement.
The reason for that specificity is a failure I have seen: A PM pushed an 8-week estimate down to 5 weeks in planning, the work took 11 weeks including 2 weeks of rework, and the team stopped giving estimates below their worst case for the following two quarters.
Where the estimate came from.
| Component | Estimate | Uncertainty |
|---|---|---|
| known CRUD work | 3 weeks | low |
| legacy data migration | 4 weeks | high |
| third-party rate limits | 1 week | high |
I would not consider it settled without evidence: Ask the team to name the largest uncertainty and agree a spike or decision that would remove it.
Argue with the scope, not with the arithmetic.
Curated: · Written: · Reviewed:
QA-80What belongs in a launch plan besides the code being live?(show answer)
What separates a strong answer on a launch plan beyond the release is knowing which number would have to move and by how much.
A launch is a coordinated change across product, support, sales, documentation, and pricing, and the release is only the part with a deploy button. Missing pieces surface as customer confusion that the product team hears about last.
Concretely, list every function that touches the customer and what each needs before the date, including enablement, documentation, pricing, and the support macro. Name an owner and a readiness check for each. Hold a go or no-go where each owner confirms rather than assumes.
The reason for that specificity is a failure I have seen: A feature launched to every account with no support documentation and no pricing decision, generating 380 tickets in 4 days, of which 140 asked whether it would be charged for, which nobody could answer for 9 days.
Readiness by function at go time.
| Function | Artefact | Ready |
|---|---|---|
| support | macro and article | no |
| pricing | decision recorded | no |
| sales | enablement session | yes |
| docs | published page | no |
I would not consider it settled without evidence: Run a go or no-go where every named owner confirms readiness against their own checklist.
The deploy is the cheapest part of a launch.
Curated: · Written: · Reviewed:
QA-81How do you roll out a risky change?(show answer)
I would frame staged rollout and the stop condition around the segment it affects, because an average across segments hides the decision.
A staged rollout is only a safety mechanism if there is a metric being watched and a threshold that stops it, otherwise it is a slow way to reach everyone. The stop condition must be agreed before the first stage.
Concretely, define the stages, the metric watched at each, and the threshold that halts the rollout. Give each stage enough time and volume to detect the effect you are worried about. Keep the rollback path tested rather than assumed.
The reason for that specificity is a failure I have seen: A change rolled to 1 percent, then 50 percent within a day because the first stage looked fine, and the error it introduced appeared only for accounts with more than 10,000 records, which were 0 percent of the first stage and 12 percent of the second.
Stage composition against the risk.
| Stage | Accounts | Large accounts included | Hours before next |
|---|---|---|---|
| 1 percent | 210 | 0 | 4 |
| 50 percent | 10,400 | 1,240 | 0 |
| revised plan | 5 percent | 60 | 48 |
I would not consider it settled without evidence: Confirm each stage includes the segment most likely to expose the risk and runs long enough to detect it.
A rollout without a stop condition is a deploy with extra steps.
Curated: · Written: · Reviewed:
QA-82A launch is causing problems but not an outage. How do you decide whether to roll back?(show answer)
Before committing to anything on deciding to roll back, I would state the smallest test that could change my mind.
The decision compares the cost of the damage while it continues against the cost of reverting and relaunching, and it is made faster by agreeing the threshold before the launch. Debating it live is how a small defect becomes a long one.
Concretely, agree the rollback threshold during launch planning in terms customers feel, such as failed operations per hour. Measure against it rather than against opinion. Roll back on the threshold and investigate afterwards rather than investigating first.
The reason for that specificity is a failure I have seen: A team debated a rollback for 6 hours while a pricing display defect showed the wrong currency to 3,200 sessions, and 40 orders were placed at the displayed amount, which cost 11,000 in honoured pricing.
Cost while the debate ran.
| Hour | Affected sessions | Orders at wrong price |
|---|---|---|
| 1 | 480 | 4 |
| 3 | 1,700 | 19 |
| 6 | 3,200 | 40 |
I would not consider it settled without evidence: Compare the live error rate against the threshold agreed before launch and act on it without reopening the debate.
Decide the threshold before you are the person who has to defend the launch.
Curated: · Written: · Reviewed:
QA-83A committed date will be missed. What do you tell the customer?(show answer)
I would handle communicating a slip to customers by making the trade-off explicit rather than promising both sides of it.
The cost of a slip is mostly the surprise, so the communication happens as soon as the slip is known rather than when it is certain. A new date given without the reason and the new confidence invites the same conversation again.
Concretely, tell the affected customers as soon as the estimate changes, give the revised date with what changed, and say what they can rely on in the meantime. Offer a partial or a workaround where one exists. Follow the new date with a check-in whether or not it holds.
The reason for that specificity is a failure I have seen: A team held a known 5-week slip until the original date passed, and 2 of the 6 affected customers had already scheduled internal training, one of which cost them 14 staff days of wasted time and triggered a contract renegotiation.
Notice period against escalation.
| Notice given | Customers | Escalations |
|---|---|---|
| 5 weeks ahead | 6 | 0 |
| on the date | 6 | 4 |
| after the date | 6 | 6 |
I would not consider it settled without evidence: Confirm every affected customer heard about the slip from you before the original date arrived.
Customers plan around your dates, so they should hear about changes before your calendar does.
Curated: · Written: · Reviewed:
QA-84How do you review a launch that went badly without blaming a person?(show answer)
I would ground running a launch postmortem in what customers do after launch rather than in what they said in the room.
A review is useful when it produces a change to how the work is done, which requires the contributing factors rather than a responsible individual. Naming a person ends the inquiry at the point where the system's defects begin.
Concretely, build the timeline from records rather than memory, identify the decisions and the information available at each. Produce a small number of changes with owners and dates. Check those changes at the next launch rather than filing the document.
The reason for that specificity is a failure I have seen: A review concluded that one engineer had missed a step, produced no process change, and the same failure occurred at the next two launches because the step was not in any checklist either time.
Two reviews of similar failures.
| Outcome | Actions with owners | Recurrence |
|---|---|---|
| named an individual | 0 | twice |
| named contributing factors | 3 | none |
I would not consider it settled without evidence: Check at the next launch whether the changes from the last review were actually in place.
A review that names a person leaves the cause in place.
Curated: · Written: · Reviewed:
QA-85Your largest customer escalates to the executive team. What do you do?(show answer)
I would start handling an escalation from a major account from the customer problem it is supposed to change, not from the feature already on the roadmap.
An escalation needs a fast acknowledgement, a named owner, and a factual account of what will change and when, which is a different job from deciding whether to build what they asked for. Conflating the two turns every escalation into roadmap capture.
Concretely, take ownership, establish the facts, and give the customer a dated plan for the part you will fix. Separately assess whether the underlying request belongs on the roadmap for everyone. Feed the pattern back into prioritisation rather than treating the account as the requirement.
The reason for that specificity is a failure I have seen: A company built 2 bespoke capabilities for its largest account over 3 quarters, which consumed 31 percent of one team's capacity and none of which was used by any other customer, while a broadly requested audit trail slipped out of the year.
Bespoke work against broader demand.
| Capability built | Accounts requesting | Accounts using after 2 quarters |
|---|---|---|
| custom export schema | 1 | 1 |
| bespoke approval flow | 1 | 1 |
| audit trail | 47 | not built |
I would not consider it settled without evidence: Check how many other accounts requested the same capability before committing roadmap capacity to it.
Fix the customer's problem now and decide the roadmap question separately.
Curated: · Written: · Reviewed:
QA-86A senior leader wants a direction you believe is wrong. How do you handle it?(show answer)
The first thing I would settle about disagreeing with a senior stakeholder is what evidence would make me drop the idea.
Disagreement is expressed as evidence and a proposed test, and once the decision is made it is executed properly whether or not you won. Quiet non-compliance is the option that destroys trust and produces the worst version of both directions.
Concretely, state the specific claim you disagree with and the evidence behind your position. Propose the cheapest test that would resolve it and offer to run it. If the decision goes against you, execute it fully and put the check in place that will show who was right.
The reason for that specificity is a failure I have seen: A PM who disagreed with a pricing change implemented it half-heartedly without the measurement they had argued for, and when revenue fell 7 percent nobody could tell whether the change or the implementation was responsible.
Three ways to lose the argument.
| Response | Decision instrumented | Learned who was right |
|---|---|---|
| comply quietly | no | no |
| execute and measure | yes | yes |
| escalate repeatedly | no | no |
I would not consider it settled without evidence: Confirm the agreed measurement is in place before the decision is executed, whichever way it went.
Disagree with evidence, commit with rigour, and instrument the outcome.
Curated: · Written: · Reviewed:
QA-87Two teams are optimising metrics that pull against each other. What do you do?(show answer)
With aligning two teams with conflicting goals, I would separate what we believe from what we have actually observed.
Conflicting local metrics produce rational teams doing damaging things, and no amount of collaboration fixes an incentive that rewards the damage. The fix is a shared measure or an explicit guardrail owned jointly.
Concretely, make the conflict explicit with the numbers from both sides. Propose a shared outcome metric or add the other team's measure as a guardrail on each. Get the decision from whoever owns both, rather than asking the teams to negotiate goodwill.
The reason for that specificity is a failure I have seen: A growth team optimising signups and a monetisation team optimising conversion ran opposing prompts on the same screen for 7 weeks, and the combined result was 3 percent fewer signups and 1 percent lower conversion than before either change.
Local wins, joint loss.
| Measure | Growth view | Monetisation view | Combined |
|---|---|---|---|
| signups | plus 4 percent | not measured | minus 3 percent |
| conversion | not measured | plus 2 percent | minus 1 percent |
I would not consider it settled without evidence: Measure the combined outcome across both teams' metrics rather than each team's own.
Two teams cannot collaborate their way out of opposing incentives.
Curated: · Written: · Reviewed:
QA-88How do you keep an executive informed without an escalation every week?(show answer)
I would answer managing up without hiding problems by naming the decision it feeds and who has to live with that decision.
Executives need the small number of things that would change their decisions, delivered on a predictable cadence, with problems arriving before they arrive from someone else. A status that only contains good news trains the reader to discount it.
Concretely, send a short written update on a fixed rhythm with progress against the outcome, the decisions you need, and the risks with what you are doing about each. Escalate early and with a proposed option rather than a question. Never let a problem reach them from another function first.
The reason for that specificity is a failure I have seen: A quarterly review was the first time an executive heard that a key integration had been blocked for 9 weeks, and the recovery plan then required moving 2 engineers from another team mid-quarter.
How the risk surfaced.
| Week | Status sent | Risk named |
|---|---|---|
| 1 to 8 | green | no |
| 9 | green | no |
| 10, quarterly review | red | by the partner team |
I would not consider it settled without evidence: Check whether any risk in the last month reached leadership from another function before it reached them from you.
Bad news travels; make sure it travels from you.
Curated: · Written: · Reviewed:
QA-89Design and product disagree about a flow. How do you resolve it?(show answer)
My approach to working with design on a contested decision starts with the outcome the business is buying rather than the output the team ships.
The disagreement is resolved by deciding what evidence would settle it and getting that evidence, since both parties are usually optimising different real things. Resolving by seniority produces compliance and loses the argument's information.
Concretely, state both positions as predictions about user behaviour. Pick the cheapest instrument that discriminates between them, such as a five-user test or a small experiment. Agree in advance that the result decides, and hold to it.
The reason for that specificity is a failure I have seen: A flow disagreement was settled by the PM's preference, the flow shipped, and a usability round afterwards found 7 of 8 users taking the path design had predicted, costing 3 weeks of rework.
Cost of resolving it two ways.
| Method | Cost | Rework after |
|---|---|---|
| five-user test | 1 day | 0 weeks |
| decide by seniority | 0 days | 3 weeks |
| full experiment | 3 weeks | 0 weeks |
I would not consider it settled without evidence: Agree the instrument and the decision rule before running it, then follow the result.
Turn the disagreement into a prediction, then buy the cheapest answer.
Curated: · Written: · Reviewed:
QA-90You are raising prices. How do you handle customers who are already paying?(show answer)
For pricing changes for existing customers, I would write down the assumption that has to hold before I argue about the solution.
A price change to existing customers is a trust event as much as a revenue event, and the plan has to cover notice, grandfathering, and what the customer gets for the difference. Applying it silently at renewal converts a pricing decision into a churn event.
Concretely, give notice well ahead of the renewal date, explain what changed in the offer, and grandfather or phase the increase for long-tenured accounts. Brief support and sales with the exact numbers per segment. Track churn and downgrade by cohort after the change rather than in aggregate.
The reason for that specificity is a failure I have seen: A 22 percent increase applied at renewal with 14 days' notice produced 9 percent churn in the affected cohort against a 3 percent baseline, and 60 percent of the churned accounts had been customers for more than 3 years.
Notice period against churn in the affected cohort.
| Notice | Increase | Churn |
|---|---|---|
| 14 days | 22 percent | 9 percent |
| 90 days with phasing | 22 percent | 4 percent |
| no change | 0 | 3 percent |
I would not consider it settled without evidence: Track churn and downgrades for the affected cohort against an unaffected one for two renewal cycles.
Price changes are remembered by the people you kept.
Curated: · Written: · Reviewed:
QA-91A feature would work better with more personal data. How do you decide?(show answer)
I would treat privacy and data collection decisions as a question about behaviour, so I would look for the behaviour before the opinion.
Data collection is a product decision with a legal boundary and a trust cost, and the defensible position is the minimum data that delivers the outcome, with a stated purpose and a retention limit. Collecting because it might be useful later is the pattern that produces both breaches and regulatory findings.
Concretely, name the purpose, the minimum field set, the retention period, and the user-facing explanation before building. Get the review from whoever is accountable for privacy rather than assuming. Prefer aggregate or on-device processing where it achieves the same outcome.
The reason for that specificity is a failure I have seen: A team logged full page URLs including query strings for analytics, which captured customer record identifiers for 14 months, and the cleanup required purging 3 systems plus a disclosure to 31 enterprise accounts.
Minimum data against what was collected.
| Field | Needed for the outcome | Collected |
|---|---|---|
| page template | yes | yes |
| full URL with query | no | yes |
| account identifier | aggregate only | raw |
I would not consider it settled without evidence: Sample the collected records and confirm no field falls outside the stated purpose.
Collect for a purpose you can state to the person it is about.
Curated: · Written: · Reviewed:
QA-92How do you prepare support for a new feature?(show answer)
What separates a strong answer on deciding what support needs before a launch is knowing which number would have to move and by how much.
Support cost is a product decision made at design time, and the questions customers will ask are predictable from the flow. Preparing answers is cheaper than absorbing the contacts, and the contact pattern is also the fastest feedback the product gets.
Concretely, walk the flow with a support lead and list the likely questions and failure states. Write the macros and the article before launch and instrument the contact reasons. Feed the first two weeks of contact reasons back into the backlog.
The reason for that specificity is a failure I have seen: A launch with no prepared answers produced 380 contacts in the first fortnight at an average handling time of 11 minutes, roughly 70 support hours, and the top reason was a single unclear error message.
Contacts in the first fortnight by reason.
| Reason | Contacts | Fixed in product |
|---|---|---|
| unclear error message | 180 | week 3 |
| pricing question | 140 | article only |
| permission confusion | 60 | week 5 |
I would not consider it settled without evidence: Instrument contact reasons for the first two weeks and confirm the top reason has an owner.
Support volume is a design review you receive after shipping.
Curated: · Written: · Reviewed:
QA-93A new team is taking over an area you own. How do you hand it over?(show answer)
I would frame onboarding a new engineering team to a product area around the segment it affects, because an average across segments hides the decision.
A handover transfers the decisions and their reasons, not only the code and the backlog, because the expensive knowledge is why things are the way they are. Without it the new team re-litigates settled questions and reintroduces solved failures.
Concretely, write the area's decision record, the known failure modes, the customers who care most, and the open questions. Run the first planning cycle jointly rather than handing over a document. Stay available for a defined period with an end date.
The reason for that specificity is a failure I have seen: An area was handed over with a backlog and no decision history, and within 2 quarters the new team had rebuilt a rate limiter that had been removed deliberately, causing a repeat of an incident from 18 months earlier.
What the handover carried.
| Artefact | Included | Repeat incidents |
|---|---|---|
| backlog only | yes | 2 |
| decision record | no | n/a |
| joint planning cycle | no | n/a |
I would not consider it settled without evidence: Ask the receiving team to explain three non-obvious decisions in the area and correct the ones they cannot.
Hand over the reasons, or watch the old mistakes come back.
Curated: · Written: · Reviewed:
QA-94You need work from a team that does not report to you and has its own priorities. How do you get it?(show answer)
Before committing to anything on influencing without authority, I would state the smallest test that could change my mind.
Influence comes from making the other team's goal easier to reach, which requires knowing what they are measured on before asking for anything. A request framed entirely around your own outcome competes with everything already on their plan.
Concretely, learn their goals and constraints, then frame the request in terms of what it does for them, with the work you have already done to reduce their cost. Bring a written proposal with the specific ask and the date. Escalate to the shared owner only after the direct route has been tried and recorded.
The reason for that specificity is a failure I have seen: A PM escalated to a shared director after 2 unanswered messages, the platform team then delivered the request 4 weeks late and deprioritised two other requests from the same PM over the following quarter.
Two framings of the same request.
| Framing | Their stated goal addressed | Delivered |
|---|---|---|
| our launch needs it | no | 4 weeks late |
| reduces your incident load | yes | on time |
I would not consider it settled without evidence: State the other team's goal back to them and confirm they agree your request helps it.
Make the ask cheap for them before making it urgent for you.
Curated: · Written: · Reviewed:
QA-95Tell me about a product decision you got wrong.(show answer)
I would handle a behavioural answer about a failure by making the trade-off explicit rather than promising both sides of it.
The answer needs a real miss, the decision you owned, what the evidence actually said, and the change you made to how you decide since. A story that ends in an unqualified success removes the only part being scored.
Concretely, name the decision, the information you had, and the outcome with a number. State the cause you owned rather than the circumstances. Close with the specific change to your process and an example of it working since.
The reason for that specificity is a failure I have seen: A candidate answered with a story where a launch slipped because of another team, gave no number, and the debrief recorded no evidence of self-assessment, which cost an otherwise strong loop.
What the interviewer is scoring.
| Element | Weak answer | Strong answer |
|---|---|---|
| decision owned | another team's | mine, named |
| measured outcome | none | 22 percent below target |
| change since | general lesson | specific, applied twice |
I would not consider it settled without evidence: Check your own story names a decision you made, a measured outcome, and a change you have applied since.
The miss is the story; the change is the point.
Curated: · Written: · Reviewed:
QA-96Tell me about a time you disagreed with an engineer or a designer.(show answer)
I would ground a behavioural answer about conflict in what customers do after launch rather than in what they said in the room.
The scored content is how the disagreement was resolved and what evidence settled it, not who turned out to be right. A story where the other party is obviously foolish reads as an inability to work with people who think differently.
Concretely, state both positions fairly, including the strongest version of theirs. Describe the evidence you sought and the decision rule you agreed. Say what happened, including where their position was better than yours.
The reason for that specificity is a failure I have seen: A candidate described overruling an engineer who turned out to be right about a scaling limit that failed at 3,000 concurrent users, presented the story as a win, and 2 of the 4 interviewers recorded no evidence that the candidate had learned anything from it.
Two versions of the same story.
| Element | Version A | Version B |
|---|---|---|
| other position | dismissed | stated fairly |
| resolution | seniority | agreed test |
| outcome | I was right | they were, and here is the change |
I would not consider it settled without evidence: Confirm your story states the other position at its strongest and names what settled the question.
Steelman the other side; that is the part being assessed.
Curated: · Written: · Reviewed:
QA-97Tell me about a time you changed someone's mind.(show answer)
I would start a behavioural answer about influence from the customer problem it is supposed to change, not from the feature already on the roadmap.
The interviewer is checking whether a person outside your reporting line changed course because of something you produced, so the story needs the artefact and the decision that moved. Being persuasive in a meeting is not evidence unless something changed afterwards.
Concretely, name the person's position and why it was reasonable. Describe the artefact you built, such as an analysis, a prototype, or a customer session you arranged. State what they did differently and when.
The reason for that specificity is a failure I have seen: A candidate described a persuasive presentation with no stated outcome, and the panel recorded that nothing in the story showed a decision changing, which scored as no evidence against the competency.
What made the difference.
| Artefact | Decision changed | Time to change |
|---|---|---|
| slide argument | no | n/a |
| 3 recorded customer sessions | yes | same week |
| prototype with data | yes | next planning cycle |
I would not consider it settled without evidence: Check your story names the artefact and the decision that changed after it.
Influence is a decision that moved, not a meeting that went well.
Curated: · Written: · Reviewed:
QA-98You join a team with no roadmap, unclear goals, and a large backlog. What do you do first?(show answer)
The first thing I would settle about handling ambiguity early in a role is what evidence would make me drop the idea.
The first job is to establish what is true before proposing what to do, because a plan built on a misread of the situation costs the credibility you need to execute it. Structure comes from evidence gathered in the first weeks, not from a framework imported on day one.
Concretely, talk to the team, read the last two quarters of decisions, look at the usage and support data, and speak to customers. Publish what you found and what you propose, separating the two. Pick one visible problem to fix early so the proposals are read by people who have seen you deliver.
The reason for that specificity is a failure I have seen: A PM published a reorganised roadmap in week 2 based on the backlog alone, and 3 of the top 5 items had already been tried and abandoned for reasons recorded in documents they had not read.
First-month sequence against outcome.
| Week | Action | Result |
|---|---|---|
| 1 to 2 | interviews and document review | 3 abandoned items understood |
| 3 | publish findings | disagreement surfaced early |
| 4 | ship one visible fix | proposals read seriously |
I would not consider it settled without evidence: Check your proposal against the last two quarters of decisions and confirm you know why each abandoned item was stopped.
Find out what is true before announcing what is next.
Curated: · Written: · Reviewed:
QA-99What would you have accomplished ninety days into this role?(show answer)
With deciding what to do in your first ninety days, I would separate what we believe from what we have actually observed.
Ninety days is enough to understand the area, deliver one meaningful improvement, and establish how decisions will be made, and not enough to redirect a strategy. Promising a transformation signals that the candidate has not run the clock before.
Concretely, plan for learning in the first month, a delivered improvement in the second, and a proposed direction with evidence in the third. Name the artefacts each produces. Keep the delivered improvement small enough to be certain and visible enough to matter.
The reason for that specificity is a failure I have seen: A PM committed to a new strategy and a replatform in the first quarter, delivered neither, and spent the second quarter rebuilding credibility with a team that had absorbed 5 weeks of planning churn.
A ninety-day plan with artefacts.
| Phase | Artefact | Verifiable |
|---|---|---|
| days 1 to 30 | findings document | yes |
| days 31 to 60 | shipped improvement | yes |
| days 61 to 90 | proposed direction with evidence | yes |
I would not consider it settled without evidence: Check the plan produces one delivered improvement inside ninety days rather than only documents.
Learn, ship something real, then propose the direction.
Curated: · Written: · Reviewed:
QA-100What would you ask us at the end of the loop?(show answer)
I would answer questions a candidate should ask the interviewer by naming the decision it feeds and who has to live with that decision.
The questions reveal what the candidate intends to be accountable for, so the useful ones probe how decisions get made, what happened to the last person in the role, and how success will be measured. Questions about perks measure nothing the panel is scoring.
Concretely, ask how the last three product decisions were made and by whom. Ask what would make the first year a success and who judges it. Ask what the team disagrees about right now, and listen for whether disagreement is safe.
The reason for that specificity is a failure I have seen: A candidate used all 10 minutes on process and tooling questions, asked nothing about how decisions were made, and 3 of the 5 panellists recorded a doubt about whether the candidate would be effective in a company where that answer was contested.
Questions against what they reveal.
| Question | Reveals |
|---|---|
| how were the last three decisions made | where authority sits |
| what does success look like in a year | whether it is measurable |
| what does the team disagree about now | whether disagreement is safe |
I would not consider it settled without evidence: Check your questions would change your own decision to accept an offer, and drop the ones that would not.
Ask the questions whose answers would change your mind about joining.
Curated: · Written: · Reviewed:
