Overview
Curated: · Written: · Reviewed:
Prioritization is an accountable choice under constrained capacity and uncertainty
The mental model: a decision, not a calculation
Prioritization is the act of choosing what a constrained team will do and—more importantly—what it will not do, under uncertainty, with an owner who is accountable for the outcome. The scoring frameworks (RICE, WSJF, MoSCoW) are decision aids, not the decision. Every one of them compresses judgment into numbers so a group can argue about the same thing; none of them supplies the goal, the constraints, or the accountability.
The mechanism underneath all of them is the same:
- A goal converts "valuable" from an adjective into a measurable direction.
- Capacity converts the backlog from a wish list into a competition—only work that fits the team's actual throughput is really being ranked.
- Evidence attaches confidence to each estimate.
- A framework makes the comparison legible enough to argue about.
- An owner commits, records the trade-off, and sets the conditions for revisiting.
Skip a layer and the output is noise. A RICE score with no goal behind it ranks work against nothing. A MoSCoW session with no capacity cap produces a room where everything is Must. The framework is step four, not step one.
The interview sequence the question is fishing for
The prompt is almost always "how would you prioritize these three features?" The interviewer is testing the sequence, not arithmetic:
- Name the goal and the strategy the ranking serves. Growth, retention, reliability, a launch date—"ranking" is meaningless until you say what it's for.
- Define the comparison basis. Horizon, capacity, eligible work, outcome metric, decision owner, review cadence.
- Then apply a framework—RICE, WSJF, MoSCoW, or plain judgment.
- State the trade-off and the re-check. What you're deferring, and what evidence would change your mind next week.
A candidate who reaches for RICE before naming the objective fails the question even with correct arithmetic. Interviewers probe the frame: "what's the goal here?", "who owns this decision?", "what would make you re-order next sprint?" A weak answer defends the numbers; a strong answer defends the frame first and treats the score as a conversation starter.
One framing detail that separates senior answers: the eligible set must include product discovery, reliability, security, privacy, accessibility, compliance, maintenance, migration and retirement work. If only feature ideas get scored, features consume the capacity while essential system work stays invisible until it becomes an incident.
RICE in working detail, with a worked score
RICE = Reach × Impact × Confidence ÷ Effort.
- Reach — a count of people or events affected, over one fixed time window, without double-counting. "2,400 users hit the checkout error per month, from the last 90 days of logs"—not "all our customers, eventually."
- Impact — the effect per person on a named goal, on a defined scale (e.g. 3 = massive, 2 = high, 1 = medium, 0.5 = low, 0 = none). "Improves the experience" is not an impact.
- Confidence — a discount for evidence gaps applied to the whole estimate: 100% when you have solid data for every input, 80% when you have some, 50% when it's mostly a hypothesis. It is not how much you want the project.
- Effort — person-months of total team cost: product, design, engineering, data, research, security, operations, rollout, and ongoing maintenance—not coding time for the first release.
Worked example (hypothetical figures, one quarter, one team of 6 engineers ≈ 4.5 person-months of feature capacity after maintenance and on-call):
| Item | Reach (users/qtr) | Impact | Confidence | Effort (person-mo) | Score |
|---|---|---|---|---|---|
| A: Fix checkout error | 7,200 | 1 | 100% | 1 | 7,200 |
| B: Referral program | 30,000 | 0.5 | 50% | 4 | 1,875 |
| C: Onboarding redesign | 9,000 | 2 | 80% | 3 | 4,800 |
Trace the arithmetic: A = 7,200 × 1 × 1.00 ÷ 1 = 7,200. B = 30,000 × 0.5 × 0.50 ÷ 4 = 1,875. C = 9,000 × 2 × 0.80 ÷ 3 = 4,800. Ranking: A, C, B. Capacity check: A + C = 4 person-months, which fits the 4.5 available; B does not fit this quarter regardless of its score. The score ranks; capacity decides.
Sensitivity check—the part interviewers actually probe
Made-up inputs produce a confident-looking decimal rank. When plausible ranges overlap, the ordering is noise, and you should say so out loud. In the table above, the B-versus-C ordering rests on two judgment inputs, and neither is data. Drop C's impact from 2 to 1 and C scores 9,000 × 1 × 0.80 ÷ 3 = 2,400—the gap over B narrows from 2,925 to 525. Raise B's confidence from 50% to 80%—a sponsor's optimism rather than new evidence—and B scores 30,000 × 0.5 × 0.80 ÷ 4 = 3,000, which passes C-at-Impact-1 (2,400) and closes most of the gap on C-at-Impact-2 (4,800). Two one-point swings on judgment inputs reorder the bottom of the list; the ranking is only as good as the inputs. The senior move is to state the ranges, not just the point estimates, and to improve the information (a small experiment) when ranges overlap on a decision that matters.
Other failure points to have ready:
- Effort swamping. Effort is often the only input anyone can estimate credibly, so it dominates the division and the backlog fills with small, low-value work. Sanity-check the top of the list against the stated goal, not just the sort order.
- Reach/Impact double-counting. If Impact already encodes "affects more people," multiplying by Reach counts the same effect twice. Keep Reach a pure count and Impact a pure per-person effect.
- Mixed time windows. Reach per quarter for one item and per year for another makes the comparison meaningless. Fix one window before scoring anything.
- Confidence as smuggled opinion. Anchor it to the actual evidence behind each input and record provenance: observed usage data for reach, representative research for impact, estimates from the people doing the work for effort. Document ranges, assumptions, exclusions and date, and re-score when evidence, scope, cost or context changes.
MoSCoW and WSJF: the other two questions
MoSCoW—negotiating scope when the date is fixed
MoSCoW classifies work for a defined scope and timebox. Its real use case is negotiating with stakeholders in the room when time or budget is fixed and the deadline is not movable.
- Must — indispensable to a viable, safe or compliant outcome; if absent, the release fails. Not merely something a stakeholder strongly prefers.
- Should — important, with a tolerable workaround or deferral consequence.
- Could — desirable contingency.
- Won't — explicitly not in this horizon. Not "never."
The discipline: agree objective category tests in advance ("the release fails PCI validation without it" is a test; "marketing needs it" is not), cap Musts so the plan retains contingency, and name who accepts the consequences of each deferral. If everything comes back Must, the method has revealed there is no actual choice—say that out loud; it's the point the interviewer is fishing for.
Worked mini-example: a fixed-date release with 6 weeks of capacity and a stakeholder list of 11 items. The pass condition is "the release can process a customer order end-to-end, and the audit log is complete." That test makes 3 items Must (order pipeline, payment retry, audit log), 4 Should, 4 Won't-this-release. The pre-agreed test is what stops the room from relitigating every item against enthusiasm.
WSJF—sequencing when delay itself has a cost
WSJF = Cost of Delay ÷ job size, usually with consistent relative estimates. Cost of Delay may combine user/business value, time criticality, and risk reduction or opportunity enablement. The denominator makes smaller valuable work advance sooner, but splitting must preserve coherent value and safety—you can't split a mandatory control into "half a penetration test."
Worked example (relative units, hypothetical):
| Job | Value | Time criticality | Risk reduction | CoD | Size | WSJF |
|---|---|---|---|---|---|---|
| Tax-rule update | 5 | 8 | 2 | 15 | 2 | 7.5 |
| Search revamp | 8 | 2 | 1 | 11 | 5 | 2.2 |
| Deprecate legacy auth | 3 | 2 | 8 | 13 | 3 | 4.3 |
The tax update wins not because it's the most valuable but because its value decays fastest—delay costs more per week. That's the insight WSJF exists to surface, and naming it is worth more than the table.
WSJF assumes comparable jobs and credible relative inputs. It does not override mandatory controls, capacity constraints, fixed dependencies, or portfolio balance. Its scores change as deadlines approach, opportunities decay, and work is partially learned—re-score, don't set and forget.
Choosing between frameworks—and when to skip one
The frameworks answer different questions:
| Situation | Tool | Why |
|---|---|---|
| Large backlog, one goal, partial evidence | RICE or a simple weighted score | Cheap, comparable, forces defined inputs |
| Fixed date, stakeholders in the room, scope must shrink | MoSCoW | Negotiates scope explicitly; makes deferrals visible |
| Queue of comparable jobs where delay has a cost | WSJF | Ranks by economic urgency, not raw value |
| Small set, unambiguous goal | Judgment + strategy | A spreadsheet adds ceremony, not information |
Saying a framework is optional—and explaining when you'd skip it—is the senior signal. A backlog is ordered, not partitioned into a permanently correct priority field. Combine whatever you use with constraints, a dependency graph, strategic fit, portfolio risk, team topology, change cost and learning value. Never average incompatible metrics just because a spreadsheet accepts numbers.
Mandatory work, dependencies and opportunity cost
Mandatory work still gets prioritized. Legal, security, privacy, accessibility and reliability obligations have deadlines, scope, consequences and alternative controls; "compliance" is not an exemption from evidence or design. Validate the obligation, right-size the response, and give risk acceptance a named authority, an expiry and evidence. Emergency defects may preempt planned work based on severity and user harm—that's a policy you decide in advance, not an ad-hoc override.
Dependencies constrain the feasible sequence. Separate hard technical or policy prerequisites from convenient preferences, build the graph, and expose critical paths and shared platforms. Don't hide incremental value inside one giant parent item. Prioritize enabling work by the outcomes it unlocks and the probability and timing of those outcomes. Starting many high-scoring items at once increases work in progress, queues and context switching—order for flow and completion, not maximum starts.
Opportunity cost is the value of the best alternative not chosen. Every commitment consumes capacity and usually creates ongoing support, infrastructure, data and cognitive load, so estimate lifecycle cost and option value, not build effort. Consider deletion, simplification, and not building. Sunk cost is not future value: continue an initiative only when expected remaining benefit justifies remaining cost and risk. Make Won't decisions visible with rationale, so stakeholders see the trade-off and teams don't perform hidden work.
Fairness, governance, and the follow-ups you'll actually get
Aggregate reach can systematically deprioritize small groups with severe harm or legal rights. Segment outcomes, identify who benefits and who bears risk, and apply floors for accessibility, safety and equity where justified. Revenue impact does not automatically outweigh user harm. Loud customers and executives are data sources, not priority overrides—use consistent intake and show how evidence, strategy and obligations shaped the decision.
Prioritization is also governance. One accountable owner makes the call after multidisciplinary input; committees advise, but ambiguous ownership invites lobbying and shadow roadmaps. Publish criteria, scores or ranges, assumptions, dependencies, overrides, the decision and the revisit trigger. Provide an escalation path for new evidence, and protect teams from direct work injection that bypasses the ordered backlog.
Expect these follow-ups, and have a real answer for each:
- "What did you deprioritize and why?" Name the item, the alternative's value, and who accepted the deferral.
- "What would make you change this next sprint?" A concrete trigger: an experiment result, a threshold breach, a new obligation.
- "Tell me about a time the data was wrong." Reach that didn't materialize, impact that didn't hold, effort that doubled—what you changed in your estimation process afterward.
Validate priorities through delivery: small batches, experiments, staged releases that produce information before full investment. Measure whether reach occurred, impact materialized, guardrails held, effort and maintenance matched estimates, and opportunity costs emerged. Update the model and your calibration rather than rewriting history. A high-scoring feature that fails its hypothesis should change future decisions; a low-confidence bet may be worth taking precisely because a cheap experiment resolves a critical uncertainty.
The output is a defensible sequence and explicit non-commitments—not a formula that absolves judgment. Good prioritization exposes goals, evidence, uncertainty, constraints, affected groups, trade-offs and authority; limits work in progress; and adapts as reality changes while preserving strategic direction.
Weak answer vs. strong answer
| Probe | Weak answer | Strong answer |
|---|---|---|
| "Prioritize these three features." | Scores all three with RICE immediately, defends the decimals | Asks the goal, states capacity and horizon, scores, then shows which inputs could flip the order |
| "Everything is a Must." | Accepts it and plans overtime | Names the missing pass/fail test and the absent capacity cap; renegotiates scope |
| "The sponsor says confidence is 100%." | Records 100% | Asks what evidence backs each input; records provenance and ranges |
| "What about the accessibility fix that scores low?" | Deprioritizes it on the score | Treats it as an obligation with a floor, not a candidate competing on reach |
| "What would change your mind?" | "New information" | A named trigger and a revisit date |
