Skip to content
Tech Interview Prep home
Technical interview guide

Metrics, KPIs & North Star Frameworks

Choosing the metric that actually reflects whether the product is succeeding, and avoiding vanity metrics.

Read
31 min
Practice MCQs
25
Interview QA
25
Edition
v3
Editorial status
Reviewed

Scope: Google HEART, Scrum.org Evidence-Based Management 2024, DORA metrics guidance, GOV.UK Service Manual, NIST statistical and privacy guidance, and FTC dark-pattern guidance current 2026-08-31.

Overview

Curated: · Written: · Reviewed:

Product metrics are decision instruments, not trophies

This is the topic where interviewers hand you a vague prompt — "How would you measure success for this feature?" — and watch whether you reach for a dashboard field or for a decision. A weak answer names a metric in the first sentence and spends the rest of the time defending it. A strong answer starts with the goal and the user problem, states the causal hypothesis, and only then picks instruments — and it arrives with the distinctions interviewers probe as follow-ups already loaded: leading vs lagging, input vs output, vanity vs actionable, guardrail vs counter-metric.

Start from the goal, not the dashboard

A metric is useful when a team can explain the outcome it represents, the decision it informs, how it is calculated, whose behavior it includes or excludes, and what could make it misleading. Begin with a product goal and user problem, not with available dashboard fields. Write the causal hypothesis explicitly: if the product creates value in the intended way, which observable signals should change, over what window, for which population, without violating which guardrails?

Interviewers often test this by giving you a feature with an obvious vanity metric attached — say, a share button, where "shares sent" is trivially gameable. What they want to hear: shares are an input; the value moment is the recipient acting on the share; the metric should sit at the value moment, and shares-sent belongs lower in the tree as a diagnostic. A weak answer maximizes the count that's easiest to move.

The North Star construct

A North Star Metric is a shared measure intended to represent recurring value delivered to users and to connect product work with durable organizational value. It is not automatically revenue, daily active users or one fashionable engagement event. Three properties make one metric qualify:

  • It captures value delivered to the customer, not activity the company extracts. "Messages read" is closer to value than "messages sent"; "items delivered on time" beats "orders placed."
  • It matches the business model. A marketplace's North Star involves transactions on both sides; a subscription product's involves recurring qualified usage, not signups. If the metric rises while the business model is broken, you picked the wrong metric.
  • It predicts long-term revenue rather than merely correlating with it. The test: when the metric rises for a cohort, does that cohort's later revenue actually rise? Correlation is cheap; interviewers probe for evidence that the link was tested, not assumed.

A good candidate definition has a clear value moment, eligible actor and unit, aggregation window, quality threshold, deduplication rule and relationship to outcomes. It should be difficult to increase substantially through spam, coercion, accidental activity or degradation elsewhere. One headline metric still needs input metrics, diagnostic measures and guardrails.

The input tree: breadth, depth, frequency, efficiency

A North Star on its own is a scoreboard, not a steering wheel. Decompose it into drivers a team can actually own. The classic decomposition splits the North Star into four input families:

  • Breadth — how many users experience the value (qualified active users, not registrations).
  • Depth — how much value each interaction delivers (tasks completed per session, quality-adjusted usage).
  • Frequency — how often the value moment recurs.
  • Efficiency — how much value per unit of input or cost (time-to-value, conversion per unit of acquisition spend).

Each sub-driver gets an owning team. The acquisition team owns breadth; the core product team owns depth; lifecycle or notification teams own frequency. This is what makes the North Star operational rather than decorative: every team can point at the branch it moves, and leadership can see which branch stalled when the headline flatlines.

Decomposition stops where a team can intervene. "Increase retention" is not a lever; "reduce first-week drop-off between signup and first completed task" is. If a sub-metric has no owner and no plausible intervention, it's a report, not a driver — cut it or push it one level down.

Frameworks: HEART as a menu, not a checklist

HEART provides categories — Happiness, Engagement, Adoption, Retention and Task Success — and a goals-signals-metrics process: state the goal, identify the signals that would show progress toward it, then choose metrics for those signals. It is a menu, not a requirement to maximize every category. Engagement can be inappropriate for a product whose value is finishing quickly or needing it rarely — a tax-filing app that maximizes engagement has misunderstood its own value proposition; Task Success or Happiness fits better. Combine behavioral evidence with research, accessibility evaluation, support and qualitative evidence; telemetry explains what happened more readily than why. Interviewers sometimes ask "which HEART categories apply here?" — the strong answer picks the two or three that match the product's value model and says why the others don't apply, rather than listing metrics for all five.

The distinctions interviewers probe as follow-ups

Whichever metric you propose, expect the interviewer to test whether you know which distinction it sits on. Have these ready:

  • Leading vs lagging. Leading indicators move earlier in the hypothesized chain and may help a team respond sooner; lagging indicators confirm later results. Neither label proves causality. Feature adoption may lead retention only for the intended cohort and behavior, while revenue may lag both value and harm. A weak answer treats "leading" as "causes."
  • Input vs output. Outputs are results you don't directly control (retention, revenue); inputs are the levers you move (activation rate, time-to-first-value). Teams are assigned inputs; the output tells you whether the inputs were the right ones.
  • Vanity vs actionable. A vanity metric can only go up and isn't tied to a decision — cumulative downloads, total signups. Actionable metrics have a denominator, a segment story, and a decision they inform. The giveaway in an interview is when a candidate's metric has no failure mode: if nothing the team could do would move it down, it measures nothing.
  • Health/guardrail vs counter-metric. Guardrails protect outcomes the primary metric might sacrifice: reliability, latency, error rate, complaints, cancellations, refund rate, safety, privacy, fairness, accessibility and cost. They need thresholds, owners and action rules. A counter-metric is the specific inverse your feature could plausibly damage — if you raise "replies per thread," the counter-metric is unsubscribes from threads. A feature that raises clicks through deceptive choice architecture has not increased value; naming the counter-metric unprompted is what senior answers do.

Operational definitions and metric hygiene

Distinguish measurements, metrics, indicators and targets. A count of completed tasks is a measurement; completion rate divides eligible completions by legitimate starts under defined rules; a KPI is a selected indicator considered important to an objective; a target states a desired future level and date. Do not change denominator, event definition or eligibility while presenting a continuous trend. Version metric definitions and restate history or annotate breaks when instrumentation changes.

Every metric needs an operational definition: name, purpose, owner, event source, grain, numerator, denominator, eligibility, deduplication, identity resolution, time zone, window, latency, exclusions, segment rules, revision policy, freshness and quality checks. State whether late events backfill prior periods and how bots, employees, tests, retries and fraud are handled. A dashboard without this contract invites incompatible interpretations — and interviewers probe this directly: "two teams report different numbers for the same metric; what do you do?" The answer is the contract, not a reconciliation meeting.

Segments, distributions and Simpson's paradox

Segment before averaging away harm. Examine new and established users, plan or market, device, channel, geography and relevant accessibility or risk groups, while applying privacy thresholds and avoiding re-identification. Simpson's paradox can reverse an aggregate trend when cohort mix changes — a classic interview trap: overall completion rate improves while every segment got worse, because the mix shifted toward a low-rate segment. Averages hide distributions: use percentiles, rates, confidence intervals and absolute counts where appropriate, and show denominators so small samples are not mistaken for stable movement.

Targets, gaming and Goodhart

When a metric becomes a target, people and systems adapt. Periodically audit gaming paths and whether the proxy has drifted from the goal. Targets should use a baseline, time horizon, expected intervention, uncertainty and rationale. Benchmarks add context but differ by population and measurement. Avoid arbitrary round-number goals and post-hoc target changes. Forecast, target and commitment are different things; a target can motivate learning without pretending the team controls external conditions, and escalation should examine causal evidence and constraints, not merely pressure teams to manipulate the numerator.

Experiments and causal claims

Experiments strengthen causal claims when assignment, sample, exposure, outcome window and analysis are appropriate. Pre-register a primary metric and guardrails, estimate required sensitivity, monitor sample-ratio mismatch and data quality, avoid repeated uncorrected peeking, and interpret practical as well as statistical significance. Network effects, interference, novelty, seasonality and long-term outcomes may require cluster designs, holdouts or observational follow-up. A failed hypothesis is useful evidence, not a reason to redefine success after seeing results — interviewers deliberately offer you a "the experiment failed" scenario to see whether you redefine the metric or report the finding.

Delivery metrics and data quality

Operational and delivery metrics matter but should improve the system rather than rank individuals. DORA measures describe software delivery throughput and instability at an appropriate service or team scope. Velocity, story points, lines of code and hours online are especially dangerous performance targets because they are context-dependent inputs and easy to game. Use flow, quality and outcome evidence together, with teams investigating constraints rather than competing on numbers.

Data quality failures are product failures when decisions rely on them. Specify event contracts, validate schemas, monitor volume and distribution, reconcile with authoritative systems, test duplicate and missing events, track freshness and lineage, and make backfills visible. Preserve raw immutable evidence where justified while minimizing personal data, access and retention. Consent and purpose limitations are not waived because a field would improve segmentation.

The metrics review, and what a strong answer sounds like

A useful metrics review asks what decision is due, what changed, whether the definition or population changed, how uncertainty and segments behave, which guardrails moved, and what evidence would distinguish explanations. It ends with an owner and an action: continue, investigate, experiment, roll back, change instrumentation or retire the metric.

In an interview, the compressed version of all of this is a repeatable shape: goal → user problem → causal hypothesis → primary metric with an operational definition → input drivers and owners → guardrails and counter-metrics → how you'd validate the causal link. Metric success is better decisions and better user outcomes — not the number of charts or the uninterrupted rise of a chosen line.