Skip to content
Tech Interview Prep home
Technical interview guide

Requirements Gathering & User Stories

Turning a vague need into a written, testable requirement — user stories, acceptance criteria, and edge cases.

Read
30 min
Practice MCQs
25
Interview QA
25
Edition
v4
Editorial status
Reviewed

Scope: GOV.UK Service Manual, Scrum Guide 2020, ISO/IEC 25010:2023, WCAG 2.2, RFC 2119, OpenAPI, and NIST guidance current 2026-08-31.

Overview

Curated: · Written: · Reviewed:

Requirements turn uncertain needs into testable decisions without freezing discovery

The mental model: requirements are decisions, not transcripts

A requirement is a testable statement of what a system must do, constrain, or tolerate, for whom, and under what conditions. The mental model that separates a strong answer from a weak one in interviews: requirements work is decision-making under uncertainty, not transcription. A stakeholder saying "we need a dashboard" is raw input. The requirement is the outcome they want, the evidence behind it, the boundaries around it, and the test that will prove it — each of which is a decision you make and can defend.

The mechanism, end to end:

  1. Elicit — gather evidence from users, data, documents, and stakeholders.
  2. Analyze — separate needs from requested solutions, surface conflicts and assumptions.
  3. Specify — write stories, acceptance criteria, and quality constraints that are observable and testable.
  4. Validate — check comprehension with the people who will use, build, operate, and audit the result.
  5. Manage change — treat every artifact as revisable, with rationale preserved.

Interviewers probe this loop at every stage. A weak answer describes artifacts ("I write user stories in Jira"). A strong answer describes decisions: what you excluded, what you measured, what you deferred and why.

Elicitation: matching technique to the question you're answering

Each technique answers a different question. Picking the wrong one is a common failure mode — running a survey when you need depth, or interviewing when you need scale.

TechniqueBest forWeaknessTypical signal
Stakeholder interviewsBusiness goals, constraints, conflicts, decision authoritySmall n; opinion, not behavior"Legal won't approve that retention period"
Observation / job shadowingActual workflows, workarounds, invisible stepsTime-intensive; observer effectUser keeps a spreadsheet beside the tool — the real system of record
SurveysScale, segmentation, ranking known issuesCan't discover unknown needs; question wording bias40% of respondents report the same failure
WorkshopsAligning conflicting stakeholders, co-designing boundariesDominant voices; false consensusTwo teams leave with different understandings of "done"
PrototypingTesting comprehension of an uncertain interactionUsers react to polish, not substanceUsers can't complete the core task in a clickable mock
Document / competitive analysisRegulatory constraints, industry baselines, existing data flowsDescribes what is, not what should beCompetitor ships the feature; your support data shows no demand for it

The decision rule: behavioral questions get observation; opinion questions get interviews; scale questions get surveys; comprehension questions get prototypes. Always triangulate — analytics and support tickets tell you what is happening; interviews and observation tell you why.

Include the people who fail, abandon, need assistance, use alternative channels, have disabilities, operate the service, or bear downstream effects — not only the happy-path customer. Record sample sizes and method limits; an interviewer will ask "how did you know?" and "who did you not talk to?"

User stories: anatomy and quality gates

The Connextra template — As a [role], I want [goal], so that [benefit] — is a value frame, not a formality. The benefit clause is the part that earns its place: it lets the team evaluate alternatives. "As a shopper, I want a one-click reorder button" is a solution; "so that I can restock in under 30 seconds without searching my order history" points at the outcome, and a better solution might exist.

The 3 Cs (card, conversation, confirmation) capture how stories actually work: the card is a placeholder for a conversation, and the conversation is confirmed by acceptance criteria. A story is an invitation to talk, not a finished specification.

INVEST is the quality gate:

  • Independent — can be built in any order relative to its neighbors
  • Negotiable — invites discussion, doesn't dictate implementation
  • Valuable — delivers user or business value end-to-end
  • Estimable — the team understands it enough to size it
  • Small — fits comfortably in an iteration
  • Testable — acceptance criteria can prove it done

When a story fails INVEST — usually on Small or Valuable — split it vertically, not horizontally. A horizontal split ("build the API layer, then the UI, then the database") delivers nothing usable alone. Vertical splits that preserve end-to-end value:

  • Workflow step: order confirmation → payment → shipping → returns
  • Business rule variant: domestic refunds first, international refunds later
  • Data set / range: support 100 rows, then 100k rows
  • CRUD operation: create + read first; update and delete in a later slice
  • Happy path before edge cases: core flow, then concurrency and failure handling

Splitting must preserve coherent authorization and data integrity — a slice that lets users read data without the permission checks is not a safe vertical slice.

Acceptance criteria: verifiable boundaries

Acceptance criteria define the observable outcomes that establish a story is done. Two common formats:

  • Rule-based (checklist): "Refunds over $500 require manager approval. Refund appears in the customer's transaction history within 5 minutes."
  • Gherkin / Given-When-Then: "Given an order with status 'delivered', when the customer requests a refund within 30 days, then the refund is queued and a confirmation email is sent."

Gherkin improves shared understanding for scenario-heavy behavior; rule-based is faster for simple constraints. Neither is a complete requirement by itself — both need the negative cases spelled out.

The edge cases interviewers expect you to name unprompted:

  • Negative: refund on a cancelled order is rejected with reason "order not eligible"
  • Boundary: exactly 30 days — eligible or not? Exactly $500 — approval or not?
  • Empty state: customer with no order history sees an explanatory message, not a blank page
  • Permission/role: a support agent can issue refunds up to $100; anything higher requires a manager
  • Concurrency: two simultaneous refund requests for the same order — one succeeds, one is rejected as duplicate
  • Error and retry: payment gateway times out — is the refund idempotent on retry?

Write preconditions, the action or event, the expected result, and who accepts the evidence. Don't overspecify internal implementation — criteria constrain observable behavior, not class design.

Ambiguity: interrogating "the system should be fast"

This is the most common interview probe in the requirements space. "Fast," "user-friendly," "robust," "intuitive," "seamless," and "real-time" are not requirements — they are placeholders for a conversation. The interrogation pattern:

  1. Ask for the scenario. "Fast for whom, doing what?" → "A support agent pulling up a customer's order history."
  2. Attach a workload and population. "Agents with 500+ orders per customer, during peak season."
  3. Pin a percentile and measurement point. "p95 under 2 seconds, measured at the API, excluding network."
  4. Define failure behavior. "If it exceeds 5 seconds, show partial results with a load-more control."

The result: "p95 order-history query under 2 seconds for customers with up to 500 orders, measured at the API in production; degrade to partial results beyond 5 seconds." That is testable, negotiable, and defensible.

Separate assumptions from validated facts explicitly. Keep a running list: assumption, evidence, owner, decision date, expiry. "We assume 90% of refunds are under $50" — validated against last quarter's data or flagged as unvalidated? A requirement with no source, owner, or rationale can't be challenged or retired. Use normative language deliberately: MUST is a hard constraint; SHOULD is a recommendation with documented exceptions. When stakeholders conflict, surface the trade-off and the decision authority — don't write two mutually incompatible requirements and hope.

Non-functional requirements: what stories silently hide

A user story that passes INVEST can still be unsafe because quality attributes hide inside it. Every story carries implicit NFRs; the skill is making them explicit:

  • Performance: workload, population, percentile, measurement point, environment, degradation behavior — not "fast"
  • Availability: scope and objective with a window — "99.9% monthly for the checkout path" is a decision; "highly available" is not
  • Security/PII: data classification, allowed actors and purposes, encryption, retention, deletion, consent or legal basis, audit logging, incident response
  • Accessibility: applicable success criteria (e.g., WCAG 2.2 AA), plus testing with assistive technologies and people with disabilities. Automated checkers catch a real but partial slice of issues — things like missing alt text, color-contrast failures, and missing form labels — but they cannot verify meaningful keyboard operation, screen-reader usability, or whether a task can actually be completed. Treat automated scans as a floor, never as evidence of accessible use.
  • Scalability: what breaks at 10× volume — pagination, rate limits, queue depth, cost ceilings
  • Regulatory: GDPR/CCPA data subject rights, PCI scope, SOC 2 controls, records-retention law — these are non-negotiable constraints, not preferences

ISO/IEC 25010's quality characteristics (functional suitability, performance, compatibility, usability, reliability, security, maintainability, portability) are a useful completeness checklist, not a substitute for product-specific risk analysis. The interview trap: a candidate lists NFRs as a separate document nobody reads. The strong answer embeds them — performance budgets in the story's criteria, security review in the definition of done, accessibility in the acceptance tests.

Failure modes, misconceptions, and when requirements work isn't the answer

Common failure modes, each of which makes a good interview story if you've lived it:

  • Solution bias: shipping the requested feature instead of addressing the need. The classic: users request an export button; the actual need is a report nobody wants to build by hand.
  • Happy-path-only criteria: every edge case discovered in production, each one a support ticket.
  • The waterfall trap: treating the specification as frozen. New evidence, incidents, and policy changes invalidate assumptions; assess impact (users, interfaces, data, tests, migration, rollout, commitments), update linked artifacts, and communicate. A late requirement isn't automatically bad — an unassessed one is.
  • Traceability theater: a requirements matrix maintained for audit that nobody uses. Traceability should be lightweight — need → story → criteria → test → telemetry — enough to answer "what breaks if this tax rule changes?" and to spot orphan features and uncovered obligations.
  • Fictional benefits: a project plan disguised as user stories, each with an invented "so that..."

Misconceptions worth naming in an interview:

  • "Requirements gathering means collecting what people ask for." It means deciding, with evidence, what to build and what to exclude.
  • "A grammatically correct story is a valid story." It can still be solution-biased, untestable, or inaccessible.
  • "Meeting acceptance criteria proves the feature worked." It proves the specified behavior under tested conditions. Post-release, compare adoption, completion, errors, support demand, and exclusion against hypotheses — then feed that back into the backlog.
  • "More documentation = better requirements." Documents are evidence carriers; conversation and research are the substance.

When not to use heavyweight requirements work: a throwaway experiment, a hackathon, an internal tool with three users, or a reversible change behind a flag. The cost of formal elicitation has to match the cost of being wrong. A one-week spike needs a hypothesis and a kill criterion, not a specification.

What changes at scale

At small scale, a conversation and a card suffice. As the surface grows:

  • Interfaces become contracts. APIs need defined operations, schemas, status and error semantics, authentication, idempotency, pagination, versioning, and rate limits — in a reviewable contract (OpenAPI, for example), because a prose paragraph can't be contract-tested.
  • Cross-team dependencies force explicit boundaries. Data ownership, lifecycle, quotas, migration, and decommissioning become requirements in their own right.
  • Traceability stops being optional. Regulatory environments (finance, healthcare) require demonstrable links from regulation to control to test, with version and rationale preserved so a passing test can be tied to the correct requirement revision.
  • Localization, time zones, and compatibility move from afterthoughts to boundary conditions — "timestamps stored in UTC, rendered in the user's zone" is a decision someone has to make.
  • Change management gets heavier. Assessing a changed rule means impact analysis across teams, not a hallway conversation.

Likely interview follow-ups and what a weak answer sounds like

Typical probes, and the signals that sink answers:

  • "A stakeholder insists on a feature your research contradicts. What do you do?" Weak: "I'd explain the data." Strong: separate the request from the need, run a cheap test of the hypothesis, surface the trade-off and the decision authority, and record the outcome either way.
  • "How do you handle vague requirements?" Weak: "I ask for more detail." Strong: the interrogation pattern above — scenario, workload, percentile, failure behavior — with a concrete rewritten example.
  • "Walk me through a story you split." Weak: describes horizontal layers. Strong: names the INVEST failure and the vertical split axis, and what each slice proved.
  • "What did you cut, and why?" Weak: nothing comes to mind. Strong: a specific exclusion, its rationale, its owner, and what evidence would bring it back.
  • "How do you know the feature worked after launch?" Weak: "we met the acceptance criteria." Strong: the outcome metrics and guardrails instrumented before release, and what they showed.
  • "What NFRs did your last project have?" Weak: a generic list. Strong: named targets with numbers, and where they were enforced.

The through-line: good requirements make scope, uncertainty, quality, exclusions, failure, measurement, and change visible enough that a cross-functional team can build the right thing safely — and learn whether it worked.

Worked example: from vague request to testable requirement

Raw input, as it arrives in a backlog: "Support wants to bulk-refund customers faster."

Interrogate. Who performs it, how often, at what volume? Support lead: ~200 bulk-refund requests per month during promotions, each covering 50–500 orders, currently done one-by-one taking up to 4 hours per batch.

Separate need from solution. The need is timely, auditable refunds at batch scale — not necessarily a new UI. (A CSV upload, an API, or a rules engine could each be the right answer; the conversation decides.)

Write the story. As a support agent, I want to refund up to 500 orders in one action, so that a promotion-affected batch is resolved in under 15 minutes instead of 4 hours.

Acceptance criteria (rule-based, with the edge cases named):

  • Agent selects up to 500 eligible orders; ineligible orders are listed with reasons and excluded, not silently dropped
  • Every refund over $500 per order requires manager approval before submission
  • Duplicate submission within 60 seconds is rejected as a duplicate, not double-processed (idempotency key)
  • Partial failure: successfully refunded orders complete; failed orders are listed with reasons and retryable
  • All refunds appear in each customer's transaction history within 5 minutes and in the audit log with actor, timestamp, and reason
  • Empty state: no eligible orders selected shows an explanatory message
  • Role check: agents without refund permission cannot access the action

Hidden NFRs made explicit: audit log retention per finance policy (7 years); PII in the refund reason field classified and masked in exports; screen-reader accessible order selection; batch API p95 under 10 seconds for 500 orders.

Validation plan: support lead walks through the prototype with two agents before build; post-release, measure batch completion time, error rate, and duplicate-refund incidents for one month against the 15-minute hypothesis.

That's the whole discipline in one artifact: evidence, a decision, exclusions, edge cases, quality constraints, and a measurement — none of which the original sentence contained.