Overview
Curated: · Written: · Reviewed:
Architecture is a traceable set of decisions, not a diagram
Translating requirements into architecture starts by understanding the outcome, users, operating context and constraints—not by choosing a product. A requirement states a needed capability or observable quality. A constraint restricts the solution space. An assumption is an unverified belief. A design decision selects an approach with consequences. Keeping these categories separate prevents a stakeholder preference, guessed workload or vendor feature from quietly becoming an immutable requirement.
This distinction is the single most common interview probe in system design rounds. When the interviewer hands you "build me a checkout system," they are watching for whether you ask about the problem before drawing boxes. A weak answer starts with "so I'd use Kafka and a microservices architecture"; a strong answer starts with "who are the users, what volume, what's already there, and what happens when it's down?" Expect the follow-up chain: after you state an assumption, the interviewer will change it—"double the traffic," "the payment provider is down," "we can't store that data in the US"—and score whether your design absorbs the change or collapses.
Elicitation: what they asked for is not the problem they have
Begin with the business and user journeys. Identify actors, desired outcomes, critical flows, data and trust boundaries, current-system dependencies, lifecycle, geographic and regulatory context, delivery deadline, budget and organizational capabilities. Define what is in and out of scope.
When a customer says "we need X," treat X as a hypothesis about the solution, not the requirement. The clarifying questions that separate the two:
- What happens today, and what breaks or costs too much about it? (the problem)
- Who is blocked or losing money, and how much? (the outcome worth paying for)
- What did they try before, and why did it stop working? (history that constrains you)
- What must this not break? (the invariants)
- What would you cut first if the deadline moved up a month? (the real priority ordering)
"We need WebSockets" often means "our dashboard feels stale"—which polling at 5s intervals might fix for a tenth of the cost. "We need microservices" often means "two teams keep blocking each other"—which module boundaries and code ownership can fix without a distributed system. In an interview, when the prompt names a technology, saying "before I accept that, what problem is it solving?" is not stalling; it is the answer.
For brownfield work, establish the current architecture, contracts, data ownership, operational pain, migration constraints and tolerable coexistence period. The architecture must include build, deploy, operate, recover, change and retire—not only the steady-state request path.
Two translation paths: functional and non-functional
Functional and non-functional requirements translate into architecture through different mechanisms, and conflating them is a classic weak answer.
Functional requirements—"a user can place an order," "an admin can refund a charge"—map to components and flows. You decompose them into services, APIs, data entities and state transitions. They are testable with examples: given this input, this state, this actor, expect this output and this side effect.
Non-functional requirements—latency, availability, durability, throughput, consistency—map to topology, redundancy, storage and consistency choices. "99.9% available" is not a feature you add; it is a statement about how many failure domains you can survive, which dictates multi-AZ deployment, health-checked failover, and whether a single-region database is even on the table. "No data loss on region failure" dictates synchronous cross-region replication, which adds write latency, which feeds back into your latency budget.
The interview trap: candidates enumerate features fluently, then say "and it should be scalable and reliable" and move on. The interviewer's next question is almost always "how scalable, exactly?"—and the design's shape depends entirely on the number. A system for 50 requests per second and one for 50,000 differ in whether a single Postgres instance, read replicas, or sharding is the right call. If you haven't sized it, you haven't designed it; you've sketched a genre.
Quantifying quality attributes: scenarios and back-of-envelope numbers
Make functional requirements testable through examples, inputs, outputs, business rules, error cases and acceptance criteria. Resolve ambiguous terms such as "real time," "secure," "highly available," "global" and "scalable." Ask how much, for whom, under which conditions and measured where.
Express architecturally significant quality attributes as scenarios with six parts: stimulus source, stimulus, environment, affected artifact, expected response, and measurable response measure. "Fast" becomes, for example, a named user action under a stated concurrent load whose end-to-end p95 latency is below a threshold at a defined observation point. Reliability includes SLO, RTO, RPO, maximum tolerable outage, degraded behavior and recovery verification. Security includes assets, actors, threats, authorization boundaries and response. Cost includes demand assumptions and a decision-relevant limit, not merely "cost effective."
Back-of-envelope sizing turns these from adjectives into numbers you can defend:
- Throughput: 1M daily orders, peak hour carries ~15% of daily traffic → 150k orders/hour → ~42 orders/second average, design peak at 3–5× → ~150–200 orders/s. Each order writes 3 rows and emits 2 events → ~1k writes/s at peak.
- Storage: 1M orders/day × 12KB per order (2KB record plus 10KB of event payload) → ~12GB/day → ~4.4TB/year before indexes, replicas and event retention. That still fits a single well-sized node's disk, but add a 7-year regulatory retention requirement and it becomes ~31TB—now replication strategy, archival tiers and restore times are architecture decisions, not operations details.
- Availability: 99.9% allows 43 minutes of downtime per month; 99.99% allows 4.3 minutes—which no manual failover process achieves. The number you commit to determines whether you need automated failover and multi-AZ from day one.
- Latency: a 2s p95 browser budget minus ~300ms network and TLS minus ~200ms frontend leaves ~1.5s for the backend, which you then allocate across service hops and database reads.
Interviewers probe the derivation, not the arithmetic: where did 15% come from (state it as an assumption), why 3–5× peak (label it a rule of thumb), and what changes if you're wrong by 10×. Saying "I'm assuming X, and here's what in the design flips if X is off" is the strong move; a precise number with no stated assumption is a weak answer wearing a costume.
Prioritize requirements with accountable stakeholders. Separate mandatory obligations from preferences and negotiate conflicts explicitly: stronger consistency may increase latency and reduce availability; lower recovery targets may raise cost and operational complexity; faster delivery may defer capabilities but not silently waive security obligations. Record the rationale, decision authority and accepted residual risk. A requirement matrix is useful only if every entry has an owner, source, priority, verification method and lifecycle status.
The constraint inventory beyond software
Constraints restrict the solution space before any design choice is made, and the weak answer ignores half of them. The full inventory:
- Budget and timeline: a $50k build and a $500k build do not produce the same architecture; a 3-month deadline rules out anything requiring a 6-month data migration.
- Team size and skills: a 4-person team with Python experience cannot operate a 12-service Kubernetes mesh, however elegant. Hiring is a constraint on design.
- Existing stack and vendor commitments: an enterprise agreement with a cloud vendor, an existing Kafka cluster with spare capacity, a Postgres standard that security already approved. Fighting these needs a reason, not a preference.
- Data residency and compliance: HIPAA for health data, SOC 2 for B2B sales, GDPR's right to erasure and cross-border transfer rules. These dictate where data can live, whether you can use a managed service, and whether "delete" must reach every replica and backup.
- Deployment environment: on-prem vs cloud vs air-gapped. An air-gapped deployment kills any architecture that assumes a managed control plane, automatic patching, or calling home for licenses.
In an interview, the prompt often omits these deliberately. "Design a medical records system" without mentioning HIPAA is testing whether you raise it. Raising a constraint the interviewer didn't mention—and being willing to accept "assume it's fine"—earns more than a flawless diagram built on unstated constraints.
Deriving the design: boundaries and views
Derive candidate architecture from the domain and flows. Define system context, external actors and dependencies; decompose around ownership, change and failure boundaries; identify data of record, consistency rules, APIs/events, identity, policy enforcement, failure modes, capacity, observability and operations. Use multiple views for different questions: context and containers for scope and responsibility, sequence or dynamic views for runtime behavior, deployment views for topology and failure domains, and data models for ownership and lifecycle. A single boxes-and-arrows picture cannot carry all semantics.
For brownfield integration, the decomposition must respect what already exists: legacy APIs with their own rate limits and error semantics, an auth model you must join rather than replace, data formats you cannot change unilaterally, and the choice of sync vs async boundaries—calling the legacy ERP synchronously couples your availability to its nightly batch window; consuming its changes asynchronously decouples you but adds eventual consistency you must design for. Plan the migration and coexistence path explicitly: strangler routing, dual-write or backfill windows, and a defined rollback point. "We'll rewrite it all" is the weak answer; "we route new traffic to the new service, keep both consistent during the transition, and retire the old path after N weeks of clean SLOs" is the one that survives contact with an operations team.
Maintain bidirectional traceability between business outcomes, requirements, risks, architecture decisions, components, interfaces, tests and operational evidence. Traceability should expose gaps and change impact, not become a ceremonial spreadsheet. Do not force every low-level implementation detail into the architecture model. Focus on decisions that affect structure, key qualities, external contracts, risk or reversibility.
Evaluating alternatives and recording decisions
Evaluate alternatives against weighted or prioritized criteria using evidence. Include the status quo and simpler options. Record assumptions, unknowns, dependencies, skills, delivery and migration effort, ongoing operations, security, reliability, performance, cost, portability and exit. Scores organize discussion but do not manufacture objectivity: document scale definitions, sensitivity to weights and veto constraints. Prototype or benchmark only the uncertainty that could change the decision, with representative workloads and explicit success/stop criteria.
Capture significant choices in append-only architecture decision records: context, decision, alternatives, requirements, evidence, tradeoffs, consequences, confidence, owner and status. When facts change, supersede rather than rewrite history. Track unresolved questions and risks separately with owners, due dates, triggers, mitigations and contingency. Treat reversible two-way-door choices differently from expensive one-way-door commitments.
The interview version of this is the trade-off narration: "I'm choosing Postgres over DynamoDB here because we need transactions across two entities and the write volume is ~200/s; if volume grew 50×, the partitioning story would flip that decision." That last clause—what would change the decision—is what distinguishes a staff-level answer from reciting a stack.
Validation and iteration
Validate architecture before and during implementation. Walk critical and failure scenarios with engineering, security, operations, data, product and affected users. Threat-model trust boundaries, model capacity and cost, test dependency and recovery assumptions, and check operational readiness. Convert requirements to automated tests, SLOs, policy, deployment checks and runbooks where practical. Reviews are constructive risk discovery, not compliance theater or a substitute for testing.
Architecture is iterative. Implementation, incidents, load tests, vendor changes and user feedback reveal new facts. Maintain current diagrams, specifications, decisions and trace links close to the workload. Revisit assumptions and quality priorities at meaningful change points. A successful architecture is not the most elaborate design; it is the simplest evolvable set of choices with credible evidence that it meets the prioritized outcomes and makes its residual risks visible to the people authorized to accept them.
Keep the translation artifacts next to the delivery work so they do not rot in a slide archive. When a requirement, SLO or constraint changes, update the trace to the affected decisions, tests and operational checks in the same change, or record why the lag is acceptable. An architecture that cannot show which user journey, quality scenario and residual risk a component exists to serve is no longer a decision record; it is decoration.
Worked example: "the checkout API must be real time"
Product writes "real time." Three teams hear three different licenses.
| statement | what it licenses | measured checkout | pass/fail |
|---|---|---|---|
| "must be real time" | WebSockets, no SLO | browser p95 2.4s | nobody can fail it |
| p95 < 2s at the API, 500 concurrent | API timeouts only | API 1.8s, browser 4.1s with JS | API green, users not |
| p95 < 2s at the browser, 500 concurrent, named journey | cache, timeouts, CDN, SLO | browser p95 1.7s | pass; 2.1s would fail |
The interview is the third row: stimulus, load, percentile, threshold, and observation point. "Real time" is not a requirement.
