Skip to content
Tech Interview Prep home
Technical interview guide

A/B Testing

Applying hypothesis testing to compare two product/design variants — sample size, statistical power, and the traps of stopping early.

Read
46 min
Practice MCQs
25
Interview QA
25
Edition
v6
Editorial status
Reviewed

Scope: Practitioner and standards references current as of 2026-09: the experimentation-platform literature, NIST/SEMATECH e-Handbook, the 2016 ASA statement on p-values, SciPy, and statsmodels.

Interview QA

Treat each question like a live interview question: answer out loud first (structure, assumptions, tradeoffs), then open the model answer to spot gaps and rehearse a tighter follow-up.

Curated: · Written: · Reviewed:

QA-1

Explain why randomisation is the point and how it changes a production decision.

QA-2

How would you reason about intention to treat in a system you own?

QA-3

Walk through trigger-based analysis, including where engineers most often get it wrong.

QA-4

What does the randomisation unit guarantee, and what does it deliberately leave open?

QA-5

Describe sample ratio mismatch and the evidence you would collect before relying on it.

QA-6

A teammate proposes a design that hinges on a/A tests. How do you evaluate it?

QA-7

Where does choosing the primary metric matter, and where is it irrelevant?

QA-8

Teach guardrail metrics to an engineer who has only seen it as a rule of thumb.

QA-9

How do you construct and validate surrogate metrics to measure long-term effects in short-duration experiments?

QA-10

How would you test a claim that the system depends on long-term holdouts?

QA-11

What would you measure before treating novelty and primacy effects as settled?

QA-12

How does weekly seasonality change if the workload grows by two orders of magnitude?

QA-13

How do you calculate variance and analyze confidence intervals for ratio metrics when the randomization unit differs from the analysis unit?

QA-14

When would you refuse a design because of variance reduction?

QA-15

How would you explain the unit-of-analysis problem in practice without using the usual slogan?

QA-16

How do you detect and handle extreme outliers in continuous metrics like revenue per user, and how does capping affect bias versus variance?

QA-17

How should on-call treat an alert that names outliers and metric capping as the cause?

QA-18

What belongs in a runbook section on interference between units, and what does not?

QA-19

How would you review a pull request whose risk is switchback experiments?

QA-20

What trade-off does segment analysis force that a junior answer usually skips?

QA-21

How would you brief product on why ramping and staged rollout delays a ship date?

QA-22

What is the smallest experiment that would change your mind about when not to experiment?

QA-23

How does instrumentation as the dominant risk interact with rollback, and where do people ignore that?

QA-24

What would you ask a candidate who recites the decision rule but cannot apply it?

QA-25

How would you document reporting an experiment so the next owner can operate it?