Skip to content
Tech Interview Prep home
Technical interview guide

Hypothesis Testing

The framework for deciding whether an observed effect is likely real or could plausibly be noise — null hypotheses, p-values, and the two ways a test can be wrong.

Read
44 min
Practice MCQs
25
Interview QA
25
Edition
v6
Editorial status
Reviewed

Scope: Standards and library references current as of 2026-09: NIST/SEMATECH e-Handbook, the 2016 ASA statement on p-values, SciPy, and statsmodels.

Interview QA

Treat each question like a live interview question: answer out loud first (structure, assumptions, tradeoffs), then open the model answer to spot gaps and rehearse a tighter follow-up.

Curated: · Written: · Reviewed:

QA-1

Explain what a p-value actually is and how it changes a production decision.

QA-2

How would you reason about the null and alternative hypotheses in a system you own?

QA-3

Walk through type I and type II errors, including where engineers most often get it wrong.

QA-4

What does statistical power guarantee, and what does it deliberately leave open?

QA-5

Describe the minimum detectable effect and the evidence you would collect before relying on it.

QA-6

A teammate proposes a design that hinges on effect size versus significance. How do you evaluate it?

QA-7

Where does choosing a test matter, and where is it irrelevant?

QA-8

Teach the assumptions behind the t-test to an engineer who has only seen it as a rule of thumb.

QA-9

How does a mismatch between the unit of randomization and unit of analysis affect hypothesis testing, and how do you correct for it?

QA-10

How would you test a claim that the system depends on non-parametric alternatives?

QA-11

When does the non-parametric bootstrap fail for estimating confidence intervals, and how do you diagnose those failures in production data?

QA-12

What are the core differences between permutation tests and bootstrapping, and when should you choose one over the other?

QA-13

How does repeated peeking inflate Type I error in an A/B test, and what mathematical methods prevent it?

QA-14

When would you refuse a design because of bonferroni and false-discovery control?

QA-15

How would you explain optional stopping without using the usual slogan?

QA-16

What failure would you inject to check a team's understanding of sequential testing done correctly?

QA-17

How should on-call treat an alert that names one-sided versus two-sided tests as the cause?

QA-18

What belongs in a runbook section on the significance threshold is a convention, and what does not?

QA-19

How would you review a pull request whose risk is confidence intervals as the primary report?

QA-20

What trade-off does practical equivalence testing force that a junior answer usually skips?

QA-21

How would you brief product on why assumption checking delays a ship date?

QA-22

What is the smallest experiment that would change your mind about the garden of forking paths?

QA-23

How does replication interact with rollback, and where do people ignore that?

QA-24

What would you ask a candidate who recites bayesian alternatives but cannot apply it?

QA-25

How would you document reporting a test honestly so the next owner can operate it?