Skip to content
Tech Interview Prep home
Technical interview guide

Hypothesis Testing

The framework for deciding whether an observed effect is likely real or could plausibly be noise — null hypotheses, p-values, and the two ways a test can be wrong.

Read
44 min
Practice MCQs
25
Interview QA
25
Edition
v6
Editorial status
Reviewed

Scope: Standards and library references current as of 2026-09: NIST/SEMATECH e-Handbook, the 2016 ASA statement on p-values, SciPy, and statsmodels.

Overview

Curated: · Written: · Reviewed:

Hypothesis testing

Hypothesis testing is misused more often than it is misunderstood, and the misuse follows a predictable pattern: a p-value read as the probability the null is true, a non-significant result reported as no difference, a threshold treated as a law of nature, and an analysis chosen after the data arrived. This guide works through what a p-value is and is not, following the American Statistical Association's 2016 statement, and then treats the design questions that determine whether a result means anything — power, minimum detectable effect, the match between randomisation unit and analysis unit, multiple comparisons, and the false-positive inflation caused by optional stopping. It ends on what a complete, replicable test report contains.

what a p-value actually is

A p-value is the probability of data at least as extreme as observed, assuming the null hypothesis is true.

It is computed under the null and says nothing about the probability that the null is true, nor about the size of any effect.

P(data this extreme | null true), not P(null true | data).

from scipy import stats
stats.ttest_ind(a, b).pvalue       # 0.03
# Means: if there were truly no difference, data this extreme would occur 3%
# of the time. It does not mean a 3% chance the null is true.

Interview trap. Reading a p-value of 0.03 as a three percent chance the null is true reverses the conditioning, which is the error the ASA statement singles out.

Engineering practice. State the p-value as evidence against the null under a specific model, and report the effect size and interval alongside it.

the null and alternative hypotheses

Hypothesis testing asks whether the data is compatible with a specific null, and failing to reject is not evidence that the null is true.

Absence of a detectable effect is compatible both with no effect and with an underpowered study, and the test cannot distinguish them.

Non-significant is not evidence of no effect.

# n = 40 per arm, observed difference 2%, 95% CI [-6%, +10%].
# "No significant difference" is true and useless: a 10% lift is still plausible.

Interview trap. Reporting 'no difference' after a non-significant result claims something the test does not support.

Engineering practice. Report the confidence interval, which shows what effect sizes remain plausible, rather than only the verdict.

type I and type II errors

The significance level bounds the false-positive rate, and power is one minus the false-negative rate at a specified effect size.

The two trade off for a fixed sample: tightening the significance level reduces false positives and increases false negatives.

Fix both rates before choosing a sample size.

alphapowern per arm (d = 0.2)
0.050.80393
0.050.90526
0.010.80586

Interview trap. Setting a significance level without computing power leaves the false-negative rate unknown and often extremely high.

Engineering practice. Fix both error rates and the smallest effect worth detecting, and derive the sample size from all three.

statistical power

Power is the probability of detecting an effect of a given size, and it must be computed before the experiment rather than after.

Power depends on the effect size, the variance, the sample size, and the significance level, and one of these is always the free variable.

Compute it before, from the effect you would act on.

from statsmodels.stats.power import TTestIndPower
TTestIndPower().solve_power(effect_size=0.2, alpha=0.05, power=0.8)   # 393.4

Interview trap. Computing power after the fact from the observed effect and reporting it as evidence adds nothing, because that quantity is a deterministic function of the p-value already reported.

Engineering practice. Compute power in advance for the minimum effect of interest, and report the study's detectable effect rather than post-hoc power.

the minimum detectable effect

Every experiment has a smallest effect it can reliably detect, and stating it is more useful than reporting a p-value.

It follows from the sample size, the variance, and the chosen error rates, so it is known before any data arrives.

Known before any data arrives.

# Baseline 5%, n = 10,000 per arm, alpha 0.05, power 0.8:
# MDE is about 0.6 percentage points, i.e. a 12% relative lift.
# If no plausible change is that large, the experiment cannot resolve anything.

Interview trap. Running an experiment whose detectable effect exceeds any plausible true effect guarantees a non-significant result regardless of reality.

Engineering practice. Compute the detectable effect during design, and decline to run the experiment when it exceeds what the business would act on.

effect size versus significance

Statistical significance says an effect is detectable; effect size says whether it matters.

With a large enough sample, any non-zero difference becomes significant, so significance alone carries no information about importance.

Large n makes any non-zero difference significant.

from scipy import stats
import numpy as np
rng = np.random.default_rng(0)
a = rng.normal(0, 1, 1_000_000); b = rng.normal(0.005, 1, 1_000_000)
stats.ttest_ind(a, b).pvalue      # ~0.0004 - significant
# The effect is 0.005 standard deviations. Nobody would act on it.

Interview trap. Shipping a change because a result was significant, without checking the magnitude, is how a statistically real but commercially irrelevant effect drives a decision.

Engineering practice. Set a practical significance threshold before the experiment, and compare the interval against it rather than against zero.

choosing a test

The test follows from the data type, the design, and the distributional assumptions, and each test states its own conditions.

Comparing means of two independent groups, paired measurements, and proportions all use different tests with different variance formulas.

Match the design; pairing that is ignored is power thrown away.

DesignTest
Two independent groups, meansWelch's t-test
Same units before and afterpaired t-test
Two proportionstwo-proportion z-test
Counts in categorieschi-squared

Interview trap. Applying an unpaired test to paired data discards the pairing that reduced variance, which wastes power.

Engineering practice. Match the test to the design, and state the assumptions it makes about independence, variance, and distribution.

the assumptions behind the t-test

The t-test assumes independent observations and approximate normality of the sampling distribution of the mean.

It is robust to non-normal data at large samples through the central limit theorem, but not to dependence or to extreme skew at small samples.

Independence is the assumption that actually fails.

# 1,000 page views from 100 users. Treating them as n = 1,000 gives an interval
# roughly sqrt(10) = 3.2x too narrow when views within a user are correlated.

Interview trap. Independence is the assumption that fails in practice — repeated measurements per user, or correlated sessions — and it inflates false positives rather than merely reducing power.

Engineering practice. Check the unit of randomisation against the unit of analysis, and use clustered standard errors when they differ.

unit of analysis mismatch

The analysis unit must match the randomisation unit, or the effective sample size is overstated.

Randomising by user and analysing by page view treats correlated views as independent, which narrows intervals well below their true width.

Aggregate to the randomisation unit, or use clustered errors.

per_user = df.groupby("user_id")["converted"].mean()     # analysis unit = user
stats.ttest_ind(per_user[df.arm == "A"], per_user[df.arm == "B"])

Interview trap. Treating each event as an independent observation because each is a separate row overstates the effective sample size, which is a common cause of significant results that fail to replicate.

Engineering practice. Aggregate to the randomisation unit before testing, or use a method that accounts for clustering explicitly.

non-parametric alternatives

Rank-based tests make weaker distributional assumptions and answer a slightly different question.

The Mann-Whitney test compares stochastic ordering rather than means, so a significant result does not directly imply a mean difference.

Mann-Whitney tests stochastic ordering, not means.

from scipy import stats
stats.mannwhitneyu(a, b)     # significant here does not license "the mean rose"

Interview trap. Substituting a rank test for a t-test and then reporting a mean difference mixes the test's question with the reported quantity.

Engineering practice. Report the quantity the test actually addresses, or use a bootstrap when a mean difference is required without normality assumptions.

the bootstrap

Bootstrapping produces intervals for almost any statistic by resampling the data, without a distributional assumption.

Resampling with replacement approximates the sampling distribution, and percentile or bias-corrected intervals are read from it.

Resample at the independent unit.

import numpy as np
rng = np.random.default_rng(0)
users = per_user.to_numpy()
boot = [rng.choice(users, len(users), replace=True).mean() for _ in range(10_000)]
np.percentile(boot, [2.5, 97.5])

Interview trap. The bootstrap requires the resampling unit to match the independence structure, so resampling rows of clustered data reproduces the mismatch problem.

Engineering practice. Resample at the independent unit, and use enough replications for the interval endpoints to be stable.

permutation tests

A permutation test computes the null distribution directly by reshuffling group labels, which makes its assumptions unusually transparent.

Under the null of exchangeability, every labelling is equally likely, so the observed statistic is compared against the distribution of shuffled statistics.

Shuffle within the structure the design actually has.

import numpy as np
obs = a.mean() - b.mean()
pool = np.concatenate([a, b]); rng = np.random.default_rng(0)
null = []
for _ in range(10_000):
    rng.shuffle(pool)
    null.append(pool[:len(a)].mean() - pool[len(a):].mean())
p = (np.abs(null) >= abs(obs)).mean()

Interview trap. Exchangeability fails when observations are clustered or time-ordered, and the shuffle must respect that structure.

Engineering practice. Shuffle within the correct structure, and use permutation tests where the analytic null distribution is unclear.

multiple comparisons

Testing many hypotheses at the same significance level guarantees false positives in proportion to the number of tests.

Twenty independent tests at five percent produce on average one false positive, which is why unadjusted metric dashboards always show something significant.

Twenty tests at 5% produce one false positive on average.

1 - 0.95 ** 20      # 0.6415 - a 64% chance of at least one false positive
1 - 0.95 ** 50      # 0.9231

Interview trap. Scanning many metrics and reporting the significant ones is a procedure with a false-positive rate far above the nominal level.

Engineering practice. Pre-register a primary metric, and control the family-wise error rate or the false-discovery rate for the rest.

Bonferroni and false-discovery control

Bonferroni controls the probability of any false positive; Benjamini-Hochberg controls the expected proportion of false positives among rejections.

Bonferroni is conservative and costs power heavily with many tests, while false-discovery control is more powerful and makes a weaker guarantee.

Family-wise for a few confirmatory tests; FDR for screening.

from statsmodels.stats.multitest import multipletests
multipletests(pvals, alpha=0.05, method="bonferroni")   # conservative
multipletests(pvals, alpha=0.05, method="fdr_bh")       # more powerful, weaker claim

Bonferroni over 200 exploratory metrics tests each at 0.00025 and finds nothing.

Interview trap. Applying Bonferroni to hundreds of exploratory metrics reduces power to near zero and produces no findings at all.

Engineering practice. Use family-wise control for a small set of confirmatory tests, and false-discovery control for exploratory screening.

optional stopping

Repeatedly testing as data accumulates and stopping at significance inflates the false-positive rate far above the nominal level.

Each look is another opportunity to cross the threshold by chance, so the true error rate compounds across looks.

Daily peeking can push a 5% error rate past 30%.

LooksTrue type I rate (nominal 5%)
15%
28%
514%
1019%
unlimitedapproaches 100%

Interview trap. Peeking at an experiment daily and stopping when it turns significant can raise a five percent false-positive rate above thirty percent.

Engineering practice. Fix the sample size in advance, or use a sequential method — group sequential boundaries or always-valid inference — designed for continuous monitoring.

sequential testing done correctly

Valid continuous monitoring requires a method whose error guarantee holds at every look, not a fixed-sample test applied repeatedly.

Group sequential designs spend the error budget across pre-planned looks; always-valid confidence sequences hold uniformly over time at the cost of width.

The guarantee must hold at every look, by design.

# O'Brien-Fleming boundaries for 4 planned looks (overall alpha 0.05):
# z thresholds 4.05, 2.86, 2.34, 2.02 - strict early, near-normal at the end.
# Always-valid confidence sequences hold at every t, and are wider in exchange.

Interview trap. Describing repeated fixed-sample testing as sequential analysis misrepresents a method with no error control as one with it.

Engineering practice. Choose the sequential method during design, and state the price in width or in required sample.

one-sided versus two-sided tests

A one-sided test is only legitimate when the opposite direction would lead to the same decision as no effect.

It concentrates the error budget in one tail, which increases power at the cost of being unable to detect harm in the other direction.

Fix the sidedness before looking at the direction.

stats.ttest_ind(a, b, alternative="greater")     # decided in the plan, not after
# Switching after seeing the sign doubles the effective false-positive rate.

Interview trap. Switching to a one-sided test after seeing the direction of the result doubles the effective false-positive rate.

Engineering practice. Fix the sidedness before looking at the data, and default to two-sided where a regression would matter.

the significance threshold is a convention

The five percent threshold is a convention, not a property of nature, and the appropriate level depends on the cost of each error.

A reversible change with low deployment cost tolerates a higher false-positive rate than an irreversible one.

Choose it from the cost of each error.

# Reversible copy change, cheap deploy: alpha 0.10 may be right.
# Irreversible pricing change: alpha 0.01 and a pre-registered guardrail set.

Interview trap. Treating the threshold as a scientific constant leads to identical decision rules for changes with very different consequences.

Engineering practice. Choose the level from the relative cost of the two errors, and document the reasoning with the experiment plan.

confidence intervals as the primary report

An interval conveys everything the p-value does plus the effect size and the precision, which is why it should lead the report.

Whether the interval excludes the null value reproduces the test's verdict, and its width shows how much the study could resolve.

The interval carries the verdict plus the precision.

# "+2.1% [95% CI: +0.4%, +3.8%]" - significant, and the size is bounded.
# "+2.1% [95% CI: -9.0%, +13.2%]" - the same point estimate, and nothing is known.

Interview trap. Reporting only significance leaves the reader unable to distinguish a precisely estimated small effect from an imprecisely estimated large one.

Engineering practice. Lead with the estimate and its interval, and mention the p-value as a secondary summary.

practical equivalence testing

Demonstrating that two options are equivalent requires an equivalence test, not a failure to reject a difference.

Two one-sided tests against a pre-specified equivalence margin establish that the effect lies within the range considered unimportant.

Two one-sided tests against a pre-specified margin.

# Margin: differences within +/-1% are not meaningful.
# Equivalence is demonstrated when the whole 90% CI lies inside [-1%, +1%].
# A non-significant difference test shows nothing of the kind.

Interview trap. Concluding equivalence from a non-significant difference is the standard misuse, and it holds even for badly underpowered studies.

Engineering practice. Define the equivalence margin in advance, and use the appropriate test when the goal is to demonstrate no meaningful difference.

assumption checking

Every test's validity depends on assumptions that should be checked rather than asserted.

Independence, variance homogeneity, and distributional shape each have diagnostics, and the consequences of violation differ in severity.

Prefer a robust method over a data-dependent choice between two methods.

stats.ttest_ind(a, b, equal_var=False)     # Welch: robust to unequal variance
# Testing variances first and then choosing the test makes the overall error
# rate depend on the data, which is exactly what the alpha was supposed to fix.

Interview trap. Selecting a test based on a preliminary assumption test changes the overall error rate, because the choice is itself data-dependent.

Engineering practice. Prefer a method robust to the doubtful assumption over a data-dependent choice between two methods.

the garden of forking paths

Analytic decisions made after seeing the data inflate false positives even when only one test is finally reported.

Choices about outlier removal, segmentation, transformation, and metric definition each multiply the space of possible analyses.

Analytic choices made after seeing the data inflate the error rate.

# Outlier rule (3 choices) x segment (5) x metric definition (3) x window (2)
# = 90 analyses. Reporting the best one as a single p-value is not honest.

Interview trap. Reporting a single p-value from an analysis that was selected after exploration understates the true error rate substantially.

Engineering practice. Pre-register the analysis plan, and label anything decided afterwards as exploratory.

replication

A single significant result is weak evidence, and independent replication is what turns it into a finding.

Replication tests the whole chain — design, instrumentation, and analysis — rather than only the statistical inference.

A single significant result is weak evidence.

# Two independent replications at 80% power:
0.8 * 0.8      # 0.64 chance both reach significance when the effect is real
# A failed replication is informative, not noise.

Interview trap. Treating a first significant result as settled is how instrumentation defects survive into product decisions.

Engineering practice. Replicate surprising or high-stakes results before acting, and treat a failed replication as informative rather than as noise.

Bayesian alternatives

A Bayesian analysis answers the question people mistakenly ask of a p-value: the probability of a hypothesis given the data.

It requires an explicit prior, which makes an assumption visible that the frequentist analysis leaves implicit in the design.

The prior is explicit; the decision rule still needs checking.

# Beta(1,1) prior, 120/2,000 vs 140/2,000:
# P(treatment better) ~ 0.90 - a direct answer to the question people ask.
# Simulate the decision rule's error rates before adopting it.

Interview trap. Adopting Bayesian reporting to avoid the stopping-rule problem still requires care, since decision rules based on posteriors have their own error properties.

Engineering practice. State the prior and its sensitivity, and evaluate the decision rule's error rates by simulation whichever framework is used.

reporting a test honestly

A complete report states the hypothesis, the design, the sample size and how it was chosen, the effect with its interval, and every analytic decision.

Each of these is required to judge whether the result is credible, and omitting any one makes the report unfalsifiable.

Hypothesis, design, n and how it was chosen, effect, interval, deviations.

Primary metric: checkout conversion (pre-registered)
Design: user-level randomisation, 14 days, 2 full weeks
n: 24,180 / 24,090 (SRM check p = 0.62)
Result: 3.11% -> 3.29%, +5.8% relative, 95% CI [+0.9%, +10.9%]
Deviations: none

Interview trap. A report of the form 'the variant won, p below 0.05' cannot be evaluated, replicated, or safely acted upon.

Engineering practice. Use a fixed report template covering all of these elements, and record deviations from the pre-registered plan.