Overview
Curated: · Written: · Reviewed:
A/B testing
An A/B test is an engineering system that happens to produce a statistic, and most of its failures are engineering failures. This guide treats it that way: randomisation and why intention-to-treat analysis is what preserves it, trigger conditions that must be implemented identically in both arms, the sample-ratio check that decides whether any metric may be read at all, and A/A tests as continuous validation of the whole pipeline. The statistical sections cover sample sizing, variance reduction with pre-period covariates, ratio metrics and the delta method, and the unit-of-analysis mismatch that produces narrow intervals and results that do not replicate. The last sections cover interference in marketplaces, switchback designs, and what a report has to contain to be trusted.
why randomisation is the point
Randomised assignment is what makes an A/B test causal, because it makes the groups comparable on every variable including unobserved ones.
Any pre-existing difference is distributed by chance rather than systematically, so the difference in outcomes estimates the treatment effect.
Adopters differ from non-adopters by more than the feature.
-- Not an experiment: adopters chose to adopt.
SELECT avg(revenue) FROM users WHERE used_feature;
-- vs
SELECT avg(revenue) FROM users WHERE NOT used_feature;
-- The gap is usually dominated by who selected in, not by the feature.
Interview trap. Comparing users who adopted a feature against those who did not is not an experiment, and the selection difference is usually larger than the effect.
Engineering practice. Randomise at assignment time and analyse by assigned group, whatever the user subsequently did.
intention to treat
Analysis is by assigned group rather than by treatment received, because switching to received treatment reintroduces selection.
Users who fail to experience the treatment are systematically different, so excluding them breaks the randomisation.
Analyse by assignment; dropping the unexposed reintroduces selection.
df.groupby("assigned_arm")["converted"].mean() # ITT - valid
df[df.saw_feature].groupby("assigned_arm")... # only treatment can "see" it
Interview trap. Dropping users who never saw the new feature is the most common way a randomised experiment is turned into an observational study.
Engineering practice. Analyse by assignment, and use a trigger-based design if only exposed users should be compared.
trigger-based analysis
Restricting the analysis to users who could have been affected increases power, but the trigger must be defined identically in both arms.
The trigger condition must depend only on pre-treatment information, so it identifies the same population in control and treatment.
The trigger must fire identically in both arms.
# Right: log `eligible=True` in both arms at the point the decision is made,
# before branching on the arm.
if user_qualifies(ctx):
log_trigger(ctx) # both arms
return new_flow() if in_treatment else old_flow()
Verify equal trigger rates before reading any metric.
Interview trap. A trigger that fires only in the treatment arm — because the control has no equivalent code path — silently selects a different population.
Engineering practice. Implement the trigger in both arms and log it identically, and verify equal trigger rates before analysing.
the randomisation unit
The randomisation unit must be at least as coarse as the level at which the treatment can leak between users.
Randomising by user prevents inconsistent experience across sessions; randomising by cluster prevents interference between connected users.
Coarser than the leakage, or the arms contaminate.
| Change | Unit |
|---|---|
| Visible UI change | user |
| Ranking model in a marketplace | market or region |
| Backend latency optimisation | request is acceptable |
| Social feature | cluster of connected users |
Interview trap. Randomising by request or session for a visible interface change gives the same user different experiences and contaminates the comparison.
Engineering practice. Choose the unit from the leakage structure, and use the same unit for analysis.
sample ratio mismatch
When the observed split differs significantly from the intended one, the experiment is broken and the result must not be read.
The split is checked with a chi-squared test against the intended ratio, and a significant deviation indicates an assignment, logging, or filtering defect.
A chi-squared check that gates the whole readout.
from scipy import stats
obs = [24_180, 23_050] # expected 50/50
stats.chisquare(obs).pvalue # 3e-07 - the experiment is broken
# Do not read any metric until this is explained.
Interview trap. Adjusting the analysis for an observed ratio mismatch, or dismissing a small one as noise, treats a broken assignment as a statistical inconvenience rather than as evidence the arms hold different populations.
Engineering practice. Run the ratio check automatically before any metric is reported, and block the readout when it fails.
A/A tests
Running an experiment with identical arms validates the whole pipeline, since any significant result is by construction a false positive.
Repeated A/A tests should produce p-values distributed uniformly, which detects variance underestimation and assignment bias.
P-values should be uniform, not merely non-significant.
from scipy import stats
pvals = [run_aa_test() for _ in range(1_000)]
stats.kstest(pvals, "uniform").pvalue # a low value means the pipeline is biased
sum(p < 0.05 for p in pvals) / 1_000 # should be about 0.05
Interview trap. Skipping A/A validation leaves clustering and instrumentation defects invisible until they produce a wrong product decision.
Engineering practice. Run A/A tests continuously, and check that the p-value distribution is uniform rather than merely that nothing was significant.
choosing the primary metric
A single pre-registered primary metric is what makes the experiment's error rate meaningful.
Secondary metrics are reported as context and guardrails, but the decision rule refers to the primary metric alone.
One, pre-registered, with the decision rule.
Primary: checkout conversion, user level
Ship if: lower bound of the 95% CI > 0 and no guardrail breached
Otherwise: do not ship, regardless of secondary metrics
Interview trap. Choosing the metric after seeing which one moved converts the experiment into an exercise in multiple comparisons.
Engineering practice. Register the primary metric and the decision rule before launch, and record it where it cannot be edited afterwards.
guardrail metrics
Guardrails detect harm that the primary metric would not reveal, and they should trigger a stop rather than a discussion.
Latency, error rate, crash rate, and unsubscribe rate are the usual ones, and each needs a pre-agreed threshold.
Thresholds agreed in advance, and automated.
| Guardrail | Stop threshold |
|---|---|
| p95 latency | +10% |
| Error rate | +0.1pp |
| Crash-free sessions | -0.2pp |
| Unsubscribe rate | +5% relative |
Interview trap. Shipping a conversion win that degraded latency substantially trades a measured gain against an unmeasured loss.
Engineering practice. Define guardrails and their thresholds with the experiment plan, and automate the stop when one is breached.
surrogate metrics and long-term effects
Short-term metrics can move in the opposite direction from long-term value, so a surrogate must be validated against the outcome it stands for.
Engagement metrics are especially prone to this, since attention captured now may cost retention later.
Validate the surrogate against the outcome it stands for.
# Session count rose 6% while 90-day retention fell 1.2%.
# The surrogate moved the wrong way relative to the thing it was chosen to predict.
Interview trap. Optimising a surrogate that was never validated is how a series of individually positive experiments produces a worse product.
Engineering practice. Validate surrogates against long-term holdouts periodically, and re-derive them when the product changes.
long-term holdouts
A holdout group excluded from all shipped changes measures cumulative effects that individual experiments cannot.
It reveals whether the sum of individually positive changes actually improved the outcome over months.
The sum of wins is not the measured trajectory.
# 24 shipped experiments, claimed lifts summing to +19% conversion.
# The 1% holdout shows +6% over the same period.
# The gap is interaction, novelty, and regression to the mean.
Interview trap. Assuming individual experiment wins add up ignores interaction and novelty effects that only a holdout can detect.
Engineering practice. Maintain a small long-term holdout, and reconcile its measured trajectory against the sum of claimed effects.
novelty and primacy effects
A visible change can produce a temporary effect from novelty or from disruption of learned behaviour, in either direction.
Both fade over days to weeks, so an effect measured only in the first days can reverse sign by the end.
Inspect the effect by day, not just the pooled number.
df.groupby("day")["lift"].mean()
# day 1: +12%, day 3: +7%, day 7: +2%, day 14: +1%
# Stopping at significance on day 2 would have shipped a 12% claim.
Interview trap. Stopping an experiment as soon as it reaches significance captures precisely the period when these effects are largest.
Engineering practice. Run for a pre-planned duration covering at least a full weekly cycle, and inspect the effect trend over time.
weekly seasonality
User behaviour varies systematically by day of week, so an experiment must cover whole weeks.
A run covering only weekdays measures a population and behaviour mix that differs from the full week's.
Run whole weeks, concurrently.
# Weekday-only run: 65% of traffic is desktop.
# Full week: 52% desktop.
# A desktop-favouring change looks better in the truncated window.
Interview trap. Comparing a treatment period against a preceding control period rather than running concurrently confounds the effect with the calendar.
Engineering practice. Randomise concurrently, and run in whole-week multiples.
computing the required sample size
The sample size follows from the baseline rate, its variance, the minimum effect worth detecting, and the two error rates.
It is computable before launch, and it determines whether the experiment is worth running at the available traffic.
Derive n and the duration before launch.
from statsmodels.stats.power import NormalIndPower
from statsmodels.stats.proportion import proportion_effectsize
es = proportion_effectsize(0.032, 0.0352) # 3.2% -> 3.52%, a 10% lift
NormalIndPower().solve_power(es, alpha=0.05, power=0.8) # ~49,800 per arm
Interview trap. Launching without a size computation produces experiments that cannot detect any plausible effect and consume weeks proving nothing.
Engineering practice. Compute the size and duration first, and abandon or redesign the experiment when the traffic is insufficient.
variance reduction
Adjusting for pre-experiment behaviour reduces variance substantially without biasing the estimate.
The CUPED technique regresses the outcome on a pre-period covariate, removing variation that predates and is unaffected by the treatment.
CUPED uses a strictly pre-treatment covariate.
theta = np.cov(y, x_pre)[0, 1] / np.var(x_pre)
y_adj = y - theta * (x_pre - x_pre.mean())
# Variance typically drops 30-50% for engagement metrics, which cuts the
# required sample by the same factor. Using an in-experiment covariate biases it.
Interview trap. Using a covariate measured during the experiment can absorb part of the treatment effect and bias the result.
Engineering practice. Use only strictly pre-treatment covariates, and verify that the adjustment leaves an A/A test unbiased.
the unit-of-analysis problem in practice
When users generate many events, treating events as independent overstates the sample size dramatically.
The effective sample is closer to the user count than to the event count, and the correct interval accounts for within-user correlation.
Effective n is closer to the user count than the event count.
# 2,000,000 events from 50,000 users.
# Naive interval uses n = 2,000,000; the honest one is nearer n = 50,000.
# That is a factor of sqrt(40) = 6.3 on the width.
Interview trap. Event-level analysis of a user-randomised experiment produces intervals several times too narrow and results that do not replicate.
Engineering practice. Aggregate to the user, or apply the delta method or clustered standard errors that account for the correlation.
ratio metrics
A metric that is a ratio of two random quantities needs a variance formula that accounts for both, not the binomial formula.
The delta method gives the variance of a ratio of user-level sums, which is what click-through and per-session metrics require.
Use the delta method; the binomial formula is wrong here.
# clicks / impressions, both varying per user:
# Var(R) ~ (1/mean_d^2) * (Var(n) - 2R Cov(n, d) + R^2 Var(d))
# Treating it as a simple proportion ignores the varying denominator entirely.
Interview trap. Treating clicks per impression as a simple proportion ignores that the denominator varies by user and understates the variance.
Engineering practice. Use the delta method or bootstrapping at the user level for ratio metrics, and validate against A/A behaviour.
outliers and metric capping
Heavy-tailed metrics such as revenue are dominated by a few users, so a single outlier can determine the result.
Capping at a high percentile reduces variance greatly at the cost of a small, stated bias.
Cap at a pre-agreed percentile, identically in both arms.
cap = np.percentile(historical_revenue, 99) # decided before the experiment
y = np.minimum(revenue, cap)
# Report capped and uncapped. Choosing the cap after seeing the data biases it.
Interview trap. Choosing the cap after seeing the data, or applying it differently across arms, introduces bias in the direction of the observed effect.
Engineering practice. Set the cap in advance from historical data, apply it identically to both arms, and report the capped and uncapped results.
interference between units
Randomisation assumes one user's assignment does not affect another's outcome, which fails in social, marketplace, and shared-resource settings.
In a marketplace, treating some buyers shifts supply available to the rest, so the control group is affected by the treatment.
In a marketplace, treating some buyers changes supply for the rest.
# Treatment bids more aggressively -> control sees fewer available listings.
# The measured lift can exceed the true one, or flip sign entirely.
# Randomise by market and accept the smaller effective sample.
Interview trap. Standard analysis under interference can report an effect with the wrong sign, because the control has moved as well.
Engineering practice. Randomise at the level that contains the interference — market, region, or time slice — and accept the reduced power.
switchback experiments
When interference makes user-level randomisation invalid, randomising time periods across the whole system is an alternative.
The system alternates between arms on a schedule, and the analysis compares periods while accounting for carryover and autocorrelation.
Periods are correlated; do not treat them as independent.
# 30-minute periods, alternating arms, 2 weeks = 672 periods.
# Effective sample is far below 672 because adjacent periods correlate;
# use a design that accounts for carryover and autocorrelation.
Interview trap. Treating each period as an independent observation ignores temporal correlation and understates the variance.
Engineering practice. Choose period length from the carryover duration, and analyse with methods that account for the time-series structure.
segment analysis
Breaking a result down by segment multiplies the hypotheses and is exploratory unless the segments were pre-registered.
A heterogeneous treatment effect is a real phenomenon, but detecting it reliably requires either pre-registration or an explicit interaction test.
Pre-register segments and test the interaction.
# 12 segments at alpha 0.05: P(at least one false positive) = 1 - 0.95**12 = 46%.
# Report the interaction test, not the one segment that happened to be significant.
Interview trap. Reporting the one segment where the effect was significant is a multiple-comparison error presented as an insight.
Engineering practice. Pre-register segments of interest, test for interaction rather than per-segment significance, and label the rest exploratory.
ramping and staged rollout
Ramping traffic gradually limits the blast radius, and each stage's result must be analysed on the traffic that stage actually saw.
Pooling data across ramp stages mixes populations if the ramp was not random, so the analysis has to account for it.
Ramp for safety; analyse per stage.
# Stage 1: 1% for 2 days. Stage 2: 50% for 12 days.
# Pooling them weights the early population - and the early days - unequally.
Interview trap. Combining a one percent early stage with a fifty percent later stage without accounting for the change treats different periods as one sample.
Engineering practice. Analyse each stage on its own or model the stage explicitly, and use the ramp for safety rather than for measurement.
when not to experiment
Some changes cannot or should not be A/B tested, and recognising them is part of experimental competence.
Legal obligations, security fixes, and changes with tiny expected effects relative to available traffic all fall outside what an experiment can decide.
Compare the detectable effect against the plausible one.
# 400 weekly signups, baseline conversion 20%.
# MDE at 80% power over 4 weeks: about 8 percentage points.
# No copy change is worth 8 points, so the test cannot resolve the question.
Interview trap. Insisting on an experiment for a change whose detectable effect exceeds any plausible true effect wastes weeks and produces a non-result.
Engineering practice. Check the detectable effect against the plausible effect, and use qualitative or monitoring-based evaluation when the experiment cannot resolve it.
instrumentation as the dominant risk
Most experiment failures are engineering failures — logging gaps, assignment bugs, caching — rather than statistical ones.
Cached responses can serve the wrong variant, and a logging change can shift a metric independently of the treatment.
Check assignment, logging and caching before the statistics.
# Symptom: treatment shows 30% fewer events.
# Cause found in 20 minutes: the new page emitted the event after a redirect,
# so the analytics call was cancelled. Nothing to do with the feature.
Interview trap. Investigating a surprising result statistically before checking the instrumentation usually wastes the investigation.
Engineering practice. Check assignment balance, logging completeness, and cache behaviour first when a result is surprising.
the decision rule
The decision rule should be written before launch and should cover the significant, non-significant, and harmful outcomes.
A pre-agreed rule prevents the outcome from being renegotiated once a stakeholder sees a result they dislike.
Written before launch, covering all three outcomes.
Significant positive, no guardrail breach -> ship
Significant negative or guardrail breach -> revert, write up
Not significant -> do not ship; the CI bounds what we ruled out
Interview trap. Deciding what counts as success after the data arrives makes the whole exercise unfalsifiable.
Engineering practice. Write the rule with thresholds for the primary metric and the guardrails, and record who decides when it is ambiguous.
reporting an experiment
A trustworthy report states the hypothesis, the design, the sample and duration, the ratio check, the effect with its interval, the guardrails, and every deviation from the plan.
Each element is required to judge credibility, and the ratio check in particular determines whether anything else can be read at all.
Missing the ratio check is a blocking omission.
Hypothesis | Design | Unit | Dates (whole weeks) | n per arm | SRM p-value
Primary metric: point estimate, CI, MDE
Guardrails: each with its threshold and observed value
Deviations from the plan: listed, or "none"
Interview trap. A report stating only that the variant won by a percentage cannot be evaluated or replicated by anyone else.
Engineering practice. Use a fixed template that includes all of these fields, and treat a missing ratio check as a blocking omission.
