Overview
Curated: · Written: · Reviewed:
Probability fundamentals
Most probability errors in engineering are not arithmetic errors; they are errors about what was being computed. This guide is organised around that: defining the sample space before computing anything, keeping the two directions of conditioning apart, and recognising that independence is a claim about the process rather than a modelling convenience. It covers expectation and its linearity, variance and why standard deviations do not average, the law of large numbers and the central limit theorem with the conditions under which each actually applies, and the frequentist meaning of a confidence interval. The applied sections cover base rates in rare-event classification, Simpson's paradox and selection bias, percentiles as order statistics that cannot be averaged across shards, and simulation with its own reported uncertainty.
sample space and events
A probability statement is meaningless until the sample space is defined, because the same English question maps to different spaces.
An event is a subset of the sample space, and its probability is determined by how outcomes were generated rather than by how the question is phrased.
The generating process, not the wording, fixes the answer.
# "A family has two children, at least one is a boy. P(both boys)?"
# Sample space {BB, BG, GB} -> 1/3.
# "The older child is a boy" -> {BB, BG} -> 1/2.
# Same English intuition, different generating process, different answer.
Interview trap. Treating the two-children and Monty Hall puzzles as questions with one obvious answer skips the step where the generating process is pinned down, which is what makes them ambiguous in the first place.
Engineering practice. Write the generating process explicitly before computing anything, and check that the stated question corresponds to it.
conditional probability
Conditioning restricts the sample space, so a conditional probability is a ratio within the restricted space rather than a modified unconditional one.
The definition divides the joint probability by the probability of the conditioning event, which is undefined when that event has probability zero.
A ratio inside the restricted space, undefined when the condition has probability zero.
P_A_given_B = P_A_and_B / P_B # requires P_B > 0
# P(rain | forecast) is not P(forecast | rain). Reversing these is the
# prosecutor's fallacy.
Interview trap. Reading the probability of the evidence given the hypothesis as the probability of the hypothesis given the evidence reverses the conditioning, which is the prosecutor's fallacy.
Engineering practice. Write which quantity is being conditioned on, and state whether the question asks for one direction or the other.
Bayes' theorem
Bayes' theorem converts between the two directions of conditioning, and the prior is what makes the conversion possible.
The posterior is proportional to the likelihood times the prior, with the normalising constant summing over all hypotheses.
The base rate is what makes an accurate test inconclusive.
prevalence, sens, spec = 0.001, 0.99, 0.99
p_pos = sens * prevalence + (1 - spec) * (1 - prevalence)
posterior = sens * prevalence / p_pos
round(posterior, 4) # 0.0902 - about 9%, from a 99%-accurate test
Interview trap. Ignoring a low base rate makes a highly accurate test look conclusive when most positives are still false positives.
Engineering practice. State the base rate explicitly, and compute the posterior rather than reasoning from the test's accuracy alone.
the base-rate fallacy in production
For a rare event, even a very accurate classifier produces mostly false positives, which is a property of the prevalence rather than of the model.
Precision depends on prevalence while recall and specificity do not, so a model validated on a balanced set behaves differently in production.
Precision moves with prevalence; recall does not.
| Prevalence | Recall | Specificity | Precision |
|---|---|---|---|
| 50% | 99% | 99% | 99.0% |
| 1% | 99% | 99% | 50.0% |
| 0.1% | 99% | 99% | 9.0% |
The model is identical in all three rows. Only the population changed.
Interview trap. Reporting accuracy on a rare-event problem hides that predicting the majority class always scores well.
Engineering practice. Report precision and recall at the production prevalence, and recompute them when the prevalence shifts.
independence
Two events are independent when conditioning on one leaves the other's probability unchanged, which is a factual claim about the process rather than an assumption of convenience.
Independence permits multiplying probabilities, which is why it appears in almost every tractable model and why its failure is so consequential.
A claim about the process, and correlated failures break it.
# Three replicas, each 99.9% available.
0.001 ** 3 # 1e-09 if independent - about 32 ms of downtime per year
# Same rack, same power feed, same deploy: the correlated failure probability is
# closer to 0.001, which is 31,536 seconds per year. Nine orders of magnitude apart.
Interview trap. Assuming independence across correlated failures — machines in one rack, requests from one user — understates joint failure probability by orders of magnitude.
Engineering practice. Name the mechanism that would make events dependent, and estimate the joint behaviour empirically when one exists.
mutually exclusive versus independent
Mutually exclusive events cannot both occur, whereas independent events do not inform each other, and the two concepts are nearly opposites.
Two mutually exclusive events with non-zero probability are necessarily dependent, since observing one rules the other out entirely.
Nearly opposites: exclusive events with non-zero probability are dependent.
# A = "die shows 1", B = "die shows 2": mutually exclusive.
# P(A and B) = 0, but P(A) * P(B) = 1/36. So they are NOT independent -
# knowing A tells you B did not happen.
Interview trap. Using the terms interchangeably leads to adding probabilities that should be multiplied and the reverse.
Engineering practice. Add for mutually exclusive alternatives, multiply for independent conjunctions, and check which relationship actually holds.
random variables and their distributions
A random variable is a function from outcomes to numbers, and its distribution summarises the induced probabilities.
This separation is what lets the same underlying process support several different variables, such as count, duration, and cost.
The variable is a function; a sample is one realisation.
# One process, three variables:
# X = request latency in ms
# Y = 1 if latency > 500 ms else 0
# Z = cost of serving the request
# "The latency was 240 ms" is a realisation, not the variable.
Interview trap. Confusing the variable with a particular realisation leads to statements about 'the value' where the distribution is what matters.
Engineering practice. Distinguish the variable from a sample of it in both notation and speech, especially when discussing variance.
expectation and linearity
Expectation is linear regardless of dependence, which makes it the most reliable tool in applied probability.
The expectation of a sum is the sum of expectations whether or not the terms are independent, which is what makes indicator-variable arguments work.
Linear regardless of dependence; variance is not.
# E[X + Y] = E[X] + E[Y] always.
# Var(X + Y) = Var(X) + Var(Y) + 2 Cov(X, Y)
# Two perfectly correlated variables with variance 1 each: Var(sum) = 4, not 2.
Interview trap. Assuming the same holds for variance is wrong: variance of a sum includes covariance terms unless the terms are uncorrelated.
Engineering practice. Use linearity of expectation freely, and add covariance terms explicitly when computing the variance of a sum.
variance and standard deviation
Variance measures spread in squared units and standard deviation returns it to the original units, which is why the latter is reported.
Variance adds across independent variables, which is what makes standard deviation grow with the square root of the count rather than linearly.
Pool variances, never average standard deviations.
# Group A: n=100, sd=10. Group B: n=100, sd=20.
(10 + 20) / 2 # 15 - wrong
((99*100 + 99*400) / 198) ** 0.5 # 15.81 - pooled within-group sd
Interview trap. Averaging standard deviations across groups is not the standard deviation of the pooled data, because variances rather than deviations combine.
Engineering practice. Pool variances weighted by degrees of freedom, then take the root, and never average standard deviations directly.
the law of large numbers
The sample mean converges to the true mean as the sample grows, which is the guarantee that makes estimation possible.
Convergence is in probability and says nothing about any particular finite sample, and its rate depends on the variance.
Deviations are diluted, not corrected.
# After 100 flips: 60 heads, 10 more than expected.
# After 10,000 flips the same absolute surplus is 0.1 percentage points.
# Nothing "evens out" - the surplus is simply swamped.
Interview trap. The gambler's fallacy misreads it as a self-correcting mechanism, when in fact deviations are diluted rather than compensated.
Engineering practice. Reason about the rate of convergence, and remember that a fair process has no memory of past outcomes.
the central limit theorem
The distribution of a sample mean approaches normality as the sample grows, largely independent of the population's shape.
This is what justifies normal-based confidence intervals for means even when the underlying data is plainly not normal.
It applies to the mean, and convergence depends on skew.
# Exponential with rate 1 (skewness 2): n = 30 already gives a near-normal mean.
# Log-normal with sigma = 3 (skewness ~ 10^4): even n = 10,000 leaves the sample
# mean badly skewed, and normal-based intervals under-cover.
Interview trap. Applying it to the data rather than to the mean, or to heavy-tailed data where convergence is far slower, produces intervals that are too narrow.
Engineering practice. Check the sample size against the skewness of the data, and use bootstrapping when convergence is doubtful.
sampling distributions
A statistic computed from a sample is itself a random variable, and its distribution is what inference is about.
The standard error is the standard deviation of that sampling distribution, which is distinct from the standard deviation of the data.
Standard error is the sd of the estimator, not of the data.
sd = 20 # spread of the data
n = 400
se = sd / n ** 0.5 # 1.0 - spread of the sample mean
# Reporting 20 as the uncertainty of the mean overstates it 20-fold.
Interview trap. Reporting the data's standard deviation as though it were the uncertainty of the mean overstates the uncertainty by the square root of the sample size.
Engineering practice. Label the quantity explicitly as a standard deviation or a standard error, and derive one from the other rather than confusing them.
confidence intervals
A ninety-five percent confidence interval is a statement about the procedure's long-run coverage, not about the probability that the parameter lies inside this particular interval.
Under repeated sampling, ninety-five percent of intervals constructed this way contain the true value; the parameter is fixed and the interval is random.
A statement about the procedure's coverage, not about this interval.
# Simulate: draw 10,000 samples, build a 95% interval for each.
# About 9,500 of those intervals contain the true parameter.
# For the one interval you actually have, the parameter is either in it or not.
Interview trap. Reading it as a probability statement about the parameter is a Bayesian claim that the frequentist interval does not make.
Engineering practice. Report intervals with the coverage interpretation, and use a credible interval when the probability statement about the parameter is wanted.
the interval width and sample size
Interval width shrinks with the square root of the sample size, so halving it requires four times the data.
This square-root relationship governs the cost of precision and therefore the feasibility of any measurement plan.
Halving the width needs four times the data.
| n | Half-width (sd = 20) |
|---|---|
| 100 | 3.92 |
| 400 | 1.96 |
| 1,600 | 0.98 |
| 6,400 | 0.49 |
Interview trap. Assuming a linear return on additional data leads to experiment plans whose required sample sizes are badly underestimated.
Engineering practice. Compute the required size from the target width up front, and report when the required size is impractical.
covariance and correlation
Correlation is covariance normalised by the two standard deviations, and it measures only linear association.
A correlation near zero is compatible with a strong deterministic non-linear relationship, which is why a plot is not optional.
Correlation measures linear association only.
import numpy as np
x = np.linspace(-1, 1, 101)
y = x ** 2 # perfectly determined
np.corrcoef(x, y)[0, 1] # ~0.0
Interview trap. Reporting a correlation coefficient without inspecting the data misses non-linearity, outlier dominance, and Anscombe-style pathologies.
Engineering practice. Plot before summarising, and use rank correlation or mutual information when the relationship may not be linear.
correlation and causation
An association can arise from causation in either direction, from a common cause, or from the selection of the sample itself.
Only an intervention, or a design that mimics one, distinguishes among these explanations from data alone.
Only an intervention, or a design that mimics one, separates the explanations.
# Users who open the app daily churn less.
# Causal? Habit reduces churn.
# Reverse? Users who intend to stay open it more.
# Confounded? Both driven by how well the product fits.
# Controlling for a post-treatment variable can make this worse, not better.
Interview trap. Assuming that adding enough control variables turns an observational estimate into a causal one is wrong, and the additional controls can introduce collider bias rather than remove confounding.
Engineering practice. State the causal question, draw the assumed causal structure, and prefer a randomised experiment where one is feasible.
Simpson's paradox
An association present in every subgroup can reverse in the aggregate when group sizes differ.
The reversal is driven by a confounding variable that correlates with both the grouping and the outcome.
The aggregate can reverse every subgroup.
| Group | A | B |
|---|---|---|
| Easy cases | 93% (81/87) | 87% (234/270) |
| Hard cases | 73% (192/263) | 69% (55/80) |
| Combined | 78% (273/350) | 83% (289/350) |
B wins overall while A wins both subgroups, because A took most of the hard cases.
Interview trap. Reporting only the aggregate, or only the subgroups, can each support the opposite conclusion, and neither is inherently the right one.
Engineering practice. Decide which comparison answers the question from the causal structure, and report the subgroup breakdown alongside the aggregate.
survivorship and selection bias
Conditioning on an outcome that depends on the variable of interest biases every estimate computed from the survivors.
Users who churned are missing from the current-user table, so any average computed there describes survivors rather than users.
The available table describes survivors, not users.
-- Average tenure of "our users":
SELECT avg(now() - signup_at) FROM users WHERE status = 'active';
-- Everyone who churned is excluded, so this measures survivors and rises over
-- time even if the product is getting worse.
Interview trap. Treating an available dataset as a random sample is the most common way this bias enters production analysis.
Engineering practice. Describe how records entered the dataset, and reconstruct the excluded population before generalising.
expectation versus percentiles
A mean describes the centre and says nothing about the tail, which is why latency is reported in percentiles.
For a skewed distribution the mean can exceed the ninetieth percentile, so it neither describes the typical case nor the bad case.
For a skewed distribution the mean can exceed the 90th percentile.
# 1,000 requests: 990 at 50 ms, 10 at 30,000 ms.
# mean = 349 ms, p50 = 50 ms, p90 = 50 ms, p99 = 30,000 ms.
# The mean describes neither the typical request nor the bad one.
Interview trap. Using average latency as a service objective hides the tail that users actually experience.
Engineering practice. Report percentiles for skewed quantities, and use the mean only where the total, not the typical value, is what matters.
percentiles do not average
The ninety-fifth percentile of a combined dataset is not the average of the per-shard ninety-fifth percentiles.
Percentiles are order statistics, so combining them requires the distributions themselves, not their summaries.
Order statistics do not combine; histograms do.
# Shard A: p95 = 100 ms. Shard B: p95 = 100 ms.
# Combined p95 is NOT 100 ms - it depends on both full distributions.
# Aggregate histograms or t-digests, then compute the percentile once.
Interview trap. Aggregating percentile metrics across hosts by averaging produces a number with no distributional meaning.
Engineering practice. Aggregate histograms or sketches and compute the percentile once, rather than averaging precomputed percentiles.
the birthday problem and collisions
Collisions become likely at roughly the square root of the space size, which is far sooner than intuition suggests.
The number of pairs grows quadratically in the sample size, so the expected collision count rises much faster than the sample.
Collisions arrive near the square root of the space.
| Space | 50% collision at |
|---|---|
| 365 days | 23 people |
| 2^32 ids | ~77,000 ids |
| 2^64 ids | ~5.1 x 10^9 ids |
| 2^122 (UUIDv4 random bits) | ~2.7 x 10^18 ids |
Interview trap. Sizing an identifier space from the expected item count rather than from the collision probability produces duplicates in production.
Engineering practice. Size identifier spaces against the square-root threshold, and use the birthday bound to justify the width chosen.
simulation as a check
When an analytic probability is hard to derive or easy to get wrong, simulation gives a reliable independent answer.
A few hundred thousand replications estimate a probability to within a fraction of a percent, which is usually enough to confirm or refute a derivation.
Seed it, or the result cannot be regenerated.
import numpy as np
rng = np.random.default_rng(seed=20260907) # recorded with the result
hits = sum(trial(rng) for _ in range(200_000))
hits / 200_000
Interview trap. An unseeded simulation is not reproducible, so a result that cannot be regenerated cannot be reviewed.
Engineering practice. Seed the generator explicitly, record the seed with the result, and report the simulation's own uncertainty.
the Monte Carlo error
A simulated estimate carries its own uncertainty, which shrinks with the square root of the replication count.
The standard error of a simulated proportion follows directly from the binomial formula, so the needed replications are computable in advance.
Report the estimate with its own interval.
p, n = 0.5, 200_000
se = (p * (1 - p) / n) ** 0.5 # 0.00112
# So the estimate is good to about +/- 0.2 percentage points at 95%.
# Quoting five decimal places from 200,000 trials implies precision that is not there.
Interview trap. Reporting a simulated probability to more digits than the replication count supports implies precision that is not there.
Engineering practice. Compute the replication count from the target precision, and report the estimate with its interval.
the conjunction fallacy in estimation
A conjunction of conditions can never be more probable than any of its parts, however plausible the combined story sounds.
Every additional condition multiplies by a factor at most one, so specificity reduces probability even as it increases narrative appeal.
Adding conditions can only reduce probability.
# P(outage) >= P(outage AND caused by the new deploy AND during peak hours)
# The detailed scenario is more memorable and strictly less likely.
Interview trap. Risk assessments that describe a detailed scenario tend to be judged more likely than the broader event containing them.
Engineering practice. Decompose scenarios into their conditions and check that the estimated probabilities respect the containment ordering.
communicating uncertainty
A point estimate without an interval is not a result, because it gives the reader no way to judge whether a difference is meaningful.
The interval carries the sample size, the variance, and the estimator's behaviour, none of which the point estimate conveys.
Estimate, interval, sample size - and the smallest detectable difference.
# Bad: "Conversion rose 8%."
# Good: "3.24% -> 3.50% (+8.0%), 95% CI [-1.2%, +17.9%], n = 12,400 per arm.
# This test could only detect a change above 12%."
Interview trap. Reporting a percentage change between two noisy measurements without intervals invites decisions based on noise.
Engineering practice. Report the estimate, the interval, and the sample size together, and state the smallest difference the measurement could detect.
