Overview
Curated: · Written: · Reviewed:
Probability distributions
A distribution is a claim about how data is generated, and the useful skill is matching the mechanism rather than the histogram. This guide works through the standard families in those terms — binomial for fixed trials, Poisson for constant-rate counts, exponential for memoryless waiting, normal for additive effects, log-normal for multiplicative ones — and is direct about the assumption each one makes and how it fails on real system data. Because engineering quantities are usually skewed, it gives particular attention to heavy tails, over-dispersion, mixtures, and censoring by timeouts, and to the practical consequences: why averaging a log-normal latency describes nothing, why capacity is sized from a percentile, and why end-to-end tail latency is worse than any single stage's.
distributions as models of a generating process
A distribution is chosen because it matches how the data is generated, not because the histogram resembles its shape.
Each standard distribution corresponds to a specific mechanism — counting successes, waiting for an event, summing many small effects — and that mechanism is what justifies its use.
Match the mechanism, then check the tail rather than the centre.
# Counting successes in a fixed number of independent trials -> binomial
# Waiting for the next event at a constant rate -> exponential
# Many small effects adding -> normal
# Many small effects multiplying -> log-normal
Interview trap. Fitting a distribution by visual similarity produces a model whose tails, which are what matter for risk, are wrong.
Engineering practice. State the generating mechanism, choose the distribution it implies, and check the fit in the tails rather than in the centre.
discrete versus continuous
Discrete distributions assign probability to individual values; continuous ones assign probability only to intervals.
A probability density can exceed one and is not a probability, which is why a continuous variable has zero probability at any exact value.
A density is not a probability and can exceed 1.
from scipy import stats
stats.uniform(0, 0.5).pdf(0.25) # 2.0 - a density, not a probability
stats.uniform(0, 0.5).cdf(0.25) # 0.5 - a probability
stats.norm().pdf(0.0) # 0.3989
Interview trap. Reading a density value as a probability produces statements that are not merely imprecise but out of range.
Engineering practice. Integrate over an interval when working with densities, and reserve point probabilities for discrete variables.
the Bernoulli and binomial distributions
The binomial counts successes in a fixed number of independent trials with a constant success probability.
Its variance is largest at a success probability of one half and shrinks towards the extremes, which affects the sample size needed to detect a change.
Variance peaks at p = 0.5 and shrinks at the extremes.
n = 1000
for p in (0.5, 0.1, 0.01):
print(p, (n * p * (1 - p)) ** 0.5)
# 0.5 -> 15.81, 0.1 -> 9.49, 0.01 -> 3.15
Interview trap. Applying it when trials are dependent or the probability varies across trials understates the variance substantially.
Engineering practice. Verify the constant-probability and independence assumptions, and use a beta-binomial or a mixed model when they fail.
conversion rates and their variance
A conversion rate is a binomial proportion whose uncertainty depends on both the rate and the sample size.
The standard error is largest near one half, so rare-event rates need far more observations for the same relative precision.
Rare conversions need far more traffic for the same relative precision.
| Base rate | n for +/-10% relative, 95% |
|---|---|
| 20% | ~1,540 per arm |
| 2% | ~18,800 per arm |
| 0.2% | ~192,000 per arm |
Interview trap. Applying a fixed sample-size rule regardless of the base rate under-powers experiments on rare conversions.
Engineering practice. Compute the required sample from the observed base rate and the effect size to be detected.
the Poisson distribution
The Poisson models counts of independent events occurring at a constant average rate in a fixed interval.
Its mean and variance are equal, which is both its identifying property and the easiest assumption to test.
Mean equals variance; real counts usually exceed that.
import numpy as np
counts = np.array(hourly_errors)
counts.mean(), counts.var()
# Poisson: 12.0, 12.1 -> consistent
# Real: 12.0, 143.0 -> over-dispersed; use a negative binomial
Interview trap. Real event counts are usually over-dispersed — variance exceeding the mean — because arrivals cluster, and Poisson intervals are then too narrow.
Engineering practice. Compare the sample variance with the mean, and use a negative binomial model when over-dispersion is present.
the exponential distribution and memorylessness
The exponential models waiting time between Poisson events and is memoryless: elapsed waiting tells you nothing about remaining waiting.
Memorylessness makes the mathematics tractable and is precisely what most real waiting times violate.
Elapsed waiting tells you nothing - which is what real waits violate.
from scipy import stats
e = stats.expon(scale=10)
(1 - e.cdf(15)) / (1 - e.cdf(5)) # 0.3679
1 - e.cdf(10) # 0.3679 - identical: memoryless
A service whose failure rate rises with age is Weibull, not exponential.
Interview trap. Assuming exponential service times in a queueing model when the real distribution has a long tail understates queue lengths badly.
Engineering practice. Test the memoryless property against the data, and use a Weibull or log-normal model when hazard changes with age.
the normal distribution and where it comes from
The normal arises when many small independent effects add, which is why it fits measurement error and rarely fits counts or durations.
It is symmetric with thin tails, so extreme values are far less likely than in the heavy-tailed distributions typical of system metrics.
Thin tails: a 5-sigma event is 1 in 3.5 million.
from scipy import stats
1 - stats.norm().cdf(5) # 2.87e-07
# Latency at 5 sigma above the mean happens far more often than that,
# which is how you know latency is not normal.
Interview trap. Assuming normality for latency or revenue leads to thresholds that are exceeded far more often than the model predicts.
Engineering practice. Check the tails with a quantile plot, and use an empirical or heavy-tailed model for skewed quantities.
the log-normal distribution
When effects multiply rather than add, the result is log-normal, which is why latency, file sizes, and income are skewed with long right tails.
The logarithm of the variable is normal, so geometric means and multiplicative confidence intervals are the natural summaries.
Report the median; the mean sits well above it.
import numpy as np
x = np.random.default_rng(7).lognormal(mean=0, sigma=1.5, size=100_000)
np.median(x), x.mean() # ~1.00 and ~3.08
np.mean(x > x.mean()) # ~0.23 - P(Z > sigma/2) = P(Z > 0.75), so only 23% exceed the mean
Interview trap. Reporting the arithmetic mean of a log-normal variable describes neither the typical value nor the median, since the mean sits well above both.
Engineering practice. Report the median and percentiles, and if a mean is needed, say explicitly that it is dominated by the tail.
heavy tails and power laws
In a heavy-tailed distribution the extreme values dominate the total, and the sample mean may not converge usefully at all.
For sufficiently heavy tails the variance is infinite, so the central limit theorem does not apply and confidence intervals built on it are invalid.
With a tail index below 2 the variance is infinite and the CLT does not apply.
import numpy as np
rng = np.random.default_rng(1)
x = rng.pareto(a=1.2, size=1_000_000) + 1 # infinite variance
[x[:n].mean() for n in (10_000, 100_000, 1_000_000)]
# The running mean keeps jumping: it does not converge usefully.
Interview trap. Averaging a heavy-tailed quantity produces an estimate that changes materially with each new extreme observation.
Engineering practice. Identify tail behaviour before choosing a summary, and use percentiles, trimmed means, or explicit tail models.
the uniform distribution and its uses
The uniform assigns equal density across a range and is mostly useful as a building block rather than as a data model.
Inverse transform sampling turns uniform draws into draws from any distribution with an invertible cumulative function.
Inverse transform turns uniforms into anything invertible.
import numpy as np
u = np.random.default_rng(3).random(100_000)
exp_samples = -np.log(1 - u) / rate # inverse CDF of the exponential
Interview trap. Assuming uniformity for hash outputs or identifier distributions without checking is how shard imbalance goes unnoticed.
Engineering practice. Test uniformity where it is assumed, and use the inverse-transform relationship when generating custom distributions.
the geometric and negative binomial distributions
The geometric counts trials until the first success and the negative binomial until a fixed number of successes.
The negative binomial also serves as an over-dispersed alternative to the Poisson, which is its more common use in practice.
Two parameterisations; check which one the library uses.
from scipy import stats
stats.geom(0.5).pmf(1) # 0.5 - counts trials, support starts at 1
# Some libraries count failures before the first success, support starting at 0.
# Mixing them shifts every result by one.
Interview trap. Two parameterisations of the geometric exist — counting trials or counting failures — and mixing them shifts every result by one.
Engineering practice. Check the library's parameterisation before using it, and validate against a simulated sample.
parameterisation differences between libraries
The same distribution is parameterised differently across libraries — rate against scale, variance against standard deviation — and the difference is silent.
A rate parameter is the reciprocal of a scale parameter, so passing one where the other is expected inverts the distribution's spread.
Rate is the reciprocal of scale, and the swap is silent.
from scipy import stats
stats.expon(scale=1/0.5).mean() # 2.0 - scipy takes a scale
# numpy: rng.exponential(scale=2.0). R: rexp(n, rate = 0.5).
# Passing 0.5 where a scale is expected gives a mean of 0.5, not 2.0.
Interview trap. Porting a model between libraries without checking parameterisation produces results that are wrong but plausible.
Engineering practice. Read each library's parameter definition, and verify by comparing the sample mean against the intended one.
the cumulative distribution function
The cumulative function is the most useful representation for engineering, because it directly answers 'what fraction is below this value'.
Percentiles are its inverse, and comparing empirical cumulative functions is a robust way to compare two samples.
Compare empirical CDFs; histograms depend on the bin width.
import numpy as np
xs = np.sort(sample)
ecdf = np.arange(1, len(xs) + 1) / len(xs)
# Two samples plotted this way are comparable without choosing a bin width.
Interview trap. Comparing histograms is sensitive to bin width, which can create or hide apparent differences between the same two samples.
Engineering practice. Plot empirical cumulative functions when comparing distributions, and use histograms only for exposition.
quantile plots for assessing fit
A quantile-quantile plot compares a sample against a reference distribution and shows exactly where the fit fails.
Deviation at the ends indicates tail mismatch, which is the failure that matters for risk and capacity decisions.
The plot says where the fit fails; a test only says that it does.
from scipy import stats
stats.probplot(latencies, dist="norm", plot=plt)
# A concave-up right end is the usual signature of a heavy right tail.
Interview trap. Assessing fit with a goodness-of-fit test alone gives a single verdict without saying which part of the distribution is wrong.
Engineering practice. Use the plot for diagnosis and the test for a decision, and note that large samples reject any model.
goodness-of-fit tests at scale
With enough data, every goodness-of-fit test rejects, because no real dataset is exactly generated by a textbook distribution.
The test answers whether the deviation is detectable rather than whether it matters for the decision at hand.
At large n every model is rejected.
from scipy import stats
x = stats.norm(0, 1).rvs(1_000_000, random_state=1) + 0.002 # tiny real deviation
stats.kstest(x, "norm").pvalue # ~1e-06: rejected, and the model is still fine
Interview trap. Abandoning a useful model because a test rejects at a large sample size mistakes statistical detectability for practical relevance.
Engineering practice. Judge fit by the magnitude of the deviation in the region that matters, not by the test's verdict alone.
mixture distributions
Data drawn from several populations is a mixture, and fitting a single distribution to it produces a model that describes none of them.
A bimodal histogram is the visible case, but mixtures often appear only as a heavier tail than any single component would produce.
A single mean can describe a value neither component produces.
# 90% of requests hit a warm cache at 5 ms; 10% miss and take 300 ms.
0.9 * 5 + 0.1 * 300 # 34.5 ms - a latency essentially no request experiences
Interview trap. Reporting a single mean for a mixture describes a value that may be rare or impossible in either component.
Engineering practice. Segment by the suspected grouping variable and model the components separately when they differ materially.
the chi-squared and t distributions
These distributions describe statistics computed from samples rather than data, which is why they appear in inference rather than in modelling.
The t distribution accounts for an estimated rather than known standard deviation, and it approaches the normal as the sample grows.
The t distribution accounts for an estimated variance.
| n | t critical (95%) | normal |
|---|---|---|
| 5 | 2.776 | 1.96 |
| 10 | 2.262 | 1.96 |
| 30 | 2.045 | 1.96 |
| 100 | 1.984 | 1.96 |
Interview trap. Using normal critical values with a small sample produces intervals that are too narrow because the estimated variance is itself uncertain.
Engineering practice. Use t-based intervals for small samples, and note that the difference is negligible beyond a few dozen observations.
the beta distribution
The beta distribution models a probability itself and is the natural prior for a binomial rate.
Because it is conjugate to the binomial, the posterior after observing successes and failures is another beta with updated parameters.
A weak prior stops two trials becoming a 100% rate.
successes, trials = 2, 2
successes / trials # 1.00
(successes + 1) / (trials + 2) # 0.75 - Beta(1, 1) prior
# With Beta(1,1), the posterior mean after 200/400 is 0.5005: the prior fades.
Interview trap. Reporting a conversion rate from a handful of observations without shrinkage produces extreme estimates such as one hundred percent from two trials.
Engineering practice. Use a weakly informative beta prior to shrink small-sample rates, and state the prior explicitly.
the empirical distribution
When the generating mechanism is unknown, the empirical distribution of the data is often a better model than any parametric fit.
Bootstrapping resamples from it to obtain intervals for any statistic, without assuming a distributional form.
Bootstrap resamples what you saw; it cannot invent the unobserved tail.
import numpy as np
rng = np.random.default_rng(0)
boot = [rng.choice(sample, len(sample), replace=True).mean() for _ in range(10_000)]
np.percentile(boot, [2.5, 97.5])
Interview trap. The bootstrap inherits the data's limitations and cannot recover behaviour beyond the observed range, which matters most for tails.
Engineering practice. Use the bootstrap for central statistics, and treat extrapolation into the unobserved tail as outside its scope.
sampling from a distribution
Reproducible simulation requires an explicit generator instance with a recorded seed, not a global random function.
Modern libraries provide independent generator objects precisely so that parallel simulation streams do not interfere.
An explicit generator with a recorded seed; not the global functions.
import numpy as np
rng = np.random.default_rng(seed=42) # explicit, passable, splittable
children = rng.spawn(8) # independent streams for parallel work
# np.random.seed(42) makes results depend on call order elsewhere in the program.
Interview trap. Seeding a global generator makes results depend on library call order elsewhere in the program.
Engineering practice. Pass an explicit generator, record its seed with the results, and use the library's stream-splitting facility for parallel runs.
truncation and censoring
Truncated data omits observations outside a range, and censored data records that a value exceeded a limit without recording the value.
Timeouts censor latency measurements, so the recorded distribution understates the tail by exactly the observations that matter.
A timeout recorded as a measurement biases every tail statistic downward.
# 1,000 requests, 30 hit a 5,000 ms timeout.
# Recording those as 5,000 ms: p99 = 5,000 ms, understating the true tail.
# Recording them as censored at 5,000 ms and using survival methods keeps the bound honest.
Interview trap. Treating a timeout as a completed measurement at the timeout value biases every tail statistic downward.
Engineering practice. Record censoring explicitly, and use survival-analysis methods that account for it rather than discarding or imputing.
distributions in capacity planning
Capacity is sized from the tail of the demand distribution, not from its mean.
Peak-to-mean ratios of several times are common, so a system sized to the mean is overloaded a large fraction of the time.
Size to a percentile of demand, not the mean.
# Mean 1,200 rps, p99 4,800 rps.
# Provisioned at 1,200 x 1.5 = 1,800 rps: overloaded for roughly 5% of minutes.
# Provisioned at p99 + headroom = 5,500 rps: overloaded about 1% of minutes.
Interview trap. Multiplying the mean by a fixed safety factor implicitly assumes a distributional shape that is usually wrong.
Engineering practice. Size against a stated percentile of observed demand, and re-derive it as the distribution changes.
convolution and aggregate latency
The distribution of a sum of independent stages is their convolution, and its tail is heavier than any single stage's.
A request touching several services with independent tail latencies has a tail probability roughly summing the components' tail probabilities.
Independent stage tails add up in the composite.
# 5 services, each p99 = 100 ms, independent.
1 - 0.99 ** 5 # 0.049 - about 5% of requests hit at least one slow stage
# End-to-end p95 is therefore near the single-service p99, not near its p95.
Interview trap. Assuming a request is as fast as its slowest measured stage understates end-to-end tail latency substantially.
Engineering practice. Model end-to-end latency by composing stage distributions, and measure the composite directly rather than inferring it.
choosing a summary statistic
The right summary follows from the distribution's shape and from what the number will be used for.
A mean answers questions about totals, a median about the typical case, and a percentile about the bad case, and these diverge sharply under skew.
Say which question the number answers.
| Question | Statistic |
|---|---|
| Total cost across all requests | mean x count |
| What a typical user sees | median |
| What the worst-served users see | p95, p99 |
| Is the shape changing | full histogram |
Interview trap. Reporting a single number for a skewed distribution invites every reader to interpret it as whichever of the three they had in mind.
Engineering practice. Report several summaries with the shape, and name which one the decision depends on.
validating a distributional assumption
Any distributional assumption in a production system should be checked against data periodically, because the generating process changes.
Traffic mix, retries, and client changes all alter the shape, so a model validated at launch is not validated a year later.
Re-check on a schedule; the generating process moves.
from scipy import stats
stats.ks_2samp(reference_window, current_window).statistic
# Alert on the distance between windows, not only on the mean.
Interview trap. Treating a distributional choice as a one-time modelling decision leaves alerting thresholds calibrated to a process that no longer exists.
Engineering practice. Re-check the fit on a schedule, and alert on distributional drift rather than only on the summary statistic.
