Overview
Curated: · Written: · Reviewed:
Regression analysis
Regression is easy to run and hard to read, and nearly every applied mistake comes from confusing two goals: predicting well and estimating an effect. This guide keeps them separate. It covers what least squares actually optimises and what a coefficient means once the holding-fixed clause is taken seriously, then works through the diagnostics that decide whether any of the reported uncertainty is valid — residual patterns, heteroskedasticity, correlated errors, multicollinearity, and influential points. The causal sections treat omitted-variable and collider bias as design questions rather than modelling ones, and the applied sections cover logistic and count models, regularisation and the scaling it requires, transformed outcomes, missing data, and regression to the mean.
what least squares actually fits
Ordinary least squares finds the coefficients minimising the sum of squared residuals, which is a specific and consequential choice of loss.
Squaring makes the estimator sensitive to large residuals, and it corresponds to a maximum-likelihood fit under normally distributed errors.
Squaring makes one outlier dominate the fit.
import numpy as np
x = np.arange(10); y = 2 * x + 1
np.polyfit(x, y, 1) # slope 2.00
y[9] = 100 # one bad point
np.polyfit(x, y, 1) # slope ~6.4 - the line now chases the outlier
Interview trap. Treating least squares as the neutral default hides that a single extreme observation can dominate the fit.
Engineering practice. Inspect the influence of individual points, and use a robust loss when outliers are expected rather than exceptional.
interpreting a coefficient
A coefficient is the expected change in the outcome per unit change in that predictor, holding the other predictors fixed.
The holding-fixed clause is what makes multiple regression different from a series of simple regressions, and it is often physically impossible.
'Holding the others fixed' is often physically impossible.
# Regressing salary on years_experience and age, which correlate at 0.95:
# the coefficient on age is "the effect of a year older at the same experience",
# a comparison that barely exists in the data.
Interview trap. Reading a coefficient as a marginal effect when the predictors are correlated describes a comparison that cannot occur in the data.
Engineering practice. Report coefficients with the correlation structure of the predictors, and prefer predicted values for communicating effects.
linearity is about parameters
A linear model is linear in its parameters, not necessarily in its predictors, so polynomial and interaction terms remain linear models.
Transforming predictors keeps the estimation machinery unchanged while allowing curved relationships to be fitted.
Transform predictors; the model stays linear.
import numpy as np
X = np.column_stack([x, x**2, np.log(x + 1)]) # still ordinary least squares
Interview trap. Rejecting linear regression because a relationship is curved misunderstands what the linearity assumption constrains.
Engineering practice. Transform predictors or add basis functions when the relationship is curved, and check the residuals afterwards.
the residual plot
Residuals plotted against fitted values are the single most informative diagnostic, and a good fit shows no pattern.
Curvature indicates a missing transformation, a funnel indicates non-constant variance, and clusters indicate a missing grouping variable.
R-squared can be high while the residuals show a clear pattern.
resid = y - model.predict(X)
# Funnel shape -> heteroskedasticity
# Curvature -> a missing transformation
# Clusters -> a missing grouping variable
# All three are invisible in the R-squared value.
Interview trap. Judging a model by its coefficient of determination alone hides every one of these, since a high value is compatible with all of them.
Engineering practice. Plot the residuals before reporting anything, and treat any visible pattern as a specification finding.
heteroskedasticity
Non-constant error variance leaves the coefficients unbiased but makes the standard errors wrong.
The consequence is invalid confidence intervals and tests, not a wrong point estimate, which is why the defect is easy to miss.
Coefficients stay unbiased; the standard errors do not.
import statsmodels.api as sm
res = sm.OLS(y, X).fit()
res_robust = sm.OLS(y, X).fit(cov_type="HC3")
# Same coefficients; the robust standard errors are often 1.5-3x larger.
Interview trap. Reporting significance from a heteroskedastic model overstates precision, often substantially.
Engineering practice. Use heteroskedasticity-robust standard errors by default, and consider transforming the outcome when the variance grows with the mean.
correlated errors
Independence of errors is the assumption whose violation does the most damage, and time series and clustered data violate it routinely.
Positively correlated errors make the effective sample size smaller than the row count, so intervals are too narrow.
Positive autocorrelation makes intervals far too narrow.
res = sm.OLS(y, X).fit(cov_type="HAC", cov_kwds={"maxlags": 7})
from statsmodels.stats.stattools import durbin_watson
durbin_watson(res.resid) # ~2 is fine; near 0 means strong positive correlation
Interview trap. Fitting a regression to daily metrics without accounting for autocorrelation produces confident conclusions from a handful of effective observations.
Engineering practice. Use clustered or time-series-aware standard errors, and check the residual autocorrelation explicitly.
multicollinearity
Highly correlated predictors leave the model's predictions stable while making the individual coefficients unstable and hard to interpret.
The variance inflation factor quantifies how much a predictor's coefficient variance is inflated by its correlation with the others.
Predictions stay stable; the individual coefficients do not.
from statsmodels.stats.outliers_influence import variance_inflation_factor
[variance_inflation_factor(X, i) for i in range(X.shape[1])]
# VIF above 10 is the usual flag; the coefficient's variance is inflated 10-fold.
Interview trap. Concluding that a predictor does not matter because its coefficient is not significant, when it is collinear with another, misreads the instability.
Engineering practice. Check inflation factors, and combine or drop redundant predictors when interpretation rather than prediction is the goal.
omitted variable bias
A variable correlated with both a predictor and the outcome biases that predictor's coefficient when it is left out.
The direction of the bias follows from the signs of the two correlations, which means it can be reasoned about even when the variable is unmeasured.
The direction follows from the two correlations.
# Regressing sales on ad_spend, omitting seasonality.
# Seasonality raises both -> the ad coefficient is biased upward.
# You can reason about the sign even when the variable is unmeasured.
Interview trap. Interpreting an observational coefficient causally without addressing omitted variables is the most consequential error in applied regression.
Engineering practice. Enumerate plausible confounders, state the expected direction of bias, and prefer an experiment when the question is causal.
collider bias
Controlling for a variable that is caused by both the predictor and the outcome creates a spurious association rather than removing one.
Conditioning on a common effect induces dependence between its causes, which is the opposite of what adjustment is intended to do.
Controlling for a common effect creates an association.
Talent ----> Admission <---- Test score
Among admitted students only, talent and test score become negatively correlated even if they are independent in the population. Controlling for admission causes it.
Interview trap. Adding every available variable as a control assumes all adjustment helps, which is false and can introduce bias where none existed.
Engineering practice. Choose controls from an explicit causal diagram, and never select them by which ones improve model fit.
the coefficient of determination
The coefficient of determination measures explained variance in the fitted data and never decreases when a predictor is added.
The adjusted version penalises the parameter count, but neither measures predictive accuracy on new data. An added predictor uncorrelated with the residual leaves it unchanged rather than raising it.
Never decreases when a predictor is added.
import numpy as np
rng = np.random.default_rng(0)
X_noise = rng.normal(size=(50, 20)) # pure noise
sm.OLS(y, sm.add_constant(X_noise)).fit().rsquared # ~0.42 on 50 rows
sm.OLS(y, sm.add_constant(X_noise)).fit().rsquared_adj # ~0.02
Interview trap. Choosing between models by this value alone selects the most over-fitted one, since adding predictors can never reduce it.
Engineering practice. Compare models by out-of-sample error or by an information criterion, not by in-sample explained variance.
prediction intervals versus confidence intervals
A confidence interval bounds the mean response, and a prediction interval bounds a single new observation, which is always much wider.
The prediction interval adds the residual variance to the uncertainty of the fitted mean, and that term does not shrink with sample size.
The prediction interval adds the residual variance and does not shrink with n.
pred = res.get_prediction(X_new)
pred.conf_int() # for the mean response - narrow
pred.conf_int(obs=True) # for a single new observation - much wider
Interview trap. Presenting a confidence band as the range future observations will fall in understates the spread dramatically.
Engineering practice. Choose the interval matching the question, and state clearly which one a plotted band represents.
extrapolation
A regression is only supported within the range of predictor values it was fitted on.
Nothing in the fit constrains behaviour outside that range, so extrapolated predictions rest entirely on the assumed functional form.
Outside the fitted range the answer rests entirely on the assumed form.
# Fitted on tenure 0-24 months, predicting at 60 months:
# a linear fit says 5x the value; a saturating one says 1.3x.
# The data cannot distinguish them, because it never went there.
Interview trap. Extrapolating a fitted trend to a capacity or growth forecast is a common way a model produces confident nonsense.
Engineering practice. Record the fitted range, refuse or flag predictions outside it, and validate any extrapolation against a mechanism.
influential observations
Influence combines an unusual predictor value with a large residual, and a single influential point can change a coefficient's sign.
Leverage measures the predictor's unusualness and Cook's distance combines it with the residual to quantify actual influence.
Leverage plus a large residual, not just an unusual y.
infl = res.get_influence()
infl.cooks_distance[0].max() # > 4/n is the usual flag
# Report the fit with and without the influential points rather than deleting quietly.
Interview trap. Removing points because they are outliers in the outcome, without examining leverage, discards data for the wrong reason.
Engineering practice. Examine influence diagnostics, investigate influential points as data-quality questions, and report the fit with and without them.
categorical predictors
A categorical predictor enters as indicator variables with one level omitted as the reference, and every coefficient is relative to that reference.
Including all levels plus an intercept makes the design matrix singular, which is why one level must be dropped.
Indicator encoding with a reference level; integers impose an order.
import pandas as pd
pd.get_dummies(df["plan"], drop_first=True) # free is the reference
# df["plan"].map({"free": 0, "pro": 1, "enterprise": 2}) claims
# enterprise - pro == pro - free, which is not a fact about the data.
Interview trap. Encoding an unordered category as a single integer imposes an ordering and a constant spacing that do not exist.
Engineering practice. Use indicator encoding for unordered categories, and state the reference level whenever coefficients are reported.
interactions
An interaction term says the effect of one predictor depends on another, and its presence changes how the main effects must be read.
With an interaction, a main-effect coefficient is the effect when the other predictor is zero, which may be outside the observed range.
Centre first, or the main effect describes a point outside the data.
df["x_c"] = df.x - df.x.mean()
df["z_c"] = df.z - df.z.mean()
sm.OLS(y, sm.add_constant(df[["x_c", "z_c"]].assign(inter=df.x_c * df.z_c))).fit()
# Without centring, the coefficient on x is the effect when z == 0, which may never occur.
Interview trap. Interpreting main effects as average effects in the presence of an interaction misstates them, sometimes with the wrong sign.
Engineering practice. Centre predictors before adding interactions so main effects are interpretable, and report predicted values across the range.
regularisation
Ridge and lasso penalties trade a little bias for a large reduction in variance, which improves prediction when predictors are many or correlated.
Ridge shrinks coefficients smoothly and lasso can drive them to exactly zero, which is why lasso also performs selection.
Biased by design; do not read the coefficients as effects.
from sklearn.linear_model import RidgeCV, LassoCV
RidgeCV(alphas=np.logspace(-3, 3, 50)).fit(Xs, y) # shrinks, keeps every feature
LassoCV(cv=5).fit(Xs, y) # can set coefficients to exactly 0
Interview trap. Reading a regularised coefficient and its ordinary standard error as an effect estimate is invalid, because the penalty deliberately biases the coefficient and invalidates the usual inference.
Engineering practice. Use regularisation for prediction, choose the penalty by cross-validation, and use unpenalised estimation when the coefficients themselves are the result.
scaling before regularisation
Penalties are scale-dependent, so predictors must be standardised before a ridge or lasso fit.
Without standardisation, a predictor measured in small units receives a large coefficient and is penalised more heavily for no substantive reason.
Fit the scaler inside the fold.
from sklearn.pipeline import make_pipeline
from sklearn.preprocessing import StandardScaler
pipe = make_pipeline(StandardScaler(), LassoCV(cv=5))
cross_val_score(pipe, X, y, cv=5) # the scaler is refit on each training fold
Interview trap. Fitting a regularised model on raw features silently selects predictors by their units rather than by their importance.
Engineering practice. Standardise inside the cross-validation pipeline so the scaling is fitted on the training fold only.
logistic regression
For a binary outcome, logistic regression models the log odds as linear, which keeps predictions inside the zero-to-one range.
Coefficients are log-odds ratios, so exponentiating gives an odds ratio rather than a probability change.
Coefficients are log-odds; exponentiating gives an odds ratio.
import numpy as np
np.exp(0.7) # 2.01 - odds ratio, not relative risk
# At a 40% base rate, an odds ratio of 2 is a risk ratio of about 1.43, not 2.
Interview trap. Reporting an exponentiated coefficient as a relative risk overstates the effect whenever the base rate is not small.
Engineering practice. Report odds ratios as such, and convert to predicted probabilities at stated covariate values when communicating.
separation in logistic models
When a predictor perfectly separates the outcome classes, the maximum-likelihood coefficient is unbounded and the fit does not converge.
The optimiser drives the coefficient towards infinity, producing an enormous estimate with an enormous standard error.
A huge coefficient with a huge standard error is a symptom.
# coef = 24.7, std err = 12,400 -> perfect separation, not a huge effect.
from sklearn.linear_model import LogisticRegression
LogisticRegression(penalty="l2", C=1.0) # penalised: finite and stable
Interview trap. Reading a huge coefficient as a huge effect misses that it is a numerical symptom of separation.
Engineering practice. Check for separation, and use a penalised likelihood or a weakly informative prior to obtain a finite estimate.
count outcomes
Counts are modelled with Poisson or negative binomial regression rather than with least squares on the raw count.
These models respect non-negativity and the mean-variance relationship that count data actually has.
Poisson or negative binomial with an exposure offset.
import statsmodels.api as sm
sm.GLM(counts, X, family=sm.families.Poisson(), offset=np.log(exposure)).fit()
# OLS on counts predicts negative values and mis-states the variance.
Interview trap. Fitting least squares to counts produces negative predictions and standard errors that ignore the variance structure.
Engineering practice. Use a count model with an offset for exposure, and check for over-dispersion before trusting the intervals.
transforming the outcome
A logarithmic transform of a skewed outcome changes the model's meaning from additive to multiplicative effects.
Coefficients become approximate proportional changes, and back-transforming the fitted mean gives a geometric rather than an arithmetic mean.
Back-transforming the mean gives a geometric mean.
import numpy as np
np.exp(np.mean(np.log(y))) # geometric mean - below the arithmetic mean
# A smearing estimator is needed when the arithmetic mean is what is wanted.
Interview trap. Exponentiating a fitted log-scale mean and reporting it as the expected value understates it, because the transform is non-linear.
Engineering practice. State whether the reported quantity is a geometric or arithmetic mean, and apply a smearing correction when the arithmetic mean is required.
missing data
Dropping rows with missing values is unbiased only when the missingness is unrelated to the outcome, which is rarely true.
Complete-case analysis silently changes the population, and mean imputation understates variance while distorting relationships.
Complete-case analysis changes the population.
df.isna().mean().sort_values(ascending=False).head()
# Dropping rows where income is missing removes the people least willing to report it -
# which is usually correlated with the outcome.
Interview trap. Imputing with the mean and then reporting ordinary standard errors treats imputed values as if they had been observed.
Engineering practice. Model the missingness mechanism, use multiple imputation when appropriate, and report how many rows were affected.
regression to the mean
Extreme measurements tend to be followed by less extreme ones purely from measurement noise, with no intervention required.
Selecting units because they were extreme and then measuring them again guarantees an apparent improvement.
Selecting the extremes guarantees an apparent improvement.
import numpy as np
rng = np.random.default_rng(0)
week1 = rng.normal(100, 15, 10_000)
week2 = rng.normal(100, 15, 10_000) # no intervention at all
worst = week1 < np.percentile(week1, 10)
week1[worst].mean(), week2[worst].mean() # ~78 and ~100
Interview trap. Evaluating an intervention applied to the worst-performing units without a control group attributes this statistical artefact to the intervention.
Engineering practice. Include a control group selected the same way, and expect a shift towards the mean in both.
model selection and validation
A model is validated on data not used to fit it, and every choice made from the data — including feature selection — must be inside the validation loop.
Selecting features on the full dataset before cross-validation leaks information and produces optimistic estimates.
Every data-dependent step goes inside the fold.
pipe = make_pipeline(StandardScaler(), SelectKBest(k=20), Ridge())
cross_val_score(pipe, X, y, cv=5)
# Selecting features on the full dataset first typically inflates the score by
# 5-15 points on wide data.
Interview trap. Reporting cross-validated error from a pipeline whose feature selection ran beforehand overstates performance, sometimes greatly.
Engineering practice. Put every data-dependent step inside the pipeline, and hold out a final test set that is used exactly once.
prediction versus explanation
A model built to predict well and one built to estimate an effect are different models, and conflating them causes most applied regression errors.
Prediction rewards any correlated feature, while effect estimation requires a defensible set of controls and excludes post-treatment variables.
Different models for different goals.
| Goal | Include | Read |
|---|---|---|
| Prediction | anything that helps, including proxies | out-of-sample error |
| Effect estimation | a defensible control set, no post-treatment variables | one coefficient with its CI |
Interview trap. Reading coefficients from a predictive model as effects imports every correlation in the data into a causal claim.
Engineering practice. State the goal first, and choose the specification, the controls, and the evaluation criterion from that goal.
