Overview
Curated: · Written: · Reviewed:
Key takeaways
- Regularization changes the learning objective or effective training capacity to prefer solutions expected to generalize; it is not a repair for leakage, invalid labels, distribution shift, or missing serving-time signal.
- L2 shrinks coefficients smoothly and often stabilizes correlated features. L1 can set coefficients exactly to zero, but selected features can be unstable when predictors are correlated. Elastic net combines both behaviors.
- Penalty strength is data-, feature-scale-, sample-weight-, model-, and implementation-dependent. Fit scaling and tune regularization inside valid cross-validation; never compare raw hyperparameter numbers across incompatible objective conventions.
- Stronger regularization usually raises training error and may reduce validation error until underfitting begins. Diagnose with train/validation curves, baselines, uncertainty, calibration, and slices rather than assuming more penalty is safer.
- Early stopping, dropout, data augmentation, tree constraints, and architectural bottlenecks are capacity controls with distinct assumptions and failure modes; they are not interchangeable knobs.
- Production selection includes predictive quality, coefficient/feature stability, calibration, latency, sparsity, reproducibility, and behavior after data drift or retraining.
1. Objective, capacity, and evidence
A common objective is empirical loss plus a strength parameter times a complexity penalty. The penalty encodes a preference, not a factual claim that small weights are always correct. Changing feature units changes coefficient magnitude and therefore the effective penalty, so scale-sensitive models need fold-fitted transformations. Sample/class weighting and whether loss is summed or averaged can also alter relative penalty strength. Read the library's exact objective: for example, some APIs expose alpha as direct strength while logistic regression often exposes C as inverse strength.
Regularization is selected as part of the entire pipeline. A validation curve across log-spaced strengths shows training and validation behavior; cross-validation respects time, groups, or other deployment boundaries. The final test evaluates the locked preprocessing, penalty, optimizer, threshold, and calibration once. Strong regularization cannot make a leaked split valid or create absent information.
2. L2, L1, and elastic net
L2 penalizes the sum of squared coefficients. Its smooth shrinkage discourages reliance on extreme weights and often distributes weight across correlated predictors. It generally does not create exact zeros, so it is not automatic feature selection. L1 penalizes absolute coefficients and can yield sparse solutions. Sparsity may reduce serving cost or aid inspection, but a nonzero coefficient is not causal importance and correlated features can enter or leave arbitrarily across resamples.
Elastic net mixes L1 and L2. It can retain sparsity while improving stability for correlated groups, with both total strength and mixing ratio tuned. The intercept is usually treated separately, but library behavior must be checked. Categorical one-hot groups, interaction expansions, and polynomial features may need group-aware reasoning because independent coefficient penalties do not necessarily preserve semantic units.
3. Other effective regularizers
Early stopping limits optimization time using validation checkpoints and a declared metric, cadence, patience, and maximum budget. A noisy validation curve, bad learning rate, or convergence failure must not be mislabeled beneficial regularization. Dropout stochastically masks activations during training and requires inference-mode scaling/behavior. Data augmentation encodes target-preserving invariances; invalid transformations introduce label error. Weight decay can coincide with L2 in simple optimizers but is not universally identical under adaptive optimization.
Tree depth, minimum leaf support, pruning, feature/row subsampling, smoothing priors, low-rank bottlenecks, parameter sharing, and Bayesian priors all constrain effective fits differently. Choose the mechanism based on the failure hypothesis and deployment constraints, then verify with controlled ablations instead of stacking every regularizer at default strength.
4. Diagnostics and operations
Under weak regularization, training performance can be excellent while deployment-like validation is unstable or worse. Under excessive strength, both training and validation underperform and coefficients/predictions collapse toward a simple baseline. Learning curves, validation curves, error slices, calibration, coefficient paths, convergence diagnostics, and stability across folds/retrains distinguish these patterns from leakage or shift.
Track selected hyperparameters with data, feature, library, solver, and random-state versions. Refit the chosen procedure, verify convergence, and compare candidate/incumbent under identical evidence. In production, monitor prediction and calibration changes, feature scaling/skew, coefficient or selected-feature stability, subgroup outcomes, latency, and fallbacks. Retuning after drift is a governed experiment; stronger penalty is not an automatic response to every degradation.
Regularisation in an interview answer
The distinction interviewers probe most often is between the two standard penalties and what each one is for. An L2 penalty shrinks coefficients smoothly towards zero without reaching it, which stabilises estimates when predictors are correlated and keeps every feature in the model. An L1 penalty can drive coefficients to exactly zero, which performs selection as a side effect of fitting and produces a sparse model that is easier to serve and to explain. Elastic net combines them because correlated groups of features behave badly under a pure L1 penalty, where the choice among near-identical predictors becomes arbitrary and unstable across resamples.
The second thing being tested is whether the candidate knows that regularisation is scale-dependent. Because the penalty applies to the coefficients, a feature measured in small units receives a larger coefficient and is therefore penalised more heavily, for reasons that have nothing to do with its importance. Standardising features is not a preprocessing nicety here but a precondition for the penalty to mean anything, and the standardisation has to be fitted inside the cross-validation fold rather than on the full dataset, or the reported error is optimistic.
The third is the boundary against inference. Regularised coefficients are deliberately biased, and their ordinary standard errors and p-values do not apply. A regularised model is the right tool when the goal is prediction on new data; when the goal is estimating a specific effect and reporting its uncertainty, the honest path is an unpenalised specification with a defensible set of controls, or a method designed for post-selection inference.
Worked example: one residual, L2 versus L1
A one-feature, no-intercept fit with a single residual term makes the algebra checkable. Observation (x, y) = (1, 3). Ordinary least squares is β = 3. The two penalized objectives below are written out, not named.
L2. Minimize (3 − β)² + λ β². Differentiating: -2(3 − β) + 2λβ = 0, so β = 3 / (1 + λ).
L1, β ≥ 0. Minimize (3 − β)² + λ |β|. Differentiating: -2(3 − β) + λ = 0, so β = 3 − λ/2, then clipped at zero. The coefficient hits zero at λ = 6, not asymptotically.
| λ | OLS | L2 β = 3/(1+λ) | L1 β = max(3 − λ/2, 0) |
|---|---|---|---|
| 0 | 3 | 3 | 3 |
| 1 | 3 | 1.5 | 2.5 |
| 2 | 3 | 1 | 2 |
| 5 | 3 | 0.5 | 0.5 |
| 6 | 3 | 0.429 | 0 |
| 8 | 3 | 0.333 | 0 |
L2 at λ = 8 is still 0.333; L1 is already exactly 0. That is the sparsity difference, not a slogan. A library that exposes C as inverse strength is the other direction: larger C is the λ = 0 column.
Scale is the same arithmetic. If this feature is stored in millions instead of units, the unpenalized coefficient becomes 3e-6 and L2 with λ = 1 barely moves it, while the units-scale coefficient 3 is halved. The penalty is on coefficient magnitude, so the unit choice is the penalty.
