Overview
Curated: · Written: · Reviewed:
Key takeaways
- A feature is valid only if its value, timestamp, provenance, and computation are available at the prediction decision—not merely present in a historical table.
- Fit learned transformations and feature selection inside each training fold. All-data imputation, scaling, vocabularies, target encoding, selection, or PCA leaks validation information.
- Transformations encode assumptions. Scaling changes distance and penalty geometry; logarithms require domain and sign handling; buckets trade resolution for robustness; missingness can be signal or pipeline failure.
- Categorical encoders need explicit unknown-value behavior, vocabulary/version ownership, and high-cardinality control. Integer IDs are not automatically ordered quantities.
- Feature importance is model- and data-dependent, not causal truth. Correlated features can split or mask importance, and selection performed repeatedly can overfit validation.
- Use one versioned transformation contract for training and serving, with schema validation, point-in-time joins, lineage, offline/online parity tests, and drift monitoring.
1. Begin at prediction time
Define the entity, event, prediction timestamp, label horizon, and action before deriving features. For each candidate feature, record its event time, ingestion/availability time, source, owner, freshness, late-arrival behavior, and serving path. A 30-day purchase count for a decision at noon must use only events knowable by noon and reproduce historical availability, not today's corrected warehouse. Post-outcome fields, final case status, future aggregates, and data produced by the intervention are leakage even if correlated.
Point-in-time correctness also applies across entities and repeated events. Joins must select the latest valid feature version as of each prediction, respect slowly changing dimensions, and avoid backfilled records that were unavailable then. Split before inspecting outcome-driven transformations, keep related entities together where needed, and test with synthetic boundary timestamps.
What interviewers probe here. The classic probe is a feature that looks great and is impossible: "account_closed_reason predicts churn with 0.27 gain—keep it?" They then push on the subtler variant: an aggregate like lifetime_value whose computation window decides whether it leaks. Expect follow-ups on late-arriving data ("the event happened at 11:58 but landed in the warehouse at 14:00—can you use it?") and on joins ("the user's plan changed yesterday; which version does the 30-day-old prediction see?").
What a weak answer sounds like. "We dropped columns with high correlation to the target"—a checklist applied without asking when each value was knowable. Or conceding the boundary case immediately: a candidate who can't say why lifetime_value might leak has not internalized that availability is a property of the computation, not the column name.
2. Numerical and missing data
Inspect range, units, precision, missingness, zero/infinity, outliers, seasonality, and collection changes. Standardization learns training mean and variance; it helps distance-, gradient-, and regularization-sensitive models, but is often unnecessary for tree splits and is sensitive to outliers. Robust scaling, clipping, log-like transforms, or bins may improve stability when domain-valid. Never choose a transform solely because it makes a histogram attractive; validate the downstream metric, calibration, cohorts, and serving behavior.
Missingness has a mechanism. Absence may mean not applicable, not measured, delayed, redacted, device failure, or a meaningful user choice. Imputation statistics are learned inside folds and paired with a missing indicator when useful. A zero fill is safe only when zero has the intended semantics. Monitor missing rate by source and cohort; a model should not normalize a broken upstream feed into silent predictions.
What interviewers probe here. "You impute median age—computed on what rows?" is the standard trap; the answer must be the training fold, inside the pipeline. Follow-ups: "missingness jumped from 2% to 18% in production—what does your model do, and how do you know?" and "when is a missing indicator more valuable than the imputed value?"
What a weak answer sounds like. "We handled missing values" with no statement of mechanism, fold discipline, or monitoring. Also weak: treating every gap as noise to fill, when the absence itself (a user never opening the billing page) may be the signal.
3. Categorical, text, time, and interactions
One-hot encoding avoids false ordinal distance for low/moderate cardinality and needs an explicit unknown bucket or supported ignore behavior. Ordinal encoding is appropriate only when order is real or the downstream model treats codes categorically. Hashing bounds vocabulary memory but introduces collisions. Frequency or target encoding can handle high cardinality, but target statistics must be computed out of fold with smoothing and deployed with cold-start behavior; otherwise they leak outcomes.
Feature crosses expose interactions to simpler models, such as product(category, region), but cardinalities multiply and rare combinations overfit. Use domain-motivated crosses, hashing or minimum-frequency handling, and ablation evidence. Time features should encode calendar, elapsed time, recency, and cyclic structure only when available and relevant. Text vocabularies, tokenization, and embeddings are versioned transformers with privacy, language, out-of-vocabulary, and latency considerations.
What interviewers probe here. "A category appears in serving that never appeared in training—walk me through every encoder you'd consider and what each does." Then: "who owns the vocabulary, and what happens when it's rebuilt?" Target encoding is the escalation: a candidate who computes target means on the full dataset before the split has failed the question, and interviewers ask it precisely because the code looks innocent.
What a weak answer sounds like. "Label encoding, it works fine"—with no unknown handling, no versioning, and no awareness that integer codes impose a distance on unordered categories.
4. Feature selection and interpretation
Filter methods score features independently and can miss interactions. Wrapper methods repeatedly fit models and are computationally expensive and selection-prone. Embedded methods use model penalties or split behavior, inheriting that model's assumptions. Low variance does not mean low predictive value, and high univariate association can be leakage or a proxy. Put selector and chosen feature count inside cross-validation and compare with a no-selection baseline under the same pipeline.
Coefficient magnitude depends on scale and collinearity. Tree impurity importance can favor continuous or high-cardinality features. Permutation importance measures held-out performance change when a column is disrupted, but correlated substitutes can hide one another and unrealistic permutation can break feature relationships. These are predictive diagnostics, not causal effects or fairness approval. Use domain review, ablation, conditional/grouped variants where appropriate, stability across folds/time, and explicit sensitive/proxy analysis.
What interviewers probe here. "Your top feature by gain is also the one your business stakeholder says is downstream of the outcome—what now?" And the repeated-selection trap: "you ran forward selection 200 times against the validation set and report the best run—what is that number?" (It is an optimistically biased estimate of nothing you will see in production.) Expect a follow-up on correlated pairs: two features carrying the same signal, each showing modest importance—candidates should reach for grouped/conditional permutation or ablation, not conclude both are weak.
What a weak answer sounds like. Reading an importance ranking as a causal story, or defending a validation score obtained after the selection process itself consumed the validation data.
5. Production feature systems
Package parsing, validation, imputation, encoding, selection, and model as one versioned graph. Prefer shared transformation code or a feature service with point-in-time offline materialization and low-latency online lookup. Store feature definitions, owners, units, lineage, freshness, default/error policy, privacy classification, and training/serving versions. Contract tests run identical fixtures through offline and online paths and compare values, ordering, types, unknown categories, timestamps, and rounding.
Monitor raw and transformed schemas, availability, freshness, missingness, ranges, category vocabulary, hash/collision load, encoding sparsity, feature distributions, out-of-vocabulary rate, transformation failures, and training-serving skew. Join these with scores, outcomes, cohorts, and model versions. Feature retirement is a migration: prove consumers, remove from training and serving coherently, retain reproducibility, and watch fallback behavior. More features are not free—they add acquisition, privacy, latency, reliability, and governance surface.
What interviewers probe here. "Training AUC is 0.83, live AUC is 0.74—give me your ordered hypothesis list." Strong candidates start with train/serve skew (different transformation code, different join semantics, different freshness) before blaming drift. Follow-ups: "how do you prove the offline and online paths compute the same feature?" and "a feature's definition needs to change—what breaks?"
What a weak answer sounds like. "We'd retrain" as the first response to a skew incident, or a monitoring plan that watches model scores but never the raw and transformed feature distributions feeding them.
6. Worked example: a leaky feature that looks like the best one
A churn model reports 0.97 AUC on a random split and 0.71 on a time-blocked one. Ranked importances explain the gap in one line:
| Feature | Gain (random split) | Gain (time-blocked) | Available at prediction time? |
|---|---|---|---|
| days_since_last_login | 0.31 | 0.29 | yes |
| support_tickets_90d | 0.18 | 0.21 | yes |
| account_closed_reason | 0.27 | 0.00 | no, written at churn |
| plan_tier | 0.09 | 0.14 | yes |
| lifetime_value | 0.15 | 0.36 | only if computed from data before the cutoff |
account_closed_reason carries 0.27 of the gain on a random split and exactly 0.00 on a time-blocked one, because on the random split some churned rows leak their own outcome. It is not a weak feature to be tuned; it cannot exist at prediction time and has to be dropped.
lifetime_value is the subtler case. It survives both splits and its gain rises to 0.36, which looks like the best remaining signal. Whether it is depends entirely on how it is computed: summed over all history it includes revenue recorded after the prediction point, and the leak hides inside a legitimate-looking aggregate rather than in an obviously post-hoc column.
The same discipline applies to fitting. A scaler or encoder fitted before the split has already seen the validation rows:
# Wrong: statistics computed over train+validation, then "validated" on rows that shaped them.
X = StandardScaler().fit_transform(X)
scores = cross_val_score(model, X, y, cv=5)
# Right: every fitted step lives inside the fold.
pipe = Pipeline([("scale", StandardScaler()), ("model", model)])
scores = cross_val_score(pipe, X, y, cv=TimeSeriesSplit(n_splits=5))
The first form typically inflates cross-validated AUC by 0.01-0.03 on tabular data — small enough to pass review, large enough to reverse a model comparison.
7. Likely follow-ups and how to rehearse them
Interviewers rarely stop at the first correct answer. The follow-up chain for this topic is predictable:
- From leakage to availability: after you name a leak, expect "how would you have caught it in production, not in review?" The strong answer is point-in-time replay tests and monitoring feature availability against the prediction timestamp, not a manual audit.
- From encoding to ownership: after you pick an encoder, expect "what happens when the vocabulary changes and the model doesn't?" Version the vocabulary with the model artifact and test unknown-category behavior in contract tests.
- From importance to causality: after you present importances, expect "so if we increase days_since_last_login, churn drops?" The answer is no—importance is predictive under the observed distribution, and the correct tool for the causal question is an experiment or a carefully argued identification strategy.
- From pipeline to incident: after you describe the training pipeline, expect a skew debugging scenario. Rehearse a concrete ordering: compare offline/online feature values on the same entity, check transformation versions, check join and freshness semantics, then look at data drift.
A weak performance on follow-ups usually means the candidate memorized rules ("fit inside the fold") without the reasoning that generated them. The reasoning is always the same question, asked one level deeper: what did this number know, and when did it know it?
