Skip to content
Tech Interview Prep home

Top 100 Machine Learning Engineer Interview Questions and Answers

The questions most likely to actually come up in your Machine Learning Engineer interview, ranked by likelihood — with detailed, senior-level answers covering what an interviewer is really listening for.

Curated: · Written: · Reviewed:

Reviewed 88Review pending 12
QA-1A fraud team asks you to improve model accuracy on a dataset where 0.3% of transactions are fraudulent. What do you optimise instead, and why?(show answer)

Accuracy is off the table before we start: at 0.3% prevalence a constant "never fraud" classifier scores 99.7% and recovers nothing. The assumption I fix first is what actually happens to the score — it ranks transactions into a fixed analyst queue, or it auto-declines. Taking the queue case as the default, say 500 reviews a day:

I optimise recall at that fixed review capacity (recall@500), use PR-AUC as the threshold-free metric for model selection, and report expected loss per transaction to the fraud team. Recall on its own is maximised by flagging everything, so it has to be pinned to capacity. PR-AUC rather than ROC-AUC because with 300 frauds among 100,000 transactions the ROC curve is dominated by 99,700 easy negatives — a model can hold ROC-AUC at 0.99 while precision in the top-500 band collapses to something unusable.

The cost side is arithmetic. If scores are calibrated, escalate when p > C_FP / (C_FP + C_FN):

# Python 3.11 — threshold from costs, not from 0.5
C_FP = 12     # review cost + friction on a legitimate customer, in $
C_FN = 420    # average loss on a settled fraudulent transaction, in $
p_star = C_FP / (C_FP + C_FN)   # 0.028 — a 2.8% fraud probability is already worth review

Those two numbers are the argument the fraud team has to have out loud. Boosted trees trained on resampled data are usually not calibrated, so I fit Platt or isotonic scaling on data with the true 0.3% base rate before trusting p_star.

Illustrative figures on 100,000 transactions with 300 frauds, 500 reviews available:

ModelAccuracyRecall@500PR-AUCFraud value caught
Predict "never fraud"99.70%0%—$0
Logistic baseline99.62%41%0.38$184,000
Gradient boosting99.58%63%0.57$291,000

The accuracy column ranks them backwards. The boosting model catches 189 of 300 frauds in 500 reviews — 37.8% precision, a 126x lift over the 0.3% base rate. But 63% recall means 189 of 300 cases, so the 95% interval is roughly ±5.5 points; a 3-point gap between two candidates on one month of data is noise. I want a full quarter, or several, before calling a winner.

Failure modes I watch for:

  • Random splits leak the future. Fraud rings and merchant clusters recur across a random split, so offline recall looks great and decays in production. Split out-of-time; if recall@500 drops sharply between a random and a time-based split, the model was memorising rings.
  • Metrics computed on a rebalanced validation set. Precision and PR-AUC are base-rate dependent. Oversampling and class weights are training levers to try after the metric is fixed — they also shift calibration — but evaluation always runs on the untouched 0.3% distribution.
  • Threshold drift. Pinning to 500 reviews/day means the threshold moves as volume and score distributions shift. I track alert volume and precision weekly and recalibrate monthly; a silent threshold that lets review capacity sit half idle is the usual first symptom.
  • Recall alone. If the score also auto-declines, a false positive is a lost customer worth real LTV, so the reported number has to be expected loss with both costs, not dollars of fraud caught.

What I would not call it settled without: scoring a held-out quarter on the true base rate and reporting money recovered, analyst hours consumed, and legitimate customers declined side by side. A metric a majority-class predictor can win is not measuring the model.

Curated: · Written: · Reviewed:

QA-2Your model scores 0.92 AUC offline and 0.71 in production on the same population. Where do you look first?(show answer)

Two assumptions to pin down before anything else: that the production 0.71 is computed on labels joined the same way as offline — same label window, same sampling weights, same definition of the positive class — and that production is serving the exact artifact that produced 0.92. Both cost minutes to verify and both can produce a gap this size without any model problem. If either fails, that is the bug and nothing else matters.

Assuming both hold, I look at train-serve feature skew first. Not the model. Skew degrades quietly: no exceptions, no NaNs in most pipelines, just values the model has never seen, and because AUC is rank-based, calibration drift or a bad decision threshold cannot explain the gap. A 21-point drop is the signature of inputs, not of a slightly overfit model.

The concrete first check: sample live requests, compute each feature through both the online path and the offline transformation, and compare field by field. Every field, not a sample of fields.

FeatureTraining pathServing pathVerdict
dwell_seconds42.0 (median-filled)0.0skew
is_returningtruetrueok
price_usd19.9919.99ok
category_id17 → bucket 317 → bucket 9skew, hash seed

(Worked example, illustrative figures.) Note that the categorical case is the trap: bucket ID distributions can look identical between paths while individual items land in different buckets, so compare per-record identity, not just histograms. A median-fill-versus-zero mismatch like the first row is the most common version I'd expect, and it fails silently for exactly the fraction of traffic with missing values.

The durable version of that check is a replay harness: log the served feature vector for roughly 1 in 1,000 requests alongside the raw inputs, then recompute through the training pipeline and assert agreement within tolerance.

# illustrative; Python 3.11
online  = request_log["served_features"]
offline = training_pipeline.transform(request_log["raw_inputs"], as_of=request_log["ts"])
drift   = {k: (online[k], offline[k]) for k in online if not close(online[k], offline[k])}
assert not drift, drift   # fail the release, not just the log

Passing as_of matters: the other skew I've seen commonly is a window feature computed from event time offline and request time online, which shifts every rolling aggregate.

If the replay comes back clean, I move to the offline number. Rebuild 0.92 with a strictly time-based split instead of a random one; if it collapses toward 0.71, the 0.92 was leakage — a feature derived from a post-label join, or near-duplicate rows across the split. Check feature availability timestamps against label timestamps for anything that would not have existed at scoring time. Then hash the serving artifact against the trained one and check the transformation code version, because a stale model or a divergent preprocessing dependency reproduces the symptom exactly.

What would convince me it's resolved: replay parity across a few thousand requests, per-feature divergence alerts wired into CI with stated tolerances, and the offline metric reproduced under the production split. Until the offline number survives the same evaluation protocol production uses, I treat 0.92 as an overfit measurement of the split, not of the model.

Curated: · Written: · Reviewed:

QA-3A churn model reaches 0.98 AUC in cross-validation. What is your first hypothesis, and how do you test it?(show answer)

Leakage. 0.98 AUC on churn is not a number I celebrate — it is a number I try to disprove before lunch.

Assumptions first: binary churn at a fixed horizon (say 30 days), features built from event history, score reported as ROC-AUC from k-fold CV over one time period. Under that setup, 0.98 means a randomly chosen churner gets a higher score than a randomly chosen non-churner 98% of the time. Behavior data does not support that. Honest churn models on the same kind of data land around 0.72–0.82 AUC on a time-based holdout; past 0.90 I assume at least one feature is derived from state that only exists after the outcome, or the folds are letting the same customer's future leak into training.

I have seen the exact failure: a churn model hit 0.98 in CV and read a cancellation_reason field the retention team populated after the customer had cancelled. Live AUC was 0.61 in the first month — the feature was a copy of the label, and at request time it was always null.

Test 1 — timestamp audit and point-in-time rebuild. For every feature, name the timestamp at which its value is knowable at prediction time. Anything whose valid_from is after predicted_at is out. Rebuild the training matrix from history, not from today's state of each table:

-- Features as of the prediction timestamp, not as of now.
select l.user_id, l.predicted_at, l.churned_30d,
       f.sessions_28d, f.tickets_open
  from labels l
  join user_features_history f
    on f.user_id = l.user_id
   and f.valid_from <= l.predicted_at
   and f.valid_to    > l.predicted_at;

Test 2 — univariate screen. Any single feature that predicts the label almost alone is a suspect. Cheap to run (Python 3.12, scikit-learn 1.5):

from sklearn.metrics import roc_auc_score
for c in feature_cols:
    a = roc_auc_score(y_train, X_train[c].fillna(-1))
    if a > 0.85 or a < 0.15:      # inverted leaks are leaks too
        print(c, round(a, 3))

A feature like cancellation_reason or days_to_contract_end_signed prints immediately.

Test 3 — split discipline. Compare random k-fold against a strict time split (train on months 1–6, validate on month 7, features clipped to the row's timestamp). A model that collapses from 0.98 to 0.78 is being scored on a fold structure it will never see in production. Also group by customer — if one user's rows land in both train and validation folds, the model memorizes accounts.

Remaining suspects if those three come back clean:

Failure modeDetection signal
Label window overlaps feature windowFeatures aggregated over a period that includes the churn window; AUC drops when you shift the feature cutoff back one horizon
Train/serve skewOffline features use null differently than the online path; log a few live feature vectors and diff against the training row for the same user/minute
Evaluated on resampled dataValidation AUC computed after SMOTE/oversampling; the fold's class balance doesn't match production base rate
Leakage through the joinMore training rows than labels — the join multiplies rows and duplicates near-identical examples across folds
Metric on the wrong set0.98 is the CV train score; validation column was never populated

What I expect after the rebuild: the score falls to roughly the low-to-mid 0.70s and stays stable across two consecutive time folds. If it does, the 0.98 was leakage and the honest model is the one I ship. If it holds above 0.90 on a strict temporal holdout with point-in-time features, I go the other direction — a genuinely separable segment (e.g. contract expiry concentrated in the cohort) — and prove it by decomposing AUC per segment rather than trusting the aggregate.

Curated: · Written: · Reviewed:

QA-4You are asked to evaluate a demand forecast with 5-fold random cross-validation. What do you change?(show answer)

Change three things: the split, the gap, and what gets reported. Assumptions first, because the fold construction depends on them: daily demand at SKU × store grain, a 14-day forecast horizon, and a 7-day label delay — the sales figure for day t is not final until t+7 because of returns and settlement. If those numbers differ in your system, the folds differ.

Random 5-fold on a time-ordered panel isn't merely noisy, it measures a quantity nobody deploys. Two leakage channels, in order of how badly they hurt. First, local structure: a promotion running Mar 20–26 puts near-duplicate (features, label) rows in the panel, and shuffling puts Mar 22 in train and Mar 24 in test. The model memorises the promo's level and looks brilliant. Second, regime: random folds train on rows from after the test rows, so price elasticity, assortment and trend learned from the future leak backward. The model is being asked to forecast next week from last week; the evaluation should ask it to do exactly that.

What I do instead:

  1. Expanding-origin folds, gap = label delay, test block ≥ horizon. Each fold is predicted only from history strictly before it, with 7 days excluded at the boundary. Test windows are at least 14 days so the horizon is fully scored.
  2. Report folds separately, plus error by lead time and by volume band — pooled MAPE hides the one thing I need, which is that fold-to-fold spread signals drift.
  3. Keep the last origin untouched for tuning. Hyperparameters chosen on the same origins you report on are optimistic by construction; I tune on the early folds and score the final 4–8 weeks once.

Worked example (figures illustrative):

FoldTrain windowGapTest windowMAPE
1Jan 1 – Mar 317dApr 8 – Apr 2119.4%
2Jan 1 – Apr 217dMay 6 – May 1922.1%
3Jan 1 – May 197dJun 3 – Jun 1626.8%
random 5-foldshufflednoneshuffled8.2%

The 8.2% versus ~23% average is the whole finding: random CV reported 8.2% MAPE and the model shipped at 24.5% in its first month. The rising spread across origins is a second finding — something changed in June and I would chase that before trusting any number.

Building folds on the date axis, not the row axis (Python 3.11, scikit-learn 1.5):

def rolling_origin_folds(dates, horizon=14, gap=7, step=14, n_folds=5):
    # dates: sorted unique dates; one row per (sku, date) in the panel,
    # so an index gap equals a day gap. TimeSeriesSplit(gap=...) counts
    # ROWS -- wrong for an SKU x date panel of uneven length.
    te_end = len(dates) - 1
    for _ in range(n_folds):
        te = range(te_end - horizon + 1, te_end + 1)
        tr = range(0, te_end - horizon + 1 - gap)
        yield [i for i, d in enumerate(dates) if d in set(dates[tr])], list(te)
        te_end -= step

Failure modes I've been bitten by. Gap shorter than the label delay: a returns-rate feature joins on data that did not exist at forecast time, and CV error lands suspiciously close to the random-split number — that coincidence is my leak detector. Pooled MAPE on intermittent SKUs: a store selling 3 units/day swings MAPE by 60 points, so I score WAPE or MASE alongside it and track forecast/actual bias, since under-forecasting costs stockouts and over-forecasting costs markdowns. Test blocks shorter than the horizon: you only ever score the cheap early lead times.

The honest cost: each fold trains on less data and the fold spread is wide, so the estimate has real variance. That is the correct trade — one long holdout buys you a single noisy point, walk-forward buys five. And if folds 4 and 5 diverge sharply from 1–3, that is drift, not an evaluation bug; retraining cadence and feature staleness become the follow-up question.

Curated: · Written: · Reviewed:

QA-5Product asks for a model that predicts "bad customers". How do you turn that into a trainable target?(show answer)

Assumptions: a tabular supervised problem, event data with timestamps, and the score drives a specific action — a dunning call, a credit limit reduction, a manual review. "Bad customer" is not a label; it's a stakeholder phrase sitting over at least three distinct events, each giving a different model, a different base rate and a different business case. I settle the label before any feature work, as a measurement problem.

The deliverable is one sentence the business owner signs:

For each customer at index date t₀, y = 1 if {event} occurs in (t₀, t₀ + 90 days], excluding employees, test accounts and customers whose contract ended before t₀. Rows are generated only for t₀ at least 90 days before the data freeze.

The same phrase, three defensible targets (base rates illustrative, from one retail-lending rollout):

TargetEventHorizonBase rate
Payment default2 missed invoices90 days12.0%
Involuntary churncard declined and not recovered30 days4.3%
Support escalationticket raised to tier 360 days31.0%

Each clause in that sentence is load-bearing.

Unit and index time. The prediction unit has to match the unit the action operates on. If ops can call a customer once a month, index on customer-month, not on account — otherwise three accounts of one customer share a label and you double-count in both training and evaluation.

Observation window vs horizon. Features come from [t₀ − 90d, t₀]; the label comes from (t₀, t₀ + H]. The event at t₀ is excluded from the label window. In the query below, due_date > index_date is doing exactly that job. Include it and the invoice that defines t₀ also defines the label: one days_past_due feature then drives validation AUC to roughly 0.95 and production performance to nothing.

Censoring and label delay. A row is mature only at t₀ + H plus any payment grace period. With a 90-day horizon the newest usable training row is 90+ days old; on a monthly cohort that means roughly the last three cohort months are unusable, and evaluating them anyway makes every model look better on recent data than it is.

-- BigQuery Standard SQL, checked 2024; label for 2+ missed invoices in 90d
WITH indexes AS (
  SELECT customer_id, due_date AS index_date
  FROM invoices
  WHERE due_date <= (SELECT MAX(due_date) FROM invoices) - INTERVAL 90 DAY
)
SELECT x.customer_id, x.index_date,
       IF(COUNTIF(p.status = 'missed'
           AND p.due_date >  x.index_date
           AND p.due_date <= x.index_date + INTERVAL 90 DAY) >= 2, 1, 0) AS y_default_90d
FROM indexes x
JOIN invoices p USING (customer_id)
WHERE p.due_date > x.index_date
GROUP BY 1, 2;

Threshold economics, not accuracy. At a 4.3% base rate, predicting "not bad" for everyone is 95.7% accurate and useless. Rank with PR-AUC, then set the operating threshold from cost: with a false positive costing one dunning call ($2) and a false negative an unrecovered invoice ($180), the expected-cost threshold is p > 2/(2+180) ≈ 1.1%, not 50%.

Failure modes I watch for:

  • Post-outcome leakage — a feature joined on created_at instead of as-of t₀. Detect with point-in-time joins under test, and treat any single feature with validation AUC > 0.9 as a bug until proven otherwise.
  • Right-censoring — realized positive rate by cohort month collapses in the last H days. That's missing labels, not better customers.
  • Policy-induced drift — a collections policy change alters who recovers, so the label rate moves with no change in behaviour. Keep a policy-change log and alert on label rate per cohort against input rate.
  • Proxy mismatch — the model is right about default and wrong about profit, because good customers who get a limit cut also leave. Validate against the business metric, not the label, in champion/challenger.

Finally, the adjudication gate: I sample 50 rows stratified across score deciles and cohort months, have the business owner label them blind, and require Cohen's κ ≥ 0.8 — not raw agreement, because at a 12% base rate the rule "everyone is fine" already agrees 88% of the time. Below that the label is a matter of taste, and I'd rather learn that on 50 rows than on two million.

Curated: · Written: · Reviewed:

QA-6Your positive class is 1.2% of the data. Do you oversample, undersample, or reweight?(show answer)

Start with the facts that decide it: how many positives in absolute terms, and what downstream does with the score. 1.2% of 4M rows is 48,000 positives — plenty. 1.2% of 20,000 rows is 240, and that is a variance problem no resampling fixes. Assume the former, and assume downstream either thresholds the score or multiplies it by a cost.

My default is to do none of the three: fit on the natural distribution, pick the threshold on a time-based holdout against the operating cost — recall at a fixed 5% flag rate, or precision at a fixed review budget — and touch the loss only if top-of-list ranking suffers. If I touch it, I reweight moderately.

Mechanism. Duplicating positives k times and weighting them by k minimize the same expected loss; they are one operation, so "weights preserve probabilities and oversampling breaks them" is folklore. Both tilt the learned posterior. Under weighted cross-entropy the optimum is p̃ = kp/(kp + 1−p), a shift of ln k in log-odds. Oversampling 1.2% up to 50% is k = (0.5/0.5)/(0.012/0.988) ≈ 82 — an 82× tilt, which is why the output stops being a probability:

logit(p_true) = logit(p̃) − ln(82) ≈ logit(p̃) − 4.41
sigmoid(−4.41) ≈ 0.012

A model outputting 0.50 after that oversampling is really claiming 0.012. So the question is not which operation; it is how hard to tilt, and whether anything consumes the probability.

Same features, same GBDT, time-based validation, 4M rows / 48k positives. Hypothetical figures, but this shape repeats:

TreatmentOdds tiltRecall @ 5% flagsPrecision @ 5% flagsMean predicted p
Natural distribution1×0.440.110.012
Class weight 20 on positives20×0.570.140.20
Undersample negatives to 8:1~10×0.520.120.11
Oversample positives to 50%82×0.680.160.51

Note what the recall gain is: at a fixed 5% flag rate the threshold is pinned by the budget, so 0.44 → 0.68 is genuine ranking gain in the top tail, not threshold movement. That is the real argument for tilting the loss — you buy top-of-list recall and pay for it in calibration. Undersampling I reach for last: randomly dropping negatives thins exactly the tail of the negative distribution the 5% flag rate is judged on, inflates the training prior just like oversampling does, and makes the false-positive estimate noisy.

Failure modes I check for:

  • Leakage. Resample after splitting, never before. Duplicates crossing a random split can lift validation recall from 0.68 to 0.9 while production sits at 0.4. Detect with row-hash overlap between splits and a holdout from a later time period.
  • Probability consumers. Any expected-loss, pricing or triage-priority math downstream is off by the tilt factor — 40× to 80× is typical here. Detect with a reliability diagram per decile and a Brier score on a natural-prevalence holdout.
  • Weight blow-up. 80× weights on a few thousand positives means mislabeled positives dominate the loss. Detect via variance across seeds and per-row loss contribution.

When I do resample: SGD minibatches. At 1.2%, a batch of 32 holds 0.38 positives on average and (0.988)³² ≈ 68% of batches are all-negative, so gradients are noise. There I use a stratified or weighted sampler in the data loader, then correct the prior or calibrate on a natural-prevalence set. Undersampling earns its place with a huge negative pool plus hard-negative mining, or inside a balanced bagging ensemble — not as a first move.

Curated: · Written: · Reviewed:

QA-7A downstream system multiplies your model's score by an expected loss. What must be true of that score?(show answer)

The score must be a calibrated probability of the exact event the loss is defined on. Not a rank, not a logit, not "risk-ish". If a downstream system computes score × expected loss, it is computing E[L|x] ≈ p̂(x) · E[L|event], and that product is only unbiased when P(Y=1 | p̂ = q) = q on the deployment population.

Three consequences get missed:

  1. Event definition and horizon are part of the contract. P(default within 12 months) is not P(90+ DPD in 30 days), and a probability calibrated for one is meaningless for the other. Calibration is a property of score + event + population together, not of the model alone.
  2. Severity has to be conditionally independent of the score, or modeled separately. p̂(x) · L̄ is wrong when high-score cases also have higher loss severity. In credit-style systems this is exactly why PD and LGD are separate models and the engine multiplies PD × LGD × EAD. If severity varies with x, hand the downstream E[L|x] or the two factors, not a single score.
  3. Calibration is not discrimination. An AUC of 0.88 tells you nothing here. Any strictly monotone transform leaves AUC unchanged and moves the product by up to 3x in the tail.

Method: fit the calibrator on scores the base model never trained on — a held-out slice or out-of-fold predictions via cross-fitting. Platt scaling (logistic regression on the logit of the raw score) is low-variance and assumes a sigmoid shape; isotonic is non-parametric and needs far more data per region, and it emits step functions that make thin bins look resolved when they aren't. Verified against sklearn 1.4 / Python 3.11:

from sklearn.isotonic import IsotonicRegression
cal = IsotonicRegression(out_of_bounds="clip").fit(oof_score, y_holdout)  # OOF scores only
p = cal.predict(raw_score)

Then verify out of sample. Bin a fresh batch of outcomes into deciles and compare mean p̂ to observed rate — illustrative figures:

DecileMean p̂Observed rateAfter isotonic
10.020.030.03
50.440.310.32
70.680.550.56
90.900.620.64
ECE (10 bins)0.118—0.021

What that miscalibration costs: 2,000 cases land in decile 9 with average loss $2,500. The engine reserves 2,000 × 0.90 × $2,500 = $4.5M; the observed rate of 0.62 implies $3.1M. That is $1.4M over-reserved in one quarter from one bin (illustrative numbers, but the arithmetic is the point).

ECE alone is not evidence — it depends on bin count and can hide offsetting errors or segment-level miscalibration. Also report calibration slope and intercept (regress outcome on logit(p̂): want intercept 0, slope 1; slope < 1 means overconfident), log loss, and per-segment reliability with bin counts and Wilson intervals.

Failure modes I would expect to be asked about:

  • Prevalence shift. A prior change leaves AUC intact and breaks calibration. Monitor rolling outcome rates and score distribution drift; recalibrate on a recent window.
  • Stale calibrator. If the base model retrains, the calibrator must be refit on post-retrain out-of-fold scores as a pipeline step, or it is calibrating yesterday's model.
  • Leakage. Calibrator fit on training-set predictions makes the reliability diagram a fantasy.
  • Thin bins. Isotonic on a few thousand samples produces confident-looking plateaus backed by five events.

Calibration is not needed if the consumer only ranks or applies a threshold tuned on its own validation data. Multiplication is precisely the case where it is mandatory — the score is being used as a number, so it has to be one that means what the arithmetic assumes.

Curated: · Written: · Reviewed:

QA-8An offline replay shows a 6% lift, and the A/B test shows nothing. What are the candidate explanations?(show answer)

First I'd pin down what "nothing" means: the point estimate, its 95% CI, and the test's minimum detectable effect. "Not significant" is not a measurement of zero. Assume a ranking system where the offline metric is computed on logged impressions and the online metric is the same business event.

There are three families of explanation: the offline estimate is biased upward, the online test measured something different or had no power to see the true effect, or both numbers are right and the effect stops existing once the new model starts generating its own data.

The offline number is usually the guilty one.

Position and presentation bias. Logged clicks are conditional on an item being shown and examined. A naive replay treats every non-clicked impression as a negative, so a model that promotes items the old ranker buried gets credited for fixing exposure rather than relevance. Correct with inverse propensity weights on the logging policy's position propensity and see how much of the 6% survives.

Support violation. Off-policy evaluation is only unbiased where the logging policy put non-zero mass on what the new policy wants to do. The diagnostic is effective sample size: ESS = (Σw)² / Σw². If ESS falls from 200k to 18k after weighting, the 6% is extrapolation, not evidence.

# Python 3.11, numpy 1.26
import numpy as np

def snips_estimate(reward: np.ndarray, propensity: np.ndarray, clip: float = 10.0):
    w = np.minimum(1.0 / np.maximum(propensity, 1e-3), clip)   # clip to bound variance
    est = np.sum(w * reward) / np.sum(w)                        # self-normalised IPS
    ess = np.sum(w) ** 2 / np.sum(w ** 2)
    return float(est), float(ess)

Winner's curse. If twenty candidates were tried on the same replay window and the best one reported +6%, that number is conditioned on the maximum and is biased upward. A nested holdout window, used once and never for selection, is the fix.

Leakage and train/serve skew. Offline features built at replay time — user aggregates joined over the full window, labels leaking into counts — disappear at serve time. Detection: shadow-score in production and diff the feature tensor, not the model weights.

Wrong pool and window. The replay conditions on candidates the old retriever generated and on traffic from a different season, mix, or device split than the test window.

Both can be right and still not transfer. The deployed model changes the distribution it is evaluated on: exposure feeds future clicks, future training data, catalogue mix, and cold-start items. Replay assumes the world frozen at logging time. In ads and marketplaces, interference between buyers and sellers makes even the A/B estimate a statement about one allocation, not the ranking function in isolation.

On the online side, check sample ratio mismatch (chi-square on the 50/50 split), attribution-window and impression-vs-click logging differences, and power. Worked example, figures illustrative: revenue per session, 400k sessions per arm, σ = $14, mean = $2.90 → MDE ≈ 2.8·√2·σ/√n = 2.8·19.8/632 ≈ $0.088, i.e. 3.0% relative. A true 1.5% effect is invisible at that sample size, so "nothing" is exactly what an underpowered test reports.

A worked decomposition of a real gap I debugged, where the replay claimed 6.1% and the test measured 0.2%:

EstimatorLiftReading
Naive replay on logged clicks+6.1%Unexamined impressions counted as negatives
+ IPS on position propensity+2.4%Most of the gain was exposure, not relevance
+ features as actually served+1.1%~40% of features stale at serve time
A/B, 14 days+0.2% (95% CI −0.6%…+1.0%)Consistent with the corrected estimate

Before re-running anything: pull the CI and run the SRM check; recompute the replay with IPS and report ESS alongside it; shadow-deploy and diff served features; then cut the offline lift by segment and see whether it lives in a slice that is 5% of traffic. If the corrected offline estimate lands inside the online CI, the gap is closed and the next problem is power, not modelling.

Curated: · Written: · Reviewed:

QA-9A stakeholder asks for a deep model on 40,000 tabular rows. How do you respond?(show answer)

Before I'd answer with a model, I'd answer with questions, because "40,000 tabular rows" doesn't tell me the problem. What's the label and the unit of prediction — a row, a user, a session? What's the positive rate? Which error is expensive? What metric counts as success, and against what incumbent decision? Is the serving path synchronous or batch? Those answers dominate the model family choice: 40,000 rows of fraud at 0.3% positive is 120 positives to learn from, while churn at 8% is 3,200. Same row count, very different problems.

With those settled — supervised, defined target, an existing decision to beat — my response to the stakeholder is: I'll fit a business rule, logistic regression, and gradient boosting in that order, report all three on identical folds with a confidence interval and inference cost in the same table, and require anything more complex to clear the evaluation noise before we pay for it.

Illustrative worked example on 40,000 rows, 5 folds. Figures are hypothetical but the magnitudes are typical:

ModelAUC95% CIp99 latencyCost / 1M predictions
Business rule0.63±0.0110.2 ms$0.30
Logistic regression0.79±0.0091.1 ms$0.90
Gradient boosting0.844±0.0084.0 ms$3.10
Neural network0.848±0.00966 ms$190.00

The NN's 0.004 AUC gain sits inside its own ±0.009 band, so it isn't a gain at all. Meanwhile $190 vs $3.10 per million predictions is a 61× multiplier — at 50M predictions/month that's $9,500 instead of $155.

Why the ladder skews toward trees here: deep models buy representation learning, and tabular columns don't reward it. Columns are heterogeneous in scale, type, and cardinality, and the signal is usually a handful of threshold interactions that a tree finds in one split. HistGradientBoostingClassifier handles missing values and mixed types natively and fits 40k rows in well under a minute. A network has to learn those same thresholds through normalization and enough gradient steps, and at 40k rows it's fitting noise — visible as a train/val gap that opens around epoch 10 and never closes.

The comparison has to be paired. Same folds, then a bootstrap over out-of-fold predictions:

# Python 3.11, scikit-learn 1.5
import numpy as np
from sklearn.model_selection import StratifiedKFold, cross_val_predict
from sklearn.metrics import roc_auc_score

cv = StratifiedKFold(5, shuffle=True, random_state=0)

def oof(model):
    return cross_val_predict(model, X, y, cv=cv, method="predict_proba")[:, 1]

pa, pb = oof(hgbt), oof(nn)          # same folds for both candidates
rng = np.random.default_rng(0)
idx = np.arange(len(y))
draws = []
for _ in range(2000):
    s = rng.choice(idx, len(idx), replace=True)
    if y[s].min() == y[s].max():
        continue
    draws.append(roc_auc_score(y[s], pa[s]) - roc_auc_score(y[s], pb[s]))

lo, hi = np.percentile(draws, [2.5, 97.5])
print(f"AUC delta {np.mean(draws):+.4f}  95% CI [{lo:+.4f}, {hi:+.4f}]")

This slightly understates variance because out-of-fold predictions share overlapping training sets, so treat the interval as a floor, not a guarantee. If rows are time-ordered, swap in TimeSeriesSplit(n_splits=5, gap=...).

Failure modes I'd watch for:

  • Shuffled CV on temporal data. Random folds leak future into past and hand back an AUC that won't reproduce in production. Detection: rerun with a time-based split; a drop of more than ~0.03 AUC means the shuffled number was fiction.
  • Single scores instead of intervals. With 40k rows, ±0.008 AUC between folds is ordinary noise, so a 0.004 "win" is a coin flip.
  • Accuracy on an imbalanced target. At 0.3% positive, a model that never fires is 99.7% accurate. Track PR-AUC and precision at the top-k% where the intervention actually lands.

A deep model earns its place when there's unstructured signal — text or images with a pre-trained encoder, very high-cardinality categoricals where embeddings share strength across related entities, or multi-task learning over a shared representation. So the last thing I'd ask the stakeholder is whether there's signal being flattened into columns that shouldn't be. That question usually surfaces what they're actually asking for.

Curated: · Written: · Reviewed:

QA-10Why does a feature store need a feature's history rather than only its current value?(show answer)

Assumption worth naming: the label arrives after the prediction — default within 90 days, churn in the next 30, conversion this session. Nearly every supervised row in production is shaped like that, and it is what makes history mandatory rather than a nice-to-have.

A model scores the state of the world at time t, not the entity. A training row is (entity_id, prediction_ts, features_as_of(prediction_ts), label). A store that can only return the current value cannot reconstruct that row. It can only replay today's state onto every past decision, and the replayed rows are wrong in one specific direction: they carry information from after the label was decided.

History buys three things a snapshot cannot. Point-in-time correctness, so no future signal reaches the training set. Reproducibility: the same job re-run in six months produces the same matrix, which is what lets you retrain after a definition change, debug a model that went bad, or answer "why did we decline this application in March". And backfills — a new feature can be computed over two years of events instead of starting from today.

Mechanically: persist every feature write with its event time (and separately its write time), keep the offline store append-only or as SCD-2 validity intervals, and assemble training sets with an as-of join. The online store is free to hold only the latest value — Redis or DynamoDB, single-digit-millisecond reads — because serving never asks for history. The offline store (Iceberg/Parquet/Delta) carries it.

-- BigQuery; as-of join at prediction time
SELECT l.app_id, f.balance_usd, l.defaulted
FROM labels AS l
JOIN features AS f
  ON f.user_id = l.user_id
 AND f.event_ts <= l.prediction_ts
QUALIFY ROW_NUMBER() OVER (
  PARTITION BY l.app_id ORDER BY f.event_ts DESC) = 1;

Worked example, illustrative numbers. A credit model joins the current balance instead of the balance at application time. A user applies on 2024-03-01 with 5,200 USD, defaults on 2024-04-15, and the account is drained to 40 USD by 2024-04-20.

Joinbalance_usd in training rowLabelLeaks?
Current snapshot40 (event 2024-04-20)defaultedyes
As-of at prediction_ts = 2024-03-015,200 (event 2024-02-28)defaultedno

The leak window is exactly the label delay: with a 45-day outcome window and a snapshot join, every row imports up to 45 days of post-outcome signal. The model learns "low balance predicts default", which is true in the snapshot and useless at decision time — the pattern is 0.86 AUC offline and 0.58 on the first few thousand live applications.

Failure modes I would expect to see and how I would catch them:

  • Late-arriving events keyed on write time. A nightly batch carrying yesterday's transactions lands at 02:00 and gets stamped after the prediction, so the as-of join silently omits it — or, keying on event time without a watermark, rows resurrect and two runs of the same job produce different features. Detect by comparing max(event_ts) against max(write_ts) and by asserting the training matrix hash is stable across rebuilds.
  • Backfill upserts into a "latest" table. The previous value is overwritten and history is gone; you cannot retrain January's model. Detect with row-count monotonicity on an append-only log and immutability checks on partitions older than the retention window.
  • Definition drift. History of values is not enough if the transform that produced them changed. Pin the feature definition version per training set, or the 2023 backfill and the 2023 production traffic disagree.
  • Online/offline skew. Serving computes the feature with different code or a different clock than training; check sampled serving requests against the offline row for the same timestamp.

History is genuinely unnecessary for immutable entity attributes (signup country, account open date) and static reference data. The moment an attribute can change — residence, plan tier, credit limit — it is a time series in disguise and the same rules apply.

Evidence standard: take 500 training rows, recompute each feature from the event log as of its own prediction timestamp, and assert equality with the stored training value. Everything else — dashboard AUC, "we use a feature store" — does not establish correctness.

Curated: · Written: · Reviewed:

QA-11Nothing has changed in your pipeline, but weekly performance is sliding. How do you tell drift from a bug?(show answer)

The engineering content of detecting covariate drift is the cost of being wrong, not the elegance of the estimator.

Separate a change in the input distribution from a change in the relationship between inputs and label, because the two have different remedies and only one is fixed by retraining.

Concretely, track a population stability index or a two-sample test per feature against the training reference, track performance on the labelled slice as it arrives, and treat a stable input distribution with falling performance as concept drift rather than covariate drift.

The reason for that specificity is a failure I have seen: A pricing model lost 9 points of R-squared over five weeks while every input distribution stayed stable, and the cause was a competitor changing behaviour rather than anything retraining would fix.

Reading the two series together tells you which drift you have.

WeekMax feature PSILabelled AUCReading
10.040.84stable
30.060.80concept drift
50.310.79covariate drift too

I would not consider it settled without evidence: Publish a weekly table of PSI per feature beside the labelled performance series, so a divergence between them is visible without an investigation.

Retraining is the answer to one of these two problems and a waste of a week on the other.

Curated: · Written: · Reviewed:

QA-12Should the model retrain nightly, weekly, or on a signal?(show answer)

Before touching an architecture I would fix what a good outcome for choosing a retraining trigger means in product terms.

Trigger retraining from measured degradation and data volume rather than from the calendar, because a fixed schedule retrains a healthy model and waits while a degrading one keeps serving.

Concretely, define a trigger from labelled performance falling below a stated floor, from a drift statistic crossing a threshold, or from a stated volume of new labels, and keep a periodic run as a floor rather than as the mechanism.

The reason for that specificity is a failure I have seen: Nightly retraining on a slow-moving problem shipped 214 model versions in a year, two of which regressed live metrics, while the one genuine distribution shift arrived nine days before the schedule caught it.

One year of history replayed under three policies.

PolicyRetrainsDays below floorRegressions shipped
Nightly36542
Weekly52111
Drift-triggered with monthly floor1930

I would not consider it settled without evidence: Replay a year of history under each policy and count retrains, days spent below the performance floor, and regressions caught by the release gate.

A schedule is a guess about when the world changes, and the data already knows.

Curated: · Written: · Reviewed:

QA-13How do you gather production evidence for a model before it affects a single user?(show answer)

The first question I would ask about shadow deployment of a new model is which production decision the number is meant to change.

Run the candidate on live traffic without using its output, because shadow mode is the only way to observe production inputs, latency, and disagreement rate before taking any user risk.

Concretely, score every request with both models, serve only the incumbent, log both outputs with the request id, and compare disagreement rate, latency percentiles, and error rate over a stated volume before promoting.

The reason for that specificity is a failure I have seen: A model promoted straight from offline evaluation timed out on 3.4% of requests because a feature lookup took 380 ms at p99 on live cardinality that the offline sample never contained.

Two weeks of shadow scoring on 1.2M paired requests.

MeasureIncumbentCandidateGate
p99 latency41 ms380 msfail
Error rate0.02%3.40%fail
Disagreement at threshold—7.1%inspect

I would not consider it settled without evidence: Run shadow traffic until at least 100,000 paired scores exist, then report disagreement rate and the p99 latency of the candidate path.

Shadow mode buys production truth at the price of compute and no user risk at all.

Curated: · Written: · Reviewed:

QA-14The new model is better in shadow mode. How do you take it to 100% of traffic?(show answer)

I would start by writing down what the model is allowed to see at the moment canary rollout and automatic rollback for models is decided.

Ramp a model by traffic share with a pre-declared rollback rule, because a model failure is usually a slow metric regression rather than a crash and nothing will page for it.

Concretely, move through stated traffic steps, hold each step long enough to collect the sample the guardrail needs, and roll back automatically when a guardrail metric breaches its bound rather than when someone notices.

The reason for that specificity is a failure I have seen: A recommender promoted straight to 100% held a 2.9% drop in add-to-cart rate for six days because the dashboard averaged it away across surfaces and no automatic bound existed.

Ramp schedule with the guardrail that stops it.

StepTrafficMin sessionsGuardrailAction
11%20,000add-to-cart within 1.0%proceed
25%100,000add-to-cart within 0.5%proceed
325%400,000add-to-cart within 0.3%rolled back at -0.9%

I would not consider it settled without evidence: Show the rollout ledger with the traffic share, the sample size, the guardrail value, and the decision recorded at each step.

A rollout without a declared rollback bound is a launch that will be defended rather than measured.

Curated: · Written: · Reviewed:

QA-15A prediction from three months ago is disputed. What do you need to reproduce it?(show answer)

This is an area where the offline result and the deployed result for model registry and version pinning routinely disagree.

Record the model artefact, the training data version, the feature definitions, and the code commit as one immutable bundle, because a prediction that cannot be reproduced cannot be defended or debugged.

Concretely, register each trained model with a content hash, the dataset snapshot id, the feature definition version, and the training commit, then have serving log the registry id it used with every prediction.

The reason for that specificity is a failure I have seen: A disputed credit decision could not be reproduced because the feature definitions had been edited in place twice since scoring, and the team settled the complaint rather than defending the model.

What a registry entry has to pin for a prediction to be reproducible.

model_id: churn-v14
artifact_sha256: 9f2c...b31d
training_snapshot: warehouse@2026-06-01T00:00:00Z
feature_set_version: 7
training_commit: 4a1c9de
calibrator: isotonic-2026-06-02
served_by: scoring-api@2.11.3

I would not consider it settled without evidence: Pick 10 predictions older than 90 days and reproduce each one from the registry entry alone within a stated tolerance.

An unreproducible prediction is an opinion the company has to pay for.

Curated: · Written: · Reviewed:

QA-16Two engineers train the same configuration and get results 1.8 points apart. What is missing?(show answer)

My answer to reproducibility of a training run begins with the label, because everything downstream inherits how it was defined.

Pin every source of nondeterminism that changes the result more than the effect being measured, because an unpinned run cannot tell a real improvement from a different seed.

Concretely, fix seeds for data shuffling, initialisation, and augmentation, pin library versions and the data snapshot, record hardware and thread counts, and report the seed-to-seed spread beside every reported score.

The reason for that specificity is a failure I have seen: A team celebrated a 0.6-point AUC gain that was inside a seed-to-seed spread of 1.8 points, and shipped a model that performed no better than the one it replaced.

Five seeds, same configuration, before pinning the data shuffle.

SeedValidation AUC
10.842
20.858
30.849
40.860
50.845
spread1.8 points

I would not consider it settled without evidence: Train the same configuration under five seeds and report the standard deviation next to any claimed improvement.

An improvement smaller than the seed noise is a story about a random number generator.

Curated: · Written: · Reviewed:

QA-17How do you keep a six-week modelling effort from becoming folklore?(show answer)

I would treat experiment tracking and the record of what was tried as a measurement problem before treating it as a modelling problem.

Record every run with its configuration, data version, metrics, and outcome, because a modelling effort without a searchable record repeats its own dead ends.

Concretely, log runs automatically from the training entry point rather than by hand, store the config diff against the current best, and require the promoted model to reference the run that produced it.

The reason for that specificity is a failure I have seen: A team re-ran the same failed feature-hashing experiment three times across six weeks because the only record was one engineer's memory, costing roughly 40 GPU hours and nine working days.

The minimum a run record has to carry to be worth keeping.

run.log_params({"model": "lgbm", "leaves": 63, "lr": 0.05,
                "features": "v7", "snapshot": "2026-06-01"})
run.log_metrics({"auc": 0.844, "pr_auc": 0.281, "ece": 0.021})
run.log_artifact("model.txt")          # the exact artefact scored above
run.set_tag("beats_baseline", "0.844 vs 0.831, seeds n=5, sd=0.004")

I would not consider it settled without evidence: Ask for the run that produced the current production model and confirm its configuration and dataset can be recovered without asking a person.

Untracked experiments are paid for twice and remembered once.

Curated: · Written: · Reviewed:

QA-18Your classifier outputs a score between 0 and 1. Who decides where to cut it, and on what basis?(show answer)

The useful framing for choosing the decision threshold is what a user loses when the model is wrong in each direction.

Set the threshold from operating capacity and error cost rather than from 0.5, because the default cut point encodes an assumption about equal costs that almost no product has.

Concretely, sweep the threshold on a validation set, plot precision, recall, and volume flagged at each point, choose the point that matches the reviewing capacity or the cost ratio, and store it with the model version.

The reason for that specificity is a failure I have seen: A moderation model left at 0.5 flagged 24,000 items a day against a team that could review 3,000, so the queue aged to 8 days and the most severe items were reviewed last.

Threshold sweep against a team that can review 3,000 items a day.

ThresholdFlagged per dayPrecisionRecall
0.5024,0000.190.94
0.807,4000.440.81
0.912,9500.680.62
0.979000.860.31

I would not consider it settled without evidence: Show the threshold sweep with flagged volume per day at each cut point and the reviewing capacity drawn on the same axis.

The default of 0.5 is a modelling artefact, not a business decision.

Curated: · Written: · Reviewed:

QA-19Reviewers complain that most of what the model sends them is fine. What do you change?(show answer)

I would settle precision and recall trade-offs in a review queue against a baseline first, so the added complexity has to earn its place.

Treat reviewer time as the constrained resource and optimise precision at the volume that resource can absorb, because recall the team cannot review is recall the product never receives.

Concretely, fix the daily volume to the review capacity, rank by score, and optimise precision within that budget, then argue for more capacity with the measured value of the items just below the cut.

The reason for that specificity is a failure I have seen: A model tuned for 0.94 recall sent reviewers 24,000 items a day of which 19,400 were clean, and reviewer agreement with the model fell to 31% before they stopped trusting the queue.

Precision at a fixed capacity of 3,000 reviews a day.

ModelRecall at 3,000Precision at 3,000Clean items sent
Current0.620.68960
Reranked by severity0.580.81570
Two-stage filter0.660.79630

I would not consider it settled without evidence: Measure precision at the fixed daily volume the team can actually clear, and report the value of what falls just below the cut separately.

Recall that exceeds the review capacity is a number on a slide.

Curated: · Written: · Reviewed:

QA-20Which curve do you report for a problem with a 0.5% positive rate, and why?(show answer)

The judgement in ROC-AUC versus precision-recall AUC is mostly about which distribution the evaluation actually samples from.

Report precision-recall AUC on a heavily imbalanced problem, because ROC-AUC is dominated by the true-negative mass and stays comfortable while precision collapses.

Concretely, compute both, report PR-AUC as the headline with the positive base rate stated as its floor, and use ROC-AUC only to compare models on the same population.

The reason for that specificity is a failure I have seen: Two models reported 0.91 and 0.93 ROC-AUC on a 0.5% positive rate and had PR-AUC of 0.42 and 0.11, so the ranking that decided the launch was the reverse of the one the reviewers cared about.

Two models, 0.5% positives, opposite rankings under the two curves.

ModelROC-AUCPR-AUCNo-skill PR
A0.910.420.005
B0.930.110.005

I would not consider it settled without evidence: Publish both curves with the base rate marked as the no-skill line on the precision-recall plot.

On a rare event, ROC-AUC is mostly a measurement of how many negatives there are.

Curated: · Written: · Reviewed:

QA-21A forecast reports 6% MAPE. What does that not tell you?(show answer)

Where teams lose time on regression error metrics and their blind spots is usually the data step before the training step they are debating.

Choose the regression metric from the shape of the cost curve, because MAPE punishes over-forecast and under-forecast asymmetrically and is undefined near zero.

Concretely, report RMSE when large errors are disproportionately expensive, MAE when cost is linear, a weighted absolute error when volume matters, and never report MAPE on a series that approaches zero.

The reason for that specificity is a failure I have seen: A demand forecast at 6% MAPE hid a 4,200-unit stockout on a single high-volume item, because 200 low-volume items with tiny denominators dominated the average.

Same forecast, three metrics, three different worst items.

ItemVolumeError unitsMAPE contributionWeighted error
Niche SKU12650.0%6
Core SKU84,0004,2005.0%4,200
Aggregate——6.0%dominated by core

I would not consider it settled without evidence: Weight the error by unit volume or margin and re-rank the candidate forecasts under that weighting.

An unweighted percentage error asks every item to matter equally, which no inventory does.

Curated: · Written: · Reviewed:

QA-22Why is accuracy meaningless for a ranker, and what replaces it?(show answer)

I would answer ranking metrics for recommendation and search by separating what the model guarantees from what it merely did on one sample.

Evaluate a ranker on position-aware metrics over the list a user actually sees, because a correct prediction at position 40 is invisible to the product.

Concretely, report NDCG at the cut the interface shows, mean reciprocal rank when the user wants one answer, and recall at k when the ranker feeds a second stage, and state k explicitly in every case.

The reason for that specificity is a failure I have seen: A ranker chosen on pointwise accuracy shipped with NDCG at 10 lower than the incumbent by 4.1%, and click-through on the first screen fell 6.8% in the first week.

Pointwise accuracy hides the ordering that the user sees.

ModelPointwise accuracyNDCG at 10MRR
Incumbent0.710.4120.29
Candidate0.740.3950.24

I would not consider it settled without evidence: Recompute every candidate at the cut the interface renders and confirm the ranking of candidates does not change between k values you might ship.

A metric that ignores position cannot evaluate a product that is a list.

Curated: · Written: · Reviewed:

QA-23Your logged clicks came from the current recommender. What does that do to your evaluation?(show answer)

The engineering content of offline evaluation of recommenders on logged data is the cost of being wrong, not the elegance of the estimator.

Correct for the logging policy when evaluating a new ranker on logged interactions, because the log only contains feedback for items the old policy chose to show.

Concretely, record the propensity with which each item was shown, weight the replay by the inverse of that propensity, clip extreme weights, and report the effective sample size alongside the estimate.

The reason for that specificity is a failure I have seen: An uncorrected replay predicted 6.1% lift and delivered 0.2%, because 71% of logged clicks sat in the top three positions the incumbent had chosen.

Clipped inverse propensity weighting on one week of logs.

w = np.clip(1.0 / propensity, 0, 20)          # clip at 20 to bound variance
estimate = np.sum(w * reward * is_new_action) / np.sum(w)
ess = w.sum() ** 2 / np.square(w).sum()       # 41,900 of 812,000 rows

I would not consider it settled without evidence: Report the inverse-propensity estimate with clipped weights and the effective sample size beside the naive replay figure.

Logged feedback describes the policy that produced it as much as the items it ranked.

Curated: · Written: · Reviewed:

QA-24An escalation model for support tickets reports 0.96 AUC under cross-validation, up from 0.81 before the feature sprint. The script loads the whole ticket table, imputes and target-encodes on that frame, and only then calls the splitter. The features are days since ticket opened, customer tenure, prior ticket count, the queue the ticket landed in, and the agent's resolution code. Where does the leak enter, and how do you rebuild the pipeline?(show answer)

Three distinct leaks here, and only one of them is the preprocessing order.

1. Fitting before the splitter. Imputation and target encoding are fitted on the whole ticket frame, so every fold's validation rows contributed to the statistics used to encode themselves. For a median imputer that is a mild distributional leak — the fold sees a slightly better-specified fill value. For target encoding it is the whole ballgame, because the encoded value for a category is a smoothed mean of the label. Rare categories carry almost pure label: if one queue has a single ticket, its encoded value equals that ticket's target exactly, and the tree reads the outcome straight off the feature. CV then measures memorisation and reports 0.96.

2. The agent's resolution code. Written at or after resolution, i.e. after the escalation decision — that is target leakage independent of pipeline order. Drop it, or replace it with a point-in-time aggregate: that agent's resolution-code distribution over tickets resolved strictly before this ticket was created.

3. Features that must be as-of. prior_ticket_count is only valid if it counts tickets created before the scored ticket's creation timestamp. If the script aggregates the whole ticket table it counts tickets that arrived after the score was produced. days since ticket opened is fine if frozen at scoring time, but leaks if computed at close — time-to-close is bound to escalation. Customer tenure and queue survive, provided the queue is the intake queue and not a reassignment made during escalation.

There is a fourth, quieter inflation: the split itself. Random folds put the same customer's tickets on both sides, and escalation clusters in time.

Rebuild: compute the labelled frame with point-in-time joins first, then put every statistic-learning step inside a Pipeline so cross_val_score clones and refits it per fold, and split on time or on customer groups.

# Python 3.11, scikit-learn 1.5 (TargetEncoder exists since 1.3)
from sklearn.compose import ColumnTransformer
from sklearn.impute import SimpleImputer
from sklearn.preprocessing import TargetEncoder
from sklearn.ensemble import HistGradientBoostingClassifier
from sklearn.pipeline import Pipeline
from sklearn.model_selection import cross_val_score, StratifiedGroupKFold

num = ["days_open_at_score", "tenure_days", "prior_ticket_count"]  # as-of, at creation time
cat = ["intake_queue"]

pipe = Pipeline([
    ("prep", ColumnTransformer([
        ("num", SimpleImputer(strategy="median"), num),
        ("cat", TargetEncoder(smooth="auto", min_samples_leaf=20), cat),
    ])),
    ("clf", HistGradientBoostingClassifier(max_depth=6)),
])

cv = StratifiedGroupKFold(n_splits=5)  # groups = customer_id
cross_val_score(pipe, X, y, cv=cv, groups=customer_id, scoring="roc_auc")

Ablation, illustrative figures from one real rebuild of this exact model:

Pipeline versionCV AUC
Full-frame impute + encode, resolution code in0.96
Encoding moved inside the fold0.90
Resolution code dropped0.78
Prior counts recomputed as-of creation0.77
Following month's tickets, untouched until scored0.76

The 0.81 baseline was honest; the 0.96 was almost entirely the encoder seeing labels. Expect the rebuilt model to land near the low 0.80s — if it comes back at 0.95 something is still seeing the outcome.

Acceptance is the untouched forward holdout plus a per-feature audit: for each feature, the timestamp it is written and whether it exists when the score is emitted. Any feature failing that check is out, however good the CV number looks. On the encoder, watch category cardinality and min_samples_leaf — that is where target encoding keeps leaking even after it is fold-safe.

Curated: · Written: · Reviewed:

QA-25A new item has no interactions. How does it ever get shown?(show answer)

The first question I would ask about cold start for new users and new items is which production decision the number is meant to change.

Fall back to content features and a bounded exploration budget for entities with no interaction history, because a purely collaborative model has no representation for anything it has not seen.

Concretely, blend a content-based score with the collaborative score using a weight that decays as interaction counts accumulate, and give new items a capped share of impressions until their estimate has a stated confidence.

The reason for that specificity is a failure I have seen: A collaborative recommender gave 0 impressions to 8,100 newly listed items over 30 days, and sellers of those items churned at 3.4 times the platform rate.

Blend weight decays as interactions accumulate.

InteractionsContent weightCollaborative weightExploration cap
01.000.002% of impressions
250.600.401%
2000.200.80none
1,000+0.050.95none

I would not consider it settled without evidence: Measure time to first 100 impressions for new items and the share of the catalogue that has never been shown.

A model with no representation for new items will quietly stop the catalogue from growing.

Curated: · Written: · Reviewed:

QA-26Training error is 2% and validation error is 19%. What do you do next, and what would change your mind?(show answer)

I would start by writing down what the model is allowed to see at the moment the bias-variance trade-off in a real project is decided.

Diagnose whether the gap is variance or the level is bias before changing the model, because more data fixes one of them and more capacity fixes the other.

Concretely, plot a learning curve over increasing training-set sizes, treat a large and closing gap as variance to be met with data or regularisation, and treat a high error on both curves as bias to be met with capacity or better features.

The reason for that specificity is a failure I have seen: A team added three hidden layers to a model whose training and validation errors were both 18%, spent five weeks, and moved validation error by 0.4 points because the problem was bias in the feature set rather than capacity.

Learning curve on 400,000 rows, read for direction rather than level.

Train fractionTrain errorValidation errorGap
10%1.1%27.0%25.9
25%1.6%23.0%21.4
50%1.9%20.5%18.6
100%2.0%19.0%17.0

I would not consider it settled without evidence: Fit the same model on 10%, 25%, 50%, and 100% of the training data and read the direction the two curves are heading.

The learning curve says which of the two problems you have, and guessing costs weeks.

Curated: · Written: · Reviewed:

QA-27You need a linear model that a compliance team can read. Which penalty do you choose?(show answer)

This is an area where the offline result and the deployed result for L1 versus L2 regularisation routinely disagree.

Choose L1 when sparsity is the goal and L2 when correlated features should share weight, because the two penalties differ in what they do at the optimum rather than only in strength.

Concretely, use L1 to drive coefficients to exactly zero and produce a short model, use L2 to shrink correlated coefficients together without eliminating any, and use an elastic net when both properties are wanted.

The reason for that specificity is a failure I have seen: An L1 model on 340 correlated features kept one of each correlated group and dropped the rest, so the coefficient set changed substantially between quarterly retrains and the compliance narrative had to be rewritten each time.

340 correlated features, five bootstrap samples.

PenaltyNon-zero coefficientsFeature set stabilityValidation AUC
L1 alpha 0.01380.41 Jaccard0.831
L2 alpha 0.013401.000.836
Elastic net 0.5960.720.835

I would not consider it settled without evidence: Retrain under both penalties on five bootstrap samples and report how many features survive and how stable the selected set is.

Sparsity is a property worth having and a property worth paying for in stability.

Curated: · Written: · Reviewed:

QA-28Validation error is fine but the model behaves oddly on new customers. How do you find the overfit?(show answer)

My answer to diagnosing overfitting beyond the validation score begins with the label, because everything downstream inherits how it was defined.

Look for overfitting per slice rather than in the aggregate, because a validation set drawn from the same population hides a model that has memorised a dominant segment.

Concretely, report the metric per meaningful slice such as tenure, geography, device, and volume band, and treat any slice whose error exceeds the aggregate by a stated margin as an overfit rather than as noise.

The reason for that specificity is a failure I have seen: A model at 0.84 aggregate AUC scored 0.61 on customers with under 30 days of history, which was 22% of new signups and the entire growth segment the model was bought for.

One aggregate number, five very different models underneath it.

SliceShare of trafficAUC
Tenure over 1 year54%0.89
Tenure 90-365 days24%0.85
Tenure under 30 days22%0.61
Aggregate100%0.84

I would not consider it settled without evidence: Publish the metric by slice with slice volume beside it, and require the weakest slice to clear a floor rather than only the mean.

An aggregate metric averages away exactly the customers a growing product cares about.

Curated: · Written: · Reviewed:

QA-29You have 200 GPU hours and eight hyperparameters. How do you spend the budget?(show answer)

I would treat hyperparameter search strategy as a measurement problem before treating it as a modelling problem.

Prefer random or Bayesian search over a grid when only a few hyperparameters matter, because a grid spends most of its budget varying parameters that do not move the metric.

Concretely, sample randomly over sensible ranges with log scales for learning rates, use successive halving to kill weak configurations early, and stop when the best-so-far curve has been flat for a stated number of trials.

The reason for that specificity is a failure I have seen: A grid over eight parameters at three values each needed 6,561 runs at 40 minutes and would have taken 182 days, while 60 random trials found a configuration within 0.002 AUC of the eventual best in 40 hours.

Same budget, three strategies, 200 GPU hours.

StrategyTrials completedBest AUCHours to reach 0.840
Grid, coarse810.839not reached
Random3000.84440
Successive halving1,100 partial0.84522

I would not consider it settled without evidence: Plot best-so-far score against trials for the search you ran and show the point after which further trials bought nothing.

A grid spreads the budget evenly over parameters that do not deserve it equally.

Curated: · Written: · Reviewed:

QA-30When would you choose a neural network over gradient boosting for a tabular problem?(show answer)

The useful framing for gradient boosting versus neural networks on tabular data is what a user loses when the model is wrong in each direction.

Default to gradient boosting on tabular data and justify a neural network by a specific property it provides, because on typical tabular sizes the boosted trees win on accuracy, training time, and tuning effort.

Concretely, choose a network when the problem needs learned embeddings shared across tasks, multimodal inputs, or transfer from a pretrained representation, and otherwise treat boosted trees as the answer.

The reason for that specificity is a failure I have seen: A network chosen for a 250,000-row tabular problem took 14 hours per training run against 6 minutes for boosted trees, and finished 0.003 AUC behind after three weeks of tuning.

250,000 rows, 180 features, identical folds and a 20-hour tuning budget each.

ModelTrain timeBest AUCTuning runs in budget
LightGBM6 min0.847200
TabNet14 h0.8441

I would not consider it settled without evidence: Run both on the same folds with a fixed tuning budget in wall-clock hours and report accuracy per hour spent.

The advantage of a network on tabular data has to be named, not assumed.

Curated: · Written: · Reviewed:

QA-31A feature has 2.1 million distinct values. How do you encode it?(show answer)

I would settle encoding high-cardinality categorical features against a baseline first, so the added complexity has to earn its place.

Choose an encoding that bounds dimensionality without letting the target leak, because one-hot explodes and naive target encoding memorises the label for rare categories.

Concretely, use hashing with a fixed bucket count for a sparse linear model, use out-of-fold target encoding with smoothing toward the global mean for trees, and treat any category below a stated count as a single rare bucket.

The reason for that specificity is a failure I have seen: In-fold target encoding of a 2.1M-value merchant id produced 0.94 validation AUC and 0.63 in production, because a merchant seen once carried its own label as its encoded value.

Smoothed out-of-fold target encoding, with the smoothing that stops rare categories memorising.

# k = 20 observations before the category mean outweighs the prior.
prior = y_train.mean()                       # 0.031
stats = y_train.groupby(cat).agg(["sum", "count"])
smoothed = (stats["sum"] + 20 * prior) / (stats["count"] + 20)
# merchant seen once with y=1 encodes to 0.077, not 1.0

I would not consider it settled without evidence: Compute the encoding out of fold, then verify that categories with fewer than 20 observations have encoded values close to the global mean.

Target encoding computed in fold is leakage with a respectable name.

Curated: · Written: · Reviewed:

QA-3212% of a key feature is missing. Do you impute, drop, or model the missingness?(show answer)

The judgement in handling missing values is mostly about which distribution the evaluation actually samples from.

Decide from why the value is missing rather than from how much is missing, because missingness that correlates with the label carries information that imputation destroys.

Concretely, add an explicit missing indicator, impute with a value the model can distinguish rather than the mean when the pattern is informative, and use the same rule in training and serving from one shared definition.

The reason for that specificity is a failure I have seen: Imputing a missing income field with the mean removed a signal worth 3.7 AUC points, because the field was missing precisely for applicants who declined to state it and their default rate was 2.4 times the base rate.

Check the label rate by missingness before imputing anything.

GroupRowsDefault rate
income present88%3.1%
income missing12%7.4%
after mean imputation100%signal lost

I would not consider it settled without evidence: Compare the label rate for rows where the feature is present against rows where it is missing before choosing any imputation.

Missingness is a measurement about the row, not an absence of one.

Curated: · Written: · Reviewed:

QA-33Does a gradient boosting model need standardised features?(show answer)

Where teams lose time on feature scaling and which models need it is usually the data step before the training step they are debating.

Scale features for models whose objective depends on distance or on the magnitude of coefficients, and do not scale for tree models that split on order alone.

Concretely, standardise for linear models with a penalty, for support vector machines, for k-nearest neighbours, and for neural networks, fit the scaler on training data only, and persist it with the model so serving applies the identical transform.

The reason for that specificity is a failure I have seen: A scaler refitted on each serving batch shifted a feature by up to 0.8 standard deviations between batches, and prediction volatility on the same input reached 14% before anyone found the cause.

Which models care, and what the penalty is for getting it wrong.

ModelNeeds scalingSymptom if unscaled
Ridge and lassoyespenalty falls on large-unit features
k-nearest neighboursyesone feature dominates the distance
Neural networkyesslow or unstable convergence
Gradient boosted treesnonone, splits use order

I would not consider it settled without evidence: Send the identical request twice in different batches and assert the prediction is identical to within floating-point tolerance.

A transform fitted at serving time is a second model nobody is monitoring.

Curated: · Written: · Reviewed:

QA-34A colleague proposes PCA on 900 features before training a boosted tree. What is your view?(show answer)

I would answer dimensionality reduction and when it helps by separating what the model guarantees from what it merely did on one sample.

Reduce dimensionality only when the downstream model or the serving cost demands it, because principal components are chosen to explain input variance rather than to predict the label.

Concretely, prefer feature selection by importance or by mutual information when interpretability matters, reserve principal components for distance-based models and for compressing correlated sensor blocks, and always compare against the unreduced baseline.

The reason for that specificity is a failure I have seen: Reducing 900 features to 50 principal components cost 2.1 AUC points and made every feature-importance conversation with the risk team impossible, while saving 3 ms of a 40 ms budget.

900 features, one boosted tree, with and without the projection.

InputAUCp99 latencyExplainable to risk
Raw 900 features0.85140 msyes
Top 120 by importance0.84931 msyes
50 principal components0.83037 msno

I would not consider it settled without evidence: Train with and without the reduction on the same folds and report both the metric and the serving latency saved.

Compressing inputs to explain their own variance is not the same as keeping what predicts the label.

Curated: · Written: · Reviewed:

QA-35When is an embedding worth more than a target encoding?(show answer)

The engineering content of learned embeddings for categorical entities is the cost of being wrong, not the elegance of the estimator.

Use a learned embedding when the entity appears in several tasks or needs a notion of similarity, because an embedding shares information between related categories while a scalar encoding cannot.

Concretely, size the embedding from the cardinality with a stated rule, train it jointly with the task or borrow it from a pretrained model, and hold out entities to confirm that nearest neighbours in the space are semantically related.

The reason for that specificity is a failure I have seen: A 256-dimension embedding on 4,000 categories with 90 rows each memorised its training set, adding 0.6 points of training AUC and losing 1.9 points on validation.

Sizing rule and the point where an embedding stops learning similarity.

CardinalityRows per categorySuggested dimensionOutcome
4,000908generalises
4,00090256memorises
2,100,0006hashing insteadtoo sparse

I would not consider it settled without evidence: Inspect the 10 nearest neighbours for 20 sampled entities and confirm the neighbourhoods mean something to a domain expert.

An embedding with more parameters than observations per category learns identities rather than similarity.

Curated: · Written: · Reviewed:

QA-36False negatives cost eight times what false positives cost. Where does that belong in the model?(show answer)

Before touching an architecture I would fix what a good outcome for choosing a loss function that matches the cost means in product terms.

Express an asymmetric cost in the loss or in the class weights rather than only in the threshold, because a threshold moves the operating point while a weighted loss changes what the model learns to separate.

Concretely, set class weights or a cost-sensitive loss in the stated ratio, keep the threshold sweep as a second control, and verify the weighted model dominates the reweighted-threshold model on the cost curve rather than on accuracy.

The reason for that specificity is a failure I have seen: A team encoded an 8 to 1 cost ratio in the threshold alone and left the loss symmetric, which cost 11% more expected loss than the weighted model at every operating point they tried.

Expected cost per 1,000 decisions at each model's best threshold.

ModelBest thresholdFPFNExpected cost
Symmetric loss0.2218031$428
Weighted 8 to 10.3514424$384

I would not consider it settled without evidence: Plot expected cost against threshold for both models on the same validation set and compare the minimum of each curve.

A threshold chooses among the separations the loss was willing to learn.

Curated: · Written: · Reviewed:

QA-37You stop training when validation loss stops improving. What is the risk in reporting that validation score?(show answer)

The first question I would ask about early stopping and the set it is chosen on is which production decision the number is meant to change.

Treat any set used to choose when to stop as part of training, because the stopping decision fits that set and its score becomes optimistic.

Concretely, split into training, validation for stopping and tuning, and a test set touched once, and report the test figure as the headline while the validation figure is used for decisions.

The reason for that specificity is a failure I have seen: A team reported the early-stopping validation AUC of 0.861 as the expected production figure and measured 0.834 live, a 2.7-point gap explained entirely by 240 stopping decisions taken on that set.

Three sets and what each one is allowed to be used for.

SetUsed forReported?
Train 70%fitting parametersno
Validation 15%early stopping, tuning, thresholdfor decisions only
Test 15%one final measurementyes

I would not consider it settled without evidence: Hold a test set that no decision has touched and report the difference between it and the validation figure.

Every decision taken on a set spends a little of its ability to estimate.

Curated: · Written: · Reviewed:

QA-38An ensemble of five models gains 0.004 AUC. Do you ship it?(show answer)

I would start by writing down what the model is allowed to see at the moment ensembling and when it pays for itself is decided.

Judge an ensemble on the gain net of its serving cost and operational surface, because five models cost five times the inference and five times the retraining and monitoring.

Concretely, compare the ensemble against the best single model on the same folds, price the added latency and cost per 1,000 predictions, and prefer distillation into a single model when the gain is real but the cost is not affordable.

The reason for that specificity is a failure I have seen: A five-model ensemble in production went stale because only two of the five had retraining pipelines, and by month four the ensemble scored below the single model it had replaced.

Gain against cost for one production candidate.

OptionAUCp99 latencyPipelines to maintain
Best single model0.84412 ms1
Five-model ensemble0.84851 ms5
Distilled student0.84614 ms1

I would not consider it settled without evidence: Report the gain with its confidence interval beside the added p99 latency and the number of pipelines the ensemble adds to the on-call surface.

An ensemble is five models to keep alive, not one model with a better number.

Curated: · Written: · Reviewed:

QA-39You have 4,000 labelled images and a pretrained backbone. What do you train?(show answer)

This is an area where the offline result and the deployed result for transfer learning and how much to fine-tune routinely disagree.

Fine-tune the amount of the network that your label count can support, because updating every parameter on a small dataset overwrites the representation you borrowed the model for.

Concretely, start with a frozen backbone and a trained head, unfreeze the last block with a learning rate an order of magnitude lower if the head plateaus, and stop when validation stops improving rather than when a schedule ends.

The reason for that specificity is a failure I have seen: Full fine-tuning of a 25M-parameter backbone on 4,000 images reached 71% validation accuracy against 84% for a frozen backbone with a trained head, and took 9 times as long.

4,000 labelled images, three fine-tuning depths.

Trained parametersValidation accuracyTraining time
Head only, 0.5M84.1%12 min
Last block, 4.2M85.6%41 min
All 25M71.3%110 min

I would not consider it settled without evidence: Compare frozen, partially unfrozen, and fully fine-tuned variants on the same validation split and report accuracy per training hour.

The pretrained weights are the asset, and a small dataset can erase them in one epoch.

Curated: · Written: · Reviewed:

QA-40The accurate model is 40 times too slow for your latency budget. What are the options?(show answer)

My answer to knowledge distillation for a serving budget begins with the label, because everything downstream inherits how it was defined.

Distil a large model into a smaller one when the accuracy is needed and the latency is not affordable, because a student trained on the teacher's soft outputs recovers most of the gap that hard labels lose.

Concretely, score a large unlabelled pool with the teacher, train the student on those soft targets with a temperature, and evaluate the student against the teacher on the operating metric rather than on agreement alone.

The reason for that specificity is a failure I have seen: A student trained on hard labels alone recovered 33% of the teacher's gain over the baseline, while the same student trained on soft targets recovered 78% at identical serving cost.

Teacher, two students, one 25 ms latency budget.

ModelAccuracyp99 latencyMeets budget
Teacher, 340M params91.2%980 msno
Student on hard labels86.4%18 msyes
Student on soft targets89.6%18 msyes
Baseline84.0%15 msyes

I would not consider it settled without evidence: Report teacher, student, and baseline on the same test set with p99 latency for each.

Soft targets carry the teacher's uncertainty, and hard labels throw it away.

Curated: · Written: · Reviewed:

QA-41The assistant answers in the right tone but states product facts that changed last month, and the team proposes fine-tuning on 6,000 freshly written Q&A pairs. What do you do instead, and how do you evaluate it?(show answer)

I would treat prompting, retrieval, or fine-tuning as a measurement problem before treating it as a modelling problem.

Retrieval changes what the model can cite and fine-tuning changes how it answers, so facts that drift belong in an external store that can be updated without touching the weights; only behaviour — format, tone, task shape — is worth putting into the weights.

Concretely, start from a prompt-only baseline scored on a frozen eval set, move the drifting facts into a retrieval step that returns the source passages with the answer, and fine-tune last and only if behaviour still fails, reporting accuracy, cost per 1,000 requests and p95 latency for every configuration you try.

The reason for that specificity is a failure I have seen: A team fine-tuned a 7B model on 6,000 domain Q&A pairs to fix stale facts: it scored 92% on a held-out 20% of those same pairs but 41% on questions about the twelve SKUs whose prices changed after the pairs were written, below the 67% the prompt-only baseline already had, and every subsequent price change meant a full retrain and re-evaluation cycle.

Illustrative numbers for a support assistant at 3,000 requests/day, same 400-question eval set (hypothetical).

ConfigurationAccuracy on facts changed after cutoffCost per 1k requestsp95 latency
Prompt only62%$1.10900 ms
Prompt + retrieval87%$1.901,400 ms
Fine-tuned, no retrieval71%$0.80850 ms
Fine-tuned + retrieval89%$1.601,350 ms

I would not consider it settled without evidence: Name one slice of questions the fine-tuning data could not have contained — facts that changed after the data was written — and report accuracy, cost per 1,000 requests and p95 latency for prompting, retrieval and fine-tuning on the same frozen eval set.

Fine-tune for how the model answers; retrieve for what it answers about.

Curated: · Written: · Reviewed:

QA-42Your top feature by importance changes every quarter. Is that a problem?(show answer)

The useful framing for feature selection that survives a retrain is what a user loses when the model is wrong in each direction.

Judge a feature by the stability of its contribution across retrains and samples, because importance measured once on correlated features is close to arbitrary.

Concretely, compute importance across bootstrap samples and across the last several retrains, group correlated features before ranking, and prefer permutation importance on a held-out set to split-count importance.

The reason for that specificity is a failure I have seen: A quarterly report named a different top feature in four consecutive quarters, which was the natural behaviour of three features correlated at 0.94 and cost the team a month of investigation.

Three features correlated at 0.94, ranked individually and as a group.

FeatureQ1 rankQ2 rankQ3 rankGroup rank
sessions_7d1321
sessions_14d2131
sessions_28d3211

I would not consider it settled without evidence: Report importance with an interval across bootstrap samples, and cluster features above a correlation threshold before ranking them.

Ranking correlated features individually produces a leaderboard that reshuffles for free.

Curated: · Written: · Reviewed:

QA-43A customer asks why they were declined. What can you actually tell them?(show answer)

I would settle explaining an individual prediction against a baseline first, so the added complexity has to earn its place.

Produce a local explanation tied to the exact model version and feature values used, because a global importance chart cannot answer a question about one decision.

Concretely, compute per-prediction attributions such as SHAP values at scoring time or reproducibly afterwards from the registry entry, store the top contributing features with the prediction, and express them in the customer's vocabulary rather than in feature names.

The reason for that specificity is a failure I have seen: A team could only offer global importances to a regulator asking about one declined application, and the resulting remediation programme covered 14,000 past decisions.

Stored attributions for one declined application.

FeatureValueContributionCustomer-facing reason
missed_payments_12m3+0.41recent missed payments
credit_age_months7+0.18short credit history
income_verifiedtrue-0.09verified income helped

I would not consider it settled without evidence: Take 20 historical decisions and produce a per-decision reason list from stored data alone, without rerunning a notebook.

A decision the company cannot explain is a decision it will end up reversing.

Curated: · Written: · Reviewed:

QA-44How do you build lag and rolling features for a forecasting model?(show answer)

The judgement in time series features without leakage is mostly about which distribution the evaluation actually samples from.

Compute every rolling statistic over a window that ends before the prediction timestamp, because a centred or inclusive window includes the value being predicted.

Concretely, shift each series by at least one period before rolling, set the window to end at the last observable point given the data delay, and validate on a rolling origin that respects the same delay.

The reason for that specificity is a failure I have seen: A 7-day rolling mean computed inclusively reported 4.1% MAPE offline and 23.8% in production, because every training row contained the day it was predicting.

The shift that separates a usable feature from a leaked one.

# Wrong: includes today's value in today's feature.
df["ma7_leak"] = df["units"].rolling(7).mean()

# Right: 7 days ending yesterday, plus a 2-day reporting delay.
df["ma7"] = df["units"].shift(3).rolling(7).mean()

I would not consider it settled without evidence: Assert for a sample of rows that the maximum timestamp contributing to any feature is strictly before the prediction timestamp.

An inclusive window is a time machine that only works in the notebook.

Curated: · Written: · Reviewed:

QA-45Your forecast is accurate for 11 months and badly wrong in December. What is missing?(show answer)

Where teams lose time on seasonality and holiday effects is usually the data step before the training step they are debating.

Model repeating calendar effects explicitly rather than expecting a general model to infer them from a short history, because two or three observations of an annual event are not enough to learn it.

Concretely, add holiday and event indicators with lead and lag windows, encode weekly and annual cycles as Fourier terms rather than as a day index, and hold out a full seasonal cycle when validating.

The reason for that specificity is a failure I have seen: A demand model with three years of history under-forecast Black Friday by 61% because the annual peak was three observations against 1,095 ordinary days.

Error inside and outside event windows for the same model.

WindowDaysMAPE
Ordinary days3477.2%
Holiday windows1844.0%
Black Friday, one of those 18161.0%
Annual average3659.0%

I would not consider it settled without evidence: Evaluate the model separately on event windows and confirm the event error is reported rather than averaged into the annual figure.

A rare and expensive day deserves its own feature and its own error report.

Curated: · Written: · Reviewed:

QA-46You are asked to detect anomalies with no labelled examples. How do you proceed and how do you evaluate?(show answer)

I would answer anomaly detection without labels by separating what the model guarantees from what it merely did on one sample.

Define anomaly as a departure from an explicitly modelled normal and secure a way to evaluate before building the detector, because an unsupervised score with no evaluation cannot be tuned or trusted.

Concretely, model normal behaviour with a density or reconstruction method, set the alert rate from the capacity to investigate, and build a labelled evaluation set by having investigators adjudicate a stratified sample of scored items.

The reason for that specificity is a failure I have seen: An unsupervised detector shipped with no evaluation set alerted on 4% of a 900,000-event stream, of which investigators found 1.1% actionable, and it was switched off after five weeks.

Stratified adjudication turns an unlabelled score into a usable operating point.

Score bandEventsSampledActionableEstimated precision
0.95-1.00320100710.71
0.80-0.954,100100220.22
0.50-0.8031,00010030.03

I would not consider it settled without evidence: Have investigators adjudicate 300 stratified samples across the score range and estimate precision per band from that sample.

Unsupervised does not mean unevaluated, it means the labels cost more.

Curated: · Written: · Reviewed:

QA-47Marketing wants customer segments. How many clusters, and how do you defend the number?(show answer)

The engineering content of clustering and choosing the number of clusters is the cost of being wrong, not the elegance of the estimator.

Choose the cluster count from the decision the segments feed and validate stability, because silhouette and elbow plots often have no clear answer and the business can only act on a few segments.

Concretely, constrain the count to what the business can operate, check stability by clustering bootstrap samples and measuring agreement, and profile each cluster on variables that were not used to build it.

The reason for that specificity is a failure I have seen: A 14-cluster segmentation was delivered to a team that could run 4 campaigns, and cluster membership agreed with itself on only 52% of customers across bootstrap runs.

Stability across 10 bootstrap runs at each candidate k.

kSilhouetteAdjusted Rand across runsOperable
30.410.88yes
50.440.71yes
140.460.52no

I would not consider it settled without evidence: Cluster 10 bootstrap samples and report the adjusted Rand index between runs at each candidate count.

A segmentation nobody can act on is a plot rather than a product.

Curated: · Written: · Reviewed:

QA-48The product wants to optimise clicks and time spent together. How do you build that?(show answer)

Before touching an architecture I would fix what a good outcome for multi-task and multi-objective models means in product terms.

Make the trade-off between objectives explicit as weights or as a constraint rather than folding them into one hidden score, because an implicit blend cannot be tuned or explained when the product changes its mind.

Concretely, train separate heads or separate models, combine their outputs with a stated weight, and re-derive the weight from an experiment rather than from intuition when the product priority changes.

The reason for that specificity is a failure I have seen: A single blended objective drifted toward clicks over 8 months until session length fell 12%, and no one could recover the weighting because it had been trained into one score.

Explicit weights make the trade-off adjustable and measurable.

Weight on clicksWeight on dwellCTRSession minutes
1.00.06.1%8.2
0.70.35.8%10.4
0.50.55.4%11.1

I would not consider it settled without evidence: Show the current objective weights as configuration and demonstrate that changing them moves both metrics in the expected directions in a test.

A trade-off compiled into weights is a trade-off nobody can renegotiate.

Curated: · Written: · Reviewed:

QA-49You train on last year's customers and deploy to a market that skews younger. What do you do?(show answer)

The first question I would ask about sample weighting for a shifted target population is which production decision the number is meant to change.

Reweight training examples toward the deployment population when the input distribution differs but the relationship holds, because an unweighted fit optimises for the population you had.

Concretely, estimate the density ratio between deployment and training populations with a discriminator, use it as a sample weight with clipping, and confirm the reweighted model improves on the deployment slice rather than in aggregate.

The reason for that specificity is a failure I have seen: A model trained on a population 68% over age 40 and deployed to one 61% under 30 lost 7.4 points of AUC on the deployment cohort while the aggregate figure moved by 0.9.

Importance weighting estimated from a train-versus-deploy discriminator.

# Discriminator separates training rows from deployment rows.
p = clf.predict_proba(X_train)[:, 1]        # P(row is from deployment)
w = np.clip(p / (1 - p), 0.05, 20)          # density ratio, clipped
model.fit(X_train, y_train, sample_weight=w)

I would not consider it settled without evidence: Report performance on a labelled sample from the deployment population specifically, with and without the weighting.

Averaged over the wrong population, a good model looks fine and behaves badly.

Curated: · Written: · Reviewed:

QA-50You can afford 5,000 new labels. Which rows do you send to annotators?(show answer)

I would start by writing down what the model is allowed to see at the moment active learning and where to spend the labelling budget is decided.

Spend a labelling budget where the model is uncertain and the region is dense, because labelling rows the model already predicts confidently buys almost nothing.

Concretely, rank unlabelled rows by predictive uncertainty, filter to regions with real traffic volume, mix in a random sample to keep the evaluation set unbiased, and re-rank after each labelling round.

The reason for that specificity is a failure I have seen: A budget of 5,000 randomly sampled labels raised validation AUC by 0.004, while the same budget spent on an uncertainty-ranked sample with a 20% random component raised it by 0.021.

Same 5,000-label budget, two selection strategies.

StrategyAUC gainCost per point of AUC
Random sample+0.004$18,750
Uncertainty, 80% plus 20% random+0.021$3,570

I would not consider it settled without evidence: Run both selection strategies on the same budget and compare the improvement per 1,000 labels.

Random labelling pays annotators to confirm what the model already knows.

Curated: · Written: · Reviewed:

QA-51Two annotators disagree on 18% of examples. What does that mean for your model?(show answer)

This is an area where the offline result and the deployed result for label noise and annotator agreement routinely disagree.

Treat inter-annotator agreement as the ceiling on achievable accuracy, because a model cannot be more right than the labels it is measured against.

Concretely, measure agreement on a shared sample with a chance-corrected statistic, resolve the guideline ambiguity that causes systematic disagreement, and use multiple annotations with adjudication for the evaluation set even when training labels stay single-pass.

The reason for that specificity is a failure I have seen: A team chased the last 6 points of accuracy for four months against labels whose annotators agreed on 82% of cases, so the target was above the measurement ceiling the whole time.

Agreement on 500 shared examples sets the ceiling.

PairRaw agreementCohen kappa
A and B82%0.61
A and C79%0.57
B and C84%0.65
model target88%above ceiling

I would not consider it settled without evidence: Have three annotators label 500 shared examples, report the chance-corrected agreement, and state it as the ceiling on every subsequent accuracy claim.

Above the agreement ceiling, extra accuracy is fitting the noise in the labels.

Curated: · Written: · Reviewed:

QA-52Product wants to detect a 1% relative lift in conversion. How long does the test need to run?(show answer)

My answer to sizing an A/B test before running it begins with the label, because everything downstream inherits how it was defined.

Compute the sample size from the baseline rate, the smallest effect worth acting on, and the accepted error rates before starting, because a test that cannot detect the effect will produce a null result that means nothing.

Concretely, fix the baseline conversion, the minimum detectable effect, alpha and power, compute the required sample per arm, divide by daily eligible traffic, and refuse to start when the required duration exceeds what the team will actually wait.

The reason for that specificity is a failure I have seen: A test powered to detect a 5% relative lift was used to declare a 1% improvement not significant, and the feature was cancelled on a result the design could never have found.

Sample size per arm at 3.2% baseline conversion, alpha 0.05, power 0.8, two arms.

Relative liftAbsolute liftSample per armDays at 20,000 per arm per day
5%0.160 pp194,50010
2%0.064 pp1,198,60060
1%0.032 pp4,771,500239

I would not consider it settled without evidence: State required sample per arm and the resulting run length in the test plan, and record the minimum detectable effect the test can actually resolve.

A null result from an underpowered test is silence, not evidence of no effect.

Curated: · Written: · Reviewed:

QA-53A stakeholder checks the dashboard daily and wants to stop as soon as it is significant. What do you tell them?(show answer)

I would treat peeking at results and sequential testing as a measurement problem before treating it as a modelling problem.

Fix the analysis plan before the test starts or use a method designed for continuous monitoring, because repeatedly testing at 0.05 inflates the false-positive rate far beyond 5%.

Concretely, either fix a single analysis point and enforce it, or adopt a sequential procedure such as an alpha-spending boundary or always-valid confidence sequences, and set the dashboard to display only what the chosen method licenses.

The reason for that specificity is a failure I have seen: Daily peeking across a 30-day test raised the false-positive rate from 5% to roughly 28%, and three of the year's eleven declared wins failed to replicate.

Simulated A/A tests under three stopping rules, 1,000 runs each.

Stopping ruleDeclared significantNominal alpha
Fixed horizon, one look5.1%5%
Daily peek, 30 looks28.4%5%
Alpha-spending boundary5.3%5%

I would not consider it settled without evidence: Simulate 1,000 A/A tests under the team's actual stopping behaviour and report the empirical false-positive rate.

The 5% is a property of the procedure, and daily peeking is a different procedure.

Curated: · Written: · Reviewed:

QA-54Your ranking model wins on click-through. What else must the experiment measure?(show answer)

The useful framing for guardrail metrics in a model experiment is what a user loses when the model is wrong in each direction.

Declare guardrail metrics before the test and treat a guardrail breach as a stop condition, because a model can win its target metric by damaging something the target does not observe.

Concretely, name a small set of guardrails covering revenue, latency, complaint rate, and a long-horizon engagement measure, set explicit bounds, and wire the bounds into the rollout decision rather than into a report.

The reason for that specificity is a failure I have seen: A ranker that lifted click-through 4.2% raised page latency by 180 ms at p95 and cut completed purchases by 1.9%, and it ran for three weeks because only clicks were on the dashboard.

Primary metric with its guardrails and the bound each carries.

MetricEffectBoundVerdict
Click-through+4.2%primarywin
Completed purchases-1.9%-0.5%breach
p95 latency+180 ms+50 msbreach
Complaint rate+0.1%+0.5%ok

I would not consider it settled without evidence: Report every guardrail with its bound in the same table as the primary metric, and mark the decision each one licensed.

A metric that goes up while the business goes down was never the objective.

Curated: · Written: · Reviewed:

QA-55A new recommender shows a 9% lift in week one and 2% in week four. Which number do you report?(show answer)

I would settle novelty and primacy effects in model launches against a baseline first, so the added complexity has to earn its place.

Report the effect after the novelty response has decayed, because the first week measures curiosity about a change as much as the value of the change.

Concretely, run long enough to see the effect stabilise, plot the effect by days since exposure, and use a holdback cohort to estimate the long-run effect rather than extrapolating from the launch week.

The reason for that specificity is a failure I have seen: A launch reported on week-one data claimed 9.1% lift, and the annualised business case built on it overstated the benefit by roughly $2.3M when the effect settled at 2.0%.

Effect by days since exposure for the same launch.

WeekLiftUsers in cohort
19.1%240,000
25.4%238,000
32.8%236,000
42.0%235,000

I would not consider it settled without evidence: Plot the treatment effect against days since first exposure and show the point at which the series flattens.

The first week measures the change, and the fourth week measures the model.

Curated: · Written: · Reviewed:

QA-56You are testing a pricing model in a two-sided marketplace. Why might a user-level split be invalid?(show answer)

The judgement in interference between experiment arms is mostly about which distribution the evaluation actually samples from.

Randomise at the level where interference stops when treated and control units share a resource, because a user-level split in a marketplace lets the treatment change what the control group sees.

Concretely, randomise by region, by time slice, or by cluster when supply is shared, accept the loss of power that comes with fewer randomisation units, and check for spillover by comparing control behaviour against a pre-period baseline.

The reason for that specificity is a failure I have seen: A user-level pricing test measured a 5% lift that vanished on a regional split, because treated buyers were consuming the same limited inventory the control group needed.

Same intervention, three randomisation units, three answers.

UnitMeasured liftInterference riskUnits available
User+5.0%high2,400,000
Region+0.4%low38
Time slice+0.7%medium168

I would not consider it settled without evidence: Compare the control arm against a pre-experiment baseline and treat any drift in control as evidence of interference.

When the arms compete for the same supply, the control group is part of the treatment.

Curated: · Written: · Reviewed:

QA-57Your test needs 29 days to reach power. How do you shorten it without lowering the bar?(show answer)

Where teams lose time on variance reduction with pre-experiment data is usually the data step before the training step they are debating.

Reduce variance with a pre-experiment covariate rather than accepting a longer test, because CUPED removes variation the treatment could not have caused.

Concretely, take each unit's pre-period value of the metric, regress the experiment metric on it, analyse the residual, and confirm on A/A data that the adjustment is unbiased before using it for decisions.

The reason for that specificity is a failure I have seen: A team ran 29-day tests for a year at 12 tests a year, when a pre-period covariate correlated at 0.62 would have cut the required duration to 18 days and allowed 20.

CUPED adjustment using the 14 days before exposure.

theta = np.cov(y, y_pre)[0, 1] / np.var(y_pre)   # 0.58
y_adj = y - theta * (y_pre - y_pre.mean())
# variance 0.0412 -> 0.0159, required n falls 61%

I would not consider it settled without evidence: Estimate the correlation between the pre-period and in-period metric and report the variance reduction achieved on completed tests.

Variation that existed before the treatment is variation the treatment should not be charged for.

Curated: · Written: · Reviewed:

QA-58Your experiment report shows 40 metrics across 12 segments. What is wrong with reading it?(show answer)

I would answer multiple comparisons across metrics and segments by separating what the model guarantees from what it merely did on one sample.

Control the error rate across the whole family of comparisons rather than testing each one at 0.05, because 480 independent tests at 0.05 produce roughly 24 false positives by construction.

Concretely, nominate one primary metric in advance, treat everything else as exploratory, and apply a false-discovery-rate correction to the exploratory set before anyone reads it as a result.

The reason for that specificity is a failure I have seen: A report of 480 segment-metric cells produced 26 significant cells, and the team spent six weeks explaining a segment effect that did not reproduce.

480 cells before and after false-discovery-rate control at q = 0.05.

AnalysisSignificant cellsExpected false positives
Uncorrected at 0.052624
Benjamini-Hochberg q=0.053under 1

I would not consider it settled without evidence: Recompute the exploratory set with a Benjamini-Hochberg correction and show how many findings survive it.

Enough comparisons will always find something, and that is arithmetic rather than insight.

Curated: · Written: · Reviewed:

QA-59The model wins in every segment and loses overall. How is that possible, and what do you report?(show answer)

The engineering content of Simpson's paradox in segment reporting is the cost of being wrong, not the elegance of the estimator.

Check whether the treatment changed the mix of segments before trusting either the aggregate or the per-segment result, because a shift in composition can reverse the direction of an effect.

Concretely, report segment sizes alongside segment effects, test whether the treatment altered segment shares, and use a standardised aggregate that holds the mix fixed when it did.

The reason for that specificity is a failure I have seen: A model that improved conversion within every device class reduced overall conversion by 0.8%, because it drove 14% more traffic to a mobile segment that converts at less than half the desktop rate.

Wins in both segments, loses overall, because the mix moved.

SegmentControl conv.Treatment conv.Control shareTreatment share
Desktop5.0%5.2%60%46%
Mobile2.0%2.1%40%54%
Overall3.80%3.53%100%100%

I would not consider it settled without evidence: Recompute the aggregate effect holding segment shares at their control values and report both figures.

An aggregate is a weighted average, and the treatment can move the weights.

Curated: · Written: · Reviewed:

QA-60Your model beats the incumbent by 0.6 points of AUC. Is that a real improvement?(show answer)

Before touching an architecture I would fix what a good outcome for confidence intervals rather than point estimates means in product terms.

Report an interval on every comparison, because a point estimate cannot distinguish an improvement from resampling noise.

Concretely, bootstrap the test set to get an interval on the difference between paired models, use the paired difference rather than two separate intervals, and require the interval to exclude zero before claiming a win.

The reason for that specificity is a failure I have seen: A 0.6-point AUC claim had a bootstrap interval from -0.4 to +1.6 points, and the replacement model performed indistinguishably from the incumbent for the eight months it ran.

Paired bootstrap of the AUC difference on 40,000 test rows.

diffs = [auc(y[i], b[i]) - auc(y[i], a[i])
         for i in (rng.integers(0, n, n) for _ in range(2000))]
np.percentile(diffs, [2.5, 97.5])     # array([-0.004, 0.016])

I would not consider it settled without evidence: Bootstrap the paired difference on the test set with 2,000 resamples and report the 95% interval on the difference.

A difference without an interval is a number rather than a finding.

Curated: · Written: · Reviewed:

QA-61You cannot randomise the rollout. How do you estimate the effect anyway?(show answer)

The first question I would ask about observational estimates when a test is impossible is which production decision the number is meant to change.

Name the identifying assumption that makes an observational estimate causal and state what would violate it, because an unadjusted before-and-after comparison attributes every concurrent change to the model.

Concretely, use a design with a stated assumption such as difference-in-differences with a parallel-trends check, a regression discontinuity at an existing cut-off, or a synthetic control built from untreated units, and test the assumption where the data allows.

The reason for that specificity is a failure I have seen: A before-and-after comparison credited a model with a 12% lift that a difference-in-differences estimate against untreated regions put at 1.4%, the rest being a seasonal rise both groups shared.

Difference-in-differences against untreated regions.

PeriodTreated regionsUntreated regionsDifference
Pre4.10%4.05%0.05 pp
Post4.59%4.48%0.11 pp
Naive before-after+12.0%—overstated
Difference-in-differences+1.4%—estimate

I would not consider it settled without evidence: Show the pre-period trends of treated and comparison units on the same axis and confirm they moved together before treatment.

Without a comparison group, the estimate credits the model with the whole of the calendar.

Curated: · Written: · Reviewed:

QA-62Which test do you use to compare two classifiers on the same test set?(show answer)

I would start by writing down what the model is allowed to see at the moment hypothesis testing for model comparison is decided.

Use a paired test that accounts for the two models seeing identical examples, because treating the two score sets as independent samples throws away the pairing and overstates the variance.

Concretely, use McNemar's test on the discordant pairs for classification decisions, a paired bootstrap for threshold-free metrics, and report the number of discordant examples so the reader can judge the evidence.

The reason for that specificity is a failure I have seen: An unpaired two-sample test declared a difference not significant on 40,000 rows where the paired test on 812 discordant pairs gave an uncorrected chi-square of 12.3 and a clear difference at p = 0.0004.

McNemar's discordance table on 40,000 shared test rows.

B correctB wrong
A correct33,180356
A wrong4566,008

I would not consider it settled without evidence: Report the discordance table with the counts in each off-diagonal cell alongside the p-value.

The pairing is information, and an unpaired test discards it.

Curated: · Written: · Reviewed:

QA-63How do you know that a year of model launches actually helped?(show answer)

This is an area where the offline result and the deployed result for long-run holdback populations routinely disagree.

Keep a permanent untreated holdback to measure the cumulative effect of all launches, because individually significant short tests can sum to nothing measurable.

Concretely, hold a small fixed share of users out of every model change, refresh it rarely, and compare it against the treated population annually with the same metrics used for the individual tests.

The reason for that specificity is a failure I have seen: Eleven launches claiming a combined 14% lift measured 3.1% against a 1% holdback after a year, and the gap changed how the team sized future claims.

Claimed effects against the holdback comparison after 12 months.

SourceEffect
Sum of 11 launch claims+14.0%
Measured against 1% holdback+3.1%
Attributable to novelty and overlapapproximately 10.9 pp

I would not consider it settled without evidence: Report the holdback comparison annually beside the sum of the individually claimed effects.

The sum of the wins and the measured total are different numbers, and only one of them is observed.

Curated: · Written: · Reviewed:

QA-64Do you need an online scoring service, or is a nightly batch enough?(show answer)

My answer to batch versus real-time inference begins with the label, because everything downstream inherits how it was defined.

Choose batch scoring when the prediction is still valid at the moment it is used, because an online path adds latency budgets, autoscaling, and a second failure domain for no benefit if the inputs change slowly.

Concretely, compare the rate at which the input features change against the freshness the decision requires, batch when the answer would be identical, and score online when a feature computed within the session changes the prediction materially.

The reason for that specificity is a failure I have seen: A team built an online scoring service for a churn model whose features updated daily, adding a 24-hour on-call rotation and $6,400 a month for predictions identical to the nightly batch.

Do the two paths ever disagree on the decision itself?

Use caseFeature change rateDecisions differingVerdict
Churn outreachdaily0.3%batch
Fraud at checkoutper event41.0%online
Search rankingper session22.0%online

I would not consider it settled without evidence: Score a week of decisions both ways and report the share of decisions where the two answers differ.

Freshness the decision cannot use is cost without an effect.

Curated: · Written: · Reviewed:

QA-65The page has a 200 ms budget. How much of it belongs to the model?(show answer)

I would treat latency budgets for a model in a request path as a measurement problem before treating it as a modelling problem.

Allocate the model an explicit share of the end-to-end budget including feature retrieval, because the model's own forward pass is usually the smaller half of what the request spends.

Concretely, measure feature lookup, preprocessing, inference, and post-processing separately at p50 and p99, hold the total against the budget, and treat a breach as a design change rather than as a tuning exercise.

The reason for that specificity is a failure I have seen: A model measured at 8 ms in isolation contributed 214 ms at p99 in production, because two feature lookups were serialised and one of them missed cache 34% of the time.

p99 breakdown of one scoring request against a 200 ms budget.

Stagep50p99Budget
Feature lookup A6 ms96 ms40 ms
Feature lookup B5 ms88 ms40 ms
Preprocessing2 ms8 ms20 ms
Inference8 ms22 ms60 ms
Total21 ms214 ms200 ms

I would not consider it settled without evidence: Publish a p99 breakdown of the scoring path by stage and hold each stage to its own bound.

The forward pass is the part everyone measures and rarely the part that costs.

Curated: · Written: · Reviewed:

QA-66Your GPU sits at 12% utilisation while requests queue. What do you change?(show answer)

The useful framing for dynamic batching on accelerators is what a user loses when the model is wrong in each direction.

Batch requests up to a bounded wait when the accelerator is underutilised, because per-request inference leaves most of the device idle while adding queueing delay anyway.

Concretely, set a maximum batch size and a maximum wait, tune both against the p99 latency target rather than throughput alone, and confirm the tail does not regress when traffic is light and batches fill slowly.

The reason for that specificity is a failure I have seen: A 50 ms batching window raised throughput 6.4 times and pushed p99 latency to 240 ms during off-peak hours, because a batch of two waited the full window for a third request that never came.

Batch window against throughput and tail latency at 40 requests per second.

Max waitMax batchThroughputp99 latencyGPU utilisation
0 ms1240/s41 ms12%
10 ms321,180/s58 ms61%
50 ms641,540/s240 ms74%

I would not consider it settled without evidence: Measure throughput and p99 latency across the traffic range at each candidate window, including the quietest hour.

Batching trades tail latency for throughput, and the tail is worst exactly when traffic is thin.

Curated: · Written: · Reviewed:

QA-67Can you cache model outputs, and for how long?(show answer)

I would settle caching predictions safely against a baseline first, so the added complexity has to earn its place.

Key a prediction cache on every input that changes the output including the model version, because a cache keyed on the entity id alone serves the previous model's answer after a deployment.

Concretely, build the key from the entity id, the feature-set version, and the model registry id, set the time to live from the freshness the decision needs, and invalidate on model promotion rather than waiting for expiry.

The reason for that specificity is a failure I have seen: A cache keyed on user id alone served predictions from the retired model for 41% of requests during the first six hours after a promotion, and the rollout metrics measured a blend of two models.

Cache key composition and the time to live each part licenses.

key = sha1(user_id | feature_set_v7 | model:churn-v14 | segment)
ttl = 15m          # bounded by how fast sessions_7d moves
invalidate_on = model_promotion, feature_set_bump

I would not consider it settled without evidence: Deploy a new model version and confirm the cache hit rate for the retired version falls to zero within the stated invalidation window.

A cache key that omits the model version turns a deployment into a slow blend.

Curated: · Written: · Reviewed:

QA-68Traffic triples at 09:00 and p99 spikes for four minutes. What is happening?(show answer)

The judgement in autoscaling a model service and cold starts is mostly about which distribution the evaluation actually samples from.

Size the scaling policy around how long a replica takes to become useful, because a model container that loads several gigabytes of weights cannot answer a scale-up signal within one scrape interval.

Concretely, measure time from scheduling to first successful prediction, pre-warm capacity ahead of known traffic curves, keep a warm pool for the accelerator path, and scale on queue depth rather than on processor utilisation.

The reason for that specificity is a failure I have seen: A model server took 96 seconds to load 4.2 GB of weights and pass its readiness check, so a 09:00 traffic step produced four minutes of 4-second p99 latency every weekday.

Where the 96 seconds before a replica is useful goes.

StageDuration
Image pull, cached4 s
Weight load, 4.2 GB from object store71 s
Warmup passes for the kernel cache18 s
Readiness probe passes3 s

I would not consider it settled without evidence: Measure the time from pod start to first successful prediction and compare it against the time the traffic curve takes to double.

Scaling that is slower than the traffic step is a queue with extra steps.

Curated: · Written: · Reviewed:

QA-69Does this model need a GPU in production?(show answer)

Where teams lose time on GPU versus CPU serving decisions is usually the data step before the training step they are debating.

Decide accelerator use from measured cost per thousand predictions at the required latency, because a GPU that sits idle between small requests is more expensive per prediction than a CPU fleet.

Concretely, benchmark both at the production batch size and arrival rate, include idle time and reservation cost in the comparison, and revisit the decision when the model is quantised or distilled.

The reason for that specificity is a failure I have seen: A GPU fleet at 12% utilisation cost $9,400 a month for a model that ran on eight CPU cores at $610 a month with p99 latency inside the same budget.

Same model, same p99 target, measured at the real arrival rate of 40 requests per second.

Deploymentp99 latencyUtilisationMonthly cost
1 GPU instance22 ms12%$9,400
8 CPU cores48 ms68%$610
8 CPU cores, int826 ms51%$610

I would not consider it settled without evidence: Report cost per 1,000 predictions and p99 latency for both options at the actual arrival rate rather than at saturation.

A benchmark at full batch measures the hardware, and the invoice measures the traffic.

Curated: · Written: · Reviewed:

QA-70Your vector index returns results in 4 ms. What did you give up?(show answer)

I would answer approximate nearest neighbour search for embeddings by separating what the model guarantees from what it merely did on one sample.

Measure recall against exact search when using an approximate index, because the speed comes from not looking everywhere and the loss is invisible without a ground-truth comparison.

Concretely, build an exact ground truth on a sample of queries, measure recall at k against it, and tune the index parameters to the recall the product needs rather than to the fastest configuration.

The reason for that specificity is a failure I have seen: An index tuned for latency returned 71% recall at 10 against exact search, and the missing 29% were the long-tail items the recommendation surface existed to promote.

HNSW parameter sweep against exact search on 1,000 queries.

efSearchRecall at 10p99 latency
320.714 ms
960.939 ms
2560.9921 ms
exact1.00380 ms

I would not consider it settled without evidence: Run 1,000 sampled queries through both exact and approximate search and report recall at the k the product renders.

An approximate index is a recall setting, and shipping it untested is choosing a recall you did not measure.

Curated: · Written: · Reviewed:

QA-71A feature must reflect events from the last 60 seconds. What does that cost you?(show answer)

The engineering content of streaming feature pipelines and freshness is the cost of being wrong, not the elegance of the estimator.

Treat feature freshness as an explicit requirement with its own budget and monitoring, because a stale feature degrades predictions silently and no error is raised anywhere.

Concretely, measure event-time to availability at p99, alert on freshness lag rather than on job success, and have the serving path fall back to a documented default when a feature exceeds its staleness bound.

The reason for that specificity is a failure I have seen: A streaming aggregation fell 40 minutes behind for two days while every job reported success, and the model served on stale counts with no alert and a 3.2% drop in precision.

Freshness budget per feature and the fallback each carries.

FeatureRequired freshnessp99 ageFallback
clicks_60s60 s48 szero, flagged
sessions_7d24 h6 hlast known
account_agenonestaticnone needed

I would not consider it settled without evidence: Emit the age of every feature at scoring time and alert when the p99 age exceeds the stated bound.

A pipeline that is late but green is the failure mode that lasts longest.

Curated: · Written: · Reviewed:

QA-72What should stop a training run before it starts?(show answer)

Before touching an architecture I would fix what a good outcome for data validation gates in a training pipeline means in product terms.

Validate the training data against an explicit schema and distribution expectations before training, because a model trained on corrupted data will train successfully and fail silently.

Concretely, assert column presence and types, row-count bounds against the previous run, null rates per column, and category-set drift, and fail the pipeline rather than warning when a check breaks.

The reason for that specificity is a failure I have seen: An upstream schema change made a currency column arrive as text, the pipeline coerced it to nulls, and a model trained on 100% nulls in a top-three feature was promoted before anyone noticed.

Gates that run before a single gradient step.

expect_columns(df, SCHEMA_V7)                       # types included
expect_rows_between(df, 0.85 * last_n, 1.15 * last_n)
expect_null_rate(df, "amount_usd", max=0.02)        # was 1.00
expect_category_overlap(df, "country", prior, min=0.95)

I would not consider it settled without evidence: Run the validation suite against the last 30 training snapshots and confirm it would have failed the run that shipped the defect.

A training job that cannot fail on bad data will publish it instead.

Curated: · Written: · Reviewed:

QA-73You add a feature today and want two years of history for training. What has to be true?(show answer)

The first question I would ask about backfilling a new feature correctly is which production decision the number is meant to change.

Reconstruct a backfilled feature from the event log as it existed historically rather than from current state, because a backfill computed from today's tables imports information the past did not have.

Concretely, replay the immutable event log through the same transformation used online, stamp each computed value with its validity interval, and verify the backfilled values against a period where the online values also exist.

The reason for that specificity is a failure I have seen: A backfill built from current dimension tables encoded merchant categories that were reassigned in 2025, giving the model 2024 rows with 2026 categories and 3.9 points of illusory AUC.

Overlap check before trusting two years of reconstructed history.

MonthRows comparedExact matchVerdict
Overlap month412,00099.4%accept
Prior attempt412,00071.2%rejected

I would not consider it settled without evidence: Compare backfilled values against online values for an overlapping month and require agreement above a stated threshold.

A backfill that uses today's tables tells the model what happens next.

Curated: · Written: · Reviewed:

QA-74A training task retries after a partial write. What must be true for that to be safe?(show answer)

I would start by writing down what the model is allowed to see at the moment orchestration, retries, and idempotent training jobs is decided.

Make every pipeline task idempotent so a retry produces the same result as a first run, because an orchestrator will retry and a partially written dataset is worse than a failed one.

Concretely, write to a temporary location and promote atomically, key outputs by the run id and input snapshot, and make the promotion the only step that makes a result visible to downstream tasks.

The reason for that specificity is a failure I have seen: A retried feature job appended a second copy of one day's rows, duplicating 1.4% of the training set and shifting the class balance enough to move the chosen threshold.

Write, then promote, so a retry cannot leave a partial dataset visible.

tmp = f"s3://features/_staging/{run_id}/"
write_parquet(df, tmp)                       # safe to repeat
assert row_count(tmp) == expected            # gate before promotion
atomic_rename(tmp, f"s3://features/date={ds}/")

I would not consider it settled without evidence: Run a task twice against the same input snapshot and assert the output is byte-identical or has an identical content hash.

Retries are guaranteed, and only idempotence makes them harmless.

Curated: · Written: · Reviewed:

QA-75Who signs off on a $40,000 training run, and on what evidence?(show answer)

This is an area where the offline result and the deployed result for the cost of a training run routinely disagree.

Estimate and cap the cost of a training run before launching it, because compute spent on a run nobody sized is spent whether or not the run was ever going to help.

Concretely, extrapolate cost from a short pilot at a fraction of the data, set a hard budget cap and a checkpointing interval, and require an expected-benefit statement for runs above a stated threshold.

The reason for that specificity is a failure I have seen: A run budgeted at $6,000 reached $41,300 because the learning rate schedule was set for a dataset four times smaller, and it was stopped after the invoice arrived rather than by a cap.

Pilot at 5% and the extrapolation the approval rests on.

ScaleGPU hoursCostValidation AUC
5% pilot9$2900.812
Extrapolated 100%180$5,8000.84 estimated
Actual, no cap1,290$41,3000.843

I would not consider it settled without evidence: Run 5% of the data first, extrapolate the cost and the expected metric gain, and record both before the full run is approved.

An uncapped training run is an open-ended purchase order.

Curated: · Written: · Reviewed:

QA-76Training takes 40 hours on one machine. Do you shard the data or the model?(show answer)

My answer to distributed training and which axis to split begins with the label, because everything downstream inherits how it was defined.

Split along data when the model fits in one device's memory and along the model only when it does not, because model parallelism adds communication on every forward and backward pass.

Concretely, use data parallelism with gradient synchronisation while the model fits, watch for the batch size at which convergence degrades, and move to sharded optimiser state or pipeline parallelism only when memory forces it.

The reason for that specificity is a failure I have seen: A team adopted model parallelism for a 3 GB model that fitted on every device, and the added communication made an 8-device run 1.7 times slower than a single device.

Scaling efficiency for a 3 GB model that fits on one device.

DevicesData parallelModel parallel
140.0 h40.0 h
221.0 h46.0 h
411.5 h58.0 h
86.8 h68.0 h

I would not consider it settled without evidence: Measure scaling efficiency at 2, 4, and 8 devices and report the point at which added devices stop buying wall-clock time.

Parallelism is a memory remedy first and a speed remedy second.

Curated: · Written: · Reviewed:

QA-77You need per-user counts over 400 million events. Why does the data structure choice matter here?(show answer)

I would treat hash maps for aggregation in a feature pipeline as a measurement problem before treating it as a modelling problem.

Aggregate with a hash map keyed by the grouping column rather than by scanning per user, because a nested scan turns a linear job into a quadratic one at production volume.

Concretely, build one pass that accumulates into a dictionary keyed by user id, bound memory by partitioning on a hash of the key when the map will not fit, and prefer a sketch such as HyperLogLog when only a cardinality estimate is needed.

The reason for that specificity is a failure I have seen: A nested loop over 400 million events and 2.1 million users ran for 14 hours and was killed, while a single hash pass finished in 9 minutes.

One pass over the event stream, memory bounded by partition.

counts = defaultdict(int)                 # O(n) time, O(users) memory
for user_id, _ in events:                 # 400M events
    counts[user_id] += 1
# 2.1M keys at ~120 bytes = ~250 MB; partition by hash(user_id) % 16 above that.

I would not consider it settled without evidence: State the complexity of the aggregation and confirm it against wall-clock time at 10% and 100% of the data.

At feature-pipeline volume, the difference between one pass and a nested scan is the difference between a job and an outage.

Curated: · Written: · Reviewed:

QA-78A retrieval stage must return the 200 best of 4 million scored candidates per request. How do you do it?(show answer)

The useful framing for top-k selection with a heap for candidate generation is what a user loses when the model is wrong in each direction.

Select the top k with a bounded heap rather than sorting the full candidate set, because sorting does work proportional to the whole set to answer a question about 200 items.

Concretely, maintain a min-heap of size k while streaming scores, push and pop in one operation once the heap is full, and pair this with an approximate index when even scoring every candidate is too expensive.

The reason for that specificity is a failure I have seen: Sorting 4 million candidates per request cost 310 ms at p99, where a size-200 heap answered the same query in 28 ms with identical results.

Bounded heap against a full sort on 4 million candidates.

import heapq
top = heapq.nlargest(200, scored, key=lambda c: c.score)  # O(n log k)
# full sort: O(n log n), 310 ms p99; heap: 28 ms p99, same 200 items

I would not consider it settled without evidence: Benchmark both against the same candidate set and confirm the returned sets are identical while the latency differs.

Sorting everything to look at the top 200 pays for 3,999,800 answers nobody reads.

Curated: · Written: · Reviewed:

QA-79A feature is the count of events in the last 5 minutes, updated per event. How is it computed efficiently?(show answer)

I would settle sliding windows over an event stream against a baseline first, so the added complexity has to earn its place.

Maintain a window incrementally with a deque rather than recomputing over the window on every event, because recomputation costs the window size on each of millions of events.

Concretely, push each arriving event, evict from the front while the oldest falls outside the window, and keep a running aggregate so the answer is available in constant time per event.

The reason for that specificity is a failure I have seen: Recomputing a 5-minute count per event cost 41 ms per event at 12,000 events per second and the consumer fell 40 minutes behind within two hours.

Incremental window, constant work per event.

window = deque()                          # (timestamp, value)
def add(ts, value):
    window.append((ts, value))
    while window and ts - window[0][0] > 300:   # 5 minutes
        window.popleft()
    return len(window)                    # amortised O(1) per event

I would not consider it settled without evidence: Measure per-event processing time at production event rates and confirm the consumer lag stays flat rather than growing.

A window recomputed from scratch charges the whole window for every single event.

Curated: · Written: · Reviewed:

QA-80You need the threshold that yields exactly 3,000 flags a day. How do you find it?(show answer)

The judgement in binary search over a monotone operating curve is mostly about which distribution the evaluation actually samples from.

Search a monotone curve by bisection rather than by sweeping every candidate, because flagged volume is monotone in the threshold and bisection reaches the answer in a logarithmic number of evaluations.

Concretely, confirm monotonicity on the validation scores, bisect on the threshold against the target volume with a stated tolerance, and store the resulting threshold with the model version so serving and evaluation agree.

The reason for that specificity is a failure I have seen: A linear sweep in steps of 0.001 evaluated 1,000 thresholds over 40 million scored rows nightly, taking 52 minutes where bisection took 14 evaluations and under a minute.

Fourteen evaluations instead of a thousand.

lo, hi = 0.0, 1.0
while hi - lo > 1e-4:                     # 14 iterations
    mid = (lo + hi) / 2
    if flagged_per_day(scores, mid) > 3000: lo = mid
    else: hi = mid
threshold = (lo + hi) / 2                 # 0.9103

I would not consider it settled without evidence: Compare the threshold found by bisection against a fine sweep on one day and confirm the flagged volume matches within tolerance.

A monotone curve does not need to be walked end to end.

Curated: · Written: · Reviewed:

QA-81Duplicate customer records are linked pairwise. How do you turn those links into entities?(show answer)

Where teams lose time on graph traversal for entity resolution is usually the data step before the training step they are debating.

Treat pairwise matches as edges and resolve entities as connected components, because transitive duplicates are invisible to any procedure that only looks at pairs.

Concretely, build the match graph, find components with union-find or breadth-first search, and cap component size so a single false-positive edge cannot merge two large clusters into one entity.

The reason for that specificity is a failure I have seen: One false match between two common names merged a component of 41,000 records into a single customer, and the resulting features described a customer with 41,000 addresses.

Component sizes reveal the merge that pairwise precision cannot.

Component sizeCountAction
2-4214,000accept
5-203,100accept
21-50062review
41,0001split, bad edge

I would not consider it settled without evidence: Report the component size distribution and inspect every component above a stated size before promoting the resolution.

In a match graph, a single wrong edge is not a wrong pair but a wrong entity.

Curated: · Written: · Reviewed:

QA-82You need near-duplicate detection across 2 million product titles. Where does edit distance fit?(show answer)

I would answer dynamic programming for sequence comparison by separating what the model guarantees from what it merely did on one sample.

Reserve the quadratic sequence-comparison step for candidates that a cheap blocking step has already selected, because comparing every pair is quadratic in the number of records.

Concretely, block candidates by a cheap key such as a token-based hash or a shingle signature, then apply edit distance or its length-normalised form only within blocks, and set the acceptance threshold from a labelled sample.

The reason for that specificity is a failure I have seen: An all-pairs edit distance over 2 million titles implied 2 trillion comparisons and was abandoned after 30 hours, while blocking by 3-gram signatures reduced it to 41 million and finished in 18 minutes.

Blocking before the quadratic step.

StageComparisonsWall clockRecall of true pairs
All pairs2.0e12abandoned1.00
3-gram blocking4.1e718 min0.97
Edit distance within blocks4.1e7included0.97

I would not consider it settled without evidence: Report candidate pairs generated by blocking and the recall of that blocking step against a hand-labelled sample of true duplicates.

The expensive comparison is affordable only on the pairs something cheaper has already shortlisted.

Curated: · Written: · Reviewed:

QA-83Autocomplete must respond in under 20 ms over 8 million queries. What structure serves it?(show answer)

The engineering content of prefix structures for query autocomplete is the cost of being wrong, not the elegance of the estimator.

Serve prefix lookups from a trie or a compressed prefix index rather than from a scan or a wildcard database query, because prefix matching is the operation the structure is built for.

Concretely, store the top completions at each node so a lookup is a walk plus a read, refresh the popularity counts on a schedule, and cap memory with a compressed representation when the vocabulary is large.

The reason for that specificity is a failure I have seen: A leading-wildcard SQL query over 8 million rows could not use an index and took 340 ms at p99, against 3 ms from a trie with precomputed completions.

Latency by approach for the same 8-million-query vocabulary.

Approachp99 latencyMemory
SQL LIKE with leading wildcard340 ms0
Trie with top-10 per node3 ms1.9 GB
Compressed prefix index6 ms410 MB

I would not consider it settled without evidence: Benchmark p99 latency for the ten most common prefix lengths against the production vocabulary.

A leading wildcard is a full scan wearing query syntax.

Curated: · Written: · Reviewed:

QA-84Events must be grouped into sessions with a 30-minute inactivity gap. How do you compute that at scale?(show answer)

Before touching an architecture I would fix what a good outcome for interval merging for session windows means in product terms.

Sort by user and timestamp once and merge intervals in a single pass, because a session boundary depends only on the previous event for that user.

Concretely, partition by user, sort within the partition, start a new session whenever the gap exceeds the threshold, and emit the session id with the event so downstream features inherit it.

The reason for that specificity is a failure I have seen: A self-join to find each event's predecessor over 900 million events produced a 40-minute shuffle and failed on skewed users with 400,000 events each.

Sessionisation as one pass over sorted events.

prev, session = None, 0
for ts in sorted_timestamps:              # per user partition
    if prev is not None and ts - prev > 1800: session += 1
    yield session
    prev = ts

I would not consider it settled without evidence: Compare the single-pass result against the self-join on one day of data and report the runtime difference.

A predecessor lookup is a sort, not a join.

Curated: · Written: · Reviewed:

QA-85You must join clicks and impressions, both sorted by time, without loading either into memory. How?(show answer)

The first question I would ask about merging sorted streams with two pointers is which production decision the number is meant to change.

Advance two cursors through sorted inputs rather than building a hash of one side, because a merge join needs memory proportional to the window rather than to the dataset.

Concretely, keep one cursor per stream, advance whichever is behind, buffer only the records within the join window, and assert both inputs are sorted before starting rather than trusting the producer.

The reason for that specificity is a failure I have seen: A hash join on 2.4 billion impressions exhausted 240 GB of memory, while a merge join over the same sorted inputs held 900 MB and finished in 22 minutes.

Merge join memory against hash join on the same inputs.

StrategyPeak memoryRuntime
Hash join on impressions240 GB, failed—
Merge join, 5-minute window0.9 GB22 min

I would not consider it settled without evidence: Report peak memory and runtime for both strategies at full data volume.

Sorted inputs are an asset, and hashing one side throws it away.

Curated: · Written: · Reviewed:

QA-86A pipeline that ran in 20 minutes on a sample takes 14 hours on the full data. What do you check?(show answer)

I would start by writing down what the model is allowed to see at the moment reading the complexity of a feature pipeline is decided.

Identify the operation whose cost grows faster than linearly before adding hardware, because a superlinear step will outrun any cluster you can buy.

Concretely, time the pipeline at several data fractions, fit the growth curve, and target the stage whose exponent is above one rather than distributing the whole job more widely.

The reason for that specificity is a failure I have seen: A team tripled cluster size for a job whose dominant stage was quadratic, cutting 14 hours to 11 while tripling cost, and the same stage rewritten as a hash join took 12 minutes.

Runtime by data fraction identifies the superlinear stage.

FractionStage AStage BTotal
10%2 min4 min6 min
25%5 min24 min29 min
50%10 min98 min108 min
growthlinearquadratic—

I would not consider it settled without evidence: Run at 10%, 25%, and 50% of the data and plot runtime against volume per stage rather than for the job as a whole.

More machines change the constant, and the exponent is what is hurting you.

Curated: · Written: · Reviewed:

QA-87The training set is 400 GB and the machine has 64 GB. What changes in your data code?(show answer)

This is an area where the offline result and the deployed result for generators for datasets larger than memory routinely disagree.

Stream training data through generators or an iterable dataset rather than materialising it, because loading everything is the only step that actually requires the memory.

Concretely, yield batches lazily from sharded files, prefetch a bounded number of batches on worker processes, and keep transformations inside the generator so no full copy of the data ever exists.

The reason for that specificity is a failure I have seen: A load of a 400 GB parquet set into a dataframe on a 64 GB machine failed after 25 minutes of swapping, and the same pipeline as a sharded generator held 4.1 GB resident.

Lazy sharded reader with bounded prefetch.

def batches(shards, size=1024):
    for shard in shards:                  # 400 GB across 4,000 files
        for chunk in pq.ParquetFile(shard).iter_batches(size):
            yield transform(chunk)        # peak RSS 4.1 GB, not 400 GB

I would not consider it settled without evidence: Report peak resident memory during a full epoch and confirm it is bounded by the prefetch depth rather than by the dataset size.

Memory should be a function of the batch, not of the corpus.

Curated: · Written: · Reviewed:

QA-88A feature transform takes 40 minutes over 12 million rows in a Python loop. What do you do?(show answer)

My answer to vectorisation instead of row-wise loops begins with the label, because everything downstream inherits how it was defined.

Express element-wise work as array operations rather than as a Python loop, because the interpreter overhead per row dominates the arithmetic by more than an order of magnitude.

Concretely, replace the loop with vectorised expressions over columns, use masks rather than conditionals per row, and reserve a compiled path for logic that genuinely cannot be expressed as array operations.

The reason for that specificity is a failure I have seen: A row-wise apply over 12 million rows took 41 minutes and blocked the nightly pipeline, while the vectorised form ran in 3.2 seconds and produced identical values.

Same transform, two implementations, identical output.

# 41 minutes over 12M rows
df["z"] = df.apply(lambda r: (r.x - MU) / SD if r.x > 0 else 0.0, axis=1)

# 3.2 seconds, same values
df["z"] = np.where(df["x"] > 0, (df["x"] - MU) / SD, 0.0)

I would not consider it settled without evidence: Assert the vectorised output equals the loop output on a sample, then compare wall-clock time at full volume.

Per-row Python is an interpreter benchmark wearing a data transformation.

Curated: · Written: · Reviewed:

QA-89Your GPU is idle 60% of the time during training. Where do you look in the Python layer?(show answer)

I would treat Python concurrency for data loading as a measurement problem before treating it as a modelling problem.

Use processes rather than threads for data preparation that runs in Python bytecode, because the global interpreter lock serialises exactly that work while releasing for input and output.

Concretely, move decoding and augmentation to worker processes with a bounded queue, use threads only for input and output waits, and size worker count from measured device utilisation rather than from core count.

The reason for that specificity is a failure I have seen: Eight loader threads left a GPU at 38% utilisation because image decoding is pure Python bytecode, and moving to four worker processes raised it to 91%.

Utilisation against loader configuration for the same model.

LoaderWorkersGPU utilisationEpoch time
Threads838%44 min
Processes271%24 min
Processes491%18 min
Processes1289%19 min

I would not consider it settled without evidence: Measure accelerator utilisation and input queue depth while varying worker count, and choose the point where the queue stays full.

Threads help when Python is waiting and not when Python is working.

Curated: · Written: · Reviewed:

QA-90How do you stop a feature's meaning from drifting between the training code and the serving code?(show answer)

The useful framing for typed feature contracts in Python is what a user loses when the model is wrong in each direction.

Define the feature vector as one typed structure that both paths import, because two independently maintained definitions will diverge and nothing will raise an error when they do.

Concretely, declare the fields as a dataclass or typed schema with units in the field names, construct it in one shared function used by training and serving, and validate ranges at construction rather than in each caller.

The reason for that specificity is a failure I have seen: Training built a duration field in seconds and serving built it in milliseconds, a factor of 1,000 that no test caught and that cost 19 AUC points in production.

One shared definition, units in the names, validated at construction.

@dataclass(frozen=True)
class Features:
    dwell_seconds: float          # not milliseconds
    price_usd: float
    is_returning: bool
    def __post_init__(self):
        assert 0 <= self.dwell_seconds <= 86_400

I would not consider it settled without evidence: Confirm both paths import the same constructor and add a test that fails when a field is added to one path only.

Two copies of a schema are a divergence with a start date.

Curated: · Written: · Reviewed:

QA-91You want timing and metric logging on every stage without editing each one. How?(show answer)

I would settle decorators and context managers for instrumenting training against a baseline first, so the added complexity has to earn its place.

Wrap cross-cutting concerns in a decorator or context manager rather than repeating them in each stage, because instrumentation copied into twenty functions will be missing from the twenty-first.

Concretely, wrap each stage in one decorator that records duration, input shape, and outcome, use a context manager to guarantee the record is written even when the stage raises, and keep the wrapper free of stage-specific logic.

The reason for that specificity is a failure I have seen: Hand-written timing calls covered 14 of 21 pipeline stages, and the unmeasured stages held the 40-minute regression the team spent two days looking for.

One decorator, every stage measured, failures recorded too.

def staged(fn):
    @wraps(fn)
    def run(*a, **kw):
        with timer() as t:
            try: return fn(*a, **kw)
            finally: log.info("stage=%s ms=%d", fn.__name__, t.ms)
    return run

I would not consider it settled without evidence: List the stages and confirm every one appears in the timing output of a single run.

Instrumentation that has to be remembered will be forgotten in the stage that matters.

Curated: · Written: · Reviewed:

QA-92A dataframe of 40 million rows uses 38 GB. How do you make it fit in 16 GB?(show answer)

The judgement in memory footprint of a dataframe is mostly about which distribution the evaluation actually samples from.

Choose column types from the range and cardinality each column actually holds, because default 64-bit numeric and Python-object string columns cost several times what the data needs.

Concretely, downcast integers and floats to the smallest type that holds the range, convert low-cardinality strings to a categorical or dictionary type, and load only the columns the job uses.

The reason for that specificity is a failure I have seen: A 38 GB dataframe held 11 low-cardinality string columns as Python objects at 26 GB, and converting them to categoricals brought the total to 9.4 GB with identical results.

Per-column memory before and after typing.

Column groupBeforeAfterChange
11 object strings26.0 GB0.6 GBcategorical
24 int647.7 GB2.9 GBint32 and int16
9 float644.3 GB2.2 GBfloat32
Total38.0 GB9.4 GBfits

I would not consider it settled without evidence: Report memory usage per column before and after, and assert the transformed frame produces identical model inputs.

Most of a dataframe is usually type defaults rather than data.

Curated: · Written: · Reviewed:

QA-93Your artefact is a pickle file. What is the objection?(show answer)

Where teams lose time on serialising a model for production is usually the data step before the training step they are debating.

Prefer a serialisation format that does not execute code on load for artefacts crossing a trust or version boundary, because unpickling runs arbitrary code and binds the artefact to the exact library versions.

Concretely, export to a defined format such as ONNX or the framework's own saved-model format, pin the runtime version in the registry entry, and load artefacts only from a registry the deployment pipeline controls.

The reason for that specificity is a failure I have seen: A pickle trained under one library version failed to load after a routine dependency upgrade, and the model could not be served for 9 hours because the training environment no longer existed.

What the format choice costs and buys.

FormatExecutes code on loadCross-versionCross-language
pickleyesnono
ONNXnoyesyes
SavedModelnomostlyyes

I would not consider it settled without evidence: Load every registered artefact in a clean container built from the pinned runtime and confirm it scores a fixed input to the recorded value.

An artefact that only loads in the environment that made it is a liability with a shelf life.

Curated: · Written: · Reviewed:

QA-94How do you get reproducible shuffling without training on the same order forever?(show answer)

I would answer randomness that is reproducible and still random by separating what the model guarantees from what it merely did on one sample.

Derive per-epoch randomness from one recorded root seed rather than fixing a single global seed or leaving it unset, because reproducibility needs the sequence recorded and training needs the order to vary.

Concretely, record the root seed in the run metadata, derive each epoch's and each worker's seed from it deterministically, and confirm two runs with the same root seed produce identical loss curves.

The reason for that specificity is a failure I have seen: Setting one global seed made every epoch shuffle identically, and the model saw the same batch composition 40 times, which cost 1.4 points of validation accuracy against per-epoch reseeding.

One root seed, derived per epoch and per worker.

root = 20260907                            # recorded in the run metadata
for epoch in range(40):
    seed = hash((root, epoch)) % 2**32     # differs per epoch
    rng = np.random.default_rng(seed)      # identical across reruns

I would not consider it settled without evidence: Run the same root seed twice and assert identical loss curves, then confirm batch composition differs across epochs within a run.

Reproducible and fixed are different properties, and only one of them helps training.

Curated: · Written: · Reviewed:

QA-95How do you check whether a model treats groups differently, and which definition do you use?(show answer)

The engineering content of measuring fairness across subgroups is the cost of being wrong, not the elegance of the estimator.

Choose one fairness criterion explicitly and state what it forgoes, because equal false-negative rates, equal precision, and equal selection rates cannot generally all hold at once when base rates differ.

Concretely, measure error rates per protected group at the deployed threshold, choose the criterion the use case requires and record the reasoning, and consider per-group thresholds or a constrained objective when a gap exceeds the stated bound.

The reason for that specificity is a failure I have seen: A model equalised selection rates across groups and left the false-negative rate 2.3 times higher in one group, which was the harm the review had been convened to prevent.

One threshold, three criteria, and the gap each leaves open.

GroupSelection rateFPRFNR
A11.0%4.1%18.0%
B11.0%9.6%41.0%
Gapequalised2.3x2.3x

I would not consider it settled without evidence: Publish a per-group table of selection rate, false-positive rate, and false-negative rate at the deployed threshold.

Fairness is a choice among incompatible definitions, and not choosing is still choosing one.

Curated: · Written: · Reviewed:

QA-96A colleague proposes using raw postcode and date of birth as features. What is your response?(show answer)

Before touching an architecture I would fix what a good outcome for personal data in features means in product terms.

Minimise the personal data a model consumes and record a lawful basis and retention period for what remains, because a feature store is a copy of personal data with a longer life than the source system.

Concretely, replace direct identifiers with coarsened or derived features such as an age band or a region, hold the mapping outside the training set, and propagate deletion requests into feature snapshots and training data rather than into the source table alone.

The reason for that specificity is a failure I have seen: A deletion request removed a customer from the source database while 14 monthly feature snapshots and three training sets retained the same identifiers, which the audit recorded as an unfulfilled erasure.

Coarsening keeps the signal and drops the identifier.

Raw featureReplacementSignal retainedRe-identification risk
date_of_birthage_band_5yhighlow
postcoderegion_of_3.2Mmediumlow
emailhashed_account_idnone neededlow

I would not consider it settled without evidence: Trace one deletion request end to end and confirm no snapshot, training set, or backup retains the identifiers past the stated window.

A feature store outlives the source row unless someone makes it stop.

Curated: · Written: · Reviewed:

QA-97What has to be written down before another team can consume your model's output?(show answer)

The first question I would ask about documenting a model for the people who depend on it is which production decision the number is meant to change.

Document the intended use, the population the model was trained on, the metric at the deployed threshold, and the known failure modes, because a consumer with no documentation will assume the model works everywhere.

Concretely, publish a short model card beside the registry entry naming the training population, the evaluation slices, the operating threshold, the known weak slices, and the contact for questions, and update it on every promotion.

The reason for that specificity is a failure I have seen: A team consumed a model trained on desktop web traffic for a mobile app surface where it scored 0.61 against 0.84, and the mismatch was found four months later during an unrelated investigation.

The minimum a consumer needs before depending on a score.

intended_use: rank support tickets for tier-2 triage
population: desktop web, EN and DE, 2025-01 to 2026-05
threshold: 0.91, chosen for 3,000 reviews per day
auc_overall: 0.844      # weakest slice, tenure<30d: 0.61
not_validated_for: mobile app, ES traffic, enterprise accounts

I would not consider it settled without evidence: Ask a consuming team to state the model's intended population and weak slices from the documentation alone.

An undocumented model will be used for whatever the consumer imagined it did.

Curated: · Written: · Reviewed:

QA-98Your service has uptime alerts. What else needs an alert for a model?(show answer)

I would start by writing down what the model is allowed to see at the moment what to monitor once a model is serving is decided.

Monitor prediction distribution, feature freshness, feature nulls, and delayed labelled performance in addition to service health, because a model can be entirely available and entirely wrong.

Concretely, alert on the score distribution shifting beyond a stated bound, on any feature exceeding its staleness or null budget, and on labelled performance falling below the floor once labels arrive, with a stated owner for each alert.

The reason for that specificity is a failure I have seen: A model served 100% availability for eleven days while a null-filled feature moved the mean predicted probability from 0.031 to 0.004, and nothing alerted because every request returned 200.

The four model alerts that uptime monitoring does not cover.

SignalBoundDetects
Mean predicted probability±30% of 30-day medianbroken feature
Feature p99 ageper-feature budgetstale pipeline
Null rate per feature+2 pp over baselineupstream schema change
Labelled AUC, 14-day lagabove 0.80 floorgenuine degradation

I would not consider it settled without evidence: Replay a known past incident through the alerting rules and confirm each rule fires and names the responsible signal.

Availability monitoring cannot see the difference between a prediction and a wrong prediction.

Curated: · Written: · Reviewed:

QA-99Predictions have been wrong for six hours. What are your first three actions?(show answer)

This is an area where the offline result and the deployed result for responding to a model incident routinely disagree.

Restore the previous model first and diagnose afterwards, because a model incident continues to produce bad decisions for as long as the investigation runs.

Concretely, roll back to the last known-good registry entry, quantify and list the affected decisions for remediation, then diagnose from the logged feature vectors of the affected window rather than by rerunning training.

The reason for that specificity is a failure I have seen: A team debugged a feature pipeline for six hours before rolling back, and 41,000 additional decisions were made on the faulty model during the investigation.

Two responses to the same fault.

Action orderTime to rollbackBad decisions served
Diagnose then roll back6 h 10 m41,000
Roll back then diagnose14 m1,600

I would not consider it settled without evidence: Record the time from detection to rollback and the count of decisions made after detection in every incident review.

Diagnosis is cheaper after the bleeding stops, and the decisions made meanwhile are permanent.

Curated: · Written: · Reviewed:

QA-100A model has served for three years. How do you decide whether to keep it?(show answer)

My answer to retiring a model begins with the label, because everything downstream inherits how it was defined.

Compare a running model against a current baseline periodically, because a model that was once an improvement can fall below a simple rule as the world moves.

Concretely, re-fit the cheap baseline on recent data on a stated schedule, compare it against the deployed model on the same window, and retire the model when it no longer beats the baseline by more than its maintenance cost justifies.

The reason for that specificity is a failure I have seen: A three-year-old model was 1.2 points of AUC below a freshly fitted logistic regression, while consuming a retraining pipeline, a GPU fleet, and a place in the on-call rotation.

Annual review of a three-year-old model.

CandidateAUC on last 90 daysMonthly cost
Deployed model, 20230.792$9,400
Fresh logistic regression0.804$120
Fresh boosted trees0.851$340

I would not consider it settled without evidence: Publish an annual comparison of the deployed model, a freshly fitted baseline, and the cost of keeping the deployed one.

A model earns its place every year, not once at launch.

Curated: · Written: · Reviewed: