Skip to content
Tech Interview Prep home
Technical interview guide

Model Monitoring & Drift Detection

Detecting when a production model's performance degrades because the world changed since it was trained.

Read
45 min
Practice MCQs
25
Interview QA
25
Edition
v4
Editorial status
Reviewed

Scope: Azure Machine Learning, Vertex AI, AWS SageMaker AI, NIST AI RMF, and Google SRE guidance current 2026-09-01; AWS states Model Monitor is unavailable to new customers.

Overview

Curated: · Written: · Reviewed:

Monitor the decision system, not just a distance statistic

Production model monitoring establishes whether the deployed system is observable, technically healthy, receiving valid inputs, behaving within expected distributions, meeting outcome and trustworthiness objectives, and remaining appropriate for its context. Drift detection is one part. A dashboard of divergence scores without coverage, ground truth, owners and response paths creates alert theater rather than risk control.

In an interview, this topic is where senior candidates separate from junior ones. Junior candidates name PSI and say "retrain when drift is detected." Staff-level candidates are expected to say which drift, measured against what baseline, with what label situation, and what action follows — and to know that most drift alarms do not justify retraining. The questions below are the ones interviewers actually ask, and the failure modes they are listening for.

The taxonomy, and the questions that discriminate between its members

Five things get called "drift" and each implies a different response. Interviewers probe whether you can tell them apart from the signals available:

  • Data-quality breakage — malformed or unusable inputs: missingness, type mismatch, invalid range, category or cardinality changes, duplicates, stale features, schema errors. Detectable without any statistical test; it's a contract check.
  • Training-serving skew — the same feature means different things at train time and serve time (unit mismatch, a different aggregation window, a preprocessing version mismatch). The discriminating question: recompute the feature from production raw data using the training pipeline, and compare against what the serving path produced. A mismatch means skew — the pipelines disagree. If the pipelines agree, the raw data itself changed, and you are looking at covariate drift.
  • Covariate (data) drift — the input distribution moved. Marginals can shift with no performance harm when the shift lands in a region the model handles well.
  • Label/prediction shift — the output distribution moved. Can be pure consequence of covariate drift, or the only visible sign of concept drift when inputs look stable.
  • Concept drift — the relationship between inputs and the true outcome changed, so equivalent inputs warrant different predictions. Generally requires ground truth or a credible outcome signal to assert; input statistics alone cannot prove it.

The cross-cases are what interviewers use to test real understanding: covariate drift without performance harm (a marketing campaign changes the age mix and the error rate doesn't move), performance loss without marginal drift (the input–outcome relationship moved underneath unchanged marginals), and stable predictions concealing a stuck feature pipeline. A weak answer treats every distance threshold as "retrain"; a strong answer diagnoses first.

Drift is not a performance verdict

Conflating drift with degradation produces both wasted retraining and missed incidents. Treat a divergence score as a prompt to look for outcome evidence, not a verdict. Pair every drift alarm with the best available performance signal for the affected slice, and record how often alarms were followed by measured quality loss — a drift monitor whose alerts have never corresponded to a degradation is measuring the business, not the model.

What you do in each case:

  • Drift, no accuracy loss — expected traffic change (campaign, season, new market). Update the reference window, document it, don't page.
  • Drift with accuracy loss — the classic case: retrain or restrict the model's decision scope for the affected slices while the retrain runs.
  • Accuracy loss, no drift — suspect concept drift, a label pipeline break, or a silent serving change (model version, feature store staleness). This is the case pure drift monitoring misses, which is why outcome monitoring is non-negotiable.

Choosing and criticizing a drift measure

Interviewers frequently ask "which drift metric would you use and why." The honest answer is that the metric matters less than the baseline, threshold, and multiple-testing discipline, but you should know the menu:

MeasureBest forFailure modes
PSITabular features, binned; industry default; interpretable "rule of thumb" bandsBin choice dominates the number; unstable on low-count bins
KS (two-sample)Continuous features, moderate samplesUseless for categorical; sensitive to sample size
Chi-squareCategorical / low-cardinalityExplodes with high cardinality; needs expected counts per cell
JS / KL divergenceGeneral distributions, embedding-friendlyKL is unbounded and undefined for zero-probability bins; needs smoothing
WassersteinContinuous, ordinal; gives "how far" in feature unitsCostly for high dimensions; binning still needed in practice
MMD, centroid/cosine distance on embeddingsText and image inputs, multivariate drift in latent spaceThresholds have no rule-of-thumb bands; needs a reference embedding set; expensive

Univariate measures miss correlated shifts — each feature's marginal can move within tolerance while the joint distribution moves a lot. Multivariate measures (MMD, Mahalanobis-type, or a domain classifier trained to distinguish reference from current) catch that but are harder to threshold and explain. A strong answer names both and says which signals route to which.

Statistical discipline matters because monitoring runs many tests continuously. A two-sample test at the 5 percent level applied to 200 features every hour produces roughly 10 alarms an hour from noise alone — which is how teams end up silencing the whole channel. Prefer effect-size measures with thresholds set from observed history (a PSI above 0.2 on a feature that historically sits below 0.05) over p-values that shrink automatically as traffic grows; with a million requests a day almost any test rejects. Set thresholds per feature from its own baseline variability, require persistence across consecutive windows before paging, and route the rest to a review queue. The goal is a small number of alarms an owner reads, not a complete inventory of every distribution that moved.

The label-delay problem: detecting degradation before ground truth arrives

Ground truth is difficult: define label source, join key, observation time, maturity, corrections, missingness and selection mechanism. The model's action may influence which outcomes are observed — a fraud model declining risky users never sees whether they were actually fraudulent — creating selective labels and feedback bias. Track label coverage and delay by slice, and never substitute predictions for truth merely to keep a performance chart populated.

When labels arrive days or weeks late (credit default, churn, claims), interviewers expect the workaround list:

  • Prediction/score distribution shift — cheap, immediate, but confounds concept drift with benign traffic change.
  • Calibration drift — if the model was calibrated, a shift in the reliability of binned scores against whatever labels have arrived is an early, interpretable signal.
  • Confidence proxies — rising share of near-threshold predictions, or falling margin between top classes, flags a model operating outside its comfort region.
  • Delayed-label performance windows — run performance metrics on matured cohorts only, with an offset so late-arriving labels aren't counted as misses.
  • Performance estimation on unlabeled traffic — importance-weighted corrections or conformal-style bounds; useful but assumption-laden, and a strong candidate says so.
  • Sampled human review — a small stratified sample of decisions reviewed by humans gives a real, if noisy, outcome signal within hours instead of weeks.

A weak answer here is "we monitor input drift until labels arrive." A strong answer says which of the above it trusts, why, and how it validates the proxy against matured labels.

Baselines, windows, and how the answer moves

"Is it drifting?" has no answer until you name the reference:

  • Training data tests departure from the development distribution — the right question for "does the model still fit its training assumptions?"
  • Validation data suits prediction-distribution comparison.
  • A recent healthy production window detects sudden change quickly but normalizes slow drift (the boiling-frog problem).
  • Seasonally or cohort-matched references prevent expected cycles from paging — retail traffic compared against last Tuesday, not last month.

Window length, offset, frequency, minimum sample and event-time policy all change the verdict: a 7-day window smooths over a step change that a 1-hour window catches; an offset lets delayed capture and late labels arrive before the window closes. Segment by model version, route, region, tenant, device and critical slices so mixed traffic doesn't hide regressions. Version the reference, metric implementation, bins, features and thresholds with the deployed model — a baseline contaminated by an incident normalizes failure. When an interviewer asks "how would your drift answer change if I moved the reference from training data to a trailing week?" they are testing exactly this.

System health, coverage, and the rest of the machine

Begin with deployed identity and telemetry coverage: model version/digest, serving code, features, configuration, endpoint/batch job, traffic allocation and business policy. Measure which requests and predictions are captured, sampled, dropped or redacted; reconcile counts to serving logs. Monitor collection delay, schema, clock, join keys and consent. No violations while capture is broken means unknown, not healthy.

Separate system health from model behavior: traffic, latency, errors, saturation, dependency health, timeouts, feature retrieval and fallback rates. A latency incident can bias the observed population if requests fail before capture. An endpoint can be operationally green while producing harmful decisions — business and safety outcomes need their own monitoring.

Response is not always retraining. Data quality may require upstream repair; skew may require transformation alignment; traffic shift may need routing or capacity fixes; drift may be expected product change; performance or fairness harm may require rollback, decision limits or human review. Retraining on corrupted or biased feedback amplifies the problem. Monitoring can trigger investigation or a candidate pipeline, but promotion needs independent evaluation.

Protect the monitoring data itself: minimize fields, tokenize join identifiers, sample proportionately, encrypt, restrict access, define retention, and keep raw values out of metric labels and alerts. Keep monitoring and alert configuration changes auditable — an attacker who disables capture or widens thresholds can conceal harm. Test the monitor itself by injecting schema breaks, missing features, shifted distributions, stale capture, delayed labels, join failures, silent model-version changes and known regressions, and reconcile actual endpoints with monitoring jobs so every serving version is covered.

One dated vendor note: as of this material, AWS documents SageMaker Model Monitor as unavailable to new customers while existing customers remain supported — verify current availability rather than assuming a greenfield default.

Measure the monitoring program by telemetry and model coverage, capture completeness and freshness, ground-truth coverage and delay, actionable alert precision and recall, detection and response time, unresolved critical slices, and false retraining triggers. Monitoring succeeds when it detects meaningful loss of system value or trust early enough for an accountable owner to take proportionate, verified action.

What interviewers probe, and what a weak answer sounds like

Likely follow-ups, and the traps in each:

  • "Your PSI alarm fired on 30 features at 3 a.m. What do you do?" — Weak: retrain. Strong: check persistence across windows, check whether the affected features share a pipeline (one upstream break, not thirty drifts), pull the best available outcome signal for the affected slices, and only then decide.
  • "Labels take 30 days. How do you know the model degraded today?" — Weak: "we wait." Strong: the proxy stack above, with an explicit statement of which proxy is validated against matured labels.
  • "Accuracy dropped but no feature drifted. What happened?" — Concept drift, label pipeline break, or silent serving change. The discriminating move is checking whether the serving model version and feature pipeline are what you think they are.
  • "How do you set a drift threshold?" — Weak: "PSI > 0.25 means significant drift" recited as law. Strong: thresholds from each feature's own baseline variability, backtested against historical incidents and expected seasonality, with warning and critical bands, persistence rules, hysteresis and cooldown, shadowed before paging.
  • "Every alert needs what attached?" — Model/version, window, baseline, affected features and slices, effect size, coverage, likely causes, owner, runbook, linked evidence. An alert without an owner and a runbook is a log entry, not a control.