Skip to content
Tech Interview Prep home

Top 100 MLOps Engineer Interview Questions and Answers

The questions most likely to actually come up in your MLOps Engineer interview, ranked by likelihood — with detailed, senior-level answers covering what an interviewer is really listening for.

Curated: · Written: · Reviewed:

Reviewed 73Review pending 27
QA-1Your model scores 0.91 AUC offline and behaves far worse in production. Where do you look first?(show answer)

The first thing I would pin down about train/serve skew is which population the evidence was measured on.

The first hypothesis is that the features computed at serving time are not the features the model was trained on. Skew comes from two implementations of the same transformation, different source freshness, or a different default for a missing value, and it degrades quality without producing a single error.

Concretely, log the exact feature vector served with each prediction, recompute the same entities through the training transformation, and compare value by value rather than comparing summary statistics, which hide sign errors and unit changes inside a matching mean.

The reason for that specificity is a failure I have seen: A revenue feature was scaled to thousands in training and passed in units at serving; offline AUC of 0.91 became 0.62 in production, and three weeks were spent tuning the model before anyone compared a single vector.

Same entity, two paths.

FeatureTraining valueServed valueMatch
revenue_30d4.214210.0no
tenure_days318318yes
region_id77yes

I would not consider it settled without evidence: Compare served and recomputed feature vectors for a sample of live entities and require an exact match before promotion.

A skew bug produces no error, only a worse model.

Curated: · Written: · Reviewed:

QA-2A model has been serving for six months. What do you need to rebuild it exactly?(show answer)

I would start reproducing a training run from reproducibility, because an unreproducible result cannot be argued with.

Reproduction needs the code revision, the data snapshot, the feature transformation version, the dependency lock, the runtime image, the hyperparameters, and the seeds. A missing element does not degrade reproduction gracefully; it ends it.

Concretely, write every one of those identifiers into the registry entry at registration time as part of the artifact record, rather than relying on the training job's logs, which are retained for weeks while models serve for years.

The reason for that specificity is a failure I have seen: A model in production for 14 months could not be rebuilt because the training dataset was a query against a mutable table; the reconstructed set differed by 4 percent of rows and the rebuilt model disagreed with the deployed one on 1 in 30 predictions.

What the registry entry has to pin.

ElementRecordedSufficient
git commityesyes
dataset snapshot idno, query text onlyno
dependency lockyesyes
seedsnono

I would not consider it settled without evidence: Rebuild a promoted model from its registry record alone, on a clean machine, and compare its predictions with the deployed artifact.

Reproducibility is proved by rebuilding, not by intending to.

Curated: · Written: · Reviewed:

QA-3You rebuild a model from its recorded inputs and the weights differ. Is that a failure?(show answer)

This is an area where a model that scores well and a model that behaves well are different events.

Bit-identical weights are one claim among three, and usually the most expensive. Non-deterministic kernels, atomic floating-point reductions, and a changed worker count move the result even when every recorded input matches, so the useful claim is metric equivalence within a stated tolerance.

Concretely, state per model which claim the pipeline supports — byte equality, metrics within a tolerance, or documented inputs — and test the one you claim, so a difference is either a violation or an expected variation rather than an argument.

The reason for that specificity is a failure I have seen: A team promised byte-identical rebuilds, spent six weeks chasing a 0.002 AUC difference caused by cuDNN kernel selection, and stopped rebuilding models at all afterwards.

Three claims, three costs.

ClaimRebuild passesEffort
identical bytesnovery high
metrics within 0.005 AUCyesmoderate
inputs documentedyeslow

I would not consider it settled without evidence: Assert the rebuilt model's evaluation metrics fall within the declared tolerance on the frozen set, rather than asserting equal weights.

Claim the bound you can test.

Curated: · Written: · Reviewed:

QA-4An auditor asks which model produced a prediction made in March. How do you answer?(show answer)

My answer to logging the resolved model version with each prediction begins with the contract between training and serving.

An alias such as champion is a mutable pointer, so it identifies a model only at the instant it is read. The prediction log has to carry the resolved version and artifact digest, because the pointer will have moved by the time anyone asks.

Concretely, resolve the alias once at load time, attach the version and digest to every prediction record along with the serving image tag, and treat a prediction written without them as an incomplete record rather than as a smaller one.

The reason for that specificity is a failure I have seen: A dispute over a March decline required identifying the model; the logs held only the alias, which had pointed to three versions that month, and the investigation could not say which weights ran.

What the alias hid.

WindowAliasActual version
Mar 1-9championv31
Mar 10-18championv33
Mar 19-31championv34

I would not consider it settled without evidence: Query a prediction from a past incident and confirm the log names an immutable version and digest, not an alias.

An alias records intent; a digest records what ran.

Curated: · Written: · Reviewed:

QA-5Someone wants to overwrite version 12 with a retrained model. What is your answer?(show answer)

I would treat immutability of a registered artifact as a claim about production that has to survive a delayed label.

A registered version is an immutable package, and overwriting it destroys the link between every past deployment record and the bytes those records describe. The retrained model is a new version, whatever the convenience argument for reusing the number.

Concretely, enforce immutability in the registry's permissions rather than by convention, so an upload against an existing version fails rather than succeeding quietly, and make the digest part of every deployment record.

The reason for that specificity is a failure I have seen: A hotfix overwrote a registered version in place; three weeks later a rollback restored the overwritten artifact, and the incident that the rollback was meant to reverse reappeared immediately.

What overwriting broke.

RecordPoints atAfter overwrite
deployment 2026-01-14v12 digest a91cnow bytes of e77d
rollback targetv12restores wrong model

I would not consider it settled without evidence: Attempt to re-upload an existing version in a test environment and require the registry to reject it.

A version that can change is not an identifier.

Curated: · Written: · Reviewed:

QA-6Your training pipeline finishes and the model is better. Should it deploy?(show answer)

The useful question for promotion as a decision separate from training is what the system does while nobody is looking at it.

Training produces a candidate; promotion is a separate decision with its own gates, and collapsing them means a training job can deploy to production on its own. That is a change to a live system authorised by a cron schedule.

Concretely, have the pipeline register the candidate and stop, then run promotion as its own job with the integrity, evaluation, and approval gates that a production change requires, under credentials the training job does not hold.

The reason for that specificity is a failure I have seen: A retraining job promoted automatically on any improvement in aggregate accuracy; a corrupted label batch produced a model that raised the aggregate by 0.6 points while recall on the fraud slice fell from 0.74 to 0.31, and it served for 9 days.

Where the gate belongs.

StepActorCan serve traffic
trainpipeline roleno
register candidatepipeline roleno
promoterelease roleyes

I would not consider it settled without evidence: Confirm the training role lacks permission to move the production alias, and that promotion requires the separate gate to pass.

Training proposes; promotion decides.

Curated: · Written: · Reviewed:

QA-7Your training set is defined by a SQL query. What is wrong with that?(show answer)

I would settle data snapshots rather than live queries against a replay of real requests before trusting the offline number.

A query against mutable tables defines a different dataset every time it runs, so two runs of the same pipeline train on different data and the difference is invisible. A snapshot identifier makes the dataset a fixed object that can be referenced later.

Concretely, materialise the training set to a versioned, immutable location with a recorded row count and content hash, and record that identifier in the model's lineage rather than the query text that produced it.

The reason for that specificity is a failure I have seen: Two runs of the same pipeline a day apart trained on sets differing by 60,000 rows because a late-arriving backfill landed between them; the metric difference was attributed to a hyperparameter change.

Same query, two runs.

RunRowsHashVal AUC
Tue1,204,3383f9a0.883
Wed1,264,401c17e0.891

I would not consider it settled without evidence: Record a row count and content hash with each training set and assert they match when a run is repeated.

A query is a recipe, not a dataset.

Curated: · Written: · Reviewed:

QA-8The model file is unchanged but predictions moved. What could have happened?(show answer)

The judgement in pinning the serving environment is which version is pinned and who approved it.

The model is only one input to a prediction. A library upgrade that changes a tokenizer, a rounding behaviour, or a default preprocessing argument changes outputs with the weights untouched, which is why the serving image belongs in the release identity.

Concretely, pin the runtime image by digest and the dependency set by lock file, deploy them as one bundle with the model, and reject a serving deployment whose image digest is not the one evaluated.

The reason for that specificity is a failure I have seen: A base-image rebuild moved a library minor version and changed a text normalisation default; predictions shifted on 4 percent of inputs while the model artifact and its metrics were unchanged.

What changed between deployments.

ComponentBeforeAfter
model digesta91ca91c
image digest55b190ff
outputs differing—4.1%

I would not consider it settled without evidence: Replay a fixed request set through the new image and require identical outputs before the image is allowed to serve.

The model file is not the model.

Curated: · Written: · Reviewed:

QA-9How do you know your training features were available at prediction time?(show answer)

Where candidates lose the interview on point-in-time correctness in training data is treating a training-time result as a serving guarantee.

Every feature value in a training row must have been knowable at the row's decision timestamp. Joining the latest value instead is leakage, and it inflates offline metrics precisely because the model is reading the future.

Concretely, retrieve features with an as-of join on the decision timestamp, honouring event time rather than ingestion time, and test it with fixtures whose events deliberately straddle the boundary.

The reason for that specificity is a failure I have seen: A churn model joined support tickets as latest-value; it read tickets filed during cancellation and scored 0.94 AUC offline against 0.71 served, and the gap was attributed to drift for two quarters.

The join that leaked.

JoinOffline AUCServed AUC
latest value0.940.71
as-of decision time0.790.78

I would not consider it settled without evidence: Build a fixture with events before and after the decision time and assert the retrieved value ignores everything after it.

Leakage looks like accuracy until it ships.

Curated: · Written: · Reviewed:

QA-10Your label is defined over 60 days and you retrain weekly. What breaks?(show answer)

I would answer label maturity and the training window by separating what the pipeline automated from what it verified.

A label that takes 60 days to settle is not available for the most recent 60 days of data, so any window that includes them trains on partial outcomes. The bias is systematic rather than noisy: it favours whatever resolves fastest.

Concretely, declare the maturation delay per label, exclude any period shorter than it from training and evaluation, and treat the delay as a property of the label rather than a tuning parameter.

The reason for that specificity is a failure I have seen: A 60-day churn label was trained on a window ending 14 days before the run; the model learned early cancellers and under-predicted slow churn by a factor of three, while offline metrics improved every week.

What the recent window contained.

Window ageLabels settledChurn rate seen
0-14 days21%1.9%
15-59 days68%4.1%
60+ days100%6.3%

I would not consider it settled without evidence: Assert that no training row's decision date falls within the maturation delay of the run date.

An immature label is a different label.

Curated: · Written: · Reviewed:

QA-11Your recommender retrains on click logs. What is the risk?(show answer)

The engineering content of feedback loops in retrained models is the gate and the rollback, not the model architecture.

Training on data the model itself selected narrows the model toward its own choices. Items never shown generate no clicks, so the next model learns they do not convert, and the blind spot compounds with every cycle.

Concretely, reserve a small randomised exploration slice, log the model version and the scores behind each served decision so selection can be corrected for, and monitor training-set diversity against the catalogue rather than against the previous training set.

The reason for that specificity is a failure I have seen: A recommender's served catalogue narrowed from 12,400 to 900 items over eight retraining cycles while every offline metric improved, because the metric was computed on the items it had already chosen.

Catalogue coverage by cycle.

CycleDistinct items servedOffline CTR
112,4000.041
44,1000.048
89000.052

I would not consider it settled without evidence: Track distinct items served per week and compare against catalogue size, not against the previous week's training data.

A pipeline fed by its own output converges on its blind spot.

Curated: · Written: · Reviewed:

QA-12Your holdout set is two years old. What does its score tell you?(show answer)

Before promoting anything through evaluation on the served population I would write down what silent degradation would look like.

A metric describes the population it was measured on. A holdout drawn from a distribution the product has since moved away from measures how the model would do on customers it no longer has.

Concretely, refresh the evaluation set on a schedule tied to how fast the population moves, keep a fixed historical set alongside it for comparability, and report both, so a change in either is attributable.

The reason for that specificity is a failure I have seen: A fraud model was gated on a 2024 holdout through 2026; the merchant mix had changed entirely, and the model that scored best on the old set was the worst on live traffic.

Ranking flips by evaluation set.

Candidate2024 holdoutRecent traffic
A0.88 (best)0.74 (worst)
B0.850.82
C0.840.83 (best)

I would not consider it settled without evidence: Compare the evaluation set's feature distribution against last month's served traffic before trusting the gate.

An old holdout measures an old business.

Curated: · Written: · Reviewed:

QA-13Accuracy improved by two points. Why is that not enough to promote?(show answer)

The first thing I would pin down about slice metrics rather than aggregates is which population the evidence was measured on.

An aggregate is a weighted average, so a gain on the majority can hide a collapse on a minority that carries most of the business risk. The decision needs the slices the model is actually used on.

Concretely, define the slices before the evaluation — by segment, geography, device, tenure, and any protected attribute in scope — and gate promotion on no slice regressing beyond a stated threshold rather than on the aggregate.

The reason for that specificity is a failure I have seen: A model improved aggregate accuracy from 91.3 to 93.6 while recall on the fraud-positive slice fell from 0.71 to 0.44; the slice was 0.3 percent of rows and most of the loss.

Where the two points came from.

SliceShareBeforeAfter
ordinary99.7%91.393.6
fraud positive0.3%0.71 recall0.44 recall

I would not consider it settled without evidence: Report per-slice metrics with sample sizes alongside the aggregate, and block promotion on a slice regression.

The aggregate is where a slice failure goes to hide.

Curated: · Written: · Reviewed:

QA-14You deploy the same weights with a different threshold. Is that a new model?(show answer)

I would start the decision threshold as a deployed parameter from reproducibility, because an unreproducible result cannot be argued with.

The threshold determines every decision the system makes, so changing it changes behaviour as surely as changing weights. It belongs inside the versioned release rather than in a configuration file edited during an incident.

Concretely, version the threshold with the model, derive it from the operating cost of a false positive against a false negative on recent data, and require the same promotion gates for a threshold change as for a weights change.

The reason for that specificity is a failure I have seen: A threshold was lowered in a config edit to reduce missed fraud; the review queue grew from 400 to 9,000 cases a day and the change was not attached to any model version, so nobody could say when it happened.

One model, two thresholds.

ThresholdCaughtReviewed daily
0.8062%400
0.4581%9,000

I would not consider it settled without evidence: Include the threshold in the release bundle and evaluate the candidate at the threshold it will actually serve at.

The threshold is half the model's behaviour.

Curated: · Written: · Reviewed:

QA-15Your model ranks well but its probabilities are wrong. Does it matter?(show answer)

This is an area where a model that scores well and a model that behaves well are different events.

Ranking is enough when the system only picks a top-k, and insufficient when a downstream decision multiplies the score by a cost. A model that ranks perfectly but reports 0.9 where the true rate is 0.4 will misprice every decision built on it.

Concretely, measure calibration explicitly with a reliability curve on recent data, recalibrate on a held-out set when the score feeds an expected-value calculation, and re-check after every retrain because calibration decays faster than ranking.

The reason for that specificity is a failure I have seen: A propensity score fed a bidding rule; AUC held at 0.86 while predicted probabilities drifted 2.3 times above observed rates, and the bid ceiling was breached for four months.

Reliability by decile.

DecilePredictedObserved
80.620.28
90.780.34
100.910.40

I would not consider it settled without evidence: Plot predicted against observed rates in deciles on last month's settled outcomes and check the diagonal.

Ranking orders; calibration prices.

Curated: · Written: · Reviewed:

QA-16What should stop a training run before it starts?(show answer)

My answer to data validation as a pipeline gate begins with the contract between training and serving.

A training run consumes whatever it is given, so the schema, ranges, null rates, and row counts of the input are the last place a corruption can be caught cheaply. After training it is only visible as a metric nobody can explain.

Concretely, assert schema, per-column null rate and range, category cardinality, and row-count bounds against the previous run before training, and fail the run rather than the model when they are violated.

The reason for that specificity is a failure I have seen: An upstream change made a currency column null for 30 percent of rows; the run completed, the model imputed zeros, and the fault surfaced six weeks later as a regional revenue anomaly.

What the gate would have caught.

ColumnNull rate baselineThis runBound
currency0.2%30.4%2%
amount0.0%0.0%1%

I would not consider it settled without evidence: Fail the pipeline when any input column's null rate moves outside its declared bound, and page the data owner rather than the ML team.

Catch corruption at the input, not in the metric.

Curated: · Written: · Reviewed:

QA-17An upstream team renames a column and your model degrades. Whose fault is it?(show answer)

I would treat schema and data contracts with upstream teams as a claim about production that has to survive a delayed label.

Without a stated contract it is nobody's fault, which is the problem. A data contract makes the schema, semantics, and change process an agreement rather than an assumption, so a breaking change is a violation instead of a surprise.

Concretely, publish the fields a model depends on with their types and meanings, require additive-only change with a deprecation window, and enforce the contract in the producer's own tests so a breaking change fails before it lands.

The reason for that specificity is a failure I have seen: A field changed from cents to a decimal amount with no announcement; the model's dominant feature was silently divided by a hundred and quality fell for 11 days before the cause was found.

The change nobody announced.

FieldWasBecameConsumers notified
amountinteger centsdecimal units0 of 6

I would not consider it settled without evidence: Run contract assertions in the producing pipeline's build, not only in the consuming one.

An assumption is a contract nobody agreed to.

Curated: · Written: · Reviewed:

QA-18How do you prove the feature store serves what training used?(show answer)

The useful question for offline and online feature parity is what the system does while nobody is looking at it.

Parity is a measured property, not a design property. Two stores fed by different jobs will diverge through materialisation lag, late events, and type coercion, and the divergence shows up as quality loss rather than as an error.

Concretely, sample entities continuously, retrieve each feature from both the offline and online paths for the same effective time, and alert on mismatch rate rather than checking parity once at build time.

The reason for that specificity is a failure I have seen: A materialisation job had been failing silently for nine days; online values were stale while offline history was correct, and every offline evaluation kept passing.

Parity sample during the outage.

DayEntities sampledMismatched
before5,0003
during5,0004,880

I would not consider it settled without evidence: Run a scheduled parity job over sampled entities and alert when the mismatch rate exceeds its baseline.

Parity holds until a job fails quietly.

Curated: · Written: · Reviewed:

QA-19A feature is nine days old at serving time. What should the system do?(show answer)

I would settle feature freshness and TTL against a replay of real requests before trusting the offline number.

A stale feature is a different feature, and silently serving it converts a pipeline outage into a quality problem with no alarm. The freshness policy is part of the feature's definition, not an operational detail.

Concretely, declare a TTL per feature, return an explicit staleness signal rather than a value past it, and decide per model whether to fall back, degrade, or refuse — then monitor the fallback rate by slice.

The reason for that specificity is a failure I have seen: An hourly feature went stale for two days during an incident; the model kept serving on values from before the outage and the affected cohort's approval rate moved 12 points with no alert.

Served ages during the outage.

Age bucketShare of requestsTTL
under 1h4%6h
1-6h9%6h
over 6h87%violated

I would not consider it settled without evidence: Record feature age with each prediction and alert on the share served past TTL.

Staleness has to be visible to be handled.

Curated: · Written: · Reviewed:

QA-20A feature lookup fails for one request. What does the model receive?(show answer)

The judgement in missing features at serving time is which version is pinned and who approved it.

The default chosen for a missing value is a modelling decision with a distribution attached. Imputing zero when training imputed the median moves those predictions in a direction nobody evaluated, and does it silently.

Concretely, use the same imputation in serving as in training, pass an explicit missingness indicator where the model was trained with one, and count missing lookups per feature as a monitored signal.

The reason for that specificity is a failure I have seen: A serving default of zero for a feature whose training median was 340 pushed every affected request into the low-risk band; 6 percent of traffic was affected for a month.

Effect of the wrong default.

PathMissing value usedMean score
trainingmedian 3400.42
serving00.11

I would not consider it settled without evidence: Assert that serving and training imputations agree for every feature, and monitor the missing rate per feature.

A default is a prediction about the missing value.

Curated: · Written: · Reviewed:

QA-21Should a model retrain on a schedule or on a signal?(show answer)

Where candidates lose the interview on choosing a retraining trigger is treating a training-time result as a serving guarantee.

A schedule retrains whether or not anything changed, and a signal retrains only when something did. The right choice follows how fast the data moves and how expensive a stale model is, and both need measuring before the cadence is chosen.

Concretely, measure the decay curve first — evaluate the current model on each subsequent week's settled outcomes — then set the trigger where the loss crosses what the business will accept, and keep a schedule as a floor so a broken signal cannot mean never.

The reason for that specificity is a failure I have seen: A daily retrain ran for a year on data whose distribution moved quarterly; it consumed 340 GPU-hours a month and every candidate was within noise of the last.

Measured decay for this model.

Weeks since trainingAUCLoss
00.842—
40.8360.006
120.7910.051

I would not consider it settled without evidence: Plot the deployed model's weekly performance on settled outcomes and set the trigger from where it degrades.

Cadence follows the decay curve, not the calendar.

Curated: · Written: · Reviewed:

QA-22How do you decide a new model is better than the one serving?(show answer)

I would answer champion and challenger evaluation by separating what the pipeline automated from what it verified.

The comparison must be against the current production model on the same recent data, not against the previous candidate or a published benchmark. A challenger that beats last quarter's model may still be worse than what is serving today.

Concretely, score both models on the same freshly labelled window, report the difference with an interval rather than two point estimates, and require the gain to exceed the interval before promotion.

The reason for that specificity is a failure I have seen: A challenger was promoted on a 0.4-point gain measured against a stale baseline; the deployed model was 1.1 points better on the same window, and quality fell on promotion.

The comparison that was missing.

ModelStale windowShared recent window
champion0.8210.858
challenger0.8250.847

I would not consider it settled without evidence: Evaluate champion and challenger on one shared, recent, settled window and report the paired difference.

Better than what, measured when.

Curated: · Written: · Reviewed:

QA-23What does running a model in shadow actually prove?(show answer)

The engineering content of shadow deployment is the gate and the rollback, not the model architecture.

Shadow mode proves the candidate can handle production inputs at production rates and shows where it disagrees with the incumbent. It proves nothing about outcomes, because its predictions never reach a user and no counterfactual exists.

Concretely, mirror real requests to the candidate, discard its outputs from decisions, and compare latency, error rate, schema validity, and disagreement rate by slice — then use a canary or experiment for the outcome question.

The reason for that specificity is a failure I have seen: A team ran two weeks of shadow traffic and promoted on a low disagreement rate; the 3 percent where the models disagreed were the high-value cases, and revenue fell.

Where the models disagreed.

SliceRequestsDisagreementValue share
routine97%0.4%31%
high value3%22%69%

I would not consider it settled without evidence: Report disagreement rate by slice and weight it by business value rather than by request count.

Shadow answers "can it run", not "is it better".

Curated: · Written: · Reviewed:

QA-24You run a 1 percent canary for a day and see nothing. Is the model safe?(show answer)

Before promoting anything through canary sizing and detection power I would write down what silent degradation would look like.

A canary has a detection floor set by cohort size and baseline variance, and a quiet result below that floor is not evidence. Stating the smallest detectable effect before the rollout is what separates a test from a reassurance.

Concretely, compute the detectable effect from expected cohort volume and baseline rate, use the canary for the failures it can catch — errors, latency, schema, safety — and size a longer experiment for quality questions it cannot.

The reason for that specificity is a failure I have seen: A 1 percent canary on 200,000 daily requests observed 2,000 events and was reported as showing no regression; the actual loss was 0.4 points of conversion, roughly a third of what that cohort could detect.

What that cohort could see.

CohortEventsDetectable dropActual drop
1% for 1 day2,0001.2 pts0.4 pts
10% for 5 days100,0000.2 pts0.4 pts

I would not consider it settled without evidence: State the minimum detectable effect for the planned cohort and duration before starting the rollout.

An underpowered canary reassures rather than tests.

Curated: · Written: · Reviewed:

QA-25You roll the model back and the incident continues. Why?(show answer)

The first thing I would pin down about rolling back the whole release bundle is which population the evidence was measured on.

A release changes weights, preprocessing, feature definitions, thresholds, and the runtime image together, so restoring one of them leaves a combination that was never evaluated. Rollback has to restore the bundle.

Concretely, keep the previous bundle deployable as a unit, rehearse the reversal with the same automation and permissions that will be used under pressure, and measure how long it takes.

The reason for that specificity is a failure I have seen: A rollback restored the previous weights while the new feature transformation stayed live; the resulting pairing had never been tested and error rates rose above the incident that prompted the rollback.

What the partial rollback left.

ComponentRolled backLive version
weightsyesv18
transformationnov19
thresholdnov19

I would not consider it settled without evidence: Rehearse a full rollback in staging and record the elapsed time and the components restored.

Roll back the release, not the file.

Curated: · Written: · Reviewed:

QA-26Traffic is back on the old model. Is the incident over?(show answer)

I would start irreversible actions taken by a model from reproducibility, because an unreproducible result cannot be argued with.

Routing stops new harm; it does not undo decisions already made. Emails sent, accounts closed, prices quoted, and records written under the bad model persist after the traffic shift, and the compensating work is usually the longer half.

Concretely, enumerate the irreversible effects a model can cause before launch, log every action with the version that caused it so the affected set is queryable, and prepare the compensating procedure alongside the rollback.

The reason for that specificity is a failure I have seen: A pricing model quoted 41,000 offers below cost in four hours; rollback took eight minutes and honouring the quotes took six weeks.

The two halves of recovery.

PhaseDuration
traffic restored8 minutes
quotes reconciled6 weeks

I would not consider it settled without evidence: Query the actions taken under the withdrawn version and size the compensating work before declaring recovery.

Rollback stops the cause, not the consequences.

Curated: · Written: · Reviewed:

QA-27Who stops a bad rollout at three in the morning?(show answer)

This is an area where a model that scores well and a model that behaves well are different events.

A rollout that can only be stopped by a human on call is protected by response time rather than by design. Severe, unambiguous failures — error rate, schema violations, safety filters — should halt the ramp automatically.

Concretely, define halt conditions with thresholds and minimum sample before the rollout starts, wire them to the deployment controller rather than to a dashboard, and keep noisy business metrics as human-reviewed signals.

The reason for that specificity is a failure I have seen: A candidate returned malformed responses for 9 percent of requests overnight; the alert fired at 02:10, the page was acknowledged at 07:40, and the ramp had reached 50 percent.

Halt policy by signal.

SignalThresholdAction
schema invalid0.5%automatic halt
p99 latency2x baselineautomatic halt
conversion-1 pthuman review

I would not consider it settled without evidence: Trigger a halt condition deliberately in staging and confirm the controller stops the ramp without a human.

Automate the halt for what is unambiguous.

Curated: · Written: · Reviewed:

QA-28You route each request randomly between two models. What does that break?(show answer)

My answer to experiment assignment for model comparison begins with the contract between training and serving.

Per-request routing is fine for technical validation and wrong for behavioural comparison, because one user sees both models and their behaviour mixes the treatments. The unit of assignment has to match the unit of the effect.

Concretely, assign stickily by the unit whose behaviour is being measured — user or session — record the assignment separately from the outcome, and exclude bots, employees, and retried requests deliberately.

The reason for that specificity is a failure I have seen: A per-request split reported a 0.2 percent difference between models; a sticky re-run of the same comparison measured 3.1 percent, because per-request mixing had diluted the effect.

Same models, two designs.

AssignmentMeasured liftUsers seeing both
per request0.2%94%
sticky by user3.1%0%

I would not consider it settled without evidence: Confirm each unit saw exactly one variant for the whole measurement window before reading the result.

Mixed exposure measures the average of both.

Curated: · Written: · Reviewed:

QA-29Labels arrive 45 days late. How do you monitor the model today?(show answer)

I would treat monitoring under delayed ground truth as a claim about production that has to survive a delayed label.

With delayed labels the honest position is that today's accuracy is unknown, and the monitoring job is to detect the conditions that precede a loss rather than to claim a number. Proxy signals are early warnings, not measurements.

Concretely, monitor input distributions, prediction distributions, and any fast proxy that correlates with the outcome, then reconcile every proxy against the true label when it settles and record how well it predicted.

The reason for that specificity is a failure I have seen: A team reported daily accuracy computed on the small share of labels that settled immediately; those were the easy cases, and the number stayed at 0.93 while true performance fell to 0.71.

Proxy against settled truth.

MonthProxy accuracySettled accuracy
March0.930.88
April0.930.79
May0.920.71

I would not consider it settled without evidence: Backfill true performance when labels settle and compare it against what the proxy claimed at the time.

An early label is a biased sample.

Curated: · Written: · Reviewed:

QA-30Feature distributions moved but accuracy is flat. Do you retrain?(show answer)

The useful question for data drift against concept drift is what the system does while nobody is looking at it.

Input drift and quality loss are different events. A distribution can move into a region the model handles well and cost nothing, while a model can fail badly with inputs that look unchanged because the relationship to the target moved.

Concretely, treat a divergence score as a prompt to look for outcome evidence rather than as a verdict, pair each drift alarm with the best performance signal for the affected slice, and record the alarm's hit rate.

The reason for that specificity is a failure I have seen: A team retrained on every drift alarm; over a year 31 of 34 retrains were triggered by a marketing-driven shift in the age mix that had never affected error rate, and one real concept shift was lost in the noise.

Alarms against outcomes.

Alarm sourceFiredFollowed by loss
age mix310
merchant mix33

I would not consider it settled without evidence: Record, for every drift alarm, whether a measured quality loss followed, and prune the monitors that never predict one.

Drift is a question; degradation is the answer.

Curated: · Written: · Reviewed:

QA-31You test 200 features hourly at the 5 percent level. What happens?(show answer)

I would settle thresholds for drift detection against a replay of real requests before trusting the offline number.

Running many tests continuously produces alarms from noise alone — roughly ten an hour at that configuration — which is how a monitoring channel gets silenced. The statistics have to account for the number of tests and the volume of traffic.

Concretely, prefer effect-size measures with per-feature thresholds set from that feature's own historical variability, require persistence across consecutive windows before paging, and route the remainder to a review queue.

The reason for that specificity is a failure I have seen: A drift channel produced 240 alerts a day; it was muted after a week, and the muting was still in place when a genuine feature outage arrived four months later.

Alert volume by rule.

RuleAlerts per dayActionable
p-value < 0.052401
PSI > 0.2 for 3 windows43

I would not consider it settled without evidence: Measure the alert rate the configuration produces on a period with no known incident before enabling paging.

A monitor nobody reads is not monitoring.

Curated: · Written: · Reviewed:

QA-32A new model version has cleared offline evaluation and the team wants it serving. What does the rollout look like before it takes all the traffic?(show answer)

The judgement in progressive model rollout is which version is pinned and who approved it.

Every stage admits more traffic only after a pre-written gate passes, and that gate must cover quality on named slices as well as service health — error rate and p99 stay green while the model gets worse.

Concretely, mirror production requests to the challenger in shadow first, then ramp live traffic 1, 5, 25, 50, 100 percent, each step naming its comparison window, minimum request count, and an automated rollback trigger. Retire the champion only after a full business cycle.

The reason for that specificity is a failure I have seen: A challenger moved from 5 to 50 percent on a Thursday with gates on error rate and p99 latency only. Aggregate conversion held within 0.4 percent of the champion while the iOS slice fell 11 percent; no gate covered that slice, full traffic came on Friday, and the model served through the weekend before anyone looked at the slice on Monday.

Gate sheet for a fraud-scoring challenger (thresholds illustrative).

StageTrafficMinimum windowGate to advanceRollback trigger
shadow0% live, requests mirrored24 h, 50k or more scoredscore distribution within agreed band, zero crashesdrop the mirror; users never see it
canary1% live24 h, 5k or more requestserror rate and p99 within 10% of championerror rate above 2x champion for 5 min
ramp5 to 25 to 50%2 h per stepslice precision and recall within 3% of championany named slice drops more than 3%
full100%7 days including a weekend peakall slices holdauto-revert to previous model version

I would not consider it settled without evidence: Name the gate for the first live traffic step — the metric, the comparison window, the minimum request count, and the rollback trigger — and say which of those the platform evaluates without a human.

Write the rollback criteria before any traffic moves.

Curated: · Written: · Reviewed:

QA-33What does an SLO for a model look like?(show answer)

Where candidates lose the interview on model service level objectives is treating a training-time result as a serving guarantee.

A model service has availability and latency objectives like any service, plus quality objectives that need a defined measurement window and a defined population. Without the quality half, the service can be perfectly available and useless.

Concretely, state availability and latency per request class, state quality as a metric on a named slice measured over a stated window, and define what happens when each is breached — including who decides to withdraw the model.

The reason for that specificity is a failure I have seen: A model met 99.95 percent availability for eight months while its precision on the slice it existed to serve fell below the manual process it replaced; no objective covered that, so nothing triggered.

The objective that was missing.

ObjectiveTargetActual
availability99.9%99.95%
p99 latency120 ms96 ms
precision, fraud slice0.700.51

I would not consider it settled without evidence: Publish the quality objective with its slice and window alongside availability, and review breaches in the same forum.

Available and correct are separate promises.

Curated: · Written: · Reviewed:

QA-34How do you choose the serving shape for a new model?(show answer)

I would answer online, batch, and streaming inference by separating what the pipeline automated from what it verified.

The shape follows the decision's deadline and the freshness of the inputs it needs. Precomputing predictions is far cheaper when the inputs change slowly, and impossible when the decision depends on the current request.

Concretely, ask when the decision is made and what it depends on: precompute where inputs are stable and the population is enumerable, serve online where the request itself carries features, and stream where events must be scored as they arrive.

The reason for that specificity is a failure I have seen: A daily-refreshed segmentation model was served as a real-time endpoint at 140 dollars a day; batch precomputation for the same 900,000 customers cost 4 dollars and met the same requirement.

Same model, two shapes.

ShapeDaily costFreshness
online endpoint$140seconds
nightly batch$424 hours

I would not consider it settled without evidence: State the decision deadline and the input freshness requirement before choosing the serving shape.

Serve online only what cannot be precomputed.

Curated: · Written: · Reviewed:

QA-35Model inference takes 8 ms. Why is the endpoint answering in 240 ms?(show answer)

The engineering content of the end-to-end latency budget is the gate and the rollback, not the model architecture.

Inference is one term in a budget that also holds admission, queueing, feature retrieval, serialisation, and downstream calls. Optimising the smallest term is the most common wasted quarter in model serving.

Concretely, instrument each stage separately and report percentiles per stage and per request class, so the work goes where the time is rather than where the team's expertise is.

The reason for that specificity is a failure I have seen: Six weeks were spent quantising a model that accounted for 8 ms of a 240 ms p99; feature retrieval was 190 ms because it made four sequential lookups that could be issued together.

Where the 240 ms went.

Stagep99
feature retrieval190 ms
queueing26 ms
inference8 ms
other16 ms

I would not consider it settled without evidence: Break the p99 down by stage before choosing what to optimise.

Optimise the term that dominates.

Curated: · Written: · Reviewed:

QA-36Batching raises GPU utilisation. What does it cost?(show answer)

Before promoting anything through dynamic batching I would write down what silent degradation would look like.

Batching trades latency for throughput by holding requests until a batch forms, and the wait is spent from the same budget the deadline comes out of. Whether the trade is worth it depends entirely on the deadline.

Concretely, tune maximum batch size and maximum wait against a load test that uses the real payload and sequence-length distribution, and measure utilisation, throughput, and p99 together rather than one at a time.

The reason for that specificity is a failure I have seen: A batch window of 10 ms was applied to a ranking call inside a 50 ms page budget; utilisation improved from 12 to 70 percent while p99 rose from 24 to 45 ms and the page missed its deadline.

The trade, measured.

ConfigUtilisationThroughputp99
no batching12%1x24 ms
batch 16, wait 10 ms70%4x45 ms

I would not consider it settled without evidence: Load-test the batching parameters with realistic input sizes and check the p99 against the deadline, not the mean.

Batching buys throughput with latency.

Curated: · Written: · Reviewed:

QA-37Why is CPU utilisation a poor scaling signal for a GPU endpoint?(show answer)

The first thing I would pin down about autoscaling an inference service is which population the evidence was measured on.

CPU barely moves while the accelerator saturates, so scaling on it reacts late or not at all. The signal has to lead demand and reflect the resource that is actually scarce.

Concretely, scale on in-flight requests, queue depth, or estimated work such as input tokens, account for model load time in the scale-up delay, keep warm capacity for strict objectives, and use stabilisation windows to prevent oscillation.

The reason for that specificity is a failure I have seen: An endpoint scaled on 70 percent CPU never scaled; queue depth reached 400 requests and p99 passed nine seconds while CPU sat at 22 percent.

During the incident.

SignalValueScaled
CPU22%no
GPU utilisation98%—
queue depth400—

I would not consider it settled without evidence: Correlate each candidate scaling signal against queue depth under a load test and pick the one that moves first.

Scale on the resource that runs out.

Curated: · Written: · Reviewed:

QA-38A new replica takes 90 seconds to become useful. What does that change?(show answer)

I would start cold starts and model loading from reproducibility, because an unreproducible result cannot be argued with.

Model load time sits between the decision to scale and the arrival of capacity, so an autoscaler tuned as though replicas appear instantly will always be late. It also decides whether scale-to-zero is viable.

Concretely, measure load and warm-up time, hold readiness false until the correct version is loaded and warmed, pre-provision headroom sized to the load time, and cache or pre-bake artifacts into the image where the size allows.

The reason for that specificity is a failure I have seen: Scale-to-zero on a 90-second load produced 90-second first requests after every idle period; the p99 for the first minute of each hour was 40 times the objective.

First-request latency after idle.

ConfigurationFirst requestSteady state
scale to zero90 s40 ms
one warm replica41 ms40 ms

I would not consider it settled without evidence: Measure time from scale decision to first successful request, and size headroom from it.

Capacity arrives when the model is warm.

Curated: · Written: · Reviewed:

QA-398-bit quantisation halves your memory. How do you decide to ship it?(show answer)

This is an area where a model that scores well and a model that behaves well are different events.

Compression trades quality for cost, and the loss concentrates in the rare inputs rather than spreading evenly. An aggregate metric that holds can conceal a collapse on exactly the cases the model exists for.

Concretely, evaluate the compressed model per slice against the full-precision baseline, include the rare and high-value slices explicitly, and publish cost per thousand predictions next to the quality delta so the trade is visible.

The reason for that specificity is a failure I have seen: Quantisation moved aggregate accuracy by 0.2 points and cut recall on the smallest, highest-value slice from 0.68 to 0.39; it shipped on the aggregate.

Where the loss landed.

SliceFull precision8-bit
aggregate0.9140.912
high value0.68 recall0.39 recall

I would not consider it settled without evidence: Compare per-slice metrics against the full-precision model, weighting slices by business value.

Compression loses the tail first.

Curated: · Written: · Reviewed:

QA-40Can you cache model outputs?(show answer)

My answer to caching predictions begins with the contract between training and serving.

A prediction can be cached only when the key covers every input that can change the answer, including the model version and the feature values. A key that omits the version serves the previous model's answers after a deployment.

Concretely, include model version, feature snapshot identity, and any request parameter that changes behaviour in the key, set a TTL tied to feature freshness, and invalidate on deployment rather than relying on expiry.

The reason for that specificity is a failure I have seen: A cache keyed on user id alone served the previous model's scores for 40 percent of requests for six hours after a promotion; the canary measured the new model on the remainder and reported no change.

Traffic after promotion.

SourceShareModel served
cache hit40%previous
computed60%new

I would not consider it settled without evidence: Deploy a new version in staging and confirm the cache returns no pre-deployment entries.

A cache key that omits the version pins the old model.

Curated: · Written: · Reviewed:

QA-41The model service is down. What should the product do?(show answer)

I would treat fallback behaviour when inference fails as a claim about production that has to survive a delayed label.

The fallback is a product decision with a measurable cost, and it must be chosen per use case rather than defaulted. Returning a neutral score is a decision to approve or decline everything, and pretending otherwise hides it.

Concretely, decide per use case whether to fail closed, serve a previous model, apply a rules-based fallback, use a stale cached result, or route to a person, then implement it explicitly and monitor how often it is used.

The reason for that specificity is a failure I have seen: An outage returned a default score of 0.5 which sat below the approval threshold; 74,000 legitimate applications were declined in two hours and the fallback path had never been reviewed.

Outcome of the default fallback.

PathRequestsOutcome
model availablenormal82% approved
default 0.574,0000% approved

I would not consider it settled without evidence: Exercise the fallback in a game day and measure the business outcome it produces.

The fallback is a decision, not a default.

Curated: · Written: · Reviewed:

QA-42A batch scoring job is retried after a partial failure. What must hold?(show answer)

The useful question for idempotency and retries in scoring pipelines is what the system does while nobody is looking at it.

A retried job must produce the same result as one clean run, which means writes are keyed and repeatable rather than appended. Without that, a retry doubles rows and every downstream aggregate is wrong by an amount nobody can reconstruct.

Concretely, key outputs by entity and scoring run, write with an upsert or into a run-scoped partition that replaces atomically, and make the job safe to run twice by construction rather than by scheduling discipline.

The reason for that specificity is a failure I have seen: A retry after a mid-run failure appended a second set of scores for 2.1 million entities; downstream averages were wrong for nine days before the duplicate run identifier was noticed.

After the retry.

EntitiesRows expectedRows written
2,100,0002,100,0004,200,000

I would not consider it settled without evidence: Run the job twice on the same input in staging and assert the output is identical, not doubled.

A pipeline that cannot be re-run cannot be recovered.

Curated: · Written: · Reviewed:

QA-43How should a training pipeline be split into tasks?(show answer)

I would settle orchestration and task boundaries against a replay of real requests before trusting the offline number.

Task boundaries decide what can be retried without repeating expensive work and what can be inspected when something fails. A single task that ingests, transforms, trains, and registers has to be re-run whole for any failure.

Concretely, split at the points where an artifact is produced and worth keeping — validated dataset, transformed features, trained model, evaluation report — and make each step read its input by identifier rather than recomputing it.

The reason for that specificity is a failure I have seen: A monolithic task failed at registration after four hours of training; each retry repeated the whole run, and three retries cost twelve GPU-hours to fix a credentials error.

Cost of a late failure.

DesignRetry cost
one task4 GPU-hours
four tasks2 minutes

I would not consider it settled without evidence: Retry a late-stage failure and confirm the earlier artifacts are reused rather than recomputed.

Boundaries are where you can resume.

Curated: · Written: · Reviewed:

QA-44What do you unit-test in an ML pipeline?(show answer)

The judgement in testing machine learning code is which version is pinned and who approved it.

The deterministic parts — transformations, joins, schema handling, serialisation, and the serving contract — are ordinary code and testable as such. Model quality is not a unit test; it is an evaluation gate with a threshold.

Concretely, test transformations on small fixtures with known answers including nulls and boundaries, test that the serving path produces the training path's vector for the same input, and keep quality checks as separate gates with their own thresholds.

The reason for that specificity is a failure I have seen: A pipeline had no tests on transformations because the team believed ML could not be tested; a timezone bug shifted every daily aggregate by one day and was found by a customer.

The fixture that would have failed.

Input timestampExpected dayProduced
2026-03-01T23:40Z2026-03-012026-03-02

I would not consider it settled without evidence: Assert transformation outputs against hand-computed fixtures covering nulls, boundaries, and timezone edges.

Most of an ML pipeline is ordinary code.

Curated: · Written: · Reviewed:

QA-45What runs on a pull request that changes a model pipeline?(show answer)

Where candidates lose the interview on continuous integration for models is treating a training-time result as a serving guarantee.

The build should verify everything that is cheap and deterministic and defer everything that needs a GPU or a full dataset to a scheduled job. A build that trains a model on every commit stops being run.

Concretely, run linting, type checks, unit tests on transformations, a smoke training run on a tiny fixture, and contract tests for the serving interface on each change, then run full training and evaluation on merge or on a schedule.

The reason for that specificity is a failure I have seen: A pipeline whose CI trained on the full dataset took 70 minutes per commit; the team disabled it, and a broken feature transformation reached production two days later.

Split by cost.

CheckWhereDuration
unit + contractpull request3 min
smoke trainpull request90 s
full train + evalnightly70 min

I would not consider it settled without evidence: Keep the pull-request build under the time the team will actually wait, and verify the full run separately.

A build nobody waits for protects nothing.

Curated: · Written: · Reviewed:

QA-46Three teams share eight GPUs. How do you allocate them?(show answer)

I would answer GPU quota and scheduling for training by separating what the pipeline automated from what it verified.

Unmanaged sharing degrades into whoever launches first, which is neither fair nor aligned with value. Quotas and priorities make the trade explicit and let an urgent retrain preempt an exploratory sweep.

Concretely, set per-team quotas with a priority class for production retraining, require jobs to declare expected duration and checkpoint regularly so preemption is cheap, and publish wait times so the constraint is visible.

The reason for that specificity is a failure I have seen: A hyperparameter sweep occupied all eight GPUs for 40 hours; a production retrain triggered by a data incident waited behind it and the stale model served two extra days.

The queue during the incident.

JobPriorityGPUsWaited
sweepexploratory8—
production retrainhigh040 h

I would not consider it settled without evidence: Confirm a production-priority job preempts exploratory work in a drill, and measure the time it waits.

Priority is what a shared cluster encodes.

Curated: · Written: · Reviewed:

QA-47A 30-hour training run dies at hour 26. What did you lose?(show answer)

The engineering content of checkpointing long training runs is the gate and the rollback, not the model architecture.

Without checkpoints the answer is everything, and the expected cost of a run is far higher than its nominal duration once failure probability is included. Checkpointing converts a failure into a delay.

Concretely, write checkpoints at an interval derived from the failure rate and restart cost, store optimiser state and data-loader position along with weights, and verify resumption by restarting from a checkpoint as part of the pipeline's tests.

The reason for that specificity is a failure I have seen: A spot instance was reclaimed at hour 26 of 30; no checkpoint existed and the run restarted from zero, turning a 30-hour job into 56 hours of wall clock.

Expected cost with and without.

SetupNominalWith one failure
no checkpoints30 h56 h
hourly checkpoints30.5 h31.5 h

I would not consider it settled without evidence: Kill a run deliberately and confirm it resumes from the last checkpoint with matching loss.

A checkpoint is insurance priced in minutes.

Curated: · Written: · Reviewed:

QA-48Should training run on preemptible instances?(show answer)

Before promoting anything through spot and preemptible capacity for training I would write down what silent degradation would look like.

Preemptible capacity is a large discount in exchange for interruption, so it is right for work that can resume and wrong for work that cannot. The decision follows from checkpointing, not from the price alone.

Concretely, use preemptible capacity for checkpointed training and batch scoring with a fallback to on-demand for deadline-bound runs, and measure the effective cost including restarts rather than the headline rate.

The reason for that specificity is a failure I have seen: A team moved all training to spot without checkpointing; restarts consumed 40 percent more compute than was saved, and the deadline for a quarterly model slipped twice.

Effective cost per completed run.

CapacityRateRestartsEffective
on demand$1.000$1.00
spot, no checkpoints$0.351.9 avg$1.02
spot, checkpointed$0.351.9 avg$0.41

I would not consider it settled without evidence: Compare effective cost per completed run, including restarts, against on-demand for the same job.

The discount is real only if the work resumes.

Curated: · Written: · Reviewed:

QA-49What does one prediction from this model cost?(show answer)

The first thing I would pin down about cost per prediction is which population the evidence was measured on.

Cost per prediction is a first-class metric and it usually falls out of utilisation rather than model size. An accelerator held at 8 percent utilisation costs the same as one at 80 percent, so consolidation is the largest lever.

Concretely, attribute infrastructure spend per model, publish cost per thousand predictions next to latency and quality, and treat a persistent low-utilisation endpoint as a design finding rather than an accounting one.

The reason for that specificity is a failure I have seen: Eleven models each held a dedicated GPU endpoint at an average of 6 percent utilisation; consolidating onto three multi-model endpoints cut monthly serving spend from 14,200 to 3,900 dollars with no latency change.

Before and after consolidation.

SetupEndpointsUtilisationMonthly
dedicated116%$14,200
multi-model361%$3,900

I would not consider it settled without evidence: Report utilisation and cost per thousand predictions per model endpoint monthly.

Idle accelerators cost full price.

Curated: · Written: · Reviewed:

QA-50What is the risk of hosting several models on one endpoint?(show answer)

I would start multi-model endpoints and isolation from reproducibility, because an unreproducible result cannot be argued with.

Consolidation shares failure as well as hardware. One model's memory growth or slow request can evict or delay the others, so isolation has to be reasoned about rather than assumed from the packing decision.

Concretely, set per-model memory and concurrency limits, keep noisy or business-critical models on their own capacity, and monitor per-model latency rather than only endpoint-level latency.

The reason for that specificity is a failure I have seen: A newly added model with a 2 GB working set evicted three others from a shared endpoint's cache; their p99 rose from 30 ms to 900 ms and the endpoint-level average hid it.

Per-model p99 after co-tenancy.

ModelBeforeAfter
A (new)—45 ms
B30 ms900 ms
endpoint mean32 ms61 ms

I would not consider it settled without evidence: Report latency and eviction rate per model on a shared endpoint, not per endpoint.

Sharing hardware shares failure.

Curated: · Written: · Reviewed:

QA-51Why should an ML environment be defined in code?(show answer)

This is an area where a model that scores well and a model that behaves well are different events.

A hand-built environment cannot be recreated, which makes disaster recovery, a second region, and a clean test environment all impossible in the same way. It also makes the difference between environments invisible.

Concretely, define endpoints, roles, storage, quotas, and pipeline definitions declaratively, review changes as code, and rebuild a non-production environment from scratch on a schedule to prove the definition is complete.

The reason for that specificity is a failure I have seen: A production endpoint had been modified by hand seven times; recreating it in a second region took nine days of discovering settings, and two were found only by comparing behaviour.

Drift found by the rebuild.

SettingIn codeIn production
concurrency limit832
timeout30 s5 s

I would not consider it settled without evidence: Rebuild the staging environment from the definition and diff it against production.

An environment you cannot rebuild is a single copy.

Curated: · Written: · Reviewed:

QA-52Your training job needs a database credential. Where does it live?(show answer)

My answer to secrets in training and serving begins with the contract between training and serving.

A credential in a notebook, a pipeline definition, or an image layer is a credential in version control and in every artifact built from it. Model pipelines touch more data sources than most services, so the exposure is broad.

Concretely, inject secrets at runtime from a managed store using a workload identity, scope each credential to the least data the job needs, rotate on a schedule, and scan images and repositories for material that should not be there.

The reason for that specificity is a failure I have seen: A warehouse credential was baked into a training image published to a shared registry; it carried read access to eleven schemas and remained valid for five months.

Blast radius of the baked credential.

PropertyValue
schemas readable11
image pulls340
days valid152

I would not consider it settled without evidence: Scan built images and the repository for credential material as a build step, and fail on a hit.

A secret in an image ships with the image.

Curated: · Written: · Reviewed:

QA-53A team wants to load a model downloaded from a public hub. What is your concern?(show answer)

I would treat untrusted model deserialisation as a claim about production that has to survive a delayed label.

Loading a pickled artifact executes code, so an untrusted model file is untrusted code with the permissions of the serving process. The supply-chain question is identical to that for a dependency, and usually less scrutinised.

Concretely, prefer formats that do not execute on load, verify a digest against a recorded source, scan and load in an isolated environment before promotion, and record provenance in the registry entry.

The reason for that specificity is a failure I have seen: A pickled artifact from an unverified mirror was loaded in a job holding credentials for 11 schemas; the review that followed could not rule out execution because no isolation had existed.

Load-time risk by format.

FormatExecutes on loadDigest verified
pickleyesno
safetensorsnoyes

I would not consider it settled without evidence: Verify the artifact digest against its recorded source and load first in a sandbox without production credentials.

Loading a model is running its author's code.

Curated: · Written: · Reviewed:

QA-54Who should be able to move the production alias?(show answer)

The useful question for access control over the registry and endpoints is what the system does while nobody is looking at it.

The ability to move an alias is the ability to change what serves production, and it should be held by fewer identities than the ability to train. Separating those duties is what stops a pipeline from approving its own work.

Concretely, grant training roles registration but not promotion, require an approval record for alias movement, use compare-and-set so concurrent promotions cannot silently overwrite, and audit every mutation.

The reason for that specificity is a failure I have seen: Two automated promotions raced within the same minute; the later write reverted the earlier decision and the deployed version did not match either team's expectation for three days.

The race in the audit log.

TimeActorSet alias to
14:02:11pipeline Av44
14:02:38pipeline Bv41

I would not consider it settled without evidence: Attempt an alias move without approval in staging and require it to fail, and check the audit record for every production move.

Promotion rights are production rights.

Curated: · Written: · Reviewed:

QA-55You log every feature vector for debugging. What have you built?(show answer)

I would settle personal data in prediction logs against a replay of real requests before trusting the offline number.

A complete feature log is a copy of the underlying personal data with a new retention period, a new access path, and usually a weaker one. Debuggability is real but it is not free of obligations.

Concretely, log what is needed for reconstruction — identifiers, versions, and hashes — rather than raw values where the data is sensitive, set retention from the obligation rather than from disk cost, and apply the same access controls as the source.

The reason for that specificity is a failure I have seen: A debug log retained full feature vectors including health-related attributes for 24 months in a bucket readable by the whole data organisation; the source table was restricted to nine people.

Where the copy weakened control.

PropertySource tableFeature log
readers91,400
retention90 days24 months

I would not consider it settled without evidence: Compare the log's access list and retention against the source system's, and reconcile the difference.

A log of features is a copy of the data.

Curated: · Written: · Reviewed:

QA-56A user asks for deletion. What happens to the models trained on their data?(show answer)

The judgement in deletion requests and model artifacts is which version is pinned and who approved it.

Deletion propagates to training sets, feature stores, logs, and evaluation snapshots, and the honest position on trained weights is that removal is bounded rather than perfect. Deciding the policy in advance is the difference between an answer and an incident.

Concretely, record which training snapshots contain an entity so the affected artifacts are queryable, delete from stores and logs on request, and define the retraining cadence at which weights derived from deleted data are replaced.

The reason for that specificity is a failure I have seen: A deletion request was honoured in the warehouse only; the feature store, 3 training snapshots, and the prediction log kept the records, and the gap surfaced in an audit 10 months later.

Coverage of the request.

StoreDeleted
warehouseyes
feature storeno
training snapshotsno
prediction logno

I would not consider it settled without evidence: Trace one deletion request through every store and snapshot and confirm each one is covered.

Deletion is a pipeline, not a table.

Curated: · Written: · Reviewed:

QA-57How do you check a model does not harm one group more than another?(show answer)

Where candidates lose the interview on fairness slices as a release gate is treating a training-time result as a serving guarantee.

The check is a measurement on named groups against a stated metric, decided before the evaluation rather than after a result appears. Which metric matters is a policy decision, because the common definitions cannot all hold at once.

Concretely, agree the protected groups in scope and the metric — false-negative rate parity, calibration within groups, or another — record the choice with its reasoning, and gate promotion on the stated threshold with sample sizes reported.

The reason for that specificity is a failure I have seen: A model was signed off on aggregate accuracy; false-negative rate differed between two groups by a factor of 2.4, and the difference was found by an external review rather than by the pipeline.

The gate that was missing.

GroupAccuracyFalse-negative rate
A0.910.09
B0.900.22

I would not consider it settled without evidence: Report the agreed metric per group with confidence intervals at every promotion, and block on a breach.

Choose the fairness metric before you see the result.

Curated: · Written: · Reviewed:

QA-58What belongs in a model's documentation?(show answer)

I would answer model cards and intended use by separating what the pipeline automated from what it verified.

The document exists so a future consumer can tell whether their use is one the model was evaluated for. That means intended use, population, metrics with their slices, known limitations, and prohibited uses, not an architecture description.

Concretely, generate the card from the registry record so it cannot drift from the artifact, require intended use and limitations to be written by the owner, and version it with the model.

The reason for that specificity is a failure I have seen: A model trained on one country's data was reused in 3 others by a team that read only its 0.94 accuracy figure; the card existed but recorded no population.

What reuse needed and did not find.

FieldPresent
accuracyyes
training populationno
prohibited usesno

I would not consider it settled without evidence: Check the card names the population and prohibited uses, and that it is generated from the registered version.

Documentation is for the next user, not the author.

Curated: · Written: · Reviewed:

QA-59An auditor wants the exact prediction made on 3 February reproduced. What do you need?(show answer)

The engineering content of reproducing a past prediction for an audit is the gate and the rollback, not the model architecture.

The reproduction needs the model version, the serving code and image, and the feature values as they were at that moment. Recomputing features today gives a different input and therefore answers a different question.

Concretely, log the served feature vector or a snapshot reference alongside the resolved model version, retain both for the audit period, and rehearse the reproduction rather than assuming the pieces fit.

The reason for that specificity is a failure I have seen: A regulator asked for six reproductions; feature values had to be recomputed from current data, four of the six differed from the logged decision, and the discrepancy became the finding.

Recomputation against the logged decision.

CaseLogged scoreRecomputed
10.310.31
20.620.48
30.770.55

I would not consider it settled without evidence: Reproduce a sample of historical predictions quarterly and confirm they match the logged outputs.

Reproduce the inputs, not just the model.

Curated: · Written: · Reviewed:

QA-60A customer asks why the model declined them. What can you provide?(show answer)

Before promoting anything through explaining an individual decision I would write down what silent degradation would look like.

An explanation is a product obligation before it is a technique, and the technique has to be chosen so the answer is stable and defensible. An attribution that changes materially between runs on the same input is not an explanation anyone can act on.

Concretely, decide which explanation the decision requires, compute it from the same logged inputs and version that produced the decision, and test its stability by recomputing on identical inputs.

The reason for that specificity is a failure I have seen: Attributions were computed against current data rather than the logged vector; 2 requests about the same decision produced different top reasons and the discrepancy reached a complaint.

Two runs, one decision.

RunTop factorSecond
Atenurebalance
Bregiontenure

I would not consider it settled without evidence: Recompute an explanation twice from the logged inputs and require the top factors to agree.

An unstable explanation is worse than none.

Curated: · Written: · Reviewed:

QA-61Your feature depends on a hosted model you do not control. What changes?(show answer)

The first thing I would pin down about third-party and foundation model dependencies is which population the evidence was measured on.

A hosted model is a dependency that can change beneath you without a version bump, so behaviour that was evaluated last month may not hold today. The provider's versioning and deprecation policy becomes part of your risk.

Concretely, pin an explicit provider version where one is offered, maintain a fixed evaluation set that runs on a schedule to detect behaviour change, budget for deprecation windows, and keep a fallback path.

The reason for that specificity is a failure I have seen: A provider silently updated an unpinned endpoint; classification of one category shifted, and the downstream routing error was diagnosed for eleven days as an internal regression.

Probe set before and after the provider change.

CategoryAgreement with baseline
A99.8%
B61.2%

I would not consider it settled without evidence: Run a fixed probe set against the provider daily and alert on output change.

An unpinned dependency changes without telling you.

Curated: · Written: · Reviewed:

QA-62A prompt template is edited in production. Is that a deployment?(show answer)

I would start prompt and template versioning from reproducibility, because an unreproducible result cannot be argued with.

For a system built on a language model the template determines behaviour as much as the weights do, so an edit is a behaviour change and needs the same version, evaluation, and rollback as any release.

Concretely, store templates in version control, include the template version in the release identity and in every logged call, evaluate a template change on the same fixed set as a model change, and roll back the pair together.

The reason for that specificity is a failure I have seen: A one-word prompt edit made for clarity moved refusal rate on a support workflow from 2 to 19 percent; it was untracked, so the cause took nine days to locate.

Effect of the untracked edit.

TemplateRefusal rateTracked
v72%yes
v8 (edit)19%no

I would not consider it settled without evidence: Log the template version with each call and require an evaluation before a template change reaches production.

The template is part of the model's behaviour.

Curated: · Written: · Reviewed:

QA-63The same input gives different outputs. How do you evaluate that?(show answer)

This is an area where a model that scores well and a model that behaves well are different events.

With a sampled decoder a single run is a draw rather than a measurement, so a comparison needs repeated samples and a stated variance. Two systems compared on one run each are being compared on noise.

Concretely, fix the seed and temperature where the use case allows, otherwise sample repeatedly and report a mean with an interval, and size the evaluation set from the variance rather than from convenience.

The reason for that specificity is a failure I have seen: A model change was accepted on a 4-point gain measured with one sample per case; repeating the evaluation five times showed a mean difference of 0.6 points with an interval spanning zero.

The gain after repetition.

RunsMeasured gain
1+4.0 pts
5+0.6 pts, interval -1.4 to +2.6

I would not consider it settled without evidence: Repeat the evaluation enough times to report an interval, and require the gain to exceed it.

One sample from a sampler is not a measurement.

Curated: · Written: · Reviewed:

QA-64You use a model to grade another model's outputs. What must you check?(show answer)

My answer to model-graded evaluation begins with the contract between training and serving.

A model grader is itself a model with error, bias, and drift, so its agreement with human judgement is the thing being relied on and it needs measuring. An ungraded grader turns evaluation into an assumption.

Concretely, calibrate the grader against a human-labelled sample, report agreement, re-measure it whenever the grader or its prompt changes, and keep a human-labelled set for the decisions that matter most.

The reason for that specificity is a failure I have seen: A grader agreed with human labels 61 percent of the time on the disputed cases; three months of accepted improvements were reversed when a human review of the same outputs disagreed.

Agreement by case type.

Case typeAgreement
clear pass96%
clear fail93%
disputed61%

I would not consider it settled without evidence: Measure grader-human agreement on a stratified sample before trusting any result it produces.

The grader is a model you have not evaluated.

Curated: · Written: · Reviewed:

QA-65An answer is wrong. Was it the retriever or the model?(show answer)

I would treat retrieval quality in a production pipeline as a claim about production that has to survive a delayed label.

A retrieval-backed system has two failure surfaces and one output, so attribution needs the retrieved context logged with the answer. Without it every failure is debated as a model problem.

Concretely, log retrieved document identifiers and scores with each response, measure retrieval separately with recall against a labelled set, and gate index changes on that metric rather than on end-to-end impressions.

The reason for that specificity is a failure I have seen: A quality regression was attributed to a model change for three weeks; an index rebuild had dropped a document set, and retrieval recall had fallen from 0.88 to 0.51 unmeasured.

What the separate metric showed.

PeriodRetrieval recallEnd-to-end quality
before rebuild0.880.82
after rebuild0.510.63

I would not consider it settled without evidence: Track retrieval recall on a labelled query set on the same schedule as end-to-end quality.

Two surfaces need two measurements.

Curated: · Written: · Reviewed:

QA-66A model has no traffic. Can you delete it?(show answer)

The useful question for deprecating and retiring a model is what the system does while nobody is looking at it.

Traffic is one dependency among several; audit obligations, rollback plans, downstream snapshots, and open disputes can all require the artifact after the last request. Retirement is a process with checks rather than a cleanup job.

Concretely, deprecate first with an announced window, prove no endpoint, batch job, rollback target, or legal hold depends on the version, retain lineage and decision evidence, and only then delete artifacts under the retention policy.

The reason for that specificity is a failure I have seen: A cleanup deleted eleven versions with no recent traffic; one was the documented rollback target for a model that failed a week later, and recovery required retraining from scratch.

Dependencies found after deletion.

VersionTrafficRollback targetLegal hold
v22noneyesno
v19nonenoyes

I would not consider it settled without evidence: Check rollback targets and legal holds, not only traffic, before deleting a version.

No traffic is not no dependency.

Curated: · Written: · Reviewed:

QA-67What does the on-call engineer do when a model degrades at 3 a.m.?(show answer)

I would settle on-call for machine learning systems against a replay of real requests before trusting the offline number.

Model incidents differ from service incidents because the mitigations are model-shaped: withdraw the version, fall back, or raise the threshold. The runbook has to name them, because the responder is usually not the model's author.

Concretely, write per-model runbooks stating the fallback, the rollback command, the person who owns the quality decision, and the signals that distinguish a data incident from a model one, and rehearse them.

The reason for that specificity is a failure I have seen: An overnight responder had no model-specific action available, escalated at 03:20, and waited until 08:00 for the owning team; the degraded model served for five hours.

Time to mitigation.

SituationMitigation availableElapsed
no runbookescalate5 h
runbook with rollbackrollback12 min

I would not consider it settled without evidence: Have a responder outside the owning team execute the runbook in a drill and time it.

A responder can only do what the runbook names.

Curated: · Written: · Reviewed:

QA-68Predictions look wrong. How do you tell whether the model or the data broke?(show answer)

The judgement in distinguishing a data incident from a model incident is which version is pinned and who approved it.

The model is a fixed function between deployments, so if nothing was deployed then the inputs changed. Establishing that ordering first prevents the common failure of retraining in response to an upstream outage.

Concretely, check deployment history first, then input validity and freshness, then feature parity, and only then the model — and hold retraining until the input question is settled.

The reason for that specificity is a failure I have seen: A team retrained in response to a quality drop caused by an upstream feed outage; the new model was trained on the corrupted window and made the incident worse.

Order of investigation.

CheckResultConclusion
deployments in windownonenot the model
feature freshness3 days staleupstream

I would not consider it settled without evidence: Confirm inputs are valid and fresh before allowing a retrain in response to a quality alert.

An unchanged model with worse output means changed input.

Curated: · Written: · Reviewed:

QA-69You fix a feature bug. What do you do about the history?(show answer)

Where candidates lose the interview on backfilling a corrected feature is treating a training-time result as a serving guarantee.

Correcting the definition going forward leaves a discontinuity in the training data at the fix date, and a model trained across it learns the artefact. The backfill decision is part of the fix, not follow-up work.

Concretely, backfill the corrected values with the event times they should have had, record the correction as a versioned change so the training set can be reconstructed either way, and re-run the affected evaluations.

The reason for that specificity is a failure I have seen: A fix applied forward only left 14 months of history on the old definition; the next model learned the fix date as a signal and its importance ranked third.

The discontinuity the model learned.

PeriodMean value
before fix0.42
after fix1.09

I would not consider it settled without evidence: Compare feature distributions either side of the fix date and require no discontinuity before training across it.

A forward-only fix is a step change in the data.

Curated: · Written: · Reviewed:

QA-70Should training and serving share transformation code?(show answer)

I would answer training/serving code sharing by separating what the pipeline automated from what it verified.

They must share semantics; sharing a runtime is one way to get it and not the only one. What matters is that a difference is detected, because two independent implementations of the same transformation diverge on the cases nobody thought about.

Concretely, share the implementation where the runtimes allow, and where they do not, generate both from one definition or run a parity test on real entities that fails the build on mismatch.

The reason for that specificity is a failure I have seen: A Python training transformation and a Java serving reimplementation agreed on every unit test and differed on null handling for one field, affecting 3 percent of requests for seven months.

Where they differed.

CaseTrainingServing
value present0.620.62
value nullmedian0.0

I would not consider it settled without evidence: Run both implementations over sampled production entities and require exact agreement.

Two implementations need a test that compares them.

Curated: · Written: · Reviewed:

QA-71A data scientist hands you a notebook that produces a good model. What now?(show answer)

The engineering content of moving a notebook into a pipeline is the gate and the rollback, not the model architecture.

The notebook is a description of a result, not a pipeline, because it depends on hidden state: cells run out of order, local files, and manual edits. Productionising means making every input explicit rather than translating the code.

Concretely, extract parameters and data references into configuration, run the whole thing top to bottom in a clean environment, pin dependencies, and compare the resulting metrics against the notebook's before changing anything else.

The reason for that specificity is a failure I have seen: A converted pipeline scored four points below the notebook; the notebook had a manually filtered dataframe from a cell that had been deleted, and reconciling the two took a week.

Reconciliation after conversion.

SourceRows usedAUC
notebook as run812,0000.874
pipeline1,204,0000.831

I would not consider it settled without evidence: Run the notebook end to end in a fresh kernel and require it to reproduce its own reported metric before conversion starts.

A notebook that only runs in order is already broken.

Curated: · Written: · Reviewed:

QA-72Why record experiments that did not work?(show answer)

Before promoting anything through experiment tracking I would write down what silent degradation would look like.

Failed runs are the record of what has been ruled out, and without them a team re-tries the same idea every few months. Tracking is a research asset before it is a compliance one.

Concretely, log parameters, dataset version, code revision, environment, metrics, and outcome for every run automatically from the pipeline, not by hand, so the record does not depend on whether anyone remembered.

The reason for that specificity is a failure I have seen: Three engineers independently tried the same embedding approach over eight months because none of the failures had been recorded; the total cost was 46 GPU-days.

The duplicated work.

AttemptMonthOutcomeGPU-days
1Febabandoned14
2Junabandoned17
3Octabandoned15

I would not consider it settled without evidence: Search the tracker for an approach before starting it, and require every run to be logged automatically.

An untracked failure is repeated.

Curated: · Written: · Reviewed:

QA-73How long should a hyperparameter sweep run?(show answer)

The first thing I would pin down about hyperparameter search budgets is which population the evidence was measured on.

A sweep has diminishing returns that can be measured, and the honest comparison is against what the same compute would buy elsewhere — more data, better features, or faster iteration. Search is rarely the highest-value use of a cluster.

Concretely, bound the budget in advance, use early stopping on unpromising trials, record the best-so-far curve, and stop when the curve flattens rather than when the calendar says so.

The reason for that specificity is a failure I have seen: A 900-trial sweep improved validation AUC by 0.003 over the first 40 trials; the same compute spent on label collection moved it by 0.021.

Returns from the sweep.

TrialsBest AUC
400.868
3000.870
9000.871

I would not consider it settled without evidence: Plot best-so-far against trials and stop where the curve flattens.

Search until the curve flattens, then spend elsewhere.

Curated: · Written: · Reviewed:

QA-74When is distributed training worth the complexity?(show answer)

I would start distributed training and its failure modes from reproducibility, because an unreproducible result cannot be argued with.

Multiple workers add synchronisation, fault handling, and a new class of silent errors in exchange for wall-clock time. Below the point where a single machine cannot hold the model or the data, the complexity buys very little.

Concretely, establish the single-node baseline first, measure scaling efficiency rather than assuming it, keep effective batch size and learning rate consistent when comparing, and checkpoint so a worker loss is recoverable.

The reason for that specificity is a failure I have seen: An eight-worker setup delivered 2.3 times the throughput of one machine while the effective batch size changed unnoticed; the resulting model was worse and the cause took a month to find.

Scaling efficiency.

WorkersThroughputEfficiency
11.0x100%
82.3x29%

I would not consider it settled without evidence: Compare distributed and single-node runs at matched effective batch size before accepting a speedup.

Measure the scaling before paying for it.

Curated: · Written: · Reviewed:

QA-75Training fails with an out-of-memory error. What do you change first?(show answer)

This is an area where a model that scores well and a model that behaves well are different events.

Memory is a function of model size, batch size, sequence length, optimiser state, and activation storage, and each has a different cost in accuracy or time. Reaching for a bigger instance first is the most expensive of the available answers.

Concretely, reduce batch size with gradient accumulation to hold the effective batch, enable activation checkpointing, use mixed precision, and only then change hardware — measuring the accuracy effect of each step.

The reason for that specificity is a failure I have seen: A team moved to instances costing 3.4 times as much to fix an out-of-memory failure that gradient accumulation resolved with no accuracy change and no extra cost.

Options at matched effective batch.

RemedyCostAccuracy
larger instance3.4xunchanged
gradient accumulation1.0xunchanged

I would not consider it settled without evidence: Compare accuracy and cost at matched effective batch size for each memory remedy before changing hardware.

Effective batch size is what matters, not the instance.

Curated: · Written: · Reviewed:

QA-76A feature must reflect the last five minutes. How do you compute it?(show answer)

My answer to streaming feature computation begins with the contract between training and serving.

A streaming aggregate introduces windowing, watermarks, and late events, and each is a place where the training and serving values can diverge. The window's semantics have to be identical in both paths or the feature means two things.

Concretely, define window type, size, and allowed lateness once, apply the same definition when reconstructing the feature for training, and test with deliberately out-of-order and late events.

The reason for that specificity is a failure I have seen: The training path computed a five-minute window over complete history while the stream dropped events arriving more than 30 seconds late; 8 percent of values differed and the model was tuned against the wrong one.

Effect of late events.

PathLate events includedMean value
training reconstructionall12.4
streamunder 30 s11.4

I would not consider it settled without evidence: Replay a stream with late and out-of-order events through both paths and require matching values.

A window is a definition, not an implementation detail.

Curated: · Written: · Reviewed:

QA-77You change the embedding model. What happens to the index?(show answer)

I would treat embedding and vector index operations as a claim about production that has to survive a delayed label.

Embeddings from two models are not comparable, so a changed encoder invalidates every stored vector at once. Mixing old and new vectors in one index produces silently meaningless neighbours rather than an error.

Concretely, version the index with the encoder, rebuild fully before switching, keep the old index serving until the new one is validated, and stamp both query and stored vectors with the encoder version.

The reason for that specificity is a failure I have seen: An encoder upgrade re-embedded new documents only; queries returned a mixture of spaces, recall fell from 0.88 to 0.34, and no error was raised anywhere.

Recall during the mixed period.

Index stateRecall
all v10.88
mixed v1/v20.34
all v20.90

I would not consider it settled without evidence: Assert the encoder version of the query matches the index's, and reject the query when it does not.

Two embedding spaces do not share a neighbourhood.

Curated: · Written: · Reviewed:

QA-78A nightly scoring job must finish by 6 a.m. How do you hold that?(show answer)

The useful question for batch scoring service levels is what the system does while nobody is looking at it.

A batch job with a downstream deadline is a service with an objective, and the objective is the completion time rather than the average runtime. Growth in input volume erodes it silently until the first miss.

Concretely, track runtime against the deadline with headroom, alert on trend rather than on breach, partition so the job can run in parallel as volume grows, and define what downstream consumers do when it is late.

The reason for that specificity is a failure I have seen: A job that took 90 minutes at launch reached 5 h 40 m over two years; the first miss was discovered by a marketing team whose campaign sent without scores.

Runtime against the window.

QuarterRuntimeWindow used
launch1 h 30 m25%
year 13 h 10 m53%
year 25 h 40 m94%

I would not consider it settled without evidence: Alert when projected runtime crosses 70 percent of the window, not when the deadline is missed.

A deadline is missed by a trend, not an event.

Curated: · Written: · Reviewed:

QA-79A cheap model filters and an expensive one decides. What is the risk?(show answer)

I would settle model cascades and chained models against a replay of real requests before trusting the offline number.

A cascade's overall quality is bounded by the first stage's recall, because anything the filter drops is never seen by the model that could have caught it. The cheap stage sets the ceiling.

Concretely, measure end-to-end rather than per stage, tune the filter for recall rather than precision, sample a share of filtered items through the full path to estimate what is being lost, and monitor that estimate.

The reason for that specificity is a failure I have seen: A filter with 0.94 precision and 0.72 recall capped the system at 0.72 recall; per-stage dashboards showed both models performing well and the loss was invisible for a year.

Stages against the whole.

MeasureValue
filter recall0.72
decider recall on what it sees0.95
end to end0.68

I would not consider it settled without evidence: Route a random sample past the filter and measure what the full model would have decided.

The first stage sets the ceiling.

Curated: · Written: · Reviewed:

QA-80Where should a person sit in an automated decision pipeline?(show answer)

The judgement in human in the loop is which version is pinned and who approved it.

Review capacity is finite, so it should be spent where the model is least certain and the cost of error is highest. Routing a fixed percentage at random spends most of it on cases the model already handles.

Concretely, route by predicted uncertainty and business impact, size the queue to actual reviewer capacity, feed reviewed outcomes back as labels, and monitor review latency as part of the decision's end-to-end time.

The reason for that specificity is a failure I have seen: A 5 percent random sample sent for review consumed the whole team's capacity on routine cases while the uncertain band went straight through; the errors were concentrated exactly there.

Where the errors were.

BandShare of volumeError rateReviewed
confident88%0.4%5%
uncertain12%9.1%5%

I would not consider it settled without evidence: Compare error rates inside and outside the review queue to check the routing is selecting the right cases.

Spend review capacity where the model is unsure.

Curated: · Written: · Reviewed:

QA-81How do you build a labelled set from a system already running?(show answer)

Where candidates lose the interview on collecting labels from production is treating a training-time result as a serving guarantee.

Labels harvested from the deployed system inherit its selection: cases it never surfaced are absent, and cases it flagged are over-represented. Treating that set as a sample of the population biases every model trained on it.

Concretely, reserve a randomised slice that bypasses the model's selection, weight harvested labels by their selection probability, and keep an audited set labelled independently of the model's decisions.

The reason for that specificity is a failure I have seen: A fraud team labelled only reviewed cases; the model learned to detect what it already detected and its recall on unreviewed fraud was never measured for two years.

Recall by label source.

Label sourceMeasured recall
reviewed queue0.91
random audit0.48

I would not consider it settled without evidence: Label a random unselected sample periodically and measure performance on it separately.

A label set built by the model describes the model.

Curated: · Written: · Reviewed:

QA-82Managed platform or your own stack?(show answer)

I would answer build against buy for the ML platform by separating what the pipeline automated from what it verified.

The decision is about which constraints you accept and which you can afford to maintain, not about capability lists. A managed platform trades flexibility and unit cost for the engineering you do not have to staff.

Concretely, cost the option including the engineers it needs, name the constraints that would be unacceptable, prototype against the real workload, and prefer components with an exit path over ones that hold the artifacts.

The reason for that specificity is a failure I have seen: A self-built platform consumed four engineers for two years to reach parity with a managed service costing a fraction of their salaries, while model delivery stalled.

Two-year comparison.

OptionPlatform spendEngineersModels shipped
self-builtlow42
managedhigher111

I would not consider it settled without evidence: Compare total cost including staffing, and test the migration path out before committing.

Count the engineers, not the feature list.

Curated: · Written: · Reviewed:

QA-83You are moving 30 models to a new serving stack. How do you do it safely?(show answer)

The engineering content of migrating models between platforms is the gate and the rollback, not the model architecture.

A migration is thirty behaviour-preserving deployments, and equivalence has to be proved per model rather than assumed from the platform. The evidence is agreement on real requests, not a successful deployment.

Concretely, run old and new in parallel on mirrored traffic, compare outputs per model, migrate in order of risk, and keep the old path deployable until the comparison has covered a full traffic cycle.

The reason for that specificity is a failure I have seen: A big-bang cutover moved all thirty at once; two had different preprocessing defaults, and the resulting incident could not be attributed for days because everything had changed together.

Agreement found during parallel run.

ModelsAgreementCut over
28100%yes
296.2%held

I would not consider it settled without evidence: Require output agreement above a stated threshold on mirrored traffic before each model is cut over.

Migrate one model at a time with evidence.

Curated: · Written: · Reviewed:

QA-84What is different about running a model in two regions?(show answer)

Before promoting anything through multi-region serving for models I would write down what silent degradation would look like.

The model must be the same version in both, and the features it reads must be equally fresh, or the same customer gets different decisions depending on routing. Consistency of the model is easier than consistency of its inputs.

Concretely, promote versions to regions through the same pipeline with a recorded order, monitor per-region feature freshness and score distributions, and decide explicitly whether a region serves stale features or fails over.

The reason for that specificity is a failure I have seen: A promotion reached one region and stalled in another for six hours; identical customers received different decisions and the discrepancy was reported by support rather than by monitoring.

During the stalled promotion.

RegionVersionFeature age
euv524 min
usv503 h

I would not consider it settled without evidence: Alert when the serving version or feature freshness differs between regions beyond a stated window.

Two regions are two chances to serve different models.

Curated: · Written: · Reviewed:

QA-85Your model registry and artifact store are lost. What is your recovery position?(show answer)

The first thing I would pin down about disaster recovery for model systems is which population the evidence was measured on.

Recovery needs the artifacts, their lineage, and the pipeline that can rebuild them, and the weakest of the three sets the actual position. A backup of weights without lineage restores serving but not the ability to change anything.

Concretely, back up artifacts and registry metadata together, replicate across a failure boundary, and rehearse a restore into a clean account rather than trusting that the backup exists.

The reason for that specificity is a failure I have seen: A restore recovered 14 artifacts but not the registry metadata; the team could serve the current model and could not tell which dataset produced it, so the next retrain started from scratch.

What the drill recovered.

AssetRestoredEnables
artifactsyesserving
registry metadatanonothing further

I would not consider it settled without evidence: Restore into an empty environment in a drill and confirm both serving and retraining are possible.

Restore the lineage, or you restore a dead end.

Curated: · Written: · Reviewed:

QA-86Nothing was deployed but behaviour changed. Where do you look?(show answer)

I would start configuration as a source of behaviour change from reproducibility, because an unreproducible result cannot be argued with.

Feature flags, thresholds, routing weights, and provider settings change behaviour without a deployment, and are usually the least audited surface in the system. Any change history that covers only code is incomplete.

Concretely, version configuration alongside code, record every change with actor and time in the same timeline as deployments, and require the same review for a threshold change as for a model change.

The reason for that specificity is a failure I have seen: A routing weight edited in a console shifted 30 percent of traffic to an experimental model; the deployment history showed no change and the investigation looked at the model for two days.

The timeline that was missing.

TimeChangeIn deploy log
09:12routing weight 5% to 30%no
—model versionunchanged

I would not consider it settled without evidence: Produce a single change timeline covering code, models, and configuration when investigating an incident.

Behaviour changes wherever configuration does.

Curated: · Written: · Reviewed:

QA-87A scheduled pipeline succeeded but produced nothing. How would you know?(show answer)

This is an area where a model that scores well and a model that behaves well are different events.

Exit status describes the process, not the work. A job that processes zero rows because its source was empty exits successfully, and a scheduler-only view reports it as healthy.

Concretely, assert output volume and freshness as part of the job, alert on rows written and on the age of the newest output, and treat an unexpectedly empty run as a failure rather than as a quiet success.

The reason for that specificity is a failure I have seen: A feature job ran green for nine days while a permission change made its source return no rows; the feature aged out of TTL and models silently fell back.

The nine green days.

DayExit statusRows written
normalsuccess4.1m
during outagesuccess0

I would not consider it settled without evidence: Fail the job when output row count falls outside its expected band, not only when it throws.

Green means it ran, not that it worked.

Curated: · Written: · Reviewed:

QA-88You version datasets with content hashes. What still goes wrong?(show answer)

My answer to dataset versioning tools and their limits begins with the contract between training and serving.

A hash proves two datasets are identical and says nothing about whether either is correct. Versioning gives reproducibility and comparison, not validity, and teams routinely conflate the two.

Concretely, pair versioning with validation — schema, ranges, volumes, and known invariants — so a version records both what the data was and that it met its expectations at the time.

The reason for that specificity is a failure I have seen: A corrupted extract was versioned, hashed, and reproduced faithfully for 4 months; all 6 models trained on it were consistent and wrong.

What the version recorded.

PropertyRecorded
content hashyes
schema checkno
null-rate checkno

I would not consider it settled without evidence: Attach validation results to the dataset version, and refuse to train on a version without them.

A hash proves sameness, not soundness.

Curated: · Written: · Reviewed:

QA-89Where does ML-specific debt accumulate?(show answer)

I would treat technical debt specific to ML systems as a claim about production that has to survive a delayed label.

The debt is in the couplings that code review does not see: features consumed by models nobody tracks, thresholds tuned once, pipelines whose outputs feed other models, and glue that outlives its purpose. It compounds because changing an input changes every downstream model.

Concretely, maintain a dependency map from features to models to decisions, retire unused features and models on a schedule, and require a consumer list before any feature's definition changes.

The reason for that specificity is a failure I have seen: A feature deprecated by its owning team was still read by three models in other teams; removing it degraded a revenue-critical model with no warning.

Consumers found after the removal.

FeatureKnown consumersActual
tenure_bucket14

I would not consider it settled without evidence: Query the consumer list for a feature before changing or removing it, and keep the list generated rather than maintained by hand.

Every shared feature is a coupling.

Curated: · Written: · Reviewed:

QA-90Your model's AUC improved and revenue did not. What went wrong?(show answer)

The useful question for aligning model metrics with business outcomes is what the system does while nobody is looking at it.

A statistical metric is a proxy for a business outcome and the mapping is rarely linear. A gain concentrated where decisions do not change, or on a population that does not convert, moves the metric without moving anything else.

Concretely, state the decision the model drives and the business measure it should move, evaluate at the operating point rather than across the whole curve, and confirm the gain lands where decisions actually flip.

The reason for that specificity is a failure I have seen: A 3-point AUC gain came entirely from better ordering within the band that was auto-approved anyway; no decision changed and revenue was flat.

Where the gain landed.

BandAUC gainDecisions changed
auto-approve+0.040
borderline+0.000

I would not consider it settled without evidence: Measure how many decisions change under the new model, not how the ranking metric moved.

A gain that flips no decision buys nothing.

Curated: · Written: · Reviewed:

QA-91How do you tell a business owner what the model can and cannot do?(show answer)

I would settle communicating model risk to stakeholders against a replay of real requests before trusting the offline number.

The useful statement names the population it was validated on, the operating point, the expected error rate at that point, and the conditions that would invalidate it. A single accuracy figure invites use well outside what was tested.

Concretely, publish the operating point with expected volumes of each error type in business units, state the invalidating conditions explicitly, and repeat both when the model is reused for a new purpose.

The reason for that specificity is a failure I have seen: A model described as 94 percent accurate was applied to a segment absent from training; the owner had no way to know that mattered, and the error rate there was 38 percent.

The statement that was needed.

ItemValue
validated populationexisting customers, EU
operating point0.62
expected false positives210 per week

I would not consider it settled without evidence: State expected false positives and false negatives per week in business units, alongside the population validated.

Name the population, not just the number.

Curated: · Written: · Reviewed:

QA-92Who owns a model in production?(show answer)

The judgement in ownership boundaries between data science and platform is which version is pinned and who approved it.

Undefined ownership resolves at the worst possible moment, during an incident. The split that works names one owner for model quality and one for the serving system, with an agreed escalation between them.

Concretely, record the quality owner and the platform owner per model in the registry, put both in the runbook, and require the quality owner to sign the promotion decision.

The reason for that specificity is a failure I have seen: A degraded model sat between 2 teams for 4 days, each believing the other held it; the registry recorded no owner and the runbook named a team that had been reorganised.

What the registry recorded.

FieldValue
quality owner(empty)
platform ownerteam since reorganised

I would not consider it settled without evidence: Check that every serving model names a current quality owner and platform owner before it is promoted.

An unowned model is owned during the incident.

Curated: · Written: · Reviewed:

QA-93A generative feature needs a safety filter. What does that change operationally?(show answer)

Where candidates lose the interview on safety filters around generative outputs is treating a training-time result as a serving guarantee.

The filter is a second model in the request path with its own error rates, latency, and failure modes. Its false positives are a product problem and its false negatives are a safety one, so both need measuring separately.

Concretely, version and evaluate the filter like any model, measure block rate by slice, log blocked outputs for review under access controls, and define behaviour when the filter itself is unavailable.

The reason for that specificity is a failure I have seen: A filter update raised the block rate on one language from 3 to 27 percent; the aggregate block rate moved by less than a point and nothing alerted for three weeks.

Block rate after the update.

SliceBeforeAfter
aggregate3.4%4.1%
one language3.0%27.0%

I would not consider it settled without evidence: Monitor block rate by slice against the rate recorded when the filter was evaluated.

The filter is a model with its own errors.

Curated: · Written: · Reviewed:

QA-94Your provider rate-limits you at peak. What is the design answer?(show answer)

I would answer rate limits and quotas on model providers by separating what the pipeline automated from what it verified.

A provider quota is a capacity constraint that arrives as errors rather than as queueing, so the system has to decide what to shed and what to defer. Retrying blindly converts a limit into an outage.

Concretely, classify requests by priority, queue what can wait with bounded depth, shed or degrade the rest deliberately, back off with jitter, and monitor headroom against the quota rather than only the error rate.

The reason for that specificity is a failure I have seen: Uniform retries on a rate limit tripled request volume during the peak; the limit held for 40 minutes instead of the 4 minutes the burst would have lasted.

Retry behaviour during the peak.

StrategyRequests sentDuration
immediate retry3.1x40 min
backoff with shedding1.2x4 min

I would not consider it settled without evidence: Load-test against the quota and confirm the system degrades in the intended order.

A quota is capacity that fails loudly.

Curated: · Written: · Reviewed:

QA-95The ML bill doubled. How do you find out why?(show answer)

The engineering content of cost attribution across teams is the gate and the rollback, not the model architecture.

Spend without attribution cannot be argued with, because no team recognises its share. Tagging by model, team, and environment turns a total into a set of decisions someone can make.

Concretely, enforce tags at resource creation, report spend per model and per environment monthly, and set budgets with alerts at the team level rather than at the account level.

The reason for that specificity is a failure I have seen: An untagged development cluster accounted for 41 percent of a doubled bill and had been idle for two months; the investigation took three weeks because nothing was attributed.

The bill once attributed.

OwnerShareIdle
dev cluster41%2 months
production serving38%no
training21%no

I would not consider it settled without evidence: Require tags at creation and report the untagged share as a defect.

Unattributed spend is nobody's to cut.

Curated: · Written: · Reviewed:

QA-96A vendor claims their model beats yours. How do you test it?(show answer)

Before promoting anything through evaluating a vendor model against your own I would write down what silent degradation would look like.

A comparison is only meaningful on your data, at your operating point, with your costs of error. A vendor benchmark answers a question about their evaluation set, which is rarely the population you serve.

Concretely, run both on a held-out set drawn from your traffic, at matched operating points, and include integration cost, latency, and the dependency risk in the comparison rather than only the metric.

The reason for that specificity is a failure I have seen: A vendor model reported 6 points better on a public benchmark and was 3 points worse on the company's own held-out set, where the class balance was very different.

Two evaluation sets.

ModelPublic benchmarkYour traffic
vendor0.9120.844
incumbent0.8510.873

I would not consider it settled without evidence: Score both models on the same sample of your production traffic at the same operating point.

Benchmarks describe benchmarks.

Curated: · Written: · Reviewed:

QA-97When is a rule better than a model?(show answer)

The first thing I would pin down about when not to use a model is which population the evidence was measured on.

Where the logic is known, stable, and legally required to be explainable, a rule is cheaper, testable, and auditable. Choosing a model for a problem that a rule solves adds pipelines, monitoring, and drift for no gain.

Concretely, establish the rule-based baseline first, require a model to beat it by a margin that justifies its operating cost, and record the comparison so the decision can be revisited.

The reason for that specificity is a failure I have seen: A model replaced a four-line eligibility rule; it matched the rule 99.4 percent of the time, added a retraining pipeline and a monitoring surface, and the 0.6 percent disagreement was the model being wrong.

Model against the rule.

MeasureRuleModel
agreement with policy100%99.4%
ongoing costnonepipeline + monitoring

I would not consider it settled without evidence: Compare the model against the rule baseline on the same data and require a margin before adopting it.

A model needs to beat the rule, not just exist.

Curated: · Written: · Reviewed:

QA-98A caller sends a field the model has never seen. What should the endpoint do?(show answer)

I would start validating the request at the endpoint from reproducibility, because an unreproducible result cannot be argued with.

A model will produce a number for almost any input, so an unvalidated request turns a caller's mistake into a confident prediction. The endpoint's schema is the boundary where a malformed input can still be rejected rather than scored.

Concretely, validate types, ranges, required fields, and category membership against the model's recorded signature, reject with a clear error rather than coercing, and count rejections per caller as a monitored signal.

The reason for that specificity is a failure I have seen: A client sent a category string the encoder had never seen; it was silently mapped to the unknown bucket, and 22 percent of that caller's traffic scored on a default for five weeks.

What the unknown bucket absorbed.

CallerRequestsUnknown category
A1.2m0.1%
B340k22%

I would not consider it settled without evidence: Send a deliberately malformed request in a test and require an explicit rejection rather than a prediction.

A model will score nonsense without complaining.

Curated: · Written: · Reviewed:

QA-99You join a team with models in production and no platform. Where do you start?(show answer)

This is an area where a model that scores well and a model that behaves well are different events.

Start where the risk is, which is almost always the absence of a record of what is serving and the absence of a way to put it back. Registry and rollback beat any amount of pipeline automation on day one.

Concretely, inventory what is serving and how it was produced, get artifacts and lineage into a registry, establish a tested rollback, and only then automate training — because automation multiplies whatever discipline already exists.

The reason for that specificity is a failure I have seen: A team automated retraining before establishing a registry; the pipeline promoted a bad model, the previous artifact had been overwritten, and recovery meant 2 days of retraining from an uncertain snapshot.

Order of work.

StepBefore automationRisk removed
inventoryyesunknown models
registry + lineageyesunreproducible artifacts
tested rollbackyesunrecoverable promotion

I would not consider it settled without evidence: Confirm a tested rollback exists for every serving model before automating promotion.

Automate after you can reverse.

Curated: · Written: · Reviewed:

QA-100How would you know an ML platform is in good shape?(show answer)

My answer to measuring the health of an ML platform begins with the contract between training and serving.

The indicators are that every serving model is reproducible, that promotion and rollback are routine and timed, that quality is monitored under delayed labels, and that unused models are retired. Those are properties of the system rather than counts of models shipped.

Concretely, track reproducibility rate, median time to rollback, share of models with an active quality monitor, and retired-model count, and treat each as a maintained number rather than a project outcome.

The reason for that specificity is a failure I have seen: A platform measured only models deployed; two years in, no model could be rebuilt from its record and rollback had never been tested outside an incident.

Platform health over two years.

MeasureYear 1Year 2
models reproducible0 of 1414 of 14
median rollback timeuntested9 min
models with quality monitor3 of 1414 of 14
models retired06

I would not consider it settled without evidence: Report the four health measures monthly and act on the one that is worst.

Health is reproducibility and reversibility, not volume.

Curated: · Written: · Reviewed: