Skip to content
Tech Interview Prep home
Technical interview guide

Model Versioning & Registries

Tracking every trained model artifact alongside the data and code that produced it, so results are reproducible.

Read
45 min
Practice MCQs
25
Interview QA
25
Edition
v4
Editorial status
Reviewed
Relevant for
MLOps Engineer

Scope: MLflow, SageMaker AI, Vertex AI, Azure Machine Learning, and NIST AI RMF guidance current 2026-09-01.

Overview

Curated: · Written: · Reviewed:

Make every deployed model reproducible, reviewable, and reversible

A model registry is a governed catalog of deployable model versions and their lifecycle evidence. It is not merely object storage, an experiment dashboard, or a list called production. The registry connects an immutable model artifact to the training run, code, data snapshot, feature definitions, environment, evaluation, intended use, approvals, deployment history and monitoring baseline needed to understand and reproduce it.

In an interview, this topic shows up as one of two questions. The narrow one — "what's a model registry and how is it different from an artifact store?" — is a vocabulary check; a strong answer names the binding between artifact bytes and lifecycle metadata and moves on. The broad one — "how do you know what model is serving in production right now, and how would you roll it back?" — is where senior candidates are separated. Interviewers are probing whether you treat the registry as a supply-chain control plane with invariants (immutability, resolved identity, audited promotion) or as a filing cabinet where teams drop files. A weak answer describes UI stages and version numbers without ever saying what happens when an alias moves mid-incident, or why a version must never be overwritten.

What a registry is, and what it is not

In a production stack the registry is the system of record that binds one immutable artifact to its metadata: version, lineage, metrics, owner, stage or alias. An artifact store answers "where are the bytes and are they intact?" An experiment tracker answers "what did we try and what happened?" The registry answers "what is deployable, what is approved, what is serving, and why." A plain store can hold a model.pkl; only a registry can say which bytes were approved by whom, against which evaluation, for which intended use.

Define model identity precisely, because interviewers probe this. A product or registered-model name groups comparable versions; a version identifies one immutable package; a content digest verifies the bytes. The package must include or reference weights, preprocessing/postprocessing, inference code, dependency lock, runtime image, model signature and schema. Re-register any byte or contract change as a new version. Never overwrite an artifact behind an existing version — that destroys auditability and rollback safety, and it is the single fastest way to fail this question.

The versioning lifecycle and why immutability makes rollback safe

The lifecycle is: register a version under a named model → promote it between environments → archive or retire it. Registration creates an immutable version; promotion moves a pointer or records a decision; retirement removes the artifact under policy. Nothing in that sequence mutates the version itself.

Version numbers are identifiers, not quality rankings. Version 42 is newer than 41 within one registry namespace, but not necessarily better, safer, approved, or equivalent across environments. Tags and descriptions annotate facts. Approval status records a governance decision. Aliases such as champion, candidate or default are mutable pointers to versions. Resolve and record the immutable version/digest at deployment time; an alias alone cannot prove what served a historical prediction.

Immutability is what makes rollback a pointer move rather than a rebuild. If version 41's bytes can never change, "roll back" is: repoint the alias, restart or reconfigure serving, done — with the full bundle (preprocessing, inference code, dependencies) still intact. If bytes can be overwritten, rollback is archaeology.

MLflow mechanics: versions, aliases, and the deprecated Stages

If the conversation turns to tooling, it usually turns to MLflow, and there is a specific fact worth knowing cold: MLflow Model Registry's built-in stages — Staging, Production, Archived — are deprecated in favor of aliases like @champion and @challenger. An alias is a mutable, named pointer attached to a registered model that resolves to exactly one version at a time. Moving it is a single atomic reassignment — set @champion to version 7 and version 6 stops being the champion in the same operation; there is no window where the alias resolves to two versions or none.

The distinction matters in design conversations. Stages were a fixed enum baked into the registry, so every team's process had to fit three buckets. Aliases are arbitrary names, so you can encode your own topology: @champion and @challenger for the serving pair, @canary for the 5% traffic version, @shadow for the mirror. A deployment references models:/fraud-model@champion, and the registry resolves it to a concrete version and digest. If a candidate describes Staging/Production stages as the current mechanism, that is a version-sensitive tell — worth naming the deprecation and moving on.

Promotion as a guarded, auditable transition

Promotion is a controlled decision, not artifact mutation. A candidate passes automated integrity, signature, dependency, security, compatibility and performance gates, then risk-based human approval where required. Separate duties so a training job cannot silently approve and deploy itself to production. Use environment-specific registries or permissions where appropriate, and promote provenance or code through a repeatable pipeline. Record who approved which immutable version with which evidence, and expiry/review triggers.

The gates a candidate must clear before an alias moves: validation metrics against thresholds on a frozen evaluation set, model quality tests (slices, robustness, fairness, latency/resource), signature and schema compatibility with the serving contract, security scans of the serialized artifact and dependencies, and sign-off where the risk profile demands it. Metrics without intended-use thresholds are observations, not approval. Keep test data independent of training and guard against evaluation leakage. Document known limitations, prohibited uses and human-oversight requirements.

Capture lineage at registration: source repository and commit, dirty-tree state, pipeline/run identifier, training and validation dataset version or query plus snapshot time, feature transformation version, random seeds, hyperparameters, base model, licenses, environment and responsible owner. Data lineage must be reproducible rather than a mutable table name. For foundation-model adaptations, include base checkpoint digest, tokenizer, prompt/template, adapters and merge method.

The registry in the serving path

Deployments should pin immutable identity even when operators choose through an alias. The deployment record links version/digest, serving code and image, feature service version, configuration, endpoint, traffic fraction, region, time and change identity. Move an alias through an audited compare-and-set operation so concurrent promotion cannot silently replace a newer decision. A mutable default alias is convenient for exploration but dangerous as an implicit production selector.

Here is the failure mode interviewers are fishing for: serving should not depend on live registry lookups. A service resolves models:/fraud-model@champion once at startup or deploy time, pins the resolved version and digest, and loads the artifact — from the registry, from a mirror, or from an image-baked copy. If every request hits the registry to resolve the alias, you have coupled your p99 to the registry's availability and made an alias move a live production change with no rollout. When the registry is down, serving continues on the last resolved artifact; what breaks is new deployments and alias moves, which is the correct blast radius. A weak answer says "the serving layer queries the registry for each request" without noticing it just made the registry a single point of failure for inference.

Rollout and rollback require more than preserving old weights. Validate input/output contracts and features, canary or shadow safely, watch technical and business guardrails, and define automatic halt. Rollback must restore the full serving bundle and compatible feature/config dependencies. Retain the last-known-good version and evidence, but investigate whether data distribution or upstream systems changed so that old model behavior is still valid.

Security, metadata drift, and deletion

Secure the registry as a software supply-chain control plane. Apply least privilege to artifact upload, metadata changes, approval, alias movement, deletion and deployment. Scan serialized formats and dependencies, prefer safe loaders, sign or attest artifacts, verify digests at promotion and load, encrypt data, restrict network paths and audit every lifecycle mutation. Treat untrusted model deserialization as code execution. Protect registry backups and signing keys separately.

Avoid metadata drift. Tags that drive automation need schemas, allowed values and authoritative writers; free text is not a reliable gate. Use append-only decision events or protected status transitions, validate required lineage before registration, and reconcile registry aliases with actual endpoints. A production alias that points to one version while endpoints serve another is an incident in the control plane, not harmless catalog lag.

Handle deletion and retention deliberately. Deprecate versions before removing them, prove no endpoint, batch job, audit record, rollback plan or legal obligation depends on them, preserve lineage and decision evidence, then delete artifacts under policy. Legal holds and incident evidence override cleanup. Cross-account or cross-region copying creates a new policy, encryption, residency and provenance boundary; verify digest and lineage after transfer.

Reproducibility as a bounded claim

Reproducibility is a bounded claim, and the registry entry should say which bound it supports. Re-running a training job from the recorded commit, dataset snapshot and seed frequently does not return identical weights: non-deterministic GPU kernel selection, atomic floating-point reductions, a different data-loader worker count and a patch-level library upgrade each move the result. Three claims are available and they cost very different amounts to support — that the artifact bytes match a recorded digest, that a re-run lands within a stated tolerance such as 0.3 points of AUC on the frozen evaluation set, or only that the inputs and process are documented well enough to explain a difference. A team that advertises the first while supporting the third treats every failed reproduction as a mystery, when most are the predictable consequence of a choice nobody recorded.

The registry earns its cost at the moment an incident asks which model produced one specific prediction, and that question is answerable only if the prediction log carries the resolved version and digest rather than the alias in force at request time. An alias is a mutable pointer: a request served at 09:14 under champion and one served at 11:40 under the same alias can have run different weights, and reconstructing which is guesswork once the pointer has moved. Log immutable identity alongside the prediction, and keep the artifact and its evidence for as long as the decisions it made remain contestable — for a credit or clinical decision that is years rather than sprints. A retention job that deletes an artifact still referenced by an open dispute, an audit or a rollback plan is an incident in the control plane, not housekeeping.

What interviewers probe, and how strong answers sound

Likely follow-ups, in rough order of frequency:

  • "How do you roll back a bad model?" Weak: "restore the previous version." Strong: repoint the alias to the pinned last-known-good version, restore the full bundle including feature and config dependencies, and check whether the data distribution that made the old model valid still holds — rollback is a decision, not a keystroke.
  • "What's in a model version besides the weights?" Weak: "the pickle file." Strong: preprocessing, inference code, dependency lock, runtime image, signature/schema — the deployable contract, not the parameters.
  • "What happens if the registry is down?" Weak: "serving fails." Strong: serving continues on resolved, cached artifacts; only new promotions and deployments block.
  • "How do you stop a training pipeline from deploying itself?" Separation of duties: the job that registers cannot be the identity that approves or moves the production alias.
  • "Version 42 vs version 41 — which is better?" Neither, necessarily. Versions are identifiers; quality lives in the attached evaluation and approval evidence.
  • "How do you know what served this prediction from three months ago?" The resolved version and digest in the prediction log, plus retention policy that keeps the artifact contestable for the decision's lifetime.

A strong answer throughout sounds like trade-offs under constraints: immutability costs storage and forces re-registration for small changes, but buys auditability and O(1) rollback; aliases cost you historical identity unless you resolve and log them at deploy time; a registry adds process weight, and the alternative is not "no process" but incident archaeology. Measure the thing that matters — lineage completeness, reproducibility rate, unsigned or mutable artifacts, alias-to-deployment divergence, unpinned deployments, time from candidate to decision, rollback success, versions serving past review expiry. Registry success is not version count. It is the ability to answer exactly what is serving, why it was approved, how it was produced, whether its assumptions still hold, and how to replace it safely.