Overview
Curated: · Written: · Reviewed:
Releasing LLM behavior: versioning, gates, and rollout for nondeterministic systems
Review status: rewritten after review — prior draft lost compatibility, supply-chain, capacity, secrets, and lifecycle coverage. Score: pending re-review.
An LLM product's behavior is not the model file. It's the model, the prompt, the retrieval config, the tool schemas, the decoding parameters, and the agent graph, composed at runtime. Interviewers probe exactly this: "Your prompt change passed review — what did you actually deploy?" A weak answer says "we pushed the prompt to prod." A strong answer names a manifest, a gate, and a rollback path.
What actually gets versioned — and why it's one unit
Everything that changes behavior is a deployable artifact:
- Prompts and system instructions (rendered, not just the template — variable interpolation can break a prompt silently)
- Model + provider + exact model revision (a provider's
gpt-4ois an alias; the revision behind it can change under you) - Decoding parameters: temperature, top-p, max tokens, stop sequences, seed policy
- Tool/function schemas — a renamed field is a behavior change for a model that was tuned on the old name
- Retrieval config: chunking strategy, embedding model, index version, top-k, reranker, filters
- Eval datasets and scoring rubrics — if these drift, your gates drift
- Agent/workflow definitions: the graph, the loop bounds, the fallbacks
The mistake interviewers fish for is versioning these independently. A prompt tuned against retrieval index v3 with top-k=10 will behave differently over index v4 with top-k=8, even if the prompt text is identical. The tested combination — model revision + prompt + retrieval config + tool schemas — is the release unit. Version each component, but promote them as one pinned manifest.
release-manifest.json (content-addressed, e.g. sha256:9f3a...)
{
"prompt": "sha256:71c2...", # rendered template + variables schema
"model": "azure/gpt-4o, rev 2024-08-06", # pinned revision, not alias
"decoding": {"temperature": 0.2, "top_p": 1.0, "max_tokens": 1024},
"tools": ["sha256:44de...", "sha256:0a91..."],
"retrieval": {"index": "kb@2024-11-30", "embedder": "text-embedding-3-large",
"top_k": 10, "reranker": "cohere-rerank-v3"},
"agent_graph": "sha256:be07...",
"eval_evidence": "run 4821: golden set 87.4%, safety set 100%, p95 1.9s, $0.011/req"
}
A mutable latest alias with no resolved manifest destroys attribution: when metrics move, you can't say which of the six components changed. Promotion should be a transactional pointer move to an already-tested digest — never a rebuild in production, because rebuilt bytes are untested bytes.
Artifact supply chain: identity is not safety
Build artifacts reproducibly: pin source commits, dependency versions, base models, adapters, tokenizer, data snapshots, and preprocessing. Generate content-addressed checksums, provenance attestations (what inputs produced this artifact, on what machine, from what commit), and vulnerability/malware scans. Then enforce the controls that stop a stolen or tampered artifact from reaching production:
- Sign promoted artifacts and verify signatures at load time; unsigned artifacts never enter the registry.
- Registry admission policy: only artifacts with a passing provenance chain, scan results, and eval evidence can be tagged for a production alias.
- Safe serialization: prefer non-executable model formats (e.g., safetensors over pickle); a pickle executes arbitrary code on load.
- Disable unreviewed remote code (e.g.,
trust_remote_code=Falsein Hugging Face transformers) and isolate model conversion/loading in a sandboxed step — the artifact that loads a model is attack surface.
The line worth repeating: a valid checksum proves identity, not safe behavior. Supply-chain checks and behavioral evals are both required; neither substitutes for the other.
Compatibility: it loads, but is it the same system?
A component that loads successfully may still produce wrong behavior. Loading checks types; compatibility checks semantics. Maintain an explicit compatibility matrix and fail loudly on mismatch rather than silently degrading:
| check | what silently goes wrong if you skip it |
|---|---|
| adapter ↔ base model pair | adapter loads against the wrong base and outputs plausible garbage |
| tokenizer + special tokens | tokens shift, prompts truncate mid-instruction, no error raised |
| chat template | model trained on ChatML gets raw text; quality drops with no crash |
| context length | silent truncation of system prompt or history |
| quantization / dtype | numerically different outputs; your eval evidence no longer applies |
| serving kernels / inference engine | same weights, different numerics and latency profile |
| retrieval vector space | embeddings from model A queried against an index built with model B — every result is wrong but plausible |
| prompt variables / tool schemas | undefined variable renders empty; renamed tool field breaks dispatch |
| cache keys / state migrations | cross-release contamination (see the trace below) |
The vector-space row is the classic interview trap: cosine similarity between mismatched embeddings still returns neighbors, so retrieval "works" while relevance quietly collapses. Gate compatibility in CI on the packaged artifact, not on developer laptops.
Where eval gates sit in CI/CD
Run on every change, before any human sees a request:
- Static checks — lint and render prompts (catch undefined variables, over-long templates), validate tool schemas, validate the manifest and its compatibility matrix, verify signatures and provenance, scan dependencies for known CVEs.
- Deterministic tests — unit tests on parsing, schema validation, tool dispatch; integration tests on the packaged artifact, loaded the way production loads it.
- Offline evals on frozen golden sets — task success, refusal correctness, safety cases, RAG faithfulness, agent trajectory checks. Freeze the sets and the rubric; a gate that rewrites its own exam is not a gate.
- Budget checks — latency percentiles and cost per request against explicit thresholds.
Which gate blocks promotion? All critical ones, non-compensably: a prompt that scores 2 points higher on quality but fails one safety case does not ship. Human approval is accountable review of the recorded evidence — who approved run 4821, on what manifest — not a substitute for the gates.
Rollback is the mirror of promotion: repoint the alias to the previous manifest digest. Because the old manifest is immutable and its eval evidence is stored, rollback needs no re-evaluation. What rollback doesn't cover automatically: in-flight agent workflows mid-tool-loop, memory writes already made, and external side effects. Those need idempotent tool operations and expand/migrate/contract schema changes so the old version can still read what the new version wrote.
Deployment strategies for nondeterministic systems
Accuracy numbers from a frozen set don't tell you what production traffic does. Nondeterminism means you need live comparison, and the standard toolbox adapts:
- Shadow/dark launch: the new stack executes on copies of real requests, results discarded, no user-visible effect — where privacy and side-effect isolation permit. Catches distribution shift the golden set missed.
- Canary with traffic split: 1% → 5% → 25% → 100%, sticky by user, with a pinned control cohort. Assignment lives server-side in a feature-evaluation system; user input never selects privileged variants.
- A/B and interleaving: for quality questions where per-user stickiness is less important than fast statistical readout.
- Blue-green: two full stacks, switch routing. Works, but you pay double capacity during cutover.
Why latency and cost gate the ramp, not just accuracy: a prompt revision that adds 2,000 tokens of context can hold accuracy constant while pushing p95 from 1.9 s to 6 s and cost per request from $0.011 to $0.031. Users churn on latency long before they notice a 1% quality delta. Predeclare the ramp criteria and the abort criteria before launch — "verified task success ≥ 85%, p95 ≤ 2.5 s, $/success ≤ $0.015, zero critical safety events" — so nobody argues about thresholds mid-incident.
Worked trace: a canary that caught a cache-key bug
Prompt v19 + pinned model revision passed all offline gates (golden set 88%). Canary launched at 5%. Verified task success in canary: 61%. The gap wasn't the prompt — production was serving cached completions keyed on prompt hash only, so v19 requests were hitting v18 cache entries.
| cache key | verified success | $/success | duplicate tool POSTs |
|---|---|---|---|
| prompt hash only (v19 canary) | 61% | 0.42 | 18 |
| tenant + manifest + index checksum | 88% | 0.11 | 0 |
Fix the key to include the manifest ID, re-canary, ramp. The lesson to repeat in an interview: a cache that ignores the release manifest is a silent rollback, and only live cohort comparison catches it — the offline gates all passed.
Choosing and pinning the model
Provider vs self-hosted is a decision about control, cost shape, and deprecation risk:
| dimension | provider API | self-hosted |
|---|---|---|
| behavior stability | revision can change under an alias; you pin what they let you pin | fully pinned, you own upgrades |
| unit cost at low volume | high per-token, no floor | high fixed GPU cost; better past roughly 10⁷ tokens/day, workload-dependent |
| deprecation | provider retires models on their timeline (typically months of notice) | your problem, indefinitely |
| data control | egress region/retention negotiated | yours |
Model tiers are the second selection axis, orthogonal to hosting. Frontier, mid-tier, and small models differ by an order of magnitude in cost and latency, and the engineering question is which tasks actually need the frontier tier. A routing design classifies requests — constrained extraction, classification, formatting, simple lookups go to a small or mid-tier model; multi-step reasoning, ambiguous tool use, and safety-sensitive judgment go to the frontier tier — and each route is its own qualified release with its own eval evidence. The failure mode is tiering by vibes: routing to a cheaper model without running the golden set against it, then discovering the cheap model fails on the long tail. The mirror failure is paying frontier prices for tasks a small model handles at parity. Interviewers like the question "how do you decide which model a request gets?" — the strong answer is a measured quality-per-dollar curve per task class, not a default.
The trap underneath both axes: a provider model under a stable-looking name changes behavior — quiet point releases, safety-tuning updates, tokenizer tweaks. Your mitigation is pinning exact revisions where the API allows, recording the revision in every log line, and running a smoke eval on a schedule against the pinned revision so you detect drift you didn't cause. If the provider retires it, you need a qualified fallback, not a silent reroute.
Multi-provider abstraction (e.g., an OpenAI-compatible gateway over several backends) buys failover, but a fallback model is a separate release: it must pass the same gates with roles, tool-schema semantics, token limits, and safety controls mapped and tested — including throttling, partial streams, ambiguous tool effects, and failback. Bound retries and circuit-break the fallback path. Silently routing to an unevaluated model changes the product contract.
Cost mechanics: where the money actually goes
Token accounting first: input tokens and output tokens price differently (output typically 3–5× input for frontier models), and cached input tokens get a discount (OpenAI's prompt caching, for example, prices cached input at 50% of base input as of late 2024 — check current rates for your provider). So spend is dominated by whichever of these your workload amplifies:
- RAG: long retrieved contexts stuffed into every turn. 8 retrieved chunks × 500 tokens = 4,000 input tokens per request before the user's question.
- Multi-turn agents: full history resent every turn — a 10-turn conversation with tool results can carry 20k+ tokens of context per call.
- Tool-call loops: each iteration pays the whole context again. A 6-step agent loop over a 5k-token context is ~30k input tokens for one task.
- Retries: a failed parse that retries the full request doubles cost and can loop.
Demand-side levers come before price shopping: trim retrieved context, cap output tokens, validate early and stop failed workflows instead of retrying blindly, reduce agent fan-out, cache governed results (with tenant authorization and every behavior version in the key). Tiered routing — the selection axis above — is itself the biggest demand-side lever. Serving-side levers — batching, quantization, speculative decoding — change behavior until evidence says otherwise, so they re-run the gates too.
Attribution is the governance layer: tag every request with product, tenant, feature, workflow, and release manifest, preserve raw provider usage records and the pricing version used, and reconcile against invoices. Budgets at request, workflow, and org level with owners and alerts; cost caps must fail with a typed outcome or safe alternative, never a truncated or unauthorized action.
Capacity, load management, and secrets
Forecast volume, input/output token distributions, concurrency, cache-hit rate, and growth, then translate that into rate limits, concurrency, token throughput, memory, and provider quota — with headroom and failure scenarios, not a point estimate sized for the average day. A rollout that saturates the pool turns a 40 ms p99 into a 180 ms p99 and makes you think the new release is slow when it's actually out of capacity.
The controls that keep a spike from becoming an outage:
- Admission control and backpressure — reject or queue early with a typed outcome rather than letting requests pile up inside the serving path until they time out en masse.
- Per-tenant limits with priority/fairness — one noisy tenant can't starve the rest.
- Timeouts and graceful degradation — fall back to a smaller qualified model or a cached result with an honest label, never a truncated answer presented as complete.
- Reserved capacity for verification, rollback, and incident response — if a canary needs 5% of capacity to run its control comparison, that capacity is not available for the ramp math.
Secrets are a release concern too: configuration (which manifest, which tier policy) is versioned and promotable; secrets (API keys, credentials) are fetched at runtime as short-lived credentials from a vault, never baked into images or manifests. A rotated provider key should not require a redeploy, and a leaked image should not contain a production credential.
End of life and incident learning
Deprecation is a release in reverse, with the same discipline: inventory every caller and stored state that references the old version, block new promotions to it, migrate callers through tests, revoke aliases and credentials, invalidate caches keyed to the old manifest, and preserve required lineage while deleting sensitive artifacts on a schedule. An alias pointing at a retired manifest is an incident waiting for a traffic spike.
When an incident happens: contain and roll back first, preserve the evidence (manifest IDs, cohort assignments, request logs) before mutating anything, scope which releases and cohorts were affected, repair one layer at a time, add the regression case to the frozen golden set so the gate that missed it never misses it again, and reenter through a smaller canary than the one that failed. The post-incident reentry is the part teams skip — going straight back to 100% after a rollback re-runs the original experiment without the learning.
What interviewers probe, and what a weak answer sounds like
- "Walk me through shipping a prompt change." Weak: "we edit it in the dashboard and watch metrics." Strong: rendered prompt artifact → CI with golden-set eval, compatibility checks, and cost/latency budget → signed manifest promotion → 5% sticky canary vs pinned control → predeclared ramp/abort criteria → alias rollback.
- "Your canary's accuracy is fine but p99 tripled. Ramp or not?" Not — latency was a predeclared gate. Diagnose: context growth, cache misses, provider throttling, or capacity saturation from the canary itself.
- "The provider deprecated your model with 30 days' notice. Go." Qualified fallback release, same gates, mapped tool semantics, tested failover and failback, migration timeline per cohort.
- "How do you roll back an agent that already made three external API calls?" Idempotent tool operations, compensation logic, stop-new-work flag, in-flight handling — rollback includes state, not just routing.
- "A model file passed your virus scan. Is it safe to deploy?" No — a checksum and a scan prove identity and hygiene, not behavior. Behavioral gates are separate and non-compensable.
- "Metrics dropped 10% Tuesday. What changed?" Compare like with like: pricing change, evaluator change, or traffic-mix shift can mimic a release effect. Immutable manifest IDs on every request make attribution a query, not an archaeology project.
The one-sentence position to defend: LLMOps makes LLM changes reproducible, attributable, reversible, and economically aligned with user value — and every one of those four words maps to a mechanism above.
