Overview
Curated: · Written: · Reviewed:
Evaluation is a release system, not a leaderboard score
Start with a decision. State what product change may ship, who uses it, which harms are unacceptable, and what evidence would approve, reject, roll back, or escalate it. Translate that decision into a versioned evaluation contract: target tasks and populations, input distribution, expected behavior, rubric, metrics, slices, baselines, uncertainty, owners, thresholds, and review date. A metric without a decision rule produces numbers rather than confidence.
The offline/online split
Evaluation happens in two places with different jobs. Offline suites run a candidate against a fixed, versioned dataset before release; their purpose is to catch regressions and gate shipping on evidence you can reproduce. Online evaluation is monitoring on live traffic after release; its purpose is to detect drift, incidents, and gaps your offline set missed. Offline scores never transfer directly to production because the offline distribution is a frozen sample of a moving population: prompts shift, users discover new behaviors, retrieval corpora change, and adversarial users probe what your dataset never contained. The bridge is deliberate — shadow traffic replayed through the candidate, a small sticky canary, and instrumentation that attaches outcomes to the exact stack that produced them. If an interviewer asks why your offline score was 58% win-rate but production satisfaction barely moved, the answer is distribution shift plus metric validity, and you should have expected it.
Building the eval set
Sample real tasks from the deployment distribution, with consent and minimization. Preserve hard and ordinary cases, and label language, domain, user cohort, input length, risk, tool or retrieval dependency, and time. Add synthetic and adversarial cases to cover rare critical failures, but never claim they represent prevalence. Deduplicate and group related prompts — near-duplicate variants and multi-turn conversations from one user — before splitting, so a dev/test leak is not just one prompt but a whole behavioral cluster. Freeze an untouched release set; prompt, model, judge, retrieval, and threshold tuning belong on development data.
Size and representativeness are budget questions, not aspirations. If your gate is "win rate ≥ 55% vs baseline," work backwards: to distinguish 55% from 50% on a paired design with 200 cases, the paired win-rate difference of 5 points sits near the edge of what n=200 resolves; slices need more cases than you think, which is why you slice on the few dimensions that carry risk, not twenty. Keep the set in sync with the product: when a new tool, language, or policy ships, add cases for it and retire cases for removed behavior, transparently. Rotate contamination-sensitive sets, hash immutable inputs, and investigate sudden score gains for leakage before celebrating them.
Deterministic checks first
Use deterministic evaluation wherever a machine-checkable contract exists. Validate JSON schemas, types, required fields, citations, tool names and argument constraints, policy decisions, executable tests, calculations, and exact task outcomes. Execute generated artifacts only in isolated sandboxes with strict time, network, secret, and resource controls. Exact match is appropriate only when equivalence really is exact; normalization can otherwise hide meaningful defects or punish harmless wording.
# pytest 8.x, jsonschema 4.x — checked against Python 3.12
import json
from jsonschema import validate
SCHEMA = {
"type": "object",
"required": ["action", "amount_cents", "account_id"],
"properties": {
"action": {"enum": ["refund", "hold", "escalate"]},
"amount_cents": {"type": "integer", "minimum": 1, "maximum": 500_000},
"account_id": {"type": "string", "pattern": "^acct_[0-9a-f]{12}$"},
},
"additionalProperties": False,
}
def test_tool_call_contract(model_output: str):
call = json.loads(model_output) # malformed JSON fails here
validate(call, SCHEMA) # wrong shape or range fails here
assert call["action"] != "refund" or call["amount_cents"] <= 10_000 or call["account_id"] in APPROVED_ACCOUNTS
This is the cheapest, fastest layer: no judge, no sampling noise, a hard pass/fail per case. Every behavior you can push into this layer is a behavior you stop arguing about.
Reference-overlap and semantic metrics
Reference-overlap metrics such as BLEU and ROUGE are useful for constrained translation or summarization comparisons, but open-ended correctness is not word overlap. Semantic metrics such as BERTScore relax surface form while still inheriting encoder and reference limitations. Report metric validity for the actual task and language. Do not average incomparable measures into one impressive score that lets a safety regression cancel a style gain — a single aggregate hides slice regressions by construction, which is exactly why slices are part of the contract.
LLM-as-judge as an engineering problem
LLM judges can scale rubric-based scoring and pairwise comparison. Pin the judge model, prompt, decoding, rubric order, candidate order, and parsing behavior. Blind system identity, randomize or swap positions, constrain evidence-backed rationales, and test verbosity, style, self-preference, prompt-injection, reference leakage, and correlated-model biases. Distinguish reference-anchored judging ("is this claim supported by the gold answer?") from reference-free judging ("is this answer helpful and correct?"); the first is more reliable but only covers behaviors your references cover.
# Position-bias probe: same pair, both orders, n=200 — checked against Python 3.12
import random
flips = 0
for case in cases:
a, b = case["cand"], case["base"]
pick1 = judge(a, b) # candidate in position 1
pick2 = judge(b, a) # candidate in position 2
if pick1 != flip(pick2):
flips += 1 # judge chose whichever answer sat in position 1
print(f"position-flip rate: {flips / len(cases):.2f}")
# 0.18 on a 200-case suite means ~1 in 6 decisions is order, not quality.
# Fix: randomize order per case and only count pairs where both orders agree,
# or report the disagreement rate as judge noise.
Calibrate every judge and threshold against blinded domain-expert labels; revalidate after any judge change or distribution shift. The judge is a measurement instrument, not ground truth. A weak answer treats judge output as truth; a strong one reports judge–human agreement and the conditions under which it was measured.
Human evaluation and agreement
Human evaluation needs the same engineering rigor. Define observable rubric anchors and abstention, train raters with worked examples, blind system identity, randomize presentation, collect independent labels, measure agreement, adjudicate critical conflicts, and retain disagreement rather than forcing false certainty. Record rater qualifications and conflicts without exposing unnecessary personal data. Experts should review high-risk factual, legal, medical, security, or policy outcomes even when automated agreement is strong.
Measure agreement with chance-corrected statistics, not raw percent. Two raters label 100 items; both say "yes" on 45, both say "no" on 45, and they split the remaining 10. Raw agreement is 90/100 = 0.90, which sounds excellent. But each rater said "yes" 50 times, so chance agreement is (0.50 × 0.50) + (0.50 × 0.50) = 0.50, and Cohen's kappa is (0.90 − 0.50) / (1 − 0.50) = 0.80 — good but not the 0.90 the raw number implied. With three or more raters use Fleiss' kappa. When humans genuinely disagree, that disagreement is data: it usually means the rubric anchor is ambiguous or the case sits on a real boundary. Rewrite the anchor, or model the case as genuinely contested — do not average the disagreement away.
RAG and task-level metrics
Evaluate a RAG system by layer, because the fix differs by layer. Retrieval measures include recall at k, precision at k, rank, coverage, freshness, permissions, and no-answer behavior. Generation measures include claim-level support (faithfulness), citation correctness, answer relevance, completeness, uncertainty, and resistance to instructions embedded in sources. End-to-end task success matters most, but component diagnostics identify whether to fix ingestion, retrieval, reranking, context assembly, prompting, model behavior, or product policy. For classification-shaped outputs — spam, intent, escalation, refusal — use precision and recall per class, and state which one the product decision weights; a 99% accurate refusal classifier with 40% recall on the harm class is a safety hole, not a rounding error.
Agent evaluation
Agent evaluation must inspect trajectories, not only polished final text. Verify tool selection, argument validity, authorization, step budget, state transitions, retries, idempotency, stopping, confirmation, side effects, recovery, and audit logs. Simulate timeouts, partial failures, stale reads, malicious tool output, conflicting instructions, revoked access, and repeated delivery. Grade the externally observed state and invariants; a convincing explanation cannot compensate for a wrong payment, deletion, or permission change.
Safety gates are non-compensable
Cover jailbreaks, direct and indirect prompt injection, sensitive-data disclosure, harmful assistance, bias, insecure output handling, denial of service, supply-chain changes, and excessive agency. Separate desired refusals from over-refusals. Preserve attack provenance and severity, rotate hidden suites, and use authorized red teams. Critical failures trigger containment and review rather than being diluted into average quality.
Statistics you can defend
Declare the unit of analysis, paired design, sampling method, estimand, uncertainty interval, multiple comparisons, and minimum practically important difference before looking at results. Keep repeated samples from one user or template grouped — a cluster of 50 paraphrases of one prompt is closer to one observation than fifty. Use paired bootstrap or an appropriate clustered method for matched systems, and report win, loss, tie and abstention rather than silently dropping ties.
# Paired bootstrap over 200 frozen cases — checked against Python 3.12
import random
wins = [1 if c else 0 for c in paired_outcomes] # 1 = candidate beats baseline
n = len(wins)
random.seed(0)
deficits = []
for _ in range(10_000):
sample = [wins[random.randrange(n)] for _ in range(n)]
deficits.append(sum(sample) / n)
deficits.sort()
lo, hi = deficits[249], deficits[9_749] # 95% interval
print(f"win rate {sum(wins)/n:.3f}, 95% CI [{lo:.3f}, {hi:.3f}]")
# With 116 wins out of 200: point estimate 0.580, CI roughly [0.510, 0.648].
# The lower bound of 0.510 sits above 0.50, so this run rejects a coin-flip
# null at the 5% level — the win is real, not sampling luck. The 55% gate
# also passes. What this run cannot do is separate a 58% candidate from a
# 55% one: the whole interval spans that 3-point gap, so "beats baseline by
# at least 5 points" would need more cases.
A tiny significant change can still be irrelevant; a noisy critical slice can still block release.
Reproducibility and the release loop
Reproducibility binds results to the exact stack: application commit, dataset snapshot and transformation, model/provider revision, prompt and policy, tools and schemas, retrieval index, judge, decoding, locale, dependencies, seeds, credentials class, and execution environment. Store per-example outputs, structured scores, judge rationale, errors, latency, tokens and cost with access controls. Hash immutable inputs, generate a model/evaluation card, and independently rerun a sample before promotion.
Offline evidence does not prove production value. Roll out through shadow or a small sticky canary, compare a pinned baseline, and monitor task completion, verified corrections, escalation, abandonment, safety, latency, errors, and cost by cohort. Instrument exposure so outcomes attach to the exact stack. User ratings are biased observations, not truth; combine them with audited samples and downstream outcomes. Define automatic rollback and preserve privacy-safe incident evidence.
Maintain the suite like production code. Every material failure becomes a minimized regression case when permitted. Track coverage and case provenance, retire obsolete cases transparently, rotate contamination-sensitive sets, and re-run critical gates on every model, prompt, retrieval, tool, policy, judge, or dependency change. Evaluation is the continuous control loop connecting requirements, evidence, release, monitoring, incident response, and improvement.
Worked example: MT-Bench 8.2, jailbreak 12/20
Support bot, n=200 frozen cases. Gate: task win ≥ 55% vs pinned baseline and jailbreak success ≤ 2/20.
| scoreboard | task win vs baseline | jailbreak success | ship? |
|---|---|---|---|
| MT-Bench 8.2 (leaderboard) | unknown | unknown | no — wrong population |
| new prompt, 58% win | 58% | 12/20 | no — safety gate fails |
| new prompt + refusal tests | 56% | 1/20 | yes |
The 8.2 is not a release number. The 12/20 is.
What interviewers probe
The common probe is a design question: "You've fine-tuned a model for our support bot. How do you know it's better?" A weak answer names a benchmark and a score — "we improved MT-Bench from 7.9 to 8.2." That answer fails because it never states the decision, the population, or the harms. A strong answer starts from the release decision and works down: contract, dataset from real traffic, deterministic layer first, judge calibrated against human labels, safety as a separate gate, statistics with a unit of analysis, then canary and rollback.
Likely follow-ups, and what each is really testing:
- "Your judge agrees with humans 85% of the time — is that good?" — whether you know agreement must be chance-corrected, measured on blinded labels, and re-measured after any judge change.
- "Win rate went up but complaints went up too. What happened?" — whether you understand slice regressions and that an aggregate can hide exactly the cohort that regressed.
- "How do you prevent your eval set from leaking into training?" — dedup, grouping, rotation, hashing, and investigating suspicious gains.
- "Offline says ship, the PM says wait. Who's right?" — whether you can articulate why offline evidence bounds risk but doesn't prove value, and what a canary would resolve.
- "A human rater disagrees with the judge on a medical answer. What now?" — whether safety-critical domains get expert adjudication regardless of automated agreement.
If you can walk a single concrete release decision end to end — numbers, gates, and the rollback trigger — you have the answer the question was fishing for.
