Overview
Curated: · Written: · Reviewed:
Key takeaways
- Choose evaluation from the decision: population, prediction unit, horizon, action, error costs, capacity, and minimum safety/fairness constraints. A convenient default metric is not the product objective.
- Separate ranking, probability, and thresholded-decision quality. ROC AUC measures pairwise ranking; log loss and Brier loss assess probability forecasts; precision and recall describe one selected operating point.
- Accuracy can look excellent under class imbalance. Precision depends on prevalence; recall and false-positive rate condition on actual class. Always report confusion counts, support, uncertainty, and relevant slices.
- Select preprocessing, model, calibration, and threshold on development evidence. Evaluate the locked procedure once on an untouched deployment-like test set; repeated test decisions consume that test.
- Regression metrics encode different costs: MAE is linear and robust relative to MSE, RMSE restores target units, pinball loss evaluates quantiles, and R² compares squared error with a mean baseline but can be negative.
- Offline model metrics are proxies. Production evaluation also needs label maturity, interventions, selective feedback, calibration and performance by cohort, operational capacity, and real-world outcomes.
1. Start with the decision and evidence boundary
Define what is predicted, for whom, when, over what horizon, and what action consumes the output. Document false-positive and false-negative costs, abstention or review capacity, latency, prevalence, regulatory/fairness constraints, and acceptable degradation by cohort. Convert these into one primary selection metric plus explicit guardrails; keep diagnostic metrics to explain failure without quietly changing the objective after seeing results.
The split must reproduce deployment. Time-dependent systems need forward validation; repeated users, devices, patients, or organizations may require group isolation; spatial/site deployment needs appropriate holdouts. Fit preprocessing, resampling, feature selection, calibration, and threshold choice only inside development folds. The final test estimates the entire locked selection procedure. Report sample counts, label maturity, confidence intervals or resampling dispersion, and comparisons with simple rules, the incumbent, and dummy/prevalence baselines.
2. Classification metrics and operating points
A confusion matrix counts true positives, false positives, true negatives, and false negatives at a declared positive class and threshold. Accuracy is (TP+TN)/N; precision is TP/(TP+FP); recall or TPR is TP/(TP+FN); specificity is TN/(TN+FP); FPR is FP/(FP+TN). Raising a score threshold usually decreases predicted positives, trading fewer false positives for more false negatives. Undefined denominators must remain visible rather than being silently converted into success.
F1 is the harmonic mean of precision and recall, while F-beta weights recall more when beta is above one. Neither includes true negatives or business costs, and a single score can hide which error changed. In multiclass evaluation, macro averaging gives each class equal weight, weighted averaging follows class support, and micro averaging pools decisions. Report per-class metrics and confusion patterns because an average can conceal an unusable rare class.
3. Ranking and probability quality
ROC plots TPR against FPR across thresholds. ROC AUC is the probability that a randomly chosen positive receives a higher score than a randomly chosen negative; it does not prove calibrated probabilities or select the deployment threshold. Under rare positives, small FPR changes can still create many false alarms, so precision-recall curves and average precision often expose operational performance more directly. Their baseline depends on prevalence, which makes cross-population comparisons hazardous.
Probability forecasts support expected-cost decisions only when calibrated for the deployment population. A reliability diagram compares predicted probability bands with observed frequency. Log loss heavily penalizes confident errors; Brier loss is mean squared probability error. Both proper scores combine calibration and discrimination, so pair them with reliability evidence. Calibration is another learned stage fitted without the evaluation fold and rechecked after prevalence, policy, or population shift.
4. Regression, ranking, and structured outputs
MAE weights errors linearly and remains in target units. MSE squares errors and emphasizes large misses; RMSE takes its square root to return to target units. Median absolute error is more resistant to extreme outcomes. MAPE is undefined or unstable near zero and asymmetric in common interpretations. R² compares squared error with predicting the evaluation-set mean; zero is mean-baseline performance and values below zero are possible. Always accompany aggregate loss with residual distributions and slices across target range, time, and important cohorts.
Quantile forecasts use pinball loss and should be checked for empirical coverage and interval width. Ranking/retrieval systems need query-grouped metrics such as precision/recall at k, mean reciprocal rank, or discounted gain aligned with the user surface. Multilabel tasks distinguish exact subset accuracy, per-label measures, and sample averaging. Each metric defines a unit of analysis; do not treat rows as independent when decisions occur per user, query, session, or organization.
5. From offline score to production decision
Choose the threshold on a validation population using explicit costs, capacity, or constraints—for example maximize recall subject to a minimum precision and daily review volume. Stress-test plausible prevalence and cost shifts, report the threshold and metric version, and snapshot the model, transform, calibration, and policy together. A threshold of 0.5 has no universal privilege.
Monitor score and feature distributions, predicted-positive volume, calibration, confusion metrics after labels mature, operational queues, abstentions, latency, overrides, downstream outcomes, and slices. Interventions change which labels become observable, creating selective feedback; randomized audits, delayed-outcome accounting, or other governed measurement may be needed. Aggregate quality never substitutes for fairness, safety, privacy, and product impact review.
6. Worked example: one model, two thresholds, two different products
A fraud model scores 100,000 transactions of which 2,000 are fraudulent (2.0% prevalence). A model that predicts "legitimate" for every row is 98.0% accurate and catches nothing, which is why accuracy is the wrong headline at this prevalence.
The same model at two operating points:
| Threshold 0.5 | Threshold 0.2 | |
|---|---|---|
| True positives | 1,180 | 1,700 |
| False negatives | 820 | 300 |
| False positives | 620 | 4,900 |
| True negatives | 97,380 | 93,100 |
| Precision | 0.656 | 0.258 |
| Recall | 0.590 | 0.850 |
| F1 | 0.621 | 0.395 |
| Accuracy | 0.986 | 0.948 |
F1 prefers the 0.5 threshold and accuracy prefers it too, but neither knows the cost. If a missed fraud averages $220 and a false positive costs $4 in review time plus some churn risk:
Alerts are every predicted positive, TP + FP -- a reviewer cannot know which is which in advance.
Threshold 0.5: 820 misses x $220 + 1,800 alerts x $4 = $180,400 + $7,200 = $187,600
Threshold 0.2: 300 misses x $220 + 6,600 alerts x $4 = $66,000 + $26,400 = $92,400
The threshold with the worse F1, worse precision, and worse accuracy is worth $95,200 more per 100,000 transactions. The ranking never changed — AUC is identical for both rows, because AUC is threshold-free — so this is a decision about the operating point, not about the model.
Two consequences worth stating in an interview. AUC-PR is the more honest summary curve here (a random model scores 0.02, not 0.5, so the baseline is visible). And the review queue is a hard constraint, sized by alerts rather than by mistakes: 1,800 alerts at the 0.5 threshold and 6,600 at 0.2. A team that can process 1,500 cannot staff either operating point, so the threshold is set by that capacity and not by any metric on this table. Note what the constraint is measured against, too: 100,000 transactions is a volume, not a day. Whether 6,600 alerts is unthinkable or trivial depends entirely on whether that volume arrives hourly or monthly, and the setup has to say which.
