Overview
Curated: · Written: · Reviewed:
Biological machine learning must generalize across relationships, batches, and populations
Machine learning in computational biology predicts phenotypes, molecular properties, variants or structure from high-dimensional, highly related data. The engineering problem is rarely fitting a flexible model — it is defining the future use population and building an evaluation that does not leak individuals, families, homologs, batches, sites, time or preprocessing information across the boundary. Interviewers in this space probe exactly that: they hand you a benchmark number and ask whether you believe it.
The mental model: the split defines the question
Every biological dataset has structure — donors, families, homologs, chromosomes, scaffolds, plates, sites, dates — and a random split quietly moves that structure onto both sides of the boundary. The model is then scored partly on recognizing near-duplicates it has already seen, and the reported metric answers a question nobody asked: "how well do you recall things similar to your training set?" The metric you deploy under is "how well do you predict on units genuinely unlike anything in training."
So the first design decision is not the architecture. It is: what is the unit of independence, and what does deployment actually look like? New patients at the same hospital? New loci in the same patients? New protein families? New sequencing platform? Each answer implies a different split, and picking the wrong one inflates every downstream number.
The mechanism of inflation is simple: biological data is redundant at many scales. Two paralogous proteins share most of their residues and their annotation. Two adjacent genomic windows share most of their bases and their linkage disequilibrium. Two cells from one donor share genotype and preparation. A random split converts that redundancy into near-duplicate train/test pairs, and near-duplicates are easy. The model is not overfitting in the usual sense — the evaluation is measuring the wrong thing.
Where ML actually shows up, and what each task is for
Interviewers often open with "where would you use ML here?" — answer with the task, the question it answers, and what the prediction feeds:
- Sequence-to-function and variant effect prediction: given a protein or a variant, predict function, pathogenicity or effect size. Downstream: variant interpretation in a clinic or a screen triage list. The failure that matters is a confident wrong pathogenicity call.
- Gene expression and single-cell annotation: label cell types or states from expression profiles. Downstream: composition analysis across conditions. Errors propagate into every differential-abundance claim made afterward.
- Structure and interaction prediction: fold proteins or predict binding. Downstream: hypothesis generation for experiments, not a substitute for them.
- Molecular property and scaffold prediction: ADMET, solubility, activity from structure. Downstream: ranking a compound library before synthesis, so the cost of a false positive is a wasted synthesis and a false negative is a missed lead.
- Imaging readouts: morphology, phenotype classification from microscopy. Downstream: screen hit calling, where plate and scanner effects sit right next to the biology.
Before modeling any of these, state three things: the unit of observation (a cell? a donor? a plate?), the label source and its error rate (curated database? clinical code? computational annotation? each has a different failure mechanism), and the asymmetric cost of a false positive versus a false negative in the downstream workflow. If you cannot say what a wrong prediction costs, you cannot pick a threshold or a metric.
Split design: the checklist that decides your numbers
Match the split to deployment:
- Homology-aware / sequence-identity splits for proteins: cluster sequences at an identity threshold matched to the novelty you claim, keep whole clusters in one fold. A model deployed on remote homologs needs a 30%-identity split, not a 50% one.
- Chromosome- or locus-held-out splits for variant and regulatory genomics: nearby windows share sequence, LD and labels, so split by large blocked regions or whole chromosomes, with buffer zones at boundaries.
- Scaffold splits for molecules: split by scaffold, not by compound, or you score memorization of decorated analogues.
- Donor/patient-held-out splits for clinical omics: multiple cells, visits or samples from one donor are one patient, not independent samples. Treating them as independent is pseudoreplication, and it inflates your effective sample size.
- Temporal splits when deployment means "the future": structures solved after a cutoff, samples collected after a date.
How much does this change conclusions? A variant-effect model trained on chromosomes 1–22 and tested on a random 20% of SNPs can post AUC 0.94 while a held-out chromosome gives 0.71, with calibration slope falling from 0.98 to 0.61. That is the difference between "ready to deploy" and "not yet," decided entirely by the split.
Worked example: a split that inflates every number
A model predicts protein function from sequence, trained on 40,000 proteins, reporting 0.93 accuracy on a random 20% holdout. The number is meaningless, and the reason is homology:
| Split | Max train-test sequence identity | Reported accuracy | What it estimates |
|---|---|---|---|
| Random 80/20 | up to 99% | 0.93 | recall of near-duplicates already seen |
| Cluster at 50% identity, split clusters | < 50% | 0.71 | generalization to related proteins |
| Cluster at 30% identity | < 30% | 0.58 | generalization to remote homologs |
| Time-split (structures solved after a cutoff) | uncontrolled but realistic | 0.62 | performance on genuinely new targets |
The fix is to cluster before splitting, never after:
# Wrong: random rows, homologs on both sides.
train, test = train_test_split(df, test_size=0.2, random_state=0)
# Right: cluster at a sequence-identity threshold, then split whole clusters.
# mmseqs easy-cluster seqs.fasta clusters tmp --min-seq-id 0.3
clusters = read_clusters("clusters_cluster.tsv") # 40,000 seqs -> 6,200 clusters
train_c, test_c = train_test_split(sorted(clusters), test_size=0.2, random_state=0)
assert not (set(train_c) & set(test_c)) # no cluster spans the split
The same trap appears wherever biological data has structure the split ignores: cell-line identity in drug response, patient identity across multiple samples, batch in single-cell data, chromosome in variant effect prediction. The unit of splitting must be the unit of independence, and in biology that is almost never the row.
Leakage beyond the split: preprocessing, labels, batches, missingness
All fit-dependent preprocessing belongs inside each training fold: scaling, imputation, batch correction, feature selection, dimensionality reduction, oversampling and representation learning. Computing them on the complete dataset exposes held-out distributions or labels. Reference databases and pretrained embeddings can also leak benchmark targets through sequence identity or publication chronology — audit training-database overlap before trusting any transfer result.
Labels are noisy proxies. Clinical codes, assay thresholds, computational annotations and curator decisions have different error mechanisms, and label uncertainty, verification bias and missing-not-at-random outcomes bound what the model can achieve. A model can accurately reproduce a biased measurement process while failing at the biological construct.
Confounding and shortcuts are the default, not the exception. Sequencing center, plate, scanner, read length, ancestry or collection period can predict the label, and random cross-validation preserves those shortcuts. Negative controls, batch-held-out tests, matched analyses, feature attribution and acquisition audits help detect them — but causal claims require a design beyond prediction.
Missing data is part of the process, not noise to fill in. An imputed value or a missingness indicator can encode care access or laboratory workflow. Fit imputation inside folds, simulate deployment missingness, decide how the system abstains when required inputs are absent, and never substitute unmeasured future variables with values only available after the outcome.
Freeze an external test set before feature choice and tuning. Cross-validation on the remaining development data estimates selection performance; nested cross-validation earns its complexity when data are too small for a fixed validation set. Repeated inspection of test performance turns the test set into training feedback and invalidates its uncertainty.
Metrics: imbalance, calibration, and what the threshold costs
Class imbalance makes accuracy misleading — always predicting negative already achieves 99% accuracy at 1% prevalence. Report confusion counts, sensitivity, specificity, precision and recall at predeclared thresholds, plus curves where appropriate. The curve choice follows from the task:
- AUPRC over AUROC when positives are rare: AUROC is dominated by the abundant true negatives and can look strong while precision at any usable threshold is poor. AUPRC exposes that directly.
- Precision at k for screening: if the lab can follow up 50 hits, the question is precision among the top 50, not average performance.
- Enrichment over a random baseline: at 1% prevalence, a random top-50 list contains ~0.5 positives; a screen that delivers 15 is a 30× enrichment. State the baseline when you state the number.
- Calibration for ranked follow-up lists: if the list drives wet-lab capacity planning, the counts must be honest — among cases assigned risk near 0.2, roughly 20% should experience the outcome in the target setting.
Discrimination and calibration are distinct, and precision depends on prevalence, so validation under enriched cases will not give deployed positive predictive value without recalibration. Thresholds encode asymmetric costs and workflow capacity; select them on development data, freeze before test, and evaluate subgroup consequences. Calibrators require held-out data, and prevalence or practice changes can invalidate them — evaluate calibration slope, intercept and curves with uncertainty and subgroup sample sizes.
Small n, large p: when the deep model is the wrong answer
Biological datasets routinely have thousands of features and dozens to hundreds of independent samples. In that regime a regularized linear model or gradient-boosted tree on engineered features often beats a deep network, and it trains in minutes with interpretable coefficients — which matters when the reviewer asks "why did the model say that?" Reach for deep models when the raw signal is long and compositional (sequence, spectrum, image) and you have either enough independent data or a good prior.
That prior is what transfer learning buys. A pretrained protein language model (e.g., ESM- or ProtT5-style embeddings) encodes what hundreds of millions of sequences say about residue co-occurrence and structural constraints; fine-tuning on 500 labeled proteins then works because most of the representation was learned elsewhere. Be precise about what pretraining does not buy: it does not add labeled examples of your task, it does not remove the need for homology-aware splits (the pretraining corpus may contain your test set's homologs — audit it), and its objective (masked-token reconstruction) is not your objective, so the fine-tuning data still sets your ceiling. When someone reports that a pretrained model beat a from-scratch model on a random split, the honest comparison is on a cluster-held-out split with a logistic-regression-on-embeddings baseline included — that baseline is cheap and frequently competitive.
Interpretation, scale, and what changes in production
Feature attribution describes model behavior locally or globally under assumptions; it does not prove biological causality. Correlated genomic features can exchange importance, and in silico perturbations may create impossible sequences. Interpretability analysis needs stable baselines, biologically valid perturbations, sensitivity checks and experimental validation.
Scale increases data and provenance obligations, not decreases them. Deep models can learn sequence motifs, interactions and structures, but at scale you must evaluate homology-aware or chromosome-held-out splits, training-database overlap, architecture ablations and simpler baselines, and report local confidence and uncertainty at the granularity of each prediction rather than one benchmark average.
Domain shift occurs across sequencing technologies, laboratories, hospitals, ancestries, species and time. Internal cross-validation estimates only variation represented in development data. External validation should use independent acquisition and preserve the intended workflow. Subgroup performance requires uncertainty — a tiny subgroup's perfect score is not evidence of equity.
Deployment requires prospective workflow evaluation, human factors, latency and failure behavior, not only retrospective AUC. Monitor input schema, missingness, acquisition shifts, calibration, subgroup outcomes, abstentions and downstream interventions. Labels often arrive late and are affected by the model's own use, so monitoring design must address feedback loops.
Reproducibility requires immutable cohort definitions, consent and licenses, raw and derived feature versions, split assignments, code, containers, random seeds, hardware and model artifacts. Determinism helps, but repeated seeds and uncertainty estimates reveal optimization variance. Document intended uses, excluded uses, training composition, metrics, subgroup limitations and update triggers.
What interviewers probe, and what a weak answer sounds like
Likely follow-ups, and the failure mode each one catches:
- "You report AUC 0.94 on this variant benchmark. Do you trust it?" Weak answer: "yes, cross-validated." Strong answer: ask what the split was, whether homologs or chromosomes cross it, whether preprocessing was fit inside folds, and what happens on a held-out chromosome.
- "Why did accuracy drop from 0.93 to 0.58 when you clustered?" Weak answer: "overfitting." Strong answer: the model didn't change — the earlier number measured near-duplicate recall; the new one measures the generalization the task actually needs.
- "Prevalence is 1% and the lab can chase 50 hits. What metric?" Weak answer: "AUC." Strong answer: precision at k against a stated random baseline (~0.5 expected hits), plus calibration if the list feeds capacity planning.
- "When would you not use a pretrained protein model?" Weak answer: never. Strong answer: when tabular features already carry the signal, when the pretraining corpus overlaps your test set, or when the deployment split is close-homology and the pretrained representation was learned on exactly that similarity regime.
- "Your model flags a variant pathogenic. What does the attribution mean?" Weak answer: "these positions caused it." Strong answer: model dependence under assumptions, correlated features can swap importance, perturbations may be biologically invalid, and confirmation requires experiment.
The production invariant to close with: leakage-free, calibrated generalization — every prediction tied to a versioned cohort, feature and label definition, a biologically independent split and frozen evaluation, with uncertainty, subgroup and domain-shift evidence sufficient for its exact intended decision, and an abstention or rollback path when those conditions fail.
