Overview
Curated: · Written: · Reviewed:
Statistical genomics estimates associations under structure, measurement, and selection
Statistical genomics connects genetic variation with traits across many samples and loci. A genome-wide association is evidence that an allele tags trait variation in the studied population under a model; it is not automatically causal, clinically useful or portable. Linkage disequilibrium, ancestry, relatedness, phenotype definition, environment and selection shape every result.
How this comes up in interviews. You will rarely be asked to define GWAS. You will be handed a scenario — a hit that won't replicate, a lambda of 1.2, a PRS that collapses in a second cohort — and asked to diagnose it. Interviewers probe three depths: (1) can you state what the regression actually estimates, (2) can you trace how structure and QC decisions change that estimate, (3) can you say what a result licenses you to claim and to whom. A weak answer recites vocabulary ("you correct for population stratification with PCA") without saying what goes wrong numerically if you don't, or why the fix is imperfect. A strong answer names the failure mode, the fix, and the fix's own failure mode.
What a GWAS actually estimates
Every per-variant result comes from a regression of phenotype on genotype dosage plus covariates. For a quantitative trait, linear regression gives beta: the expected phenotype-unit change per additional effect allele, holding covariates fixed. For a binary trait, logistic regression gives log-odds, exponentiated to an odds ratio. The effect allele, other allele, units, phenotype transform, model and sample count are inseparable from the number — a beta of 0.03 in standardized BMI units is not the same claim as an OR of 1.03 per allele, and neither is a probability the variant is causal.
The standard error and p-value come with the effect: SE reflects sample size, allele frequency and residual variance, and the p-value tests the null of no association in this model, in this sample. Case-control sampling makes odds ratios efficient but non-portable: ORs depend on ascertainment, and prediction calibration drifts with the case mix. Quantitative traits may need transformation or robust models; survival, longitudinal, family and admixed data need methods matching their dependence structure.
Interviewer probe: "You have beta = 0.05, SE = 0.01, MAF = 0.3, N = 50,000 for a height GWAS. What can you claim?" A weak answer stops at "significant." A strong answer states: per-effect-allele change in the phenotype's units, the population and model it was fit in, that the variant may be tagging a causal variant through LD, and that replication requires the same effect allele, build and units.
Multiple testing and what significance licenses
Millions of correlated tests demand multiplicity control, and the conventional genome-wide threshold is arithmetic, not tradition:
A GWAS tests 1,000,000 SNPs. At uncorrected p < 0.05, expected false positives under a complete null are 1,000,000 x 0.05 = 50,000. Bonferroni divides the family-wise error rate by the test count: 0.05 / 1,000,000 = 5 x 10^-8 — the genome-wide line. The choice of correction changes what you control:
| Method | Controls | Expected false positives | Expected true hits found (of 40 real) |
|---|---|---|---|
| Uncorrected p < 0.05 | nothing useful | ~50,000 | 40 |
| Bonferroni, 5 x 10^-8 | P(>=1 false positive) <= 0.05 | 0.05 | 12 |
| Benjamini-Hochberg, FDR 0.05 | 5% of reported hits are false | ~1.6 of 32 reported | 31 |
| Permutation, 1,000 shuffles | family-wise, empirically | 0.05 | 15 |
(Hit counts are a hypothetical worked example, not a guarantee.) Bonferroni finds 12 of 40; BH finds 31 while accepting ~1.6 wrong among its 32. Neither is more correct — Bonferroni suits an expensive validation step, FDR suits a ranked follow-up list. Permutation recovers 15 because it estimates the effective number of independent tests rather than assuming it, at the cost of 1,000 GWAS runs.
The 5 x 10^-8 line itself is a pragmatic standard derived for common variants in European-ancestry GWAS density; it is not universal across ancestries, sequencing studies or variant sets. Report exact p-values with effect and uncertainty. Genomic inflation (lambda) can reflect confounding, relatedness, batch effects or genuine polygenicity — it is a diagnostic to investigate, not a knob to divide every statistic by.
What significance licenses: an association in the studied population under the stated model. It does not license causality, individual prediction, or portability to another ancestry. Weak-answer tell: claiming a hit "is the gene for" the trait.
Linkage disequilibrium: why tests aren't independent and hits aren't causal
Nearby variants correlate because recombination breaks haplotypes only so often. Two consequences follow, and interviewers test both:
- The threshold. SNPs in LD are not independent tests, so Bonferroni over 1,000,000 genotyped SNPs is conservative:
1,000,000 genotyped SNPs
~ 1,000,000 tests assumed by Bonferroni
~ 500,000 effectively independent tests, European ancestry, common variants
-> the true family-wise threshold is nearer 1 x 10^-7 than 5 x 10^-8
- The interpretation. The most significant marker may only tag the causal allele. Clumping picks approximately independent lead signals; conditional analysis reveals multiple signals in a region; fine-mapping assigns posterior support under LD and causal-number assumptions but cannot resolve statistically indistinguishable variants. Use ancestry-matched LD throughout — a mismatched LD panel is a classic production failure.
Probe: "Your lead SNP is significant but not coding. Is it causal?" Weak answer: no. Strong answer: it may tag a causal variant via LD; fine-mapping with matched LD, conditional analysis and functional annotation narrows the credible set, and indistinguishable candidates stay in it.
Phasing estimates haplotypes; imputation uses a reference panel to infer untyped allele dosages. Imputation quality, reference ancestry and marker overlap determine reliability — retain dosage uncertainty in association, filter with frequency-aware quality, and never hard-call low-confidence dosages without acknowledging the lost uncertainty.
Population stratification, relatedness, and the standard fixes
Ancestry confounds association when allele frequencies and trait prevalence both vary across ancestry-related groups: genotype and trait correlate for reasons that have nothing to do with the variant's biology. Relatedness violates independence the same way — relatives share genotypes and often environments. Neither is a race label; ancestry is continuous and context-dependent.
The standard toolkit, and what each one does and doesn't fix:
- PCA / ancestry covariates: principal components summarize the major axes of correlated genetic variation in the sample. Include the leading PCs as covariates. Limitation: captures coarse structure; fine-scale structure, environment and participation confounding survive.
- Genomic control: divide test statistics by the inflation factor. Blunt — it also shrinks genuine polygenic signal.
- Linear/logistic mixed models with a kinship (genetic relationship) matrix: model genome-wide phenotypic covariance from genetic similarity, retaining related samples. Limitations: kinship must be estimated from suitable variants, and the tested region can contaminate its own relationship estimate (proximal contamination) unless handled.
- Removing relatives: simplifies the analysis but discards information and can change representativeness; family designs or cluster-robust approaches retain samples when assumptions match.
Probe: "Lambda is 1.15 in your GWAS. What do you do?" Weak answer: "add more PCs." Strong answer: distinguish confounding from polygenicity (QQ plot shape, LD-score intercept where assumptions fit), check batch and ancestry stratification, compare models, and note that overcorrection can also erase true signal.
QC before analysis
Sample and variant QC precede any association, with thresholds predeclared and checked for differential removal by outcome or ancestry:
- Samples: call rate, heterozygosity, sex-chromosome consistency with recorded metadata, duplicates and relatedness, ancestry structure, contamination. A sex check or fingerprint that conflicts with metadata is a labeling failure to resolve, not a p-value problem.
- Variants: missingness (measured by batch and outcome — differential genotyping failure creates spurious association), allele frequency, reference orientation, imputation quality, and Hardy-Weinberg equilibrium.
Hardy-Weinberg deviation deserves its own sentence because it is the most misused filter: deviation can indicate genotyping error, population mixture, inbreeding or true association. It is a diagnostic, not a universal deletion command — test it in appropriate unrelated ancestry groups (often controls separately), with sample-size-aware thresholds and documented exceptions.
MAF affects power and asymptotic calibration; rare-variant analysis aggregates by gene or functional group and must handle annotation, direction and allele-count uncertainty. Imputation INFO/quality scores need frequency-aware filtering: rare-allele dosages can be unstable even when broad metrics look acceptable — one cutoff hides rare failure.
Probe: "Cases were genotyped on a different plate than controls. What could that produce?" Weak answer: "batch effects." Strong answer: differential missingness or differential genotype error correlated with outcome — spurious genome-wide hits — and the fix is measuring missingness by batch and outcome, inspecting cluster plots, and sensitivity analysis, not just adding batch covariates.
From single hits to scores: meta-analysis and downstream uses
Meta-analysis combines compatible studies while retaining cohort-level privacy. Harmonize build, variant identity, effect allele, trait units and model before pooling; strand-ambiguous A/T and C/G variants require allele-frequency or reference evidence, not blind flipping. Fixed-effect meta-analysis estimates one common effect; random-effects models heterogeneity but does not explain its source — report cohort effects and investigate predeclared moderators.
Summary statistics support LD-score regression (partitioning polygenicity from confounding under assumptions), genetic correlation, Mendelian randomization and polygenic scores — all sensitive to sample overlap, ancestry-matched LD, allele harmonization and selection. MR additionally needs relevance, independence and exclusion restriction; pleiotropy breaks the causal reading.
Polygenic scores sum risk alleles weighted by estimated effects. Evaluate discrimination, calibration and clinical utility in independent target populations. Scores commonly transfer poorly across ancestries because LD, allele frequency, environment, phenotype and training representation differ. A worked illustration: a European-trained coronary PRS at the 90th percentile showed ~3.2-fold risk in the discovery ancestry and ~1.4-fold with miscalibrated absolute incidence in an African-ancestry validation set of 4,100 people (hypothetical figures, but the pattern is well documented). A strong population association does not determine an individual's outcome. Ship a score with locked variants, alleles and weights, per-population calibration plots, and the training ancestry disclosed — or don't ship it.
Production checklist and what interviewers listen for
The production invariant is calibrated inference: every association or score traceable to a defined population, phenotype, reference build, QC flow, model and effect allele, with confounding, multiplicity, uncertainty, transfer limits and privacy obligations quantified rather than implied away. Genomes identify people and their relatives: minimize access, encrypt, audit, and release only under compatible consent.
The recurring pass/fail pattern across every topic above:
| Topic | The claim you must defend | The first production miss |
|---|---|---|
| Estimand | Population, phenotype, effect scale and model stated | Only the smallest p-value |
| Sample QC | Genetics conflicts with metadata, resolved before exclusion | Mislabeling |
| Variant QC | Differential missingness creates spurious association | Case-control batch confounding |
| HWE | Diagnostic for error, structure, inbreeding or real biology | Deleting a true association |
| Allele orientation | Effect and sign refer to a named allele on a named build | Sign flip after merge |
| Stratification | Ancestry-correlated frequencies and trait differences | Fine-scale structure surviving PCs |
| Relatedness | Correlated genotypes, modeled or excluded deliberately | Wrong kinship in admixed groups |
| Mixed models | Genome-wide covariance modeled; tested region handled | Proximal contamination |
| Binary traits | Separation and sparse cells bias or inflate estimates | Infinite OR reported as a finding |
| Multiple testing | Threshold matched to the question (FWER vs FDR) | Phenotype fishing |
| LD | Lead SNP tags a block; fine-mapping keeps tied candidates | Wrong LD panel |
| Imputation | Dosages with uncertainty, frequency-aware filtering | Reference ancestry mismatch |
| Meta-analysis | Build, allele, units and model harmonized | Strand flip |
| Rare variants | Masks and annotations defined before testing | Annotation circularity |
| LDSC / MR | Assumptions stated; intercept and pleiotropy checked | Sample overlap ignored |
| PRS | Validated per intended population, calibration reported | Data leakage between train and test |
| Portability | LD, frequency, environment and representation differ | Disparity amplification at deployment |
| Provenance | Full schema: alleles, counts, model, QC, software | Allele ambiguity in released summary stats |
| Operations | Consent, checksums, immutable results, disclosure review | Sample reidentification |
Likely follow-ups after any answer: "what if that assumption fails," "how would you detect it in the data you have," and "what would you tell a clinician this result means." Have a concrete detection step for every fix you name — that is the difference between a senior answer and a recited one.
