Overview
Curated: · Written: · Reviewed:
A variant call is a probabilistic claim tied to a reference and representation
Variant calling infers how sequenced sample molecules differ from a reference assembly. The pipeline observes noisy reads, chooses alignments among repeats and homologous regions, models sequencing and sample biology, and emits genotype or allele hypotheses with evidence. A VCF record is not the variant itself and a PASS label is not a clinical conclusion.
The mental model: one molecule's sequence, degraded three times
Start from the molecule. A DNA fragment existed in the sample with a definite sequence. Sequencing degraded it once: the instrument emitted base calls with per-base error probabilities. Alignment degraded it a second time: the read was placed somewhere in a reference that may not contain the true source locus, and the placement carries its own confidence. Calling degrades it a third time: from noisy, ambiguously placed reads the caller infers a genotype, which is a probability distribution over hypotheses, collapsed into one reported answer plus a quality score.
Every field in the final VCF is a summary of this chain, and every field is conditional on choices made upstream: which reference, which aligner, which error model, which ploidy. When an interviewer asks "is this variant real?", the strong answer traces the claim back through these three degradations and states what would have to be true for the call to be wrong. The weak answer treats the VCF row as ground truth and argues from annotations.
The data path: FASTQ → BAM → VCF as a representation story
Each format answers a different question, and the transitions between them are where information is lost or assumed.
FASTQ holds reads as sequence plus per-base quality. A quality character encodes an estimated error probability on a log scale under a declared ASCII offset (Phred+33 for the common Sanger/Illumina convention). These are estimates, not measurements: instrument calibration, cycle position, sequence context and chemistry can make them systematically optimistic or pessimistic. Base-quality recalibration learns empirical residual error patterns using trusted known sites, so genuine variation is not learned as error — but it cannot repair wrong sample identity or bad alignment, and on non-human platforms the known-sites resources may not transfer at all.
SAM/BAM holds alignments: where each read was placed, how (CIGAR), with what confidence (MAPQ), in which library (read group), and with what mate relationship. CRAM can encode alignments relative to an external reference, which compresses well but means decoding requires recovering the exact reference sequences used for encoding — record reference URIs and checksums, or your archive is unreadable in practice.
Two confidences live here and get conflated constantly. Base quality is confidence in a called residue; mapping quality is confidence in the read's placement. Repeats, paralogs, alternate loci and reference incompleteness reduce MAPQ, and MAPQ calibration is aligner-specific — a MAPQ of 30 from one aligner is not the same claim as 30 from another. A mapped read is not necessarily uniquely or correctly placed.
Duplicate marking identifies reads likely derived from the same original molecule. Optical duplicates, PCR duplicates and independent fragments that coincidentally share endpoints have different interpretations, and the right treatment depends on library design: molecular barcodes define molecules directly, while coordinate-based marking is a heuristic. Removing duplicates blindly damages high-depth, amplicon and single-cell analysis, where apparent duplicates are often genuine molecules.
VCF holds the inference: per-site hypotheses with evidence. The transition from BAM to VCF is where the pipeline stops reporting observations and starts making claims.
VCF anatomy, column by column
A VCF record is dense, and nearly every filtering decision is made from the fields on one line:
#CHROM POS ID REF ALT QUAL FILTER INFO FORMAT SAMPLE1
chr7 117559593 . G A 842.7 PASS DP=87;AF=0.494;MQ=59.2 GT:AD:DP:GQ 0/1:44,43:87:99
- CHROM/POS: 1-based position of the record's first reference base. Not zero-based half-open like BAM — the single most common off-by-one in production code.
- REF: must match the reference sequence at that position. This is the field that catches reference-build mismatches; a REF that doesn't match the FASTA means the file was made against a different reference, whatever the header claims.
- ALT: ordered alternate alleles. Their order defines the allele indexes used everywhere else in the record.
- QUAL: phred-scaled probability that the site is not variant — it concerns the site, not any individual genotype. A site QUAL of 842.7 is roughly a 10^-84 chance the site is non-variant.
- FILTER: PASS means the record passed the filters declared by this file's producing workflow — nothing more. A dot means filters were not applied or are unspecified; never coerce dot to PASS.
- INFO: site-level annotations (here DP=87 total depth, AF=0.494 allele fraction, MQ=59.2 mean mapping quality).
- FORMAT/SAMPLE: per-genotype fields, in the order the FORMAT column declares.
Genotype fields: how a call is actually represented
The sample column 0/1:44,43:87:99 decodes as:
- GT
0/1: heterozygous, unphased. Indexes refer to REF (0) and ordered ALT alleles (1, 2, …). A slash means unphased; a pipe means phased order, valid only within a phase set — check the PS field before claiming phase continuity across a region. - AD
44,43: allelic depths — 44 reference reads, 43 alternate reads. Evidence, not confidence. - DP
87: total depth. Note DP and sum(AD) can legitimately differ (low-quality reads excluded from AD); they are not interchangeable. - GQ
99: genotype quality — phred-scaled confidence that the selected genotype is right relative to the next-best alternative, not that the site is variant. GQ is capped at 99 in many callers; a capped value means "at least this confident."
The 44/43 split is what makes this call trustworthy, and its absence is what makes others suspect:
| AD (ref,alt) | Allele balance | Reading |
|---|---|---|
| 44,43 | 0.494 | clean heterozygote |
| 78,9 | 0.103 | likely artifact, or somatic at low fraction, or contamination |
| 3,2 | 0.400 | balance fine, depth too low to trust: DP=5 |
| 0,91 | 1.000 | homozygous alternate, expect GT 1/1 not 0/1 |
| 61,58 | 0.487 | clean, but check MQ — balanced noise in a repeat looks like this |
A heterozygous call at allele balance 0.103 is the common one to catch. In a germline sample there is no mechanism producing 10% alternate reads at a true heterozygous site; the usual causes are alignment artifacts in a repetitive region, PCR error, index hopping, or sample contamination. MQ=59.2 here is near the maximum of 60, so these reads map uniquely and a mapping artifact is unlikely.
Missing genotype (./.), reference genotype (0/0) and an absent sample are three different statements. A gVCF preserves the distinction by recording reference-confidence blocks — evidence that a sample was observed and appears reference-like — which is exactly what lets joint genotyping distinguish "reference at this position" from "no data at this position." Merge plain VCFs and that difference silently becomes a false negative.
Ploidy determines the genotype space: sex chromosomes, organelles, tumors and pooled samples all violate a flat diploid assumption. Genotype likelihoods (PL in VCF) are the relative probabilities of the observed read data under each possible genotype, ordered per the header's ploidy and allele count — decode the indexes from the header definition, don't assume diploid ordering.
How callers actually work, and where they break
Modern small-variant callers (GATK HaplotypeCaller being the archetype) don't evaluate isolated mismatches. They detect active regions, assemble candidate haplotypes locally from reads, realign reads against those haplotypes, and compute genotype likelihoods over the haplotype graph. This reduces reference-alignment artifacts around indels — a deletion realigns the downstream bases, so a sloppy aligner's artifact cluster disappears — but the output remains conditional on active-region detection, the candidate graph, ploidy and the error model. A true haplotype pruned from the graph is unrecoverable downstream; no filter setting recovers it.
Deep-learning callers learn representations from training data and must be revalidated for platform, chemistry, ancestry, coverage and genomic context shifts. Their failure modes are distributional, not random.
Germline calling assumes inherited alleles under a ploidy model. Somatic calling seeks variants in a mixture of tumor and normal cells where purity, subclonality, copy number and contamination all shift expected allele fractions. Low allele fraction is not synonymous with error, and a germline caller's diploid heterozygosity expectation is not a valid tumor model — a 10% allele fraction is contamination in one context and a real subclone in the other, and only the matched normal and purity model tell them apart.
Normalization: why the same variant has two spellings
VCF represents alleles relative to REF, with insertions and deletions normally including an anchor base. Equivalent haplotypes can have multiple textual representations, especially in repeats: a 2 bp deletion inside a homopolymer can be written at several positions with the same resulting sequence. Two callsets of the same variant can therefore disagree at the string level.
Normalization fixes this with three operations against a pinned reference: split multiallelic records per explicit policy, trim shared sequence from the ends of REF/ALT, and left-align the alleles. After normalization, comparison should still be haplotype-aware, because complex clusters can remain ambiguous.
Multiallelic decomposition is where pipelines quietly corrupt data. GT allele indexes, and allele-indexed INFO/FORMAT fields like AD and PL, must be transformed when a record splits: 1/2 with ALT A,T becomes two records with remapped genotypes, and AD arrays must be subset per allele. Get the remap wrong and a heterozygote becomes a wrong homozygote. Use specification-aware tools and regression-test the round trip.
Reference identity underlies all of it. Assembly name, accession version, contig dictionary, decoys, alternate sequences and masking influence alignment and coordinates. Records from different builds cannot be merged by matching chromosome and position. Liftover is an alignment-based transformation that can fail, reverse orientation or change allele representation; always verify REF against the destination reference.
A merge that looked clean on chromosome and position was the usual disaster: 14,812 "shared" SNPs between a GRCh37 callset and a GRCh38 callset, of which 2,041 had REF alleles that did not match the destination FASTA. Liftover without REF verification invented 2,041 variants. Compare contig dictionaries and checksums first; then normalize with a haplotype-aware comparator inside confident regions, not by string-matching POS.
Filtering and quality: what a raw callset is, and how truth is established
A raw callset is the caller's full hypothesis set with likelihoods attached. Filtering turns it into a deliverable, and the honest pipeline keeps both: filter genotypes and sites at the appropriate level without erasing raw evidence.
Hard filters encode explicit thresholds — MQ, DP, allele balance, strand bias — and their failure mode is that thresholds tuned on one assay or cohort don't transfer. Statistical recalibration (VQSR-style) ranks variants using annotations plus truth and training resources; it fails for small cohorts, shifted assays or poorly represented variant classes, because the model has nothing reliable to learn from. Either way, FILTER semantics are per-file: PASS means this workflow's declared filters passed, and every FILTER ID must resolve against the header.
Truth is established by benchmarking with a defined scope: truth set, confident regions, reference, sample, variant classes and comparison semantics. False positives and false negatives count only where truth is asserted — a call outside the confident regions is neither. Precision is the fraction of your calls that are true; recall is the fraction of benchmark truth variants your callset recovered; they are not reciprocals of each other. Both change with representation matching and stratification, and a single whole-genome F1 can hide failures in difficult repeats or clinically important regions. Stratify by variant type, size and genomic context, and investigate discordant variants rather than the aggregate score.
Structural variants extend the representation problem: symbolic alleles, END, confidence intervals and breakend notation describe imprecise or complex rearrangements, and one VCF row may describe only part of an event — event identity and mate relationships must be preserved or the callset is uninterpretable. Copy-number calls depend on depth, allelic balance and segmentation, with GC, mappability, purity and ploidy all moving observed depth, and often use different coordinate semantics from small variants.
Misconceptions, scale, and what interviewers probe next
Misconceptions that fail interviews:
- "PASS means real." PASS means this file's declared filters passed. A dot means filters weren't applied — the opposite claim.
- "QUAL is the genotype confidence." QUAL is about the site being variant; GQ is about the selected genotype versus alternatives.
- "DP and AD measure confidence." They measure evidence. Five reads and 500 reads can both be wrong for different reasons.
- "Low allele fraction means artifact." In a tumor it can be a real subclone; in a germline sample it usually is an artifact. The ploidy and mixture model decides.
- "Same position means same variant." Only after normalization against the same reference, and even then compare haplotypes in repeats.
What changes at scale: joint genotyping across a cohort produces consistently represented sites, but adding samples changes genotypes and annotations — cohort release and sample set are part of provenance. Deep-learning callers degrade under distribution shift. Benchmarking stratification matters more as callsets grow, because aggregate metrics improve while difficult regions stay broken.
Likely follow-ups, in rough order: which reference and how do you verify it (checksums, contig dictionaries, REF validation); how you'd tell artifact from low-fraction variant (allele balance, MQ, matched normal, purity); how you'd merge two callsets (normalize, haplotype-aware compare, never string-match POS); how you'd benchmark (scope, stratification, discordant review); and how the whole thing is reproducible — read groups, pedigree, checksums, tool and model versions, parameters, commands, logs and raw calls, so every call can be reconstructed from exact inputs. Genomic data is identifying: consent-aware identifiers, access controls, encryption, audit and retention policy are part of the pipeline, not an afterthought.
The production invariant is traceable evidence: every call can be reconstructed from exact sample reads, reference, alignment and caller configuration, and its representation, genotype likelihoods, filters, benchmark scope and annotation versions remain separable from downstream biological or clinical interpretation.
