Top 100 Computational Biologist Interview Questions and Answers
The questions most likely to actually come up in your Computational Biologist interview, ranked by likelihood — with detailed, senior-level answers covering what an interviewer is really listening for.
Curated: · Written: · Reviewed:
QA-1A hit is 92% identical over 38 residues. Why is that not 92% homology, and what else must be on the report?(show answer)
The first thing I would pin down about homology versus percent identity is which claim the number is actually allowed to support.
Homology is a yes-or-no statement about common ancestry; percent identity is a length-dependent similarity statistic under a scoring model. Neither identity nor an E-value by itself is orthology or function.
Concretely, report alignment length, query and subject coverage, bit score, E-value with database size, gaps, and the exact BLAST program. Require domain architecture before transferring a name. A 38-residue match on a 510-residue query is a motif, not a full-length homolog.
The reason for that specificity is a failure I have seen: An annotation pipeline transferred "kinase" from a 38-residue BLAST hit at 92% identity; 410 of 510 residues were unmatched and the protein was a dehydrogenase. Twelve hundred records inherited the wrong name.
Same 92%, two different claims.
| Hit | Identity | Aln length | Query coverage | Allowed claim |
|---|---|---|---|---|
| A | 92% | 38 | 7% | motif only |
| B | 92% | 498 | 98% | full-length candidate |
| used | 92% | 38 | 7% | wrongly named kinase |
I would not consider it settled without evidence: Show query coverage, subject coverage, and bit score beside identity; reject name transfer below a predeclared coverage and domain-architecture rule.
Identity is a local score; homology is an evolutionary claim.
Curated: · Written: · Reviewed:
QA-2When must you use Needleman-Wunsch instead of Smith-Waterman, and what does the wrong choice invent?(show answer)
I would start global versus local alignment from the samples, reference, and assay, not from the first plot that looked clean.
Global alignment scores end-to-end correspondence and is for sequences expected to match across most of their lengths. Local alignment finds the best subsequence and will happily ignore unrelated termini.
Concretely, use Needleman-Wunsch for two complete homologous proteins of similar length; use Smith-Waterman or BLAST for shared domains, fragments, or database search. Terminal-gap policy is part of the model, not a display option.
The reason for that specificity is a failure I have seen: A fusion protein was globally aligned to a 180-residue domain; the unmatched 620 residues were gapped and the "identity" looked like 88% over the domain while the rest of the gene was fiction.
An 800-residue fusion versus a 180-residue domain.
| Method | Score region | Unmatched ends | Misleads? |
|---|---|---|---|
| Smith-Waterman | 180 aa domain | free | no |
| Needleman-Wunsch | forced end-to-end | 620 gapped | yes |
| reported identity | 88% over domain | fusion hidden |
I would not consider it settled without evidence: State expected homology scope before choosing the algorithm; if termini are not homologous, local alignment is the method, not a post-hoc trim.
The algorithm is a statement about which residues are allowed to be unmatched.
Curated: · Written: · Reviewed:
QA-3Why open a gap at 11 and extend at 1 rather than charging every indel residue the same, and when does that still fail?(show answer)
This is a place where a green QC report and a correct handling of affine gap penalties are not the same event.
An affine gap charges opening more than extension because one multi-residue indel is usually one event. Linear penalties invent many independent events in a repeat.
Concretely, keep match, gap-in-query, and gap-in-subject states. Document the 11/1 (or program default) pair with the matrix. Test a 12-residue indel versus twelve 1-residue gaps; the scores must differ.
The reason for that specificity is a failure I have seen: A linear gap of -4 per residue split a 12-residue indel into 12 scattered gaps and called several events where the biology was one exon skip.
One 12-residue indel, two penalties.
| Model | Open | Extend | Gaps drawn | Score |
|---|---|---|---|---|
| affine | 11 | 1 | 1 | 22 |
| linear | 4/res | 4/res | 12 | 48 |
| biology | one event | 1 | affine wins |
I would not consider it settled without evidence: A unit test on a known 12-residue indel where affine 11/1 keeps one gap costing 22 and linear -4 does not.
Gap costs are an evolutionary hypothesis, not a display preference.
Curated: · Written: · Reviewed:
QA-4Someone says "higher matrix number means more distant sequences." For BLOSUM, why is that backwards, and how do you pick?(show answer)
My answer to BLOSUM numbering versus PAM distance begins at the confounder: if I cannot name what would fake the signal, I do not have a design.
BLOSUM62 is built from blocks clustered at 62% identity, so a higher BLOSUM number is closer sequences. PAM250 is a longer evolutionary distance than PAM30. The two families number in opposite directions.
Concretely, start from the program default (often BLOSUM62 for proteins). Move toward BLOSUM80 or PAM30 for close homologs and BLOSUM45 or PAM250 for remote ones, then benchmark sensitivity and false hits rather than picking by folklore.
The reason for that specificity is a failure I have seen: A remote-homolog search used BLOSUM80 because "higher is more sensitive"; true remote hits dropped 31% versus BLOSUM45 on the same 400-query set.
Which way is distant?
| Family | Close | Distant |
|---|---|---|
| BLOSUM | 80 | 45 |
| PAM | 30 | 250 |
| default blastp | 62 | not "max" |
I would not consider it settled without evidence: A table of recall at 1% false discovery for BLOSUM45, 62, and 80 on a labeled remote-homolog set.
The number is a clustering threshold or a distance, not a universal "strength" dial.
Curated: · Written: · Reviewed:
QA-5The same alignment has E-value 3e-112 against nr and about 1.7e-114 against a 12 Mb genome. What changed, and which number travels?(show answer)
I would treat E-value search space as an inference under a model, not as a file the pipeline emitted.
An E-value estimates chance hits at least this strong in the effective search space. Database size is part of the statistic. Bit score normalizes the raw score for the scoring system and is the number that can move between searches.
Concretely, report bit score, E-value, and database name plus date together. Do not copy an E-value threshold from nr onto a small custom database. Recalculate if the database grows.
The reason for that specificity is a failure I have seen: A lab kept "E < 1e-10 means real" from nr on a 12 Mb genome; 86 low-complexity hits passed that copied threshold and became "orthologs."
One 412-bit alignment, two databases.
| Database | Letters | Bit score | E-value |
|---|---|---|---|
| nr | 2.1e9 | 412 | 3e-112 |
| 12 Mb genome | 1.2e7 | 412 | 1.7e-114 |
| copied rule | 86 false hits |
I would not consider it settled without evidence: Recompute E-values after any database update; compare bit scores when ranking the same alignment across databases.
E-values are search-specific; bit scores are the portable score.
Curated: · Written: · Reviewed:
QA-6You have a protein query and want unannotated coding loci in a genome. Which BLAST program, and what does swapping query and database invent?(show answer)
The useful question for tblastn versus blastx orientation is what still holds after a sample swap, a batch, or a database update.
tblastn translates the nucleotide database in reading frames and compares it to a protein query. blastx translates the nucleotide query against a protein database. Orientation is not cosmetic.
Concretely, protein query, genome database → tblastn. Nucleotide query, protein database → blastx. Account for six frames, genetic-code, and pseudogene hits. Do not treat a tblastn HSP as an expressed gene.
The reason for that specificity is a failure I have seen: A pipeline used blastx of 2 kb genome chunks against Swiss-Prot when the question was one protein versus a genome; coordinates were fragment-relative and 1,104 chunk-edge peptides were submitted as novel genes.
Program choice.
| Query | Database | Program |
|---|---|---|
| protein | genome nt | tblastn |
| mRNA | swissprot | blastx |
| 2 kb genome chunks | swissprot | blastx, wrong grain |
I would not consider it settled without evidence: Run tblastn of the protein against the genome on a known coding locus and a known intergenic control; coordinates must land on the genome, not on chunk files.
Who is translated is the experiment.
Curated: · Written: · Reviewed:
QA-7Two proteins are reciprocal best BLAST hits. Why is that not orthology, and what extra evidence do you require?(show answer)
I would settle reciprocal best hits against a negative control first, so the clever method has to survive it.
Reciprocal best hits are a heuristic for candidate orthologs. Duplication, loss, incomplete databases, and rate variation all produce RBH pairs that are paralogs or coincidences.
Concretely, require comparable coverage, then check synteny, phylogeny, or gene-tree/species-tree reconciliation before naming orthologs. Incomplete draft genomes are a known RBH trap.
The reason for that specificity is a failure I have seen: An annotation called 6,400 "1:1 orthologs" from RBH against a 70%-complete draft; 890 were lineage-specific duplicates whose true orthologs were in unassembled gaps.
RBH that was a paralog.
| Test | Result |
|---|---|
| A vs B RBH | yes |
| coverage | 94% |
| gene tree | A sister to A' duplicate |
| 1:1 emitted | 890 false |
I would not consider it settled without evidence: A gene tree on the RBH pair plus the two closest paralogs; if the species tree and gene tree conflict, do not emit 1:1.
Best hit is a search result; orthology is a history.
Curated: · Written: · Reviewed:
QA-8Yesterday's top hit is missing today. What must be stored with every BLAST result so that is not a mystery?(show answer)
The judgement in database version as part of the hit is which identifiers and versions are pinned, not which tool name is fashionable.
The database name, release or build date, sequence count, and tool version are part of the result. Accessions without versions are mutable names.
Concretely, store BLAST+ version, database date, command line, and versioned accessions. Re-run rather than comparing yesterday's E-values to today's nr.
The reason for that specificity is a failure I have seen: A clinical report cited NP_000001 without version; a weekend update replaced the record and the reported 99% identity pointed at a different isoform.
What to pin.
| Field | Example |
|---|---|
| blastp | 2.16.0 |
| db | nr 2026-08-01 |
| accession | NP_000001.3 |
| missing date | not a result |
I would not consider it settled without evidence: Checksum or date of the BLAST database file next to every stored hit; accession.version, not the display name.
A hit without a database date is not reproducible.
Curated: · Written: · Reviewed:
QA-9A search of a collagen-like protein returns thousands of hits. What is masking for, and what can it hide?(show answer)
Where candidates lose the interview on low-complexity masking is usually treating a hit list as biology.
Low-complexity sequence generates high scores that are composition, not homology. Masking reduces those seeds but can also hide genuine repetitive biology.
Concretely, use program-appropriate SEG/dust or soft masking, then inspect unmasked alignments for the candidates you intend to keep. Document whether query, database, or both were masked.
The reason for that specificity is a failure I have seen: Hard-masked coiled-coil query lost 22 true homologs in a filament family; unmasked search then dumped 8,400 composition hits into the same table without a complexity filter on reporting.
Two wrong defaults.
| Mode | True homologs | Composition hits |
|---|---|---|
| hard mask | 4 of 26 | 12 |
| no mask | 26 | 8400 |
| soft + filter | 24 | 40 |
I would not consider it settled without evidence: Compare masked and unmasked hit lists on a labeled set; keep a reporting rule that is not "whatever BLAST printed."
Masking is a search bias you chose; it is not a finding.
Curated: · Written: · Reviewed:
QA-10A banded aligner is 40× faster. When is the band not allowed to contain the true alignment, and how do you know?(show answer)
I would answer banded dynamic programming by separating what the experiment measured from what the software merely displayed.
A band around the diagonal assumes the optimal path stays near it. Long indels or length differences push the path outside a narrow band and the "optimum" is then a constrained one.
Concretely, set band width from expected indel size, not from runtime. For unknown divergence, compare a subset against unbanded Smith-Waterman. Chromosome-scale pairwise DP is the wrong tool for database search anyway.
The reason for that specificity is a failure I have seen: A 16-residue band missed a 40-codon indel in an exon; the gene was called a frameshift versus the reference when unbanded alignment recovered one gap of 120 nt.
Band versus indel.
| Band (residues) | True indel | Path inside? |
|---|---|---|
| 16 | 40 | no |
| 64 | 40 | yes |
| reported | frameshift | wrong gene |
I would not consider it settled without evidence: On a 100-gene subset, unbanded versus banded scores and coordinates; any coordinate shift larger than the band is a method error.
Speed that excludes the path is a different algorithm.
Curated: · Written: · Reviewed:
QA-11Assembly A has contig N50 38 Mb and B has 19 Mb. Why might B still be the genome you should publish?(show answer)
The engineering content of N50 is not correctness is the evidence boundary and the exclusion log, not the p-value.
N50 is the length L such that half the assembled bases are in contigs ≥ L. It measures contiguity of submitted span, not accuracy, completeness, or absence of joins.
Concretely, report N50 with total span, contig count, NG50, BUSCO, misassembly counts, and contaminant span. Investigate any N50 jump that is not matched by mapping support.
The reason for that specificity is a failure I have seen: A chimeric 12 kb satellite join pushed N50 from 18.6 Mb to 38.2 Mb; QUAST later counted 14 misassemblies and a 2.84 Gb span that was 130 Mb of duplicated contaminant.
Two assemblies, one genome.
| Metric | A | B |
|---|---|---|
| N50 | 38.2 Mb | 18.6 Mb |
| span | 2.84 Gb | 2.71 Gb |
| BUSCO S | 91% | 96.1% |
| contaminant | 130 Mb | 0.4% |
I would not consider it settled without evidence: A release table with N50, NG50, BUSCO, misassemblies, and contaminant bases; N50 cannot be the only gate.
Longer contigs can be longer mistakes.
Curated: · Written: · Reviewed:
QA-12Why can NG50 fall when N50 rises, and which genome-size estimate are you allowed to use?(show answer)
Before releasing I would write what a wrong result for NG50 genome size would look like in the same table.
NG50 uses an expected genome size as the halfway target, so missing span is visible. A wrong genome-size estimate moves NG50 without changing the FASTA.
Concretely, state how genome size was estimated (k-mer, flow cytometry, prior reference). Apply the same size and contig filter to every comparison. Do not mix haploid and diploid size.
The reason for that specificity is a failure I have seen: A team used 3.2 Gb for NG50 on a 2.4 Gb bird genome copied from a mammal page; NG50 looked catastrophic while the assembly was nearly complete.
Same FASTA, two sizes.
| Assumed size | NG50 | Truth |
|---|---|---|
| 3.2 Gb mammal | 8 Mb | wrong size |
| 2.4 Gb k-mer | 19 Mb | honest |
| N50 | 19 Mb | unchanged |
I would not consider it settled without evidence: Publish the genome-size estimator beside NG50; a sensitivity table at ±10% size.
NG50 is only as honest as the size you fed it.
Curated: · Written: · Reviewed:
QA-13Raising k from 21 to 77 fragmented a short-read assembly. What did you buy and what did you lose?(show answer)
The first thing I would pin down about k-mer size tradeoff is which claim the number is actually allowed to support.
Larger k is more unique and resolves more repeats, but needs uninterrupted coverage across k bases, so errors and low depth fragment the graph.
Concretely, use odd k below read length. Inspect the k-mer spectrum first. Multi-k assemblers exist because one k is a compromise. Do not pick k by maximum N50.
The reason for that specificity is a failure I have seen: k=77 on 2×150 reads at 18× mean depth shattered a 4.5 Mb plasmid into 220 contigs; k=33 recovered it as one circle with 2 repeats left unresolved.
One plasmid, two k.
| k | Contigs | N50 | Circles |
|---|---|---|---|
| 33 | 3 | 4.5 Mb | 1 |
| 77 | 220 | 22 kb | 0 |
| depth | 18× | low for 77 |
I would not consider it settled without evidence: k-mer spectra plus a small grid of k versus N50, BUSCO, and misassemblies on the same reads.
k is a resolution-versus-connectivity knob, not a quality score.
Curated: · Written: · Reviewed:
QA-14When do you pay for PacBio HiFi instead of cheaper raw long reads, and what does polishing not replace?(show answer)
I would start HiFi versus noisy long reads from the samples, reference, and assay, not from the first plot that looked clean.
HiFi circular-consensus reads are long and accurate; raw long reads are longer or cheaper but need error-aware assembly and polishing. Polishing does not repair a wrong join.
Concretely, match molecule length to repeat length. If haplotypes must be phased, HiFi plus a phasing graph is the default. If the DNA is already degraded to 4 kb, neither platform invents span.
The reason for that specificity is a failure I have seen: A 40 kb tandem repeat was assembled from 8 kb raw reads and polished to Q40; the copy number was still 3 instead of 11 because no read spanned the array.
Repeat array of 40 kb.
| Data | Span | Consensus Q | Copies |
|---|---|---|---|
| 8 kb raw + polish | no | Q40 | 3 |
| 45 kb HiFi | yes | Q30+ | 11 |
| need | molecule > 40 kb array |
I would not consider it settled without evidence: Read N50 and accuracy histogram versus the repeat-length spectrum of the genome; polishing QV is a separate column.
Accuracy without span still collapses repeats.
Curated: · Written: · Reviewed:
QA-15A diploid assembly has 1.12× expected span. How do you tell collapsed heterozygosity from duplicated haplotigs, and what does aggressive purge destroy?(show answer)
This is a place where a green QC report and a correct handling of haplotype collapse versus false duplication are not the same event.
Collapse merges homologous haplotypes into one mosaic; haplotigs keep both and look like extra copies. Purging by coverage can delete real segmental duplications.
Concretely, inspect coverage of primary versus alternate, bubble structure, and BUSCO duplicated counts. Trio or Hi-C phasing is evidence, not a guarantee. State the representation goal before purge.
The reason for that specificity is a failure I have seen: purge_dups removed 64 Mb of real segmental duplication that had 1× haploid coverage; a clinical CNV caller then missed 9 disease-relevant copies.
Span 1.12× haploid.
| Class | Coverage | Action |
|---|---|---|
| haplotig | ~1× each | purge or phase |
| collapse | ~2× mosaic | do not purge |
| real duplicon | ~1× extra | keep |
I would not consider it settled without evidence: Coverage histograms of purged versus retained sequence plus an independent FISH or long-read spanning count on known duplicons.
1× coverage is not automatically the extra haplotype.
Curated: · Written: · Reviewed:
QA-16BUSCO reports 97% complete. What has that not measured, and which lineage dataset are you even using?(show answer)
My answer to BUSCO is not the whole genome begins at the confounder: if I cannot name what would fake the signal, I do not have a design.
BUSCO estimates recovery of near-universal single-copy orthologs expected in a chosen lineage dataset. It is not noncoding completeness, copy-number truth, or chromosome order.
Concretely, name the lineage dataset and version. Report complete single, duplicated, fragmented, and missing. Relate duplicates to ploidy and haplotigs. Pair with k-mer completeness and mapping.
The reason for that specificity is a failure I have seen: A plant assembly scored BUSCO 98% on embryophyta_odb10 while missing 180 Mb of repeat-rich pericentromeres that held the trait locus the project funded.
97% BUSCO, missing trait.
| Check | Result |
|---|---|
| BUSCO S | 97% |
| k-mer complete | 88% |
| trait locus | absent |
| funded question | fails |
I would not consider it settled without evidence: BUSCO table plus k-mer completeness and a known-locus recovery list; the trait locus must be in the recovery list, not implied by BUSCO.
Conserved proteins can look fine while the genome you care about is absent.
Curated: · Written: · Reviewed:
QA-17A contig BLAST-hits a vector at 99%. Why is that still not enough to delete it, and what bundle of evidence is?(show answer)
I would treat assembly contamination as an inference under a model, not as a file the pipeline emitted.
Taxonomy, coverage, composition, linkage, and sample controls together classify contamination. A single BLAST hit can be a conserved gene, a horizontal transfer, or a mislabeled database record.
Concretely, compare blanks and related samples. Require concordant coverage and no spanning host reads before deletion. Quarantine uncertain contigs with reasons. Horizontal transfer is a scientific claim, not a default delete.
The reason for that specificity is a failure I have seen: Every "bacterial" contig was dropped from a gut metagenome host assembly, including a 48 kb genuine insertion with spanning HiFi reads; the paper then "discovered" it in a later long-read study.
Same 99% vector hit, two outcomes.
| Contig | Host-spanning HiFi | Blank | Decision |
|---|---|---|---|
| 12 kb | 0 | yes | delete |
| 48 kb insert | 14 | no | keep |
| BLAST-only rule | both deleted |
I would not consider it settled without evidence: A contamination review table: taxonomy, coverage, GC, host-spanning reads, blank presence, decision, and operator.
BLAST is a clue, not a delete key.
Curated: · Written: · Reviewed:
QA-18Three rounds of polishing raised QV from 28 to 42. Why might the assembly still be wrong, and when do you stop?(show answer)
The useful question for polishing a misassembly is what still holds after a sample swap, a batch, or a database update.
Polishing corrects consensus bases and small indels using aligned evidence. It reinforces whatever join you already made. Structural error needs spanning evidence, not more pileup.
Concretely, track each round's QV, BUSCO, and mapping discordance. Stop when independent metrics stall or worsen. Never polish with a reference that is the question.
The reason for that specificity is a failure I have seen: Polishing a chimeric join to Q42 made the false junction look like high-quality sequence; 14 genes were fused in the annotation.
Three polish rounds.
| Round | QV | Discordant pairs | Misassemblies |
|---|---|---|---|
| 0 | 28 | 4.1e4 | 14 |
| 3 | 42 | 5.8e4 | 14 |
| stop | QV up, structure not |
I would not consider it settled without evidence: QV, misassembly count, and discordant pairs per round; a round that improves QV while raising discordance is a stop.
A beautiful consensus can sit on a broken join.
Curated: · Written: · Reviewed:
QA-19The assembler flagged a contig as circular. What independent evidence is required before you rotate and publish it?(show answer)
I would settle circular closure evidence against a negative control first, so the clever method has to survive it.
Closure needs consistent end overlap plus reads that span the join and coverage that supports one traversal. A "circular" tag is a hypothesis. Rotation is a coordinate choice, not biology.
Concretely, confirm unique overlap, spanning molecules, and even coverage across the join. Remove only duplicated overlap bases. Document the origin (dnaA, a convention) without changing sequence content.
The reason for that specificity is a failure I have seen: A repeat-driven false closure duplicated 2.1 kb at the origin of a 5.2 Mb chromosome; every coordinate in the paper was shifted and the origin gene was annotated twice.
Proposed bacterial circle.
| Evidence | Pass? |
|---|---|
| 2.1 kb overlap | yes, but a repeat |
| spanning HiFi | 0 |
| coverage dip | 0.3× at join |
| publish as circular | no |
I would not consider it settled without evidence: Spanning-read count, overlap identity, and coverage across the proposed join in the release package.
Circles are joins you can show, not flags you can trust.
Curated: · Written: · Reviewed:
QA-20Mean coverage is 40×. Why can half the genome still be unassembled, and which plot answers that?(show answer)
The judgement in nominal coverage versus dropout is which identifiers and versions are pinned, not which tool name is fashionable.
Nominal coverage is total bases divided by haploid genome size. Mean hides GC dropout, duplicates, organelles, and inaccessible repeats. Lander-Waterman assumes uniform independent sampling that libraries violate.
Concretely, inspect k-mer and mapped-depth histograms, not the mean. Report organellar overrepresentation separately. Set targets in filtered nuclear bases.
The reason for that specificity is a failure I have seen: A GC-poor fungus was sequenced to a headline 40× of which 22× of those genome-equivalents were mitochondrial; nuclear unique sequence sat at 9× and the assembly missing 38% of BUSCO was called a library failure a month later.
Headline 40× library.
| Compartment | Genome-equivalents |
|---|---|
| mitochondrial | 22× |
| nuclear unique | 9× |
| rest / repeats | 9× |
| headline mean | 40× |
I would not consider it settled without evidence: Depth histogram with nuclear versus organellar split; mean is a footnote.
Forty-fold average can be nine-fold where it counts.
Curated: · Written: · Reviewed:
QA-21Two VCFs share chr1:12345. Why might they still name different loci, and what do you compare before merging?(show answer)
Where candidates lose the interview on reference build mismatch is usually treating a hit list as biology.
Coordinates are meaningful only inside one reference accession, contig dictionary, and version. Chromosome and position are not a universal key.
Concretely, compare assembly name, contig lengths, and checksums. Confirm every REF base matches the destination FASTA. Liftover is an alignment, not a rename; it can reverse and drop.
The reason for that specificity is a failure I have seen: A merge of 14,812 “shared” SNPs between GRCh37 and GRCh38 kept 2,041 records whose REF did not match GRCh38; those became invented variants in the joint callset.
chr1:12345 in two builds.
| Build | REF | Same locus? |
|---|---|---|
| GRCh37 | A | |
| GRCh38 | G | no |
| merged anyway | 2041 false SNPs |
I would not consider it settled without evidence: A pre-merge report of contig length mismatches and REF mismatches; zero REF failures required.
Same numbers on the page are not the same genome.
Curated: · Written: · Reviewed:
QA-22You convert a VCF SNP and a VCF insertion to BED. What off-by-one happens if you treat both like a 0-based SNP?(show answer)
I would answer VCF one-based positions by separating what the experiment measured from what the software merely displayed.
VCF positions are 1-based at the first reference base of the record. Insertions include an anchor base. BED is usually 0-based half-open. Class-specific conversion is required.
Concretely, sNP at POS 100, REF=A → BED [99,100). A 3 bp deletion with REF=ACGT ALT=A at POS 100 spans BED [99,103). Test first-base, insertion, deletion, and symbolic alleles.
The reason for that specificity is a failure I have seen: A 3 bp deletion at POS 100 was loaded as BED 100-103 instead of [99,103); 18,000 variants shifted one base and missed every ClinVar overlap.
POS=100 on a 1-based VCF.
| Class | REF/ALT | BED [0, half-open) |
|---|---|---|
| SNP | A/G | 99-100 |
| del 3bp | ACGT/A | 99-103 |
| wrong SNP rule on del | 100-103 |
I would not consider it settled without evidence: Round-trip fixtures for SNP, insertion, deletion, and BND against a 10-base toy reference.
Position arithmetic is a variant-class function.
Curated: · Written: · Reviewed:
QA-23FILTER=. and FILTER=PASS look equally “fine” in a spreadsheet. What is the difference, and which one may you treat as passing?(show answer)
The engineering content of FILTER PASS versus dot is the evidence boundary and the exclusion log, not the p-value.
PASS means the producing workflow applied its declared filters and this record passed. A dot means filters were not applied or are unspecified. They are not equivalent.
Concretely, validate FILTER IDs against the header. Retain raw calls. Never coerce missing filter status to PASS at ingest.
The reason for that specificity is a failure I have seen: An aggregator rewrote "." to PASS for 2.4 million sites from an unfiltered gVCF dump; a diagnostic panel reported 310 unfiltered artifacts as clinical candidates.
Three FILTER values.
| FILTER | Meaning | Clinical? |
|---|---|---|
| PASS | declared filters passed | maybe |
| q30 | failed a named filter | no |
| . | unknown | no |
I would not consider it settled without evidence: Ingest rejects files that contain "." if the contract required filtering, and never silently rewrites the field.
Missing filters are not a pass.
Curated: · Written: · Reviewed:
QA-24A sample is 0/2 on a record with ALT=C,T. Which allele is 2, and what breaks if you split the record without rewriting GT?(show answer)
Before releasing I would write what a wrong result for GT allele indexes would look like in the same table.
GT indexes REF as 0 then ALT in order. 0/2 is REF plus the second ALT. Allele-indexed AD and PL arrays belong to that full allele set.
Concretely, parse FORMAT per record. When splitting multiallelics, remap GT, AD, and PL with a spec-aware tool. Do not copy the row.
The reason for that specificity is a failure I have seen: A split remapped GT 0/2 to 0/1 on the T row but left AD as 12,3,40; the T allele inherited the C count of 3 instead of 40.
REF=A ALT=C,T GT=0/2.
| Index | Allele |
|---|---|
| 0 | A |
| 1 | C |
| 2 | T |
| genotype | A/T |
I would not consider it settled without evidence: Round-trip split/join on a 3-allele site with known AD and PL; indexes must match the new ALT list.
The number in GT is a pointer, not a copy number.
Curated: · Written: · Reviewed:
QA-25Two VCFs disagree on a homopolymer indel by 6 bp. Why might both be the same haplotype, and how do you compare them?(show answer)
The first thing I would pin down about left-aligning indels is which claim the number is actually allowed to support.
Equivalent indels in repeats can have different textual POS and alleles until they are left-aligned and trimmed against the exact reference.
Concretely, normalize with the pinned FASTA, then compare with a haplotype-aware tool such as vcfeval, not string equality.
The reason for that specificity is a failure I have seen: A truth-set evaluation counted 1,104 homopolymer indels as false because POS differed by 1–8 bp; precision looked 11 points worse than a normalized comparison.
Homopolymer AAAAAA, deletion of one A.
| Spelling | POS | Same haplotype? |
|---|---|---|
| left-aligned | 100 | yes |
| right-aligned | 105 | yes |
| string match | fail | wrong metric |
I would not consider it settled without evidence: Normalize both files, then vcfeval inside confident regions; quote representation-matched counts.
The letters in the VCF are a spelling of a haplotype.
Curated: · Written: · Reviewed:
QA-26A sample has no row at a site. How do you tell “called reference” from “no coverage,” and why does joint genotyping need the difference?(show answer)
I would start gVCF reference confidence from the samples, reference, and assay, not from the first plot that looked clean.
A gVCF reference-confidence block records that the sample was observed and looks reference-like. Absence of a row in a sites-only VCF is not a genotype.
Concretely, emit gVCFs with consistent intervals. Joint-genotype a defined cohort. Adding samples can change genotypes; the cohort identity is provenance.
The reason for that specificity is a failure I have seen: A rare-disease search treated missing rows as homozygous reference for 96 exomes; 14 compound-heterozygote candidates were hidden in uncovered exons.
Site with no VCF row.
| Evidence | Interpretation |
|---|---|
| gVCF block DP=32 | 0/0 OK |
| no row, DP=0 | no-call |
| treated as 0/0 | 14 missed |
I would not consider it settled without evidence: For every “0/0” used in a Mendelian analysis, show a gVCF block or DP supporting observation.
Not listed is not wild type.
Curated: · Written: · Reviewed:
QA-27You archived 40 TB as CRAM. What must travel with the files, and what fails if the FASTA moved?(show answer)
This is a place where a green QC report and a correct handling of CRAM reference recovery are not the same event.
CRAM may encode alignments relative to an external reference. Decoding needs that exact sequence, not “a GRCh38.”
Concretely, store reference URI and checksums, test decode without a local cache, retain indexes. A renamed FASTA with one contig patch is a different reference.
The reason for that specificity is a failure I have seen: A 2029 restore used a patched GRCh38 with 3 decoy changes; 1.1% of records failed decode and a 6-year study lost 800 WGS.
Archive bundle.
| Object | Required |
|---|---|
| CRAM+index | yes |
| FASTA sha256 | yes |
| “GRCh38” label | not enough |
| decode drill | yearly |
I would not consider it settled without evidence: A restore drill that decodes a random 1% of files against the checksummed FASTA only.
CRAM without its reference is compressed mystery.
Curated: · Written: · Reviewed:
QA-28A base is Q40 and the read is MAPQ 0. What are you allowed to believe, and how should a caller treat it?(show answer)
My answer to MAPQ versus base quality begins at the confounder: if I cannot name what would fake the signal, I do not have a design.
Base quality is confidence in the residue. Mapping quality is confidence in placement. Repeats can have beautiful bases in the wrong locus.
Concretely, do not convert MAPQ 0 into a confident SNP. Stratify calls by mappability. Understand aligner-specific MAPQ caps.
The reason for that specificity is a failure I have seen: A paralog SNP was reported at Q40/MAPQ 0 as a pathogenic het; the read belonged to a 99% identical gene 8 Mb away.
One pileup.
| Field | Value | Means |
|---|---|---|
| BQ | 40 | base OK |
| MAPQ | 0 | place unknown |
| called het | yes | error |
I would not consider it settled without evidence: A mappability track overlaid on called SNPs; MAPQ 0 pileups are not genotypes.
A confident letter can still be in the wrong sentence.
Curated: · Written: · Reviewed:
QA-29Why is HaplotypeCaller’s diploid model the wrong default for a 30% purity tumor, and what must the model include instead?(show answer)
I would treat diploid caller on tumor as an inference under a model, not as a file the pipeline emitted.
Germline diploid calling assumes inherited copy-2 genotypes. Tumors mix purity, subclones, and copy number, so allele fraction is not 0/0.5/1.
Concretely, use a somatic caller with matched normal or a tumor model, panel of normals, and copy-number context. Do not lower germline GQ thresholds as a somatic method.
The reason for that specificity is a failure I have seen: A 30% purity tumor run through diploid HC missed 41 subclonal SNVs at 8% AF that a somatic caller recovered at 80×.
Variant at 8% AF, 80×.
| Caller | Purity 30% | Recovered? |
|---|---|---|
| diploid HC | AF≠0.5 | no |
| somatic | mixture model | yes |
| 41 SNVs | missed |
I would not consider it settled without evidence: Sensitivity versus AF and purity on a mix-in truth set; diploid GQ is not the x-axis.
Allele fraction is the tumor’s genotype space.
Curated: · Written: · Reviewed:
QA-30Your caller’s genome-wide F1 is 0.999. Why might it still be unsafe in a repeat, and where are you allowed to count FP and FN?(show answer)
The useful question for GIAB confident regions is what still holds after a sample swap, a batch, or a database update.
Headline precision and recall are valid only where the benchmark asserts truth. Outside confident regions, unlabeled sites are not automatic false positives.
Concretely, intersect callable and GIAB confident beds. Stratify by type, size, and difficult context. Match sample and build.
The reason for that specificity is a failure I have seen: A marketing F1 of 0.999 excluded all GIAB “difficult” beds where the product would be used for MHC typing; in-region F1 was 0.81.
Two F1 numbers.
| Region | F1 |
|---|---|
| GIAB confident | 0.999 |
| MHC | 0.81 |
| advertised | 0.999 |
I would not consider it settled without evidence: Publish stratified F1 with counts, not a single genome-wide number.
You can only score what the truth set claimed.
Curated: · Written: · Reviewed:
QA-31You have 12 WGS. Why might VQSR be the wrong filter, and what do you use instead?(show answer)
I would settle VQSR on a small cohort against a negative control first, so the clever method has to survive it.
VQSR needs enough variants, representative annotations, and training resources matched to the assay. Twelve WGS can still fail that match even if SNP counts look large.
Concretely, inspect annotation distributions and truth-set recall. Do not copy GATK 1000 Genomes tranches onto a shifted assay without validation.
The reason for that specificity is a failure I have seen: VQSR on 12 WGS from a new kit placed 40% of true singleton SNPs below the 99.9 tranche because the annotation model did not match the assay; a hard-filter set recovered them.
Singletons, n=12.
| Filter | Singleton recall |
|---|---|
| VQSR 99.9 | 0.58 |
| hard filters | 0.91 |
| copied tranche | wrong N |
I would not consider it settled without evidence: Truth-set recall of singletons under VQSR versus predeclared hard filters on this assay.
A copied tranche is not a model of this chemistry.
Curated: · Written: · Reviewed:
QA-32A VCF BND looks like A[chr2:10[. What information lives in the brackets, and what does interval-only parsing lose?(show answer)
The judgement in breakend orientation is which identifiers and versions are pinned, not which tool name is fashionable.
Breakend notation encodes oriented adjacency, mates, and events. A pair of intervals without orientation is not the rearrangement.
Concretely, parse brackets, mate IDs, EVENT, and confidence intervals. Keep imprecise positions. Do not collapse BNDs into a DEL by taking min/max POS.
The reason for that specificity is a failure I have seen: A min/max parser turned a translocation into a 40 Mb “deletion” on chr1 and dropped the chr2 partner; a diagnostic report listed a huge CNV that did not exist.
Balanced t(1;2).
| Parser | Output |
|---|---|
| BND grammar | two oriented joins |
| min/max POS | 40 Mb DEL |
| partner chr2 | lost |
I would not consider it settled without evidence: A fixture with a balanced translocation; both breakends and orientations must survive ingest.
Rearrangements are oriented joins, not bounding boxes.
Curated: · Written: · Reviewed:
QA-33You sequenced the same library three times. Why does that not let you claim a population expression difference?(show answer)
Where candidates lose the interview on biological versus technical RNA-seq replicates is usually treating a hit list as biology.
Biological replicates are independent units from the target population. Technical replicates measure a narrower process and do not replace them.
Concretely, define the experimental unit (animal, donor, culture dish). Allocate budget to independent units once genes are adequately observed. Do not treat lanes of one library as n=3.
The reason for that specificity is a failure I have seen: A paper reported n=3 at FDR 0.05 from three lanes of one treated mouse versus one control mouse; 600 “DE genes” were lane noise.
Two designs, both labeled n=3.
| What was repeated | Population n |
|---|---|
| 3 mice | 3 |
| 3 lanes | 1 |
| DE genes | 600 false |
I would not consider it settled without evidence: The sample sheet’s experimental unit column must match the design matrix rows.
Repeated measurement of one mouse is still one mouse.
Curated: · Written: · Reviewed:
QA-34Every treated sample is on plate A and every control on plate B. Can DESeq2 unconfound treatment, and what do you do instead?(show answer)
I would answer RNA-seq batch confounding by separating what the experiment measured from what the software merely displayed.
If condition and batch are the same design column, their effects are not identifiable. No likelihood can split them.
Concretely, build the design matrix before sequencing. Cross-tab condition versus plate. If already confounded, state non-identifiability; do not publish a gene list.
The reason for that specificity is a failure I have seen: DESeq2 on 8 vs 8 fully confounded plates reported 1,842 FDR<0.05 genes; a later balanced rerun reproduced 61.
Plate layout.
| Plate | Treated | Control |
|---|---|---|
| A | 8 | 0 |
| B | 0 | 8 |
| identifiable? | no |
I would not consider it settled without evidence: A rank check of the design matrix; a singular batch+condition model must halt the pipeline.
Software cannot invent a contrast the experiment did not make.
Curated: · Written: · Reviewed:
QA-35Why are TPM values the wrong matrix for DESeq2, and which unit belongs in the likelihood?(show answer)
The engineering content of TPM as DESeq2 input is the evidence boundary and the exclusion log, not the p-value.
DESeq2’s negative-binomial likelihood expects integer sampling counts and estimates its own size factors. TPM already divided by length and library size and changed the mean-variance relationship.
Concretely, use raw fragment counts or tximport’s count-from-abundance for gene-level DE. Use TPM for descriptive within-sample abundance, labeled as such.
The reason for that specificity is a failure I have seen: Rounding TPM×10 to integers and feeding DESeq2 produced 400 “DE” genes driven by length, not condition; counts on the same samples recovered a 40-gene list that validated.
Same samples, two units.
| Input | FDR<0.05 genes |
|---|---|
| counts | 40 |
| rounded TPM | 400 |
| validated | 40 |
I would not consider it settled without evidence: A methods sentence that names the unit; tests refuse non-integer or length-normalized input for that path.
The likelihood has a unit; TPM is a different quantity.
Curated: · Written: · Reviewed:
QA-36One globin gene ate 40% of reads in the treated group. Why does dividing by total counts create fake DE elsewhere, and what estimator do you use?(show answer)
Before releasing I would write what a wrong result for composition bias in RNA-seq would look like in the same table.
Composition bias happens when a few genes consume the denominator. Median-of-ratios or TMM estimate size factors from a robust gene subset.
Concretely, inspect size-factor versus total-count plots. If biology is a global shift, say so and consider spike-ins. Do not assume total RNA is constant.
The reason for that specificity is a failure I have seen: Total-count normalization after globin surge called 900 repressed genes; median-of-ratios left 12, matching qPCR.
Globin 40% of treated reads.
| Normalization | “Repressed” genes |
|---|---|
| total counts | 900 |
| median-of-ratios | 12 |
| qPCR | 12 |
I would not consider it settled without evidence: Size-factor table next to the top abundant genes’ fraction of reads.
The denominator is a biological assumption.
Curated: · Written: · Reviewed:
QA-37A stranded library counted with the opposite convention lost 70% of assigned reads. What did strandedness decide, and how do you catch it?(show answer)
The first thing I would pin down about stranded RNA-seq counting is which claim the number is actually allowed to support.
Strandedness decides which overlapping antisense feature a fragment supports. The wrong convention discards or misassigns signal.
Concretely, record the protocol, infer strandedness from alignments as a check, and fail on disagreement. Do not copy a previous sample sheet.
The reason for that specificity is a failure I have seen: Running a reverse-stranded dUTP library with the forward-stranded featureCounts setting discarded or misassigned fragments across 1,200 antisense gene pairs.
dUTP library.
| Counter strand | Assigned | Wrong antisense |
|---|---|---|
| correct | 88% | 0 |
| opposite | 18% | 1200 genes |
| copied sheet | opposite | fail |
I would not consider it settled without evidence: A strandedness inference report versus metadata; mismatch is a hard fail.
Strand is part of the count, not a FASTQ comment.
Curated: · Written: · Reviewed:
QA-38A gene has adjusted p=0.04 at FDR 5%. Why is that not a 4% chance the gene is false, and what else belongs on the plot?(show answer)
I would start FDR is not a gene posterior from the samples, reference, and assay, not from the first plot that looked clean.
An FDR procedure bounds an expected false-discovery proportion in a set under assumptions. An adjusted p-value is not the posterior probability that this gene is a false positive.
Concretely, predeclare the tested family and FDR. Report effect, uncertainty, and normalized counts. Validate hits independently.
The reason for that specificity is a failure I have seen: A slide said “this gene is 96% true” from padj=0.04; replication failed, because the set-level 5% FDR said nothing of the kind.
One gene, padj=0.04.
| Claim | Valid? |
|---|---|
| list at FDR 5% | yes |
| P(false)=0.04 | no |
| replication | required |
I would not consider it settled without evidence: Figures show log2 fold change with interval and padj, not a posterior cartoon.
FDR polices a list, not a single gene’s soul.
Curated: · Written: · Reviewed:
QA-39A tumor RNA-seq study shows immune genes up. How do you tell regulation inside cells from more immune cells in the mix?(show answer)
This is a place where a green QC report and a correct handling of bulk tissue composition are not the same event.
Bulk counts combine per-cell expression with cell-type proportions. A tissue-level change is not automatically a within-cell regulatory event.
Concretely, inspect marker genes and pathology. Use deconvolution or matched single-cell when the claim is cellular. Phrase results as tissue-level unless origin is supported.
The reason for that specificity is a failure I have seen: A “IFN activation signature” looked 8-fold in bulk until cell fractions were measured; a 12%→24% T-cell increase with 1.3-fold within-cell induction predicts about 2.6-fold, so the rest was mixture, not regulation.
IFN signature in bulk.
| Driver | Fold |
|---|---|
| T-cell fraction | 12% → 24% (2.0) |
| within T-cell IFN | 1.3 |
| mixture prediction | 2.6 |
| uncorrected bulk | 8 |
I would not consider it settled without evidence: A composition table beside the signature; claims of regulation require a cellular assay.
More of a cell type is not the same as turning a gene on.
Curated: · Written: · Reviewed:
QA-40Why must count matrices keep Ensembl IDs with versions instead of joining on HGNC symbols?(show answer)
My answer to gene symbols as join keys begins at the confounder: if I cannot name what would fake the signal, I do not have a design.
Symbols are duplicated, renamed, and retired. Stable IDs with versions join to a specific gene model. Display names are labels.
Concretely, store ENSG with version and GTF release. Attach symbols as attributes. Define merge and retired-ID policy.
The reason for that specificity is a failure I have seen: A join on symbols merged C1orf106 with a retired alias and dropped 37 IDs that shared a display name across Ensembl 110 and 113.
Two releases, one symbol.
| Release | Symbol | ENSG |
|---|---|---|
| 110 | C1orf106 | ENSG000001000.12 |
| 113 | IL12 | different ID |
| join on symbol | 37 rows lost |
I would not consider it settled without evidence: Joins are on versioned IDs; a CI test feeds a known alias collision.
A pretty name is not a primary key.
Curated: · Written: · Reviewed:
QA-41Pre/post samples from 20 donors were analyzed as 40 independent libraries. What did that do to the variance, and how should the design look?(show answer)
I would treat paired RNA-seq subject blocking as an inference under a model, not as a file the pipeline emitted.
Paired designs estimate within-subject change. Ignoring the subject term treats correlated observations as independent and misstates uncertainty.
Concretely, include a subject blocking factor and a condition coefficient. Confirm each subject has the required observations. Do not drop incomplete pairs silently.
The reason for that specificity is a failure I have seen: Unpaired DESeq2 on 20 pairs reported 310 DE genes; a subject-blocked model left 44, and 9 of the 310 were donor main effects.
20 donors, pre/post.
| Model | DE genes |
|---|---|
| ~ condition | 310 |
| ~ subject + condition | 44 |
| donor effects in 310 | 9 |
I would not consider it settled without evidence: The design formula contains subject; a residual plot by donor is part of QC.
Two time points on one person are not two people.
Curated: · Written: · Reviewed:
QA-42A gene with 4 versus 0.3 expected counts shows a huge raw log2FC. Why shrink it, and when is a large ratio still real?(show answer)
The useful question for shrunken log fold change is what still holds after a sample swap, a batch, or a database update.
Low counts yield unstable ratios. Shrinkage trades bias for variance so ranking is not dominated by noise. It does not prove a gene is uninteresting.
Concretely, use method-compatible shrunken effects for ranking and volcano plots. Keep the unshrunken test for inference. Inspect raw counts for any highlighted gene.
The reason for that specificity is a failure I have seen: A ranking by raw log2FC put 80 of the top 100 genes below 10 counts; qPCR failed 71 of them. Shrunken ranking recovered a 12-gene list that held.
Gene with 4 vs 0.3 expected counts.
| Estimate | log2FC |
|---|---|
| raw MLE | 3.7 |
| shrunken | 1.1 |
| qPCR | 1.0 |
I would not consider it settled without evidence: A scatter of raw versus shrunken log2FC with count as size; the top list is the shrunken one.
A ratio of two tiny numbers is not a mechanism.
Curated: · Written: · Reviewed:
QA-43A GWAS of height is significant at thousands of SNPs before PCs and almost none after. What did ancestry do, and when can PCs overcorrect?(show answer)
I would settle GWAS population stratification against a negative control first, so the clever method has to survive it.
If allele frequency and the trait both track ancestry, unmodeled structure looks like association. PCs summarize major genotype axes; they can also capture batch or the trait itself.
Concretely, inspect PC-versus-geography and PC-versus-phenotype. Include justified PCs or a mixed model. Validate in family or stratified analyses.
The reason for that specificity is a failure I have seen: A height GWAS without PCs produced 4,200 loci at 5e-8 that tracked the first PC along a north-south recruitment cline; after 10 PCs, 28 loci remained and 22 replicated.
Height, n=80k.
| Model | Loci at 5e-8 |
|---|---|
| no PCs | 4200 |
| 10 PCs | 28 |
| replicated | 22 |
I would not consider it settled without evidence: A QQ plot and PC-trait table before and after correction; genomic inflation is interpreted with LD-score intercept when assumptions fit.
A cline is not a gene.
Curated: · Written: · Reviewed:
QA-44A cohort includes 1,200 sibling pairs. Why are ordinary logistic SEs too small, and what are two legal remedies?(show answer)
The judgement in relatedness and GWAS standard errors is which identifiers and versions are pinned, not which tool name is fashionable.
Relatives share genotypes and often phenotypes, so independent-error regression overstates information.
Concretely, estimate kinship, then use an unrelated subset or a mixed model that matches the design. Document cryptic relatedness too.
The reason for that specificity is a failure I have seen: Ignoring 1,200 sib pairs inflated chi-square by ~1.18 and added 60 loci that vanished in a leave-one-sib mixed model.
n=10,000 with 1,200 sib pairs.
| Analysis | lambda | Extra loci |
|---|---|---|
| ignore relatedness | 1.18 | 60 |
| mixed model | 1.02 | 0 |
| drop one sib | 1.03 | 0 |
I would not consider it settled without evidence: A kinship histogram and a comparison of mixed-model versus unrelated-subset top hits.
Shared parents are not two independent genomes.
Curated: · Written: · Reviewed:
QA-45A SNP has p=8e-7. Why is that not genome-wide evidence in a common-variant European GWAS, and when would the threshold change?(show answer)
Where candidates lose the interview on genome-wide significance threshold is usually treating a hit list as biology.
Millions of correlated tests make p=0.01 common. 5e-8 is a pragmatic common-variant human standard, not a law for every ancestry, sequencing study, or variant set.
Concretely, predeclare the family (imputed common SNPs vs rare burden). Use permutation or a study-appropriate threshold. Report effect and SE for every highlighted SNP.
The reason for that specificity is a failure I have seen: A sequencing study of 2 million rare variants used 5e-8 and published 14 “loci” that a Bonferroni on independent tests would have left at 0.
p=8e-7.
| Search | Pass? |
|---|---|
| common GWAS 5e-8 | no |
| 20 predeclared SNPs | maybe |
| 2e6 rare tests | no |
I would not consider it settled without evidence: The methods name the tested family and the threshold’s justification.
The threshold is a function of the search, not a lucky number.
Curated: · Written: · Reviewed:
QA-46The lead SNP is in an intron of gene X. Why might gene X still be the wrong story, and what does fine-mapping actually return?(show answer)
I would answer lead SNP versus causal allele by separating what the experiment measured from what the software merely displayed.
LD means the most significant marker may only tag a causal allele. Fine-mapping returns posterior support under LD and causal-number assumptions, often a set, not a star.
Concretely, use ancestry-matched LD, conditional analysis, and functional data. Keep statistically indistinguishable candidates.
The reason for that specificity is a failure I have seen: A press release named gene X from a lead intron SNP; the 95% credible set of 18 SNPs included an enhancer 80 kb away that later eQTL mapped to gene Y.
Credible set of 18.
| Object | Count |
|---|---|
| lead SNP in X intron | 1 |
| 95% set | 18 |
| later gene Y eQTL | 1 in set |
I would not consider it settled without evidence: Publish the credible set, not only the lead rsid.
The smallest p-value is a tag, not a mechanism.
Curated: · Written: · Reviewed:
QA-47A meta-analysis beta is near zero after combining two cohorts with large opposite effects. What must you have aligned, and which strand class is the usual villain?(show answer)
The engineering content of effect allele and sign flips is the evidence boundary and the exclusion log, not the p-value.
Beta and odds ratio refer to a named effect allele. Unaligned REF/ALT or strand-ambiguous A/T and C/G variants cancel or fabricate effects.
Concretely, harmonize build, alleles, frequency, and units. Resolve A/T and C/G with frequency or reference, not a coin flip.
The reason for that specificity is a failure I have seen: An A/T SNP was flipped in 4 of 9 cohorts; the pooled beta was 0.01 while every cohort’s |beta| exceeded 0.12.
Nine cohorts, one A/T SNP.
| Set | beta |
|---|---|
| 5 aligned | +0.14 |
| 4 flipped | -0.13 |
| naive pool | +0.01 |
I would not consider it settled without evidence: Allele-frequency correlation plots per cohort against a reference; outliers are flips.
A sign is an allele, not a vibe.
Curated: · Written: · Reviewed:
QA-48An imputed SNP has INFO 0.4. Why is a hard 0/1/2 genotype the wrong input, and how should association use it?(show answer)
Before releasing I would write what a wrong result for imputation dosage versus hard calls would look like in the same table.
Imputation is a dosage with uncertainty. Hard-calling throws the uncertainty away and is worst at low frequency and low INFO.
Concretely, associate on expected dosage. Filter with frequency-aware INFO. Match the reference panel to ancestry.
The reason for that specificity is a failure I have seen: Hard-calling INFO 0.4 rare SNPs turned uncertain dosages into overconfident 0/1/2 genotypes; a recessive model that needed homozygotes lit up 40 false loci.
INFO 0.4 rare SNP.
| Input | What it is |
|---|---|
| dosage | expected copies with uncertainty |
| hard 0/1/2 | overconfident call |
| false recessive loci | 40 |
I would not consider it settled without evidence: A scatter of dosage versus hard call by INFO bin; recessive tests require INFO floors and genotype probabilities, not hard calls.
A fractional dosage is the observation.
Curated: · Written: · Reviewed:
QA-49A European-trained PRS is applied to an African-ancestry clinic. What must you show besides the discovery AUC, and why can it harm?(show answer)
The first thing I would pin down about PRS ancestry transfer is which claim the number is actually allowed to support.
LD, allele frequency, environment, phenotype, and training representation differ. Discrimination and calibration transfer separately. A percentile is not a diagnosis.
Concretely, validate AUC, calibration, and decision utility in each intended population. Do not use race as a proxy for genetic ancestry or environment.
The reason for that specificity is a failure I have seen: A coronary PRS at the 90th percentile was 3.2-fold in Europeans and 1.4-fold with miscalibrated absolute risk in 4,100 African-ancestry patients; high-risk labels over-referred 800 people.
90th percentile label.
| Population | Relative risk | Calibration |
|---|---|---|
| EUR discovery | 3.2 | OK |
| AFR clinic n=4100 | 1.4 | off |
| extra referrals | 800 |
I would not consider it settled without evidence: Calibration plots per population in the model card; no plot, no deployment.
Weights are not portable just because DNA is DNA.
Curated: · Written: · Reviewed:
QA-50A SNP fails HWE at p=1e-12 in cases but not controls. Do you drop it, and what else can HWE mean?(show answer)
I would start Hardy-Weinberg as a delete key from the samples, reference, and assay, not from the first plot that looked clean.
HWE deviation can be genotyping error, structure, inbreeding, or true association. It is a diagnostic, not a universal delete.
Concretely, test appropriate unrelated ancestry groups, often controls separately. Inspect clusters. Use sample-size-aware thresholds.
The reason for that specificity is a failure I have seen: Dropping every case-only HWE failure removed the true HLA signal in an autoimmune GWAS (p=4e-18) because cases really were out of HWE at a causal locus.
Autoimmune GWAS.
| Filter | HLA locus |
|---|---|
| drop all HWE 1e-12 | gone |
| controls-only HWE | kept |
| true p | 4e-18 |
I would not consider it settled without evidence: Cluster plots for HWE failures; case-only failures are reviewed, not bulk-deleted.
A tiny HWE p-value is a question, not an answer.
Curated: · Written: · Reviewed:
QA-51lambda_GC is 1.15. How do you tell polygenicity from confounding with LD-score regression, and when is the tool the wrong one?(show answer)
This is a place where a green QC report and a correct handling of LD-score intercept are not the same event.
LD-score regression asks whether test statistics rise with LD score as polygenicity predicts. The intercept can flag confounding. It needs matched LD and well-calibrated summary stats.
Concretely, harmonize summaries, use ancestry-matched LD scores, inspect intercept and attenuation. Do not run LDSC on a tiny or strongly selected study as gospel.
The reason for that specificity is a failure I have seen: An intercept of 1.18 was called “polygenic” without LDSC; LDSC intercept 1.16 and slope near 0 showed batch confounding, not heritability.
lambda 1.15.
| Tool | Intercept | Meaning |
|---|---|---|
| lambda only | ambiguous | |
| LDSC | 1.16 slope~0 | confounding |
| LDSC | 1.02 slope>0 | polygenic |
I would not consider it settled without evidence: Report intercept, slope, and LD reference; lambda alone is not a verdict.
Inflation has more than one parent.
Curated: · Written: · Reviewed:
QA-52A publication figure is a fully resolved tree with 100% support on every node. What must still be in the methods, and what can the picture hide?(show answer)
My answer to tree as a model not a picture begins at the confounder: if I cannot name what would fake the signal, I do not have a design.
A phylogeny is an inference under alignment, model, and search. A clean graphic can hide alternative equally scoring histories.
Concretely, publish sequences, alignment, model, runs, seeds, and support trees. Report discordance. Do not treat a consensus as the only history.
The reason for that specificity is a failure I have seen: A 100% UFBoot tree grouped a contaminant 16S with a pathogen; removing one chimeric read moved the clade 9 nodes.
One chimeric read.
| Dataset | Topology |
|---|---|
| all reads | pathogen+contaminant |
| drop chimera | contaminant basal |
| figure | 100% support |
I would not consider it settled without evidence: The supplement contains the alignment checksum and at least two independent runs that agree.
A drawing is not a posterior.
Curated: · Written: · Reviewed:
QA-53Two distant parasites group together with high support. Why might that be long-branch attraction, and which test do you run?(show answer)
I would treat long-branch attraction as an inference under a model, not as a file the pipeline emitted.
Long-branch attraction is a systematic error where rapidly evolving lineages group regardless of history, especially under overly simple models or parsimony.
Concretely, try better-fitting models, more taxa that break branches, and composition tests. High support does not rule LBA out.
The reason for that specificity is a failure I have seen: A mitochondrial tree put two long-branched insects as sisters at 99% bootstrap under JC69; a partitioned GTR+G tree with 12 extra taxa split them by 40 million years of accepted taxonomy.
Two long branches.
| Model | Together? | Support |
|---|---|---|
| JC69 | yes | 99% |
| GTR+G +12 taxa | no | 87% split |
| taxonomy | split |
I would not consider it settled without evidence: A table of topology under JC versus a mixture model plus taxon addition.
Support can be confidence in the wrong model.
Curated: · Written: · Reviewed:
QA-54UFBoot is 95% and the Felsenstein bootstrap is 62% on the same node. Which number do you quote, and why do they differ?(show answer)
The useful question for ultrafast versus Felsenstein bootstrap is what still holds after a sample swap, a batch, or a database update.
Ultrafast bootstrap is a different, faster approximation with its own bias correction. Felsenstein’s nonparametric bootstrap resamples characters. They are not interchangeable percentages.
Concretely, report which support method and version. Do not convert UFBoot 95 into “Felsenstein 95.” Use UFBoot2 correction if that is the software.
The reason for that specificity is a failure I have seen: A paper quoted “bootstrap 95%” from UFBoot without saying so; a reviewer rerun with 1,000 Felsenstein replicates got 58% and the clade was dropped.
Same node.
| Method | Support |
|---|---|
| UFBoot2 | 95% |
| Felsenstein 1000 | 58% |
| quoted as “bootstrap” | misleading |
I would not consider it settled without evidence: Methods name UFBoot versus Felsenstein and the number of replicates.
Percent support is a method, not a universal unit.
Curated: · Written: · Reviewed:
QA-55Gene trees disagree with the species tree. How do you tell ILS from introgression, and why is a concatenated tree the wrong first move?(show answer)
I would settle incomplete lineage sorting versus introgression against a negative control first, so the clever method has to survive it.
ILS is random gene-tree discordance from ancestral polymorphism under the multispecies coalescent. Introgression is a historical gene flow event. Concatenation can yield a resolved wrong tree.
Concretely, use coalescent methods, check branch lengths and D-statistics or related tests, and do not force one concatenated topology as truth.
The reason for that specificity is a failure I have seen: A 400-gene concatenation produced a 100% bootstrap species tree that contradicted 70% of gene trees; an MSC analysis matched the gene-tree majority and D-statistics found no evidence of introgression at the study’s power.
400 genes.
| Method | Topology | Concordance |
|---|---|---|
| concat | A | 100% BS |
| MSC | B | 70% genes |
| D-stat introgression | not detected |
I would not consider it settled without evidence: A gene-tree concordance factor beside the species tree; concatenation-only is not enough.
Disagreement is data, not noise to concatenate away.
Curated: · Written: · Reviewed:
QA-56A viral alignment has detectable recombination. Why is one phylogeny the wrong object, and what do you estimate instead?(show answer)
The judgement in recombination and a single tree is which identifiers and versions are pinned, not which tool name is fashionable.
Recombination means different regions have different histories. A single tree averages incompatible signals and can invent branches.
Concretely, detect breakpoints, infer trees per non-recombinant region, or use explicit network methods. Do not “build a tree anyway” for a recombining alignment.
The reason for that specificity is a failure I have seen: A whole-genome HIV tree placed a recombinant as a new subtype with 100% support; region-specific trees showed it was a mosaic of two known subtypes.
HIV mosaic.
| Region | Parent |
|---|---|
| 1-2000 | subtype A |
| 2001-end | subtype B |
| whole-genome tree | “new subtype” |
I would not consider it settled without evidence: A recombination screen before tree search; if positive, no single-tree claim.
One tree is one history; mosaics have several.
Curated: · Written: · Reviewed:
QA-57Trimming gappy columns raised bootstrap from 50% to 99%. When is that cleaner signal, and when did you just delete the disagreement?(show answer)
Where candidates lose the interview on alignment trimming bias is usually treating a hit list as biology.
Trimming can remove noise or remove the characters that supported an alternative topology. Choosing columns to favor a tree is circular.
Concretely, predeclare a trimming rule. Compare trees on trimmed and untrimmed alignments. Do not trim until the tree looks like the hypothesis.
The reason for that specificity is a failure I have seen: A trim that dropped 40 gappy columns moved a parasite clade 11 nodes and created 99% support for a host grouping the authors preferred.
40 columns removed.
| Alignment | Topology | UFBoot |
|---|---|---|
| full | parasites together | 51% |
| trimmed | with host | 99% |
| predeclared? | no |
I would not consider it settled without evidence: Both alignments and both trees in the supplement; a predeclared trim setting.
Deleted columns cannot vote against you.
Curated: · Written: · Reviewed:
QA-58Two domains each have pLDDT 90. Why might their relative orientation still be a guess, and which AlphaFold output answers that?(show answer)
I would answer pLDDT is not domain orientation by separating what the experiment measured from what the software merely displayed.
pLDDT is a local per-residue confidence. Predicted aligned error (PAE) describes confidence in relative positions. High local score does not fix a flexible linker.
Concretely, use PAE for inter-domain claims. Do not dock a ligand into a high-pLDDT domain using an untrusted orientation to the rest.
The reason for that specificity is a failure I have seen: A docking paper used residues 12–340 at pLDDT 91 against a second domain with PAE 28 Å; the interface was an orientation guess.
Two domains, linker of 40 aa.
| Metric | Value | Allows interface? |
|---|---|---|
| pLDDT each | 91 | local only |
| PAE | 28 Å | no |
| docked anyway | paper | wrong |
I would not consider it settled without evidence: PAE plot quoted for any inter-domain distance; pLDDT only for local backbone.
Colorful domains can still float.
Curated: · Written: · Reviewed:
QA-59You need a ligand-bound conformation. Why is a default AlphaFold monomer the wrong structure, and what would you use instead?(show answer)
The engineering content of AlphaFold is not a holo complex is the evidence boundary and the exclusion log, not the p-value.
A default AF2 monomer omits the ligand and does not establish which conformational state is represented. Partners and alternative states may be absent or wrong.
Concretely, prefer experimental holo structures, or methods that condition on partners, and validate with density or biochemistry. Do not assume the pocket is the drug pose.
The reason for that specificity is a failure I have seen: A virtual screen against AF2 apo ranked 200 compounds that did not bind; the crystal holo pocket differed by 4.2 Å RMSD at the gate.
Kinase pocket.
| Model | Pocket RMSD to holo |
|---|---|
| AF2 monomer | 4.2 Å |
| PDB holo | 0 |
| screen hits | 0 of 200 |
I would not consider it settled without evidence: Pocket RMSD to any available holo PDB before screening; AF2-only is a risk flag.
A predicted chain is not a binding assay.
Curated: · Written: · Reviewed:
QA-60A paper measures an interface on chain A–B from the PDB file as downloaded. Why might that interface not exist in biology?(show answer)
Before releasing I would write what a wrong result for biological assembly versus asymmetric unit would look like in the same table.
The asymmetric unit is a crystallographic convenience. The biological assembly may be a different set of chains. Crystal contacts are not always interfaces.
Concretely, load the author or software assembly from mmCIF, not the first two chains in the PDB. Check assembly evidence.
The reason for that specificity is a failure I have seen: An “interface” of 1,800 Ų was a crystal contact; mutating it did nothing in solution because the biological unit was a different dimer.
PDB entry.
| Object | Chains | Interface real? |
|---|---|---|
| ASU | A,B | crystal contact |
| assembly 1 | A,A' | yes |
| mutation | no effect |
I would not consider it settled without evidence: The methods name assembly ID and how it was chosen.
The file’s first chains are not the molecule.
Curated: · Written: · Reviewed:
QA-61Two models have RMSD 8 Å and TM-score 0.82. Which one says they share a fold, and why can RMSD mislead?(show answer)
The first thing I would pin down about TM-score versus RMSD is which claim the number is actually allowed to support.
RMSD is dominated by the worst-aligned regions and by length. TM-score is length-normalized and more sensitive to global topology. A high RMSD can still be the same fold.
Concretely, quote TM-score for fold claims and local RMSD for a site. Do not average RMSD over unaligned tails.
The reason for that specificity is a failure I have seen: A reviewer rejected a model at 9 Å RMSD that had TM-score 0.81; the tails were disordered and the core superimposed at 1.8 Å.
Same pair.
| Metric | Value | Fold? |
|---|---|---|
| all-atom RMSD | 9 Å | unclear |
| TM-score | 0.81 | yes |
| core RMSD | 1.8 Å | yes |
I would not consider it settled without evidence: Report TM-score, aligned length, and core RMSD together.
One bad tail can drown a fold in RMSD.
Curated: · Written: · Reviewed:
QA-62When must you prefer a 3.5 Å crystal over a pLDDT-92 AlphaFold model for a catalytic residue claim?(show answer)
I would start experimental versus predicted coordinates from the samples, reference, and assay, not from the first plot that looked clean.
Experimental models carry density or restraints for that crystal or sample. Predictions encode learned regularities and can be locally wrong at functional sites even when globally pretty.
Concretely, for chemistry at a site, use experimental density and local validation. Use AF2 when no structure exists, with local pLDDT/PAE stated.
The reason for that specificity is a failure I have seen: A mechanism paper used AF2 to place a catalytic Ser that the 3.5 Å map showed pointing the other way; the mutation table matched the crystal, not AF2.
Catalytic Ser.
| Source | Orientation | Matches mutants? |
|---|---|---|
| AF2 pLDDT 92 | away | no |
| 3.5 Å map | toward substrate | yes |
I would not consider it settled without evidence: If a PDB exists for the construct, the claim cites local density, not only AF2.
A high pLDDT does not outrank density at the same residue.
Curated: · Written: · Reviewed:
QA-63A model is in the Ramachandran favored region but fits density poorly. Which check is missing, and why is geometry not enough?(show answer)
This is a place where a green QC report and a correct handling of Ramachandran versus density are not the same event.
Geometry can look ideal while the chain is registered wrong or over-refined. Density or restraint fit is independent evidence.
Concretely, read MolProbity and map-model metrics for the region used. Do not gate release on Ramachandran alone.
The reason for that specificity is a failure I have seen: A one-residue register error sat in favored Ramachandran and failed local map correlation 0.41; the “binding motif” was shifted by 3.8 Å.
Register error.
| Check | Value |
|---|---|
| Ramachandran | favored |
| local map CC | 0.41 |
| shift | 3.8 Å |
I would not consider it settled without evidence: Local map CC or equivalent beside geometry for any site-level claim.
A pretty backbone can still be the wrong sequence placement.
Curated: · Written: · Reviewed:
QA-64A classifier of disease from single-cell RNA-seq hits AUC 0.97 with random cells in train and test. Why is that leakage, and what is the split unit?(show answer)
My answer to grouped splits by donor begins at the confounder: if I cannot name what would fake the signal, I do not have a design.
Cells from one donor are not independent patients. Random cell splits put the same person on both sides.
Concretely, split by donor, family, site, or time to match deployment. Freeze the grouping before features.
The reason for that specificity is a failure I have seen: Random cell splits gave AUC 0.97; a donor-grouped split on the same 48 people dropped to 0.61.
48 donors, 200k cells.
| Split | AUC |
|---|---|
| random cells | 0.97 |
| by donor | 0.61 |
| leaked people | 48 |
I would not consider it settled without evidence: A split table with zero donor overlap; AUC is reported on that split.
The independence unit is the person, not the barcode.
Curated: · Written: · Reviewed:
QA-65A protein function model is tested on random 20% of UniProt. Why can that still leak, and how do you split?(show answer)
I would treat homolog leakage in sequence models as an inference under a model, not as a file the pipeline emitted.
Homologous sequences share labels. A random row split puts family members in train and test.
Concretely, cluster at a sequence-identity threshold or by family, split clusters, and report the threshold. Chromosome or locus grouping for genomic models.
The reason for that specificity is a failure I have seen: Random UniProt split AUC 0.94 fell to 0.71 when clusters at 30% identity were held out.
Function prediction.
| Split | AUC |
|---|---|
| random rows | 0.94 |
| 30% identity clusters | 0.71 |
| claimed generalization | overstated |
I would not consider it settled without evidence: Test proteins have max identity to train below the stated threshold.
A random row is not a random protein family.
Curated: · Written: · Reviewed:
QA-66You tuned 40 hyperparameters on the same 5 folds you report. Why is that number optimistic, and where does the test set live?(show answer)
The useful question for nested cross-validation is what still holds after a sample swap, a batch, or a database update.
Tuning on the evaluation folds uses the test performance as training feedback. Nested CV or a frozen external test is required.
Concretely, inner loop tunes, outer loop evaluates. Touch the external test once.
The reason for that specificity is a failure I have seen: A 40-parameter search on the reporting folds quoted 0.88; nested CV was 0.74, matching a later prospective set.
Same 200 samples.
| Protocol | AUC |
|---|---|
| tune+report same 5 folds | 0.88 |
| nested | 0.74 |
| prospective | 0.73 |
I would not consider it settled without evidence: The methods diagram shows inner versus outer folds; the headline metric is outer.
A fold you tuned on is not a test.
Curated: · Written: · Reviewed:
QA-67AUC is 0.91 but at the 0.5 threshold the positive predictive value is 0.11. What did AUC not tell you?(show answer)
I would settle calibration is not AUC against a negative control first, so the clever method has to survive it.
AUC measures ranking. Deployment needs calibration and a threshold tied to prevalence and cost. A rare event can rank well and still flood false positives.
Concretely, show a calibration curve and counts at the operating point. Choose the threshold on a validation set with the real prevalence.
The reason for that specificity is a failure I have seen: A 1% prevalence screen with AUC 0.91 at 0.5 threshold yielded PPV 0.11 and 8,900 extra follow-ups per year.
Prevalence 1%.
| Metric | Value |
|---|---|
| AUC | 0.91 |
| PPV at 0.5 | 0.11 |
| extra follow-ups | 8900/year |
I would not consider it settled without evidence: A reliability plot and a confusion matrix at the chosen threshold, not only AUC.
Ranking well is not being right at the cut.
Curated: · Written: · Reviewed:
QA-68A variant-effect CNN trained on chr1–22 and tested on random SNPs from those chromosomes still leaks. What split actually tests generalization?(show answer)
The judgement in chromosome-held-out sequence CNNs is which identifiers and versions are pinned, not which tool name is fashionable.
Overlapping windows and LD mean random SNPs from the training chromosomes are not independent. A held-out chromosome or locus block is the unit.
Concretely, hold out at least one chromosome or large locus. Do not tune on it. Report that chromosome’s metrics as the generalization number.
The reason for that specificity is a failure I have seen: Random SNP split AUC 0.94; chr8 holdout 0.71 and calibration slope 0.61.
Variant-effect CNN.
| Test | AUC | Slope |
|---|---|---|
| random SNPs | 0.94 | 0.98 |
| chr8 holdout | 0.71 | 0.61 |
I would not consider it settled without evidence: The test SNPs’ chromosome is disjoint from training windows.
A random SNP is still from a chromosome the model memorized.
Curated: · Written: · Reviewed:
QA-69A deep variant caller trained on NovaSeq is run on a new chemistry. What must you re-benchmark, and why is “the model is the same file” irrelevant?(show answer)
Where candidates lose the interview on caller trained on one chemistry is usually treating a hit list as biology.
Learned callers absorb platform error profiles. A new chemistry is a new domain. The artifact is not a guarantee.
Concretely, benchmark on GIAB for that chemistry and coverage. Recalibrate or abstain on unsupported platforms.
The reason for that specificity is a failure I have seen: A NovaSeq-trained model on a new kit dropped SNP F1 from 0.999 to 0.97 in homopolymers, adding 2,400 FP per genome.
GIAB SNP F1.
| Chemistry | F1 | FP/genome |
|---|---|---|
| NovaSeq train | 0.999 | 120 |
| new kit | 0.97 | 2400 |
I would not consider it settled without evidence: A per-chemistry GIAB table in the release notes.
Weights plus a new error process is a new method.
Curated: · Written: · Reviewed:
QA-70A model card lists AUC 0.93. What else must it say before a clinic may use it, and what is an excluded use?(show answer)
I would answer model card intended use by separating what the experiment measured from what the software merely displayed.
Intended use, excluded use, training population, operating threshold, subgroup metrics, and governance belong on the card. A score without a decision is not a product.
Concretely, tie the card to artifact hash and dataset versions. Disclose failure modes and update policy.
The reason for that specificity is a failure I have seen: A research AUC 0.93 model was used to deny follow-up in a population absent from training; error was 3× higher in that group.
Card excerpt.
| Field | Example |
|---|---|
| intended | triage research cohort EUR adults |
| excluded | denial of care |
| missing group error | 3× |
I would not consider it settled without evidence: The card’s intended-use sentence matches the deployment, or deployment is refused.
AUC is not a license.
Curated: · Written: · Reviewed:
QA-71Nextflow finished with exit 0. Why is that not a release, and what contract would have caught the silent 41-sample cache hit?(show answer)
The engineering content of pipeline exit code versus QC is the evidence boundary and the exclusion log, not the p-value.
Process success is not scientific validity. Each task needs an output contract: schema, identity, checksums, and domain QC. Cache keys are content hashes, not filenames.
Concretely, publish a run manifest of task hashes, input checksums, container digests, and QC. Refuse resume that cannot show re-execution after input change.
The reason for that specificity is a failure I have seen: A sample-sheet rename left 41 of 96 BAMs as cache hits on old FASTQ hashes; exit 0 in 11 minutes.
Resume after rename.
| BAMs | Cached | New FASTQ ran |
|---|---|---|
| 96 | 41 | 55 |
| exit | 0 | |
| release | no |
I would not consider it settled without evidence: Manifest lists executed versus cached tasks; input sha256 must match the sheet.
Zero is a process code, not a genome.
Curated: · Written: · Reviewed:
QA-72The workflow pins bwa:latest. Why is that not reproducible, and what do you pin instead?(show answer)
Before releasing I would write what a wrong result for container digest versus tag would look like in the same table.
Tags move. Digests identify content. latest is a moving pointer.
Concretely, pin image digest and platform. Scan licenses. Avoid privileged runs.
The reason for that specificity is a failure I have seen: bwa:latest moved over a weekend; a rerun aligned 2% fewer reads and shifted 180 indels with the same script.
Two Mondays.
| Pin | Reads aligned |
|---|---|
| bwa:latest week 1 | 99.1% |
| bwa:latest week 2 | 97.0% |
| sha256:abc | 99.1% both |
I would not consider it settled without evidence: The config contains sha256: and a pull of that digest in CI.
A name is not an image.
Curated: · Written: · Reviewed:
QA-73FASTQ headers say sample A and the sheet says sample B. What must win, and which QC is the last chance to catch a swap?(show answer)
The first thing I would pin down about sample-sheet identity is which claim the number is actually allowed to support.
Identity is a chain from tube to file to table. A sheet that disagrees with genetic sex, fingerprints, or headers is a stop.
Concretely, validate headers, barcodes, sex, and fingerprints against metadata before alignment is treated as truth.
The reason for that specificity is a failure I have seen: A swapped tumor-normal pair called 4,100 “somatic” variants that were germline in the other person.
Tumor-normal swap.
| Check | Result |
|---|---|
| sheet | T=A N=B |
| fingerprint | T matches B |
| somatic SNVs | 4100 false |
I would not consider it settled without evidence: A pre-alignment identity gate; mismatches fail the run.
The sheet is a claim until genetics agrees.
Curated: · Written: · Reviewed:
QA-74A reviewer says FAIR requires uploading the BAM to a public bucket. What does accessibility actually require?(show answer)
I would start FAIR does not mean public genomes from the samples, reference, and assay, not from the first plot that looked clean.
FAIR accessibility includes persistent identifiers and clear access conditions. It does not require unprotected human genomes.
Concretely, publish metadata and access procedures. Encrypt and audit. Controlled access can be FAIR.
The reason for that specificity is a failure I have seen: A “FAIR” public bucket exposed 120 WGS; accession numbers were enough to reidentify 8 people via a genealogy site.
Human WGS.
| Mode | FAIR? | Allowed? |
|---|---|---|
| public bucket | no | no |
| controlled accession | yes | yes |
| reidentified | 8 people |
I would not consider it settled without evidence: The data-use agreement and access log are part of the release, not a public ACL.
Findable is not the same as world-readable.
Curated: · Written: · Reviewed:
QA-75Task logs include FASTQ paths with patient names. Why is that a breach even if the BAM is encrypted, and what identifiers belong in logs?(show answer)
This is a place where a green QC report and a correct handling of work directory PHI are not the same event.
Operational metadata can reidentify. Filenames and logs are PHI if they carry names or MRNs.
Concretely, use run-scoped pseudonyms, private logs, least privilege, and verified cleanup of scratch.
The reason for that specificity is a failure I have seen: A shared HPC scratch left 14 days of paths like /scratch/JaneDoe_tumor.fq; 6 names indexed by the cluster search.
Scratch policy.
| Path | OK? |
|---|---|
| /work/run_9f3a/R1.fq | yes |
| /scratch/JaneDoe_tumor.fq | no |
| indexed names | 6 |
I would not consider it settled without evidence: A log redaction test; patient names must not appear in paths or stdout.
Encryption of the BAM does not hide the filename.
Curated: · Written: · Reviewed:
QA-76The largest 8 tumors failed out of memory and were dropped. What bias did that create, and how should resources be declared?(show answer)
My answer to OOM bias in cohorts begins at the confounder: if I cannot name what would fake the signal, I do not have a design.
Selective failure by size or complexity biases the surviving cohort. Resource declarations are hypotheses to measure and bound.
Concretely, scale memory with input size, retry with higher limits for known classes, and report who did not finish.
The reason for that specificity is a failure I have seen: Dropping 8 largest tumors removed 7 of the 8 whole-genome-duplication tumors from a “CNV landscape” paper.
96 tumors.
| Class | Finished | WGD |
|---|---|---|
| small 88 | 88 | 1 |
| large 8 OOM | 0 | 7 of 8 |
| published landscape | biased |
I would not consider it settled without evidence: A completion table by input size; incomplete samples are listed, not silent.
The genomes that finish are not a random sample.
Curated: · Written: · Reviewed:
QA-77P04637 and P04637-2 are not the same sequence. What must you store, and what happens if you join on accession without isoform?(show answer)
I would treat UniProt accession versions as an inference under a model, not as a file the pipeline emitted.
Accessions can be canonical plus isoforms and have sequence versions. Annotation changes without the accession changing.
Concretely, store accession, isoform, and sequence version. Revalidate mappings after a UniProt release.
The reason for that specificity is a failure I have seen: A phospho-site on residue 387 of isoform 2 was mapped onto canonical p53 and looked “conserved” in 40 species that lack that exon.
P53 site 387.
| Record | Length | Site exists? |
|---|---|---|
| P04637 canonical | 393 | no |
| P04637-2 | 429 | yes |
| join on P04637 | 40 false |
I would not consider it settled without evidence: Sequence checksum of the exact isoform used for coordinates.
The ID without isoform is a family, not a sequence.
Curated: · Written: · Reviewed:
QA-78A gene is annotated to a pathway with IEA. Why is that weaker than IDA, and how should enrichment treat electronic annotations?(show answer)
The useful question for GO evidence codes is what still holds after a sample swap, a batch, or a database update.
IEA is electronic, often transitive, and can be circular with the very sequences under study. IDA is experimental. Evidence codes are not optional color.
Concretely, filter or stratify by evidence. Do not treat IEA as wet-lab proof. Pin the GO release.
The reason for that specificity is a failure I have seen: An enrichment of “kinase activity” was 80% IEA from BLAST to the same family the analyst had just named; experimental codes showed no enrichment.
Kinase term.
| Evidence | Genes | Enrichment |
|---|---|---|
| all incl IEA | 400 | p=1e-12 |
| IDA/IMP only | 18 | ns |
| circular BLAST | yes |
I would not consider it settled without evidence: Enrichment tables split IEA versus experimental; the claim uses the experimental slice unless IEA is explicit.
A computed annotation is not an experiment.
Curated: · Written: · Reviewed:
QA-79Counts used Ensembl 110 GTF and variants used 113. What breaks, and how do you keep a coordinate system?(show answer)
I would settle mixing Ensembl releases against a negative control first, so the clever method has to survive it.
Genome and annotation releases are one coordinate system. Mixing silently misassigns reads and variants.
Concretely, pin FASTA+GTF from the same release. Lift with a documented chain if you must change.
The reason for that specificity is a failure I have seen: A gene-level RNA-seq to GWAS overlay used different releases; 6.2% of “eQTLs” sat in exons that did not exist in the RNA GTF.
Overlay.
| Data | Release |
|---|---|
| RNA GTF | 110 |
| VCF | 113 |
| fake eQTLs | 6.2% |
I would not consider it settled without evidence: A single release ID on counts, variants, and annotation in the manifest.
Two Ensembls are two maps.
Curated: · Written: · Reviewed:
QA-80A Python merge on the symbol column doubled 90 rows. What should the key have been, and how do you detect it?(show answer)
The judgement in pandas merge on gene symbols is which identifiers and versions are pinned, not which tool name is fashionable.
Many-to-many symbol merges explode rows. Versioned gene IDs are the key. validate in pandas is part of the contract.
Concretely, merge(..., validate="one_to_one") on ENSG. Count rows before and after. Do not merge on names.
The reason for that specificity is a failure I have seen: A symbol merge turned 18,402 genes into 19,110 rows and silently averaged duplicates; 90 genes were double-counted in a signature.
merge on symbol.
| Stage | Rows |
|---|---|
| left | 18402 |
| after merge | 19110 |
| validate 1:1 | would fail |
I would not consider it settled without evidence: Row-count assertions and validate= in the notebook that builds the matrix.
If the row count changed, the key was not unique.
Curated: · Written: · Reviewed:
QA-81Coverage 29.6× was stored as 30 and used as a gate of ≥30. What class of samples did that drop or keep wrongly?(show answer)
Where candidates lose the interview on floating coverage rounding is usually treating a hit list as biology.
Rounding before a threshold is a different rule. Depth is a distribution, not an integer headline.
Concretely, store floating mean and a percentile (e.g. 20× at 90% of bases). Apply gates on the stored quantity.
The reason for that specificity is a failure I have seen: Rounding 29.6 to 30 passed 14 samples whose 20× coverage was only 71% of bases; variant recall in those samples was 0.88 versus 0.97.
≥30× gate.
| Sample | Mean | Breadth 20× | Rounded pass? |
|---|---|---|---|
| 14 samples | 29.6 | 71% | yes |
| recall | 0.88 | should fail |
I would not consider it settled without evidence: The QC table keeps unrounded depth and a breadth metric.
The gate must see the number you measured.
Curated: · Written: · Reviewed:
QA-82A query counted rows in the call table as sample N. Why was N 8× too high, and what must DISTINCT apply to?(show answer)
I would answer SQL COUNT DISTINCT samples by separating what the experiment measured from what the software merely displayed.
One sample has many variant rows. COUNT(*) is not cohort size. COUNT(DISTINCT sample_id) is.
Concretely, aggregate at the grain of the question. Test with a known 10-sample fixture.
The reason for that specificity is a failure I have seen: A dashboard showed n=9,600 because it counted variant rows for 1,200 people; a rare-disease filter looked well powered and was not.
1,200 people, 8 variants each.
| Expression | Result |
|---|---|
| COUNT(*) | 9600 |
| COUNT(DISTINCT sample_id) | 1200 |
| power calc used | 9600 |
I would not consider it settled without evidence: A unit test that 10 samples × 100 variants yields n=10.
The grain of the table is not the grain of the study.
Curated: · Written: · Reviewed:
QA-83You need the top transcript per gene by TPM. Why does ORDER BY TPM LIMIT 1 per gene in a loop fail at 60k transcripts, and which SQL pattern is right?(show answer)
The engineering content of window functions for per-gene ranks is the evidence boundary and the exclusion log, not the p-value.
Window functions compute ranks in one pass. Per-gene loops and non-deterministic LIMIT without PARTITION are slow or wrong.
Concretely, rOW_NUMBER() OVER (PARTITION BY gene ORDER BY tpm DESC) filtered to 1. Ties need a documented tie-break.
The reason for that specificity is a failure I have seen: A Python loop picked an arbitrary transcript when TPM tied at 3 decimals; 400 genes flipped isoform between runs.
Two transcripts, TPM 12.0 both.
| Pick | Stable? |
|---|---|
| LIMIT 1 no ORDER extra | no |
| ROW_NUMBER + longer CDS | yes |
| genes flipped | 400 |
I would not consider it settled without evidence: A query with explicit tie-break (length, ID) and a test on a 3-way TPM tie.
A rank without a partition is a global sort in disguise.
Curated: · Written: · Reviewed:
QA-84In an interview you are asked to fill a 4×4 NW table for match +1 mismatch -1 gap -2. What does each cell mean, and which neighbor is illegal?(show answer)
Before releasing I would write what a wrong result for Needleman-Wunsch table filling would look like in the same table.
Each cell is the best score for prefixes ending at those indices. Transitions are diagonal substitution or a gap from the left or up cell. You cannot skip a cell.
Concretely, initialize first row/column with cumulative gaps. Fill by max of three. Traceback from the corner for global alignment.
The reason for that specificity is a failure I have seen: A candidate took max of the whole row, which is local-style, and reported a “global” alignment that dropped both termini.
AC vs AG, +1/-1/-2.
| - | A | G | |
|---|---|---|---|
| - | 0 | -2 | -4 |
| A | -2 | 1 | -1 |
| C | -4 | -1 | 0 |
I would not consider it settled without evidence: A filled toy table whose corner score matches an independent implementation.
The recurrence is the algorithm; the pretty path is the traceback.
Curated: · Written: · Reviewed:
QA-85Why does BLAST not run Smith-Waterman against every nr sequence, and what does a seed miss?(show answer)
The first thing I would pin down about BLAST word seeds is which claim the number is actually allowed to support.
BLAST uses word-hit lookup heuristics, not a suffix-array scan of every database residue with full DP. Heuristics miss alignments that never hit a seed, especially remote homologs.
Concretely, choose word size from sensitivity needs. For proteins, compare blastp word sizes such as 3 versus 2 on a labeled set. Confirm with a slower method on a subset. Do not claim exhaustive search.
The reason for that specificity is a failure I have seen: blastp word size 3 missed a remote homolog that word size 2 found; the family was declared absent from a pathogen.
Remote protein homolog.
| blastp word size | Found? |
|---|---|
| 3 | no |
| 2 | yes |
| exhaustive SW | yes |
I would not consider it settled without evidence: A sensitivity table versus word size on a labeled remote protein set.
A seed is a filter, not a proof of absence.
Curated: · Written: · Reviewed:
QA-86A de Bruijn graph of a tandem repeat is a cycle. Why is there no unique Eulerian traversal, and what extra evidence resolves copy number?(show answer)
I would start de Bruijn cycles and repeats from the samples, reference, and assay, not from the first plot that looked clean.
Identical k-mers collapse. A cycle can be traversed any integer number of times. Reads or spanning molecules supply multiplicity.
Concretely, use k-mer counts, long reads, or paired links to estimate copy number. Do not output one loop as copy 1 by default.
The reason for that specificity is a failure I have seen: An assembler emitted a 2 kb unit once for an 11-copy 22 kb array; annotation lost 10 copies of a toxin gene.
2 kb unit, 11 copies.
| Evidence | Copies |
|---|---|
| collapsed cycle | 1 |
| k-mer depth 11× unique | 11 |
| spanning HiFi | 11 |
I would not consider it settled without evidence: k-mer depth / unique depth ≈ copy number, checked with spanning reads.
A cycle is a count question, not a sequence question.
Curated: · Written: · Reviewed:
QA-87You hash both a k-mer and its reverse complement as different keys. What double-counting happens on a 2-strand genome, and what is the canonicalization rule?(show answer)
This is a place where a green QC report and a correct handling of canonical k-mer hashing are not the same event.
DNA is double-stranded. Counting both orientations as distinct inflates unique k-mers and breaks spectra. Canonical k-mers keep min(kmer, rc).
Concretely, define strand convention, count canonical keys, preserve multiplicity. Test a 20-mer reverse-complement palindrome and a non-palindromic 21-mer with its reverse complement.
The reason for that specificity is a failure I have seen: A genome-size estimate from non-canonical counts was 1.96× too large and the library was “undersequenced” when it was not.
k=21, haploid 3.0 Gb.
| Count mode | Unique k-mers |
|---|---|
| both strands | 5.9e9 |
| canonical | 3.0e9 |
| size guess | 1.96× high |
I would not consider it settled without evidence: A palindromic 20-mer counts once; a 21-mer and its reverse complement share one canonical key; total unique k-mers are near the haploid genome size for sufficiently large k.
Reverse complement is the same word on the other strand.
Curated: · Written: · Reviewed:
QA-88A variant at POS 100 should hit an exon 100-110. Using Python half-open [start,end) with 1-based POS unconverted, what miss happens?(show answer)
My answer to interval overlap for exons begins at the confounder: if I cannot name what would fake the signal, I do not have a design.
Interval libraries are usually 0-based half-open. VCF is 1-based. Off-by-one at exon edges is a systematic annotation error.
Concretely, convert POS to the interval convention in one function. Test the first and last base of an exon.
The reason for that specificity is a failure I have seen: Splice-site SNPs at the last exonic base missed annotation for 2,200 variants because POS=end was excluded by half-open 1-based misuse.
Exon 100-110 inclusive 1-based.
| POS | Should hit? | Buggy [100,110) 1-based |
|---|---|---|
| 100 | yes | yes |
| 110 | yes | no |
| missed splice SNPs | 2200 |
I would not consider it settled without evidence: Fixtures at start, interior, and end of a 10-base exon.
Overlap is a coordinate system, not a boolean guess.
Curated: · Written: · Reviewed:
QA-89Why is readlines() on a 80 GB FASTQ the wrong memory model, and what invariant must a streaming parser preserve?(show answer)
I would treat streaming FASTQ in Python as an inference under a model, not as a file the pipeline emitted.
FASTQ is four-line records. Streaming constant memory is required. A parser must not split a record across chunks or drop a quality line.
Concretely, iterate records, not the whole file. Validate lengths of sequence and quality. Bound memory in tests with a huge fixture.
The reason for that specificity is a failure I have seen: A QC script read an 80 GB file into RAM, was OOM-killed, and a downstream step treated the missing sample as empty — a silent zero-expression.
80 GB FASTQ.
| Method | Memory | Sample in matrix |
|---|---|---|
| readlines | OOM | missing → 0 |
| stream records | ~MB | present |
| quality length check | required |
I would not consider it settled without evidence: A 10-record fixture plus a memory cap; incomplete records raise, they do not skip.
The file is larger than the machine; the record is not.
Curated: · Written: · Reviewed:
QA-90You drop genes with mean count below 10 after looking at which ones were significant. Why is that invalid, and when is a filter allowed?(show answer)
The useful question for independent filtering in RNA-seq is what still holds after a sample swap, a batch, or a database update.
Independent filtering can raise power when the filter statistic is independent of the test under the null. Choosing the filter after seeing p-values is not that.
Concretely, predeclare a mean-count filter shown to be independent of the test under the null. Apply it before testing. Do not tune the cutoff to maximize discoveries, and do not assume a variance filter is independent.
The reason for that specificity is a failure I have seen: Raising the mean-count cutoff until 200 genes remained “significant” produced a list that failed 18 of 20 qPCR tests; a predeclared filter of 10 left 40 genes with 16 of 20 confirmed.
Cutoff hunt.
| Filter | DE genes | qPCR 20 |
|---|---|---|
| after p-values | 200 | 2 pass |
| predeclared 10 | 40 | 16 pass |
I would not consider it settled without evidence: The cutoff is in the protocol dated before unblinding.
A filter chosen from the p-values is another test.
Curated: · Written: · Reviewed:
QA-91A SNP is a strong instrument for LDL. Why might using it to claim LDL causes depression still be invalid?(show answer)
I would settle Mendelian randomization exclusion restriction against a negative control first, so the clever method has to survive it.
Exclusion restriction says the instrument affects the outcome only through the exposure. Horizontal pleiotropy violates it even when the SNP-exposure association is huge.
Concretely, state the DAG, test relevance, and run pleiotropy-robust sensitivity. Harmonize alleles. Avoid overlapping discovery samples.
The reason for that specificity is a failure I have seen: An LDL instrument that also tags a BMI locus produced a “LDL→depression” effect that vanished in a pleiotropy-robust method.
LDL instrument.
| Method | Effect on depression |
|---|---|
| IVW | significant |
| MR-Egger | null |
| extra BMI path | yes |
I would not consider it settled without evidence: A sensitivity table (IVW, MR-Egger, weighted median) with the same direction, or a stated failure.
A strong association is not a clean pipe.
Curated: · Written: · Reviewed:
QA-92Global clashscore is fine but the claimed ligand pocket has 8 clashes. Which number governs the claim?(show answer)
The judgement in MolProbity clashes at a site is which identifiers and versions are pinned, not which tool name is fashionable.
Global validation can hide a locally impossible geometry. Site-level claims need local clashes, rotamers, and density.
Concretely, inspect the residues you interpret. A pretty whole-model score is not a pocket score.
The reason for that specificity is a failure I have seen: A drug-design figure used a pocket with 8 MolProbity clashes; the ligand pose was an overlap, not a bind.
Kinase pocket.
| Scope | Clashscore | Clashes in pocket |
|---|---|---|
| whole model | 4 | |
| 8 | ||
| pose | invalid |
I would not consider it settled without evidence: Local clash count for the residues in the figure.
The site you talk about is the site you validate.
Curated: · Written: · Reviewed:
QA-93A process expected one BAM per sample but a channel mixed 2 BAMs into one task. What invariant did the workflow break?(show answer)
Where candidates lose the interview on Nextflow channel cardinality is usually treating a hit list as biology.
Channels carry typed records. Cardinality and keys must match the process. Mixing samples in one task is a silent join error.
Concretely, key by sample ID. Assert one-to-one between sheet and channel. Fail on duplicate keys.
The reason for that specificity is a failure I have seen: Two libraries of different people merged into one GATK task; the “het” calls were a mixture of two genomes.
96 samples.
| Channel items | Unique keys | Tasks |
|---|---|---|
| 96 | 95 | 95 |
| mixed BAM | 1 key 2 files | 1 chimeric |
I would not consider it settled without evidence: A channel dump with unique sample keys equal to the sheet.
A tuple in a channel is a contract.
Curated: · Written: · Reviewed:
QA-94Enrichment used all 20,000 HGNC genes as background while only 12,000 were measured. What bias does that create?(show answer)
I would answer gene-set universe by separating what the experiment measured from what the software merely displayed.
Over-representation needs the universe of genes that could have been selected. A larger impossible background inflates terms for well-annotated genes.
Concretely, use tested genes as universe. Freeze mapping and database release.
The reason for that specificity is a failure I have seen: “Cell cycle” looked enriched versus 20k genes; versus the 12k measured universe it was not, because unmeasured genes were mostly poorly annotated.
GO enrichment.
| Universe | Cell-cycle p |
|---|---|
| 20k HGNC | 1e-8 |
| 12k measured | 0.4 |
| honest | 12k |
I would not consider it settled without evidence: Universe size equals the rows that entered DE, not the ontology.
Background is the experiment, not the database.
Curated: · Written: · Reviewed:
QA-95Why store FASTA as a dataclass with id and sequence fields rather than a dict of mutable strings, in a pipeline that checksums inputs?(show answer)
The engineering content of dataclass FASTA records is the evidence boundary and the exclusion log, not the p-value.
Identity and sequence are a record. Mutable dicts let a later step rewrite a sequence without changing the ID, breaking checksums.
Concretely, frozen dataclasses or tuples. Checksum the sequence bytes. Do not mutate in place after validation.
The reason for that specificity is a failure I have seen: A trimmer mutated sequences in a shared dict while IDs stayed; the manifest hashed the original file and the aligner used trimmed reads — irreproducible BAM.
Shared dict.
| Stage | ID | Length |
|---|---|---|
| QC hash | read1 | 150 |
| after trim mutate | read1 | 140 |
| manifest | still 150 | lie |
I would not consider it settled without evidence: Objects are frozen after QC; a mutation test fails.
The ID is a name for one byte string.
Curated: · Written: · Reviewed:
QA-96You need mean coverage per sample and also each contig’s coverage. Why does GROUP BY sample alone lose contigs, and when is a window right?(show answer)
Before releasing I would write what a wrong result for GROUP BY versus window for coverage would look like in the same table.
GROUP BY collapses to the grain you group. Windows compute extra columns without collapsing. Two grains need two aggregations or a window plus a distinct.
Concretely, aggregate contig-sample first, then sample. Or use AVG(cov) OVER (PARTITION BY sample).
The reason for that specificity is a failure I have seen: GROUP BY sample averaged 24 contigs into one number and dropped a 0× contig that was a missed chromosome.
24 contigs.
| Grain | 0× contig visible? |
|---|---|
| GROUP BY sample | no |
| sample+contig | yes |
| missed chr | hidden in 30× mean |
I would not consider it settled without evidence: A query that still lists contig 0× samples.
An average that ate a chromosome is the wrong grain.
Curated: · Written: · Reviewed:
QA-97You merge overlapping exons after sorting by start. Why is a linear two-pointer merge O(n), and what bug appears if the sort is by end?(show answer)
The first thing I would pin down about sorted interval merge is which claim the number is actually allowed to support.
On intervals sorted by start, a single scan merges overlaps. Sorting by end does not guarantee the next start is comparable.
Concretely, sort by (chrom, start, end). Sweep, extending the current end. Test nested and touching intervals with a stated closed/open rule.
The reason for that specificity is a failure I have seen: Sort-by-end merge left a nested exon unmerged and double-counted 1,100 bases of CDS.
Three intervals.
| Sort key | Merged bases |
|---|---|
| start | 500 |
| end | 610 with overlap double |
| nested bug | +1100 CDS |
I would not consider it settled without evidence: A fixture with nested, adjacent, and disjoint intervals.
The sweep’s order is part of the algorithm.
Curated: · Written: · Reviewed:
QA-98A search returns 5 million HSPs. How do you keep the top 100 by bit score in memory, and what ties must you define?(show answer)
I would start heap for top-k BLAST hits from the samples, reference, and assay, not from the first plot that looked clean.
A bounded min-heap of size k tracks top-k in O(n log k). Sorting all 5 million is unnecessary. Ties need a second key (E-value, coordinates) or results shuffle.
Concretely, push HSPs, pop when size>k. Document tie-break. Do not sort the whole list in RAM.
The reason for that specificity is a failure I have seen: A full sort of 5e6 HSPs OOM-killed the node; a restart used unordered equal bit scores and the “top 100” changed 40 rows between runs.
k=100, n=5e6.
| Method | Memory | Stable ties? |
|---|---|---|
| sort all | OOM | |
| heap + E-value | O(k) | yes |
| heap no tie-break | O(k) | 40 rows flip |
I would not consider it settled without evidence: Memory stays O(k); a tie fixture is stable.
Top-k is a heap, not a sorted universe.
Curated: · Written: · Reviewed:
QA-99An overlap graph of reads has 3 components. Why might that be 3 chromosomes or 3 unjoined repeats, and what extra edge would merge them wrongly?(show answer)
This is a place where a green QC report and a correct handling of connected components in overlap graphs are not the same event.
Components are sets with no observed overlap path. That can be biology or missing links. Adding an edge from a repeat can glue chromosomes.
Concretely, require unique overlaps above a length and identity. Inspect bridging reads. Do not force a single component.
The reason for that specificity is a failure I have seen: A 200 bp repeat edge joined two chromosomes into one 180 Mb contig; Hi-C showed they were different chromosomes.
3 components.
| Extra edge | Result |
|---|---|
| none | 3 chr |
| 200 bp repeat | 1 contig |
| Hi-C | 2 chr glued |
I would not consider it settled without evidence: Component count versus expected chromosomes plus long-range confirmation of any join.
A path in the overlap graph is only as unique as the overlap.
Curated: · Written: · Reviewed:
QA-100You tested 1 million SNPs and used Bonferroni 5e-8 as 0.05/1e6. When is that conservative, and when is 5e-8 already the LD-aware convention instead?(show answer)
My answer to Bonferroni versus correlated tests begins at the confounder: if I cannot name what would fake the signal, I do not have a design.
Bonferroni on 1e6 fully independent tests is 5e-8. Common-variant GWAS SNPs are correlated, so the historical 5e-8 already approximates that SNP family. Testing 20 traits is a second family and needs a declared extra correction or joint FDR, not a silent extra Bonferroni after the fact.
Concretely, predeclare SNP and trait families together. Do not Bonferroni an already genome-wide-significant SNP list without a written multi-trait policy.
The reason for that specificity is a failure I have seen: A paper required 5e-8 per trait and then Bonferroni across 20 traits in the supplement only, hiding 8 loci that a predeclared joint FDR kept and replicated.
20 traits, same SNPs.
| Rule | Loci |
|---|---|
| 5e-8 per trait | 22 |
| plus Bonferroni 20 | 14 |
| replicated of hidden 8 | 8 |
I would not consider it settled without evidence: One multiplicity paragraph that names SNPs, traits, and the exact procedure.
You get one honest error budget written before the scan.
Curated: · Written: · Reviewed:
