Skip to content
Tech Interview Prep home
Technical interview guide

Genome Assembly & Sequencing Technologies

How raw sequencing reads become a genome — short-read vs. long-read platforms, de Bruijn graph assembly, and what coverage actually buys you.

Read
49 min
Practice MCQs
25
Interview QA
25
Edition
v3
Editorial status
Reviewed

Scope: Foundational Lander-Waterman and de Bruijn graph models; SPAdes, Canu, Flye, hifiasm, QUAST and BUSCO source publications; current Illumina, PacBio, Oxford Nanopore and NCBI guidance reviewed 2026-09-04.

Overview

Curated: · Written: · Reviewed:

Genome assembly is an inference from incomplete, biased observations

Genome assembly reconstructs longer genomic sequence from reads sampled by a sequencing process. It is not a lossless concatenation problem. Coverage is uneven, reads carry platform-specific errors, repeats create ambiguity, heterozygous haplotypes diverge, contaminants resemble genuine sequence, and library preparation removes or overrepresents regions. An assembler returns one or more graph traversals supported by a model and parameters — never "the" genome.

The interview frame to hold onto: every assembly decision is a bet under uncertainty, and the engineer's job is to make the bet explicit and checkable. Interviewers probe whether you know where the inference can go wrong and what evidence would expose it.

The mental model: sampling, not reconstruction

Coverage is total sequenced bases divided by haploid genome size, but mean depth hides dropout and duplication. Lander-Waterman reasoning relates random sampling to expected gaps; real libraries violate uniform independent sampling through GC bias, amplification, contamination and physically inaccessible sequence. K-mer spectra give an empirical view of sequencing errors, coverage modes, heterozygosity, repeats and approximate genome size before any assembly runs — a strong answer uses them as the first diagnostic, not a post-hoc check.

What coverage buys: redundancy for error correction, repeat depth signal, and k-mer spectra. What it cannot buy: resolution of repeats longer than your reads, rescue of a bad insert-size distribution, or recovery of GC-biased dropout. Realistic targets: PacBio HiFi around 25–30×, Illumina PCR-free 30–50×, ONT ultra-long at lower depth when the goal is span rather than per-base accuracy.

Platform choice and error profiles

Choose technology from the biological objective, not headline throughput. Genome size, ploidy, heterozygosity, repeat spectrum, GC extremes, structural variation, DNA molecule length, target completeness, phasing needs, budget and compute all matter.

  • Illumina short reads: high per-base accuracy, inexpensive, high volume — but PCR/GC bias and lengths (typically 2×150 bp) that cannot span most interspersed repeats. Paired-end orientation and insert-size distributions add linking information within a few hundred bases.
  • PacBio HiFi: circular-consensus reads, typically 10–25 kb at roughly Q30 per-read accuracy, with low systematic error. Long range plus accuracy, at the cost of multiple passes around each molecule and lower top-end read length.
  • ONT ultra-long: reads of 100 kb and beyond, no amplification step, but higher indel error concentrated in homopolymers. Best for spanning the largest repeats and structural variants; usually paired with polishing or a HiFi/short-read complement.

Hybrid assemblies combine accurate short reads with long-range reads but introduce cross-platform mapping and polishing assumptions you must state.

What interviewers probe here: "Which platform for this genome and why?" A weak answer names a platform and stops. A strong answer works backward from repeat lengths and heterozygosity to the read length and accuracy actually required, and names what the chosen platform still cannot resolve.

Repeats: the central obstacle

A repeat is resolvable only when reads or read pairs span it and reach distinguishable unique flanks on both sides. Segmental duplications, centromeric alpha-satellite, rDNA arrays and subtelomeres are the canonical hard cases; telomere-to-telomere assemblies such as T2T-CHM13 were what finally closed them, using ultra-long reads to bridge arrays that HiFi alone left fragmented. If your longest molecules are shorter than the repeat, no amount of coverage resolves it — it collapses or fragments, by arithmetic, not by tool choice.

Likely follow-up: "You have 30× HiFi and the assembly breaks at the centromere — what now?" The answer is more span (ultra-long ONT or optical/Hi-C maps), not more depth of the same library.

Two algorithm families, and what errors do to each

de Bruijn graph assemblers break reads into k-mers, connect prefix-suffix overlaps, and find paths. Sequencing errors create tips and bubbles; insufficient k collapses repeats that differ by less than k. Small k improves connectivity through low coverage but collapses repeats; large k resolves more repeats but fragments at errors or low depth. Multi-k approaches combine evidence across scales. This family dominates short-read assembly because k-mer decomposition makes overlap detection cheap and exact.

Overlap-layout-consensus / string graph assemblers detect read overlaps, arrange them, and compute consensus. They suit long reads, where per-read error is tolerable because overlaps survive it; the failure mode is repeat-induced overlaps — two reads from different copies of a repeat overlap perfectly, and a mis-join follows. Distinguishing true overlaps from repeat-induced ones requires overlap length, identity, unique flanks and read placement.

Graph branches have several causes: sequencing errors, heterozygous alleles, repeats, structural variation, contamination. Tips and bubbles are not automatically errors. Simplification must weigh coverage, sequence identity, topology, read support and expected ploidy without deleting real rare alleles or duplications. Retaining and inspecting the assembly graph is often the only way to explain a linear FASTA.

Weak answer signal: treating "bubble = error, remove it" as a rule. A strong answer asks what the local coverage ratio, path divergence and ploidy predict before deleting anything.

Contigs, scaffolds and metrics

A contig is contiguous assembled sequence without gaps. A scaffold orders and orients contigs using linking evidence and may contain gap placeholders — an N-run is a statement of ignorance, not sequence. Assembly representation matters: primary, alternate, haplotig, unlocalized and unplaced sequences answer different questions, and downstream analysis must not count alternate loci as extra genes or discard genuine haplotypes as duplication.

N50 is the contig length L such that at least half of assembled bases occur in contigs of length at least L. It measures contiguity, not correctness or completeness, and can be inflated by contamination or erroneous joins. NG50 uses an expected genome size, making missing sequence visible. Report total span, contig count, longest sequence and the full length distribution alongside both.

Polishing and consensus quality

Consensus polishing corrects bases and small indels by realigning reads or raw signals to the assembly — Racon/Medaka-style rounds for ONT, Pilon-style for short reads, with platform-appropriate models. Homopolymer indels dominate ONT residual error because the basecaller's hardest cases are runs where length is ambiguous in the signal. Polishing with the same evidence that built the assembly can reinforce a misassembly; excessive rounds can degrade quality. Track every round and validate on held-out or orthogonal evidence.

Consensus quality should be estimated, not assumed: k-mer-based tools such as Merqury compare assembly k-mers against read k-mers to yield a consensus QV and completeness estimate without a reference. Structural correction requires long-range support, not pileup consensus.

Diploid and polyploid assembly

Diploid assembly requires an explicit representation goal. A collapsed assembly merges homologous sequence; a primary-plus-alternate model separates some variants; haplotype-resolved assembly aims to phase chromosomes. Heterozygous bubbles can be misclassified as duplicated contigs, and aggressive purging can delete real segmental duplication — the purging step is where interviewers look for nuance. Trio, Hi-C or other long-range data add phasing evidence with their own error modes (switch errors, misassignments). Report phase blocks and switch rates, and retain unresolved regions explicitly.

Validation, contamination and special cases

Assembly validation must combine independent evidence. Map reads back and inspect coverage, orientation, clipping and discordant pairs; k-mer methods assess consensus quality and completeness without a reference. Reference-based tools identify structural disagreement but cannot tell whether the assembly or the reference is wrong. BUSCO estimates recovery of lineage-specific conserved genes — with the correct lineage dataset and version — not noncoding completeness or perfect copy number.

Contamination detection combines taxonomic similarity, coverage, GC, composition, read linkage and sample context. Removing every foreign-looking sequence can delete horizontal transfers, symbionts or conserved host regions; keeping everything can publish laboratory or index-hopping contaminants. Quarantine uncertain contigs with reasons.

Circular bacterial chromosomes and plasmids need evidence that contig ends overlap consistently, reads span the join, and coverage supports one traversal; rotating the origin is a representation change, not biological closure. Metagenome assembly adds strain mixtures and variable abundance, so single-genome assumptions about coverage, bubbles and completeness no longer hold.

Reproducibility

Production assembly is reproducible and inspectable: raw-read accessions and checksums, sample and library metadata, basecaller and chemistry, trimming, tool and database versions, commands, resource limits, random seeds, graph files, polishing rounds and quality reports. Keep failures separate from biologically fragmented results. The production invariant is evidence-preserving reconstruction: every sequence, join, correction and removed contig traces to reads and versioned parameters, while contiguity, correctness, completeness, phasing and contamination are measured separately.

What interviewers probe, and what a weak answer sounds like

  • "Walk me through choosing a sequencing strategy." Weak: names a favorite platform. Strong: starts from genome size, repeat spectrum and heterozygosity, derives required read length and depth, names what stays unresolved.
  • "Your N50 doubled after a parameter change — good news?" Weak: yes. Strong: suspect mis-joins or contamination; check read-back discordance and NG50 before celebrating.
  • "What creates a bubble in a de Bruijn graph?" Weak: "errors." Strong: errors, heterozygosity, repeats, contamination — and the evidence that distinguishes them.
  • "How do you know your assembly is right?" Weak: "BUSCO was high." Strong: BUSCO with lineage and version, Merqury-style k-mer QV, read-back coverage and discordance, independent maps — each measuring a different failure mode.
  • "Why did the centromere break?" Weak: blames the tool. Strong: longest molecule < repeat length; span, not depth, is the binding constraint.

Worked example: the N50 trap

A release that "won" on N50 illustrates the trap. Assembly A spanned 2.84 Gb with contig N50 38.2 Mb after a chimeric join across a 12 kb satellite. Assembly B spanned 2.71 Gb with N50 18.6 Mb, BUSCO 96.1% complete single-copy, and 0.4% contaminant span. The longer N50 was the worse genome. Report N50 with NG50, BUSCO lineage and version, mapped-read discordance and a contamination table — never as a single gate.

Worked checks

Each row states the claim the lesson defends and the first production miss.