Skip to content
Tech Interview Prep home
Technical interview guide

Phylogenetics & Evolutionary Analysis

Reconstructing evolutionary relationships from sequence data — tree-building methods, substitution models, and molecular clocks.

Read
50 min
Practice MCQs
25
Interview QA
25
Edition
v3
Editorial status
Reviewed

Scope: MAFFT 7; current IQ-TREE, RAxML-NG, MrBayes and BEAST documentation; ModelFinder, ultrafast bootstrap, multispecies coalescent and recombination source literature reviewed 2026-09-04.

Overview

Curated: · Written: · Reviewed:

A phylogeny is an inference under a history model, not a picture of certainty

Phylogenetic analysis estimates evolutionary relationships from characters such as nucleotide or amino-acid sequences. The result depends on sampled taxa, orthology, alignment, partitions, substitution and rate models, inference algorithm, rooting, and processes such as recombination, duplication, horizontal transfer and incomplete lineage sorting. A visually clean tree can still represent the wrong question or model.

The interview version of that claim: support values are statements about your pipeline, not certificates about the past. When an interviewer shows you a tree with 98% bootstrap on a suspicious clade, they are testing whether you ask about the alignment, the model, and the alternative histories before praising the support. A weak answer treats the tree as data. A strong answer treats it as the output of a chain of modeling decisions and names which decisions most plausibly produced the odd result.

What interviewers probe, and what a weak answer sounds like

Almost every phylogenetics question is really one of four probes:

  • "Why is this clade wrong?" — they hand you a tree with a biologically implausible grouping and want you to walk upstream: alignment forcing, long-branch attraction, hidden paralogs, contamination, recombination. Weak answer: "the bootstrap is high, so it's probably right."
  • "What does this support value mean?" — they want the exact interpretation (below). Weak answer: "95% means 95% probability the clade is true."
  • "Gene tree or species tree?" — they want the discordance processes and which method handles which. Weak answer: "concatenate everything and take the consensus."
  • "What would you change in this pipeline?" — they want model adequacy, sensitivity checks, and reproducibility. Weak answer: "use the default model, it's standard."

Likely follow-ups once you answer well: how you'd test for recombination, when relaxed clocks are justified, why your outgroup is trustworthy, and what happens to your topology if you drop the trimming step. Prepare one concrete story of a tree that changed under a defensible reanalysis — the case at the end of this guide is a template.

Alignment quality is the upstream error source

Sequence selection comes before inference. Compare homologous features with consistent boundaries and avoid silently mixing paralogs, isoforms, contaminants or non-overlapping domains. More taxa can break long branches and improve context, while biased sampling can change topology. Missing data are not automatically fatal, but their distribution and informativeness matter.

Multiple-sequence alignment proposes positional homology — every column asserts that these residues descend from a common ancestral position. That assertion is a hypothesis, and tree error propagates from it: a column that pairs non-homologous residues injects convergent signal that no downstream model can remove. Progressive and iterative methods trade speed and accuracy; protein structure or codon awareness may matter. Over-alignment forces unrelated residues together; under-alignment discards signal. Trimming ambiguous columns can reduce noise but also creates confident bias if the filter was chosen to favor a topology.

Interview probe: "Your topology changed after trimming — is the trimmed or untrimmed tree right?" The defensible answer: neither is automatically right; you compare topology and support across plausible alignment and trimming choices, predeclare objective criteria, retain the untrimmed input, and report which splits are unstable. If a reviewer asks whether trimming drove the result, you should already have run both.

Orthology, paralogy, and the gene tree versus species tree gap

A gene tree is not automatically a species tree. Duplication and loss, incomplete lineage sorting, hybridization, introgression and horizontal transfer create discordance between the two. The interview trap is treating a gene tree as a species tree without stating which processes you checked for.

  • Duplication and loss: after a duplication, each paralog's history is a subtree of the species history. Reconciliation explains a gene tree by mapping it onto a species tree with duplication and loss events; it requires defensible rooting of both trees and awareness that hidden losses can mimic absence.
  • Incomplete lineage sorting (ILS): successive speciations occur before ancestral alleles coalesce, so individual loci can group taxa in an order that contradicts the species tree. The multispecies coalescent predicts the expected discordance pattern; you need many independent, recombination-bounded loci, and treating linked loci as independent inflates confidence.
  • Horizontal transfer: a transferred gene groups donor and recipient rather than following species descent. Distinguish transfer from contamination using multiple core loci, synteny, composition and mobile elements.

Concatenating loci estimates one compromise history and can yield strong support despite conflict — high support on a concatenation is support for the average signal, not evidence that loci agree. Species-tree and coalescent-summary methods model particular discordance processes, but they do not fix gene misidentification or recombination within loci. Weak answer: "the concatenated tree has 100% support, so discordance isn't an issue." Strong answer: report gene-tree concordance across loci, quantify conflict, and present the discordance rather than hiding it behind a consensus.

Inference methods and what support actually measures

Three families dominate practice, and the trade-offs are the answer:

  • Distance methods (e.g., neighbor joining) compress each pairwise comparison into one distance before tree building. Fast, useful for exploration and huge datasets, but the compression discards site-pattern information.
  • Maximum likelihood (ML) chooses the topology and parameters that maximize the probability of the observed alignment under the model. Tree space grows super-exponentially with taxa, so practical programs use heuristic searches from multiple starts; repeating the same deterministic start is weaker evidence than independent searches that agree. Branch length is expected substitutions per site, not calendar time.
  • Bayesian phylogenetics samples a posterior over trees and parameters via MCMC. Independent chains, trace plots, effective sample size, stationarity and cross-run convergence matter; a finished MCMC run is not necessarily a converged one. Priors on topology, rates and clocks can dominate weak data, so prior-only runs and posterior predictive checks are part of the analysis.

Parsimony minimizes changes and can be inconsistent under long-branch attraction; it remains useful for exploration when its assumptions are stated.

Bootstrap support measures how consistently a split is recovered when alignment columns are resampled under a specific procedure. It is not the probability that the clade is true. Approximate bootstraps (e.g., ultrafast) have method-specific interpretations and convergence criteria — a 95% ultrafast-bootstrap value is not interchangeable with 95% standard bootstrap. Posterior clade probability instead conditions on model, priors and data, and can differ substantially from bootstrap support on the same split.

The unifying point interviewers want: a supported clade is a claim about the model as much as the data. High support under a mis-specified model is high confidence in the wrong thing. Weak answer: quoting the number. Strong answer: stating the procedure, the replicate count, and whether the split survives model and alignment changes.

Substitution model selection and adequacy

Substitution models specify equilibrium frequencies and relative change rates. Nucleotide models range from equal rates (JC69) to parameter-rich reversible processes (GTR); protein matrices summarize empirical replacements (e.g., WAG, LG) or mechanistic ones (e.g., MT matrices for mitochondria). Among-site rate heterogeneity is commonly approximated with discrete-gamma categories, sometimes plus an invariant-site class — note that invariant-plus-gamma has known identifiability issues, which is a fair follow-up question.

Model selection optimizes a criterion on candidates: likelihood-ratio tests for nested models, AIC/BIC for non-nested trade-offs of fit against parameters. Two failure modes to name in an interview:

  1. Accepting the default. The default model in a tool is a convenience, not a finding. State which candidates you compared and under what criterion.
  2. Confusing selection with adequacy. Winning an AIC comparison among ten wrong models does not make the winner right. Adequacy testing — posterior predictive simulation, comparing observed versus expected site-pattern frequencies, checking composition bias — asks whether the model fits at all. Site-heterogeneous and mixture models (e.g., CAT-style mixtures, or C60 profile mixtures for composition heterogeneity) exist precisely because a single exchangeability matrix often fails on divergent data.

Interview probe: "Your GTR+Gamma tree has a long-branch artifact — what do you try?" Expected answer: check composition bias and saturation, add taxa to break long branches, try a site-heterogeneous or mixture model, and compare topologies — not just scores — across the change.

Recombination: when one tree is the wrong description

Every standard phylogenetic method assumes the alignment has a single history. Recombination breaks that assumption: different genome segments inherit different genealogies, so a single bifurcating tree is not a compressed summary of the data — it is a wrong model of it. The failure is insidious because the inferred tree is a mosaic with no true counterpart, and support values can still be high on the chimera.

What recombination-aware inference does instead:

  • Screen first. Detect breakpoints with recombination-detection methods, keeping in mind that rate variation can produce false positives — corroborate with multiple methods.
  • Split the data. Analyze nonrecombining blocks separately and test whether key conclusions hold across segments.
  • Change the representation. Where reticulation is central, use phylogenetic networks or recombination-aware models that represent the genealogy as an ancestral recombination graph rather than a single tree.

For pathogens, sampling bias, within-host diversity and transmission bottlenecks also make a phylogeny an imperfect proxy for a transmission tree — a common interview distinction. Weak answer: "recombination adds noise." Strong answer: recombination is a model violation with a specific fix, and you can name the fix.

Rooting, clocks, and time

An unrooted tree describes splits without ancestry direction. An outgroup roots the ingroup only if it is correctly placed and not so distant that model error dominates — a distant outgroup is a classic long-branch-attraction vector. Midpoint rooting assumes roughly clock-like distances. Root choice changes statements about ancestral states and direction even when the unrooted topology is unchanged.

Molecular clocks connect genetic change to time. Strict clocks assume one rate; relaxed clocks allow branch-rate variation under a distribution. Calibration comes from fossils, sampling dates or biogeography and is represented by uncertainty distributions, not exact points — a fossil usually bounds a node age rather than dating it exactly. Temporal signal must be tested: date randomization (re-running with randomized tip dates should destroy the signal if it is real) and sensitivity to priors expose analyses driven by assumptions rather than sequences. Rate and time are confounded without calibration; say so before an interviewer does.

Tree files, provenance, and the production invariant

Tree files require careful semantics. Newick parentheses encode topology, but internal labels may be names or support values depending on the tool, and a tree can be rooted or unrooted with no obvious marker. Preserve taxon mapping, branch units, support type, partitions and annotations; validate round trips through target tools rather than parsing labels with ad hoc string operations.

Production workflows pin input accessions, alignment and partition files, software and models, seeds, commands, logs, all candidate searches and support replicates. Release the alignment with the tree and document exclusions. Sensitive pathogen or human metadata require access control and aggregation to avoid exposing identities or locations.

The production invariant is uncertainty-preserving history: every displayed split, root, branch length and date is traceable to exact taxa, homologous characters, alignment, model, search and support procedure, with discordance and alternative explanations reported rather than hidden by a consensus tree.

A worked case: support is not topology insurance

A 98% ultrafast-bootstrap clade grouped two long-branched parasites with a distant host after an alignment that forced 40 gappy columns. Repeating the search with a better-fitting mixture model and without those columns moved the clade 11 nodes. Nothing about the data changed — only the modeling decisions did.

If asked to defend a published tree, the checklist this case implies:

  1. Publish the alignment, model, and independent runs — not just the final tree.
  2. Show the support procedure and its convergence, with the method named.
  3. Report at least one alternative history (ILS, recombination, contamination) that was actually tested, and what testing it would have looked like.
  4. Flag which splits are unstable across defensible pipeline choices.

That last item is what separates a senior answer from a junior one: the junior defends the tree, the senior characterizes its uncertainty.