Skip to content
Tech Interview Prep home
Technical interview guide

Structural Bioinformatics & Protein Structure Prediction

Predicting and validating 3D protein structure from sequence — homology modeling, deep-learning methods like AlphaFold, and how confident a predicted structure actually is.

Read
49 min
Practice MCQs
25
Interview QA
25
Edition
v4
Editorial status
Reviewed

Scope: PDBx/mmCIF v5 and wwPDB validation; AlphaFold and AlphaFold DB guidance; ModelCIF; DSSP, MolProbity, TM-align and comparative-modeling source literature reviewed 2026-09-04.

Overview

Curated: · Written: · Reviewed:

A protein structure is a coordinate model with evidence and uncertainty

Structural bioinformatics connects sequence, three-dimensional coordinates, experimental observations and computed predictions. A coordinate file is not a direct photograph of one rigid molecule. Experimental models are interpretations of diffraction, density or restraints; predicted models encode learned structural regularities. Conformation, partners, ligands, environment, constructs and uncertainty determine what questions a model can answer.

Primary structure is sequence; secondary structure describes local backbone patterns; tertiary structure is the fold of one chain; quaternary structure describes assemblies of chains. Domains may fold independently and move relative to each other. A high-confidence domain does not establish the orientation of another domain, and one crystallographic conformation does not enumerate a protein's ensemble.

X-ray crystallography, cryo-electron microscopy and NMR provide different evidence and resolution behavior. Global resolution is only one quality signal. Local density or restraint support, refinement statistics, map-model fit, geometry, occupancy and alternative conformations matter. Compare entries within method-appropriate context and read validation reports rather than ranking structures by one number.

The crystallographic asymmetric unit is not automatically the biological assembly. Symmetry mates, author annotations and computational interfaces inform oligomeric state. Solution conditions and orthogonal biochemical evidence may contradict a crystal contact. Choose assembly deliberately before analyzing interfaces, ligand sites or distances.

PDBx/mmCIF carries entity, chain, residue, atom, assembly, experiment and provenance data without legacy PDB format limits. Author residue numbering can contain insertion codes and may differ from sequential label numbering. Alternate locations, occupancies, missing residues, modified residues and multiple models must not be silently collapsed by a parser.

Geometry validation checks whether bond lengths, angles, backbone dihedrals, side-chain rotamers and nonbonded contacts are plausible. Ramachandran outliers and clashes are diagnostic, not automatic proof of a wrong structure: supported strained conformations can be functional. Conversely, ideal geometry does not show that coordinates fit experimental data.

Structural similarity is alignment-dependent. RMSD is sensitive to outliers, length and chosen atoms and can look small for a short common core. TM-score weights global topology and length differently. Always report aligned residues, coverage, sequence identity, atom selection, symmetry and flexible-domain treatment alongside a score.

Comparative modeling relies on a template whose homologous coordinates guide the target. Alignment error often dominates when sequence identity is low. Template resolution, construct, ligand, state, missing loops and oligomeric context can matter more than nominal identity. Build multiple models and validate independently rather than optimizing only the modeling objective.

AlphaFold-style predictions provide per-residue confidence such as pLDDT and pairwise confidence such as predicted aligned error. High pLDDT supports local geometry, not necessarily ligand pose, oligomer, conformational state or biological activity. Low confidence can reflect intrinsic disorder, flexible linkers, lack of evolutionary information or prediction failure. PAE is critical for judging relative domain orientation.

Predicted structures should be checked for sequence coverage, signal peptides, transmembrane segments, disordered regions, cofactors, post-translational modifications and construct differences. A monomer prediction does not prove a native interface. Multimer and docking hypotheses require interface confidence, evolutionary or experimental support and negative controls.

Docking ranks poses under approximate scoring and sampling. Protonation, tautomer, waters, metals, receptor flexibility and pocket state can dominate. A docking score is not binding affinity, selectivity or efficacy. Validate enrichment, pose recovery and prospective experiments using realistic decoys and avoid training-target leakage.

Mutation interpretation combines conservation, structural environment, stability, interfaces, active sites and experimental phenotype. A predicted destabilization is not equivalent to pathogenicity, and predicted structure uncertainty should propagate into the claim. Population, splicing and regulatory evidence may be more relevant than a three-dimensional contact.

Production workflows retain exact accession and revision, assembly, chains, sequence mapping, parser and dictionary versions, selected alternate locations, modeling inputs, templates, databases, software, seeds and quality metrics. Cache by content hash and revalidate when structures are remediated or predictions updated.

The production invariant is evidence-bound geometry: every distance, site, interface and predicted effect can be traced to exact sequence and coordinate versions, biological assembly and conformation, with experimental support or model confidence and coverage reported at the same local scale as the conclusion.

A docking claim used AlphaFold residues 12–340 (mean pLDDT 91) to place a ligand against a second domain whose PAE to the first was 28 Å. The interface was a guess about domain orientation, not a structure. Use PAE for relative domain placement, pLDDT for local backbone, and an experimental assembly when the question is a contact. Color is not evidence.

A second production check: CASP-style global metrics can look excellent while a ligand-contact residue is locally wrong. If PDB 1ABC has density for that side chain, quote local map correlation and MolProbity clashes for those atoms; if only AlphaFold exists, quote pLDDT and PAE for that pair of residues and refuse a contact claim when PAE exceeds the distance you report.

Worked example: reading pLDDT before trusting a predicted structure

A predicted structure arrives with a per-residue confidence score in the B-factor column. Treating the model as uniformly reliable is the error; the score is what tells you which parts to use:

pLDDTInterpretationSafe to use for
> 90very high, backbone and side chains reliableactive-site geometry, mutation modelling
70 - 90confident backbone, side chains less sofold assignment, domain boundaries
50 - 70low, treat as a hypothesisrough topology only
< 50very low; often an intrinsically disordered regionnot structure at all — a disorder signal

The last row is the one that is misread most often. A pLDDT of 32 across a 60-residue stretch is usually not a failed prediction; it is a positive prediction of disorder, and those regions are frequently functional — linkers, degrons, and phosphorylation sites live there.

A worked read of one protein:

Residues   1- 24   pLDDT 41   signal peptide, cleaved -- disordered by design
Residues  25-198   pLDDT 94   catalytic domain -- use it
Residues 199-241   pLDDT 38   linker -- position relative to other domains is meaningless
Residues 242-410   pLDDT 91   binding domain -- use it
PAE(150, 300) = 24 A          the two domains are each confident, their RELATIVE position is not

Both domains score above 90 and are individually trustworthy. The predicted aligned error between residue 150 and residue 300 is 24 Angstroms, which says the model does not know how they sit relative to each other. Any conclusion about the inter-domain interface, or docking a ligand into a cleft formed between them, is unsupported — even though every residue involved has a high pLDDT.

That distinction — per-residue confidence versus relative-position confidence — is the whole reason PAE is published alongside pLDDT, and it is the most common way a predicted structure is over-interpreted.