Skip to content
Tech Interview Prep home
Technical interview guide

Biological Databases & Ontologies

The major public repositories a computational biologist works against daily — NCBI/GenBank, Ensembl, UniProt, and Gene Ontology — what each holds and how they cross-reference each other.

Read
49 min
Practice MCQs
25
Interview QA
25
Edition
v3
Editorial status
Reviewed

Scope: Current NCBI E-utilities and data policies; Ensembl release 116; UniProt, RCSB PDB, Gene Ontology, OBO Foundry, Sequence Ontology, HGNC, Identifiers.org and OLS guidance reviewed 2026-09-04.

Overview

Curated: · Written: · Reviewed:

Biological knowledge systems are versioned claims, identifiers, and relations

Biological databases organize sequences, genes, proteins, structures, variants, phenotypes, publications and experimental annotations. An identifier lookup is not merely fetching a fact. The result has a namespace, release, entity model, evidence, provenance, license and update history. Reproducible analysis must preserve those semantics.

Primary archives accept submitted experimental data; knowledge bases curate, integrate or compute interpretations. One record can contain submitter assertions, automated annotation and expert review with different confidence. Database presence is not independent validation. Read evidence and status fields and cite the exact record version used.

An accession identifies an entity within a namespace, while a version identifies a particular sequence or record state. Some accessions remain stable as annotations change; sequence-version suffixes increment when sequence changes. Display names and gene symbols are convenient labels but can be duplicated, renamed or reused. Store canonical namespace plus accession and version, attaching labels as attributes.

Identifier mapping is a many-to-many, time-dependent transformation. One gene may have several transcripts and proteins; merged or split records create retired identifiers; orthologs are not synonyms. A mapping service output must record source and target namespaces, release, mapping method and unresolved or ambiguous cases. Never silently choose the first row.

Genome annotations bind identifiers to a reference assembly and coordinate model. Ensembl releases, NCBI accessions, gene models and transcript versions evolve. A gene identifier without species and release is incomplete, and a genomic coordinate without reference accession is unsafe. Archive releases or content-addressed exports when results must be reproduced.

Ontologies define classes and relations with machine-readable meaning. Gene Ontology separates molecular function, biological process and cellular component. Sequence Ontology describes genomic features. An ontology is usually a directed graph, not a simple folder tree: a term can have multiple parents and relation types such as is_a and part_of have different semantics.

The true-path expectation allows appropriate annotations to propagate to broader ancestors along defined transitive relations. Propagation must be relation-aware; not every edge is transitive and qualifiers can negate or contextualize a statement. Annotation extensions add context such as anatomical location or target and must not be flattened into an unconditional term.

GO annotations connect a biological entity to a term with reference, evidence code, assigned-by source and often qualifiers. Evidence codes describe how the assertion was supported, not a universal quality ranking. An electronically inferred annotation can be useful at scale, while an experiment may be indirect or context-specific. Preserve evidence rather than collapsing all annotations into booleans.

Ontology terms evolve. Labels and definitions change, classes are obsoleted, replaced or merged and relations are revised. Never delete an obsolete identifier from historical data; use replacement and consideration metadata with review for one-to-many changes. Pin ontology release for analysis and test the effect of updates.

Synonyms improve retrieval but have scopes such as exact, broad, narrow or related. A broad synonym is not interchangeable identity. Cross-references and mappings between ontologies can express exact, close or partial alignment and require provenance. Naive string matching confuses homonyms and different species or biological granularities.

Semantic similarity measures shared ontology information, often using ancestor structure and annotation frequencies. Results depend on ontology and corpus versions, propagation policy and evidence filters. Similarity supports hypothesis generation or retrieval; it does not prove shared mechanism or disease causality.

APIs are live services with pagination, rate limits, transient errors and evolving schemas. Resolve identifiers, encode queries, follow documented batching and use contact metadata where required. Validate status codes and response schema, retry bounded transient failures and distinguish empty result from incomplete pagination. For large or reproducible analysis, prefer release-pinned bulk downloads with checksums.

Licenses and consent determine reuse. Public molecular archives can still include contributed materials with third-party terms, and controlled human data cannot be made open merely for reproducibility. Record source licenses, attribution and access conditions at ingestion and propagate them to derived releases.

Production integrations keep a source registry with namespace, authoritative provider, release, retrieval date, checksum, license and schema. Stage updates, diff entity and relationship counts, test sentinel identifiers and downstream outputs and support rollback. Cache by version and query semantics, not mutable URL alone.

The production invariant is resolvable meaning: every stored identifier, annotation and ontology-derived conclusion can be resolved to an authoritative namespace and version, interpreted through explicit entity and relation semantics, traced to evidence and provenance and reproduced from a governed snapshot even after the live database changes.

Joining 18,402 RNA-seq rows on gene symbol merged C1orf106 with a retired alias and silently dropped 37 Ensembl IDs that share the display name across releases 110 and 113. Store ENSG IDs with version, pin the GTF release, and treat symbols as labels. A live API 200 is not a frozen identifier.

A second check: GO enrichment against “all genes in the ontology” while only 12,000 genes were measured inflates well-annotated terms. The universe is the assayed set in that GTF release, and IEA annotations are electronic. Split enrichment by evidence code; if the experimental-only slice disappears, the story was a computed annotation, not an experiment.