Skip to content
Tech Interview Prep home
Technical interview guide

Data Catalog & Lineage

Making data discoverable and traceable — what a dataset means, where it came from, and what depends on it.

Read
33 min
Practice MCQs
25
Interview QA
25
Edition
v4
Editorial status
Reviewed
Relevant for
Data Architect

Scope: W3C DCAT 3 and PROV-O, OpenLineage 1.53 object model/facets, Google Cloud Knowledge Catalog lineage, dbt manifest artifacts, Apache Atlas lineage API, and UK metadata guidance current 2026-08-31.

Overview

Curated: · Written: · Reviewed:

Catalog and lineage turn metadata into operational evidence

Two artifacts, deliberately separated. A data catalog answers what a dataset is: definition, ownership, schema, classification, quality expectations, certification status, how to request access. Lineage answers where it came from and what breaks if it changes: which jobs produced it, which tables fed those jobs, which dashboards and downstream models read it. Interviews conflate these constantly — “we have a catalog” as an answer to a lineage question is a reliable weak-answer signal. A search box and a graph are interfaces, not the outcome. The outcome is faster, safer decisions about reuse, change, incident scope, quality, privacy and retirement based on current, attributable evidence. An interviewer probing this topic is usually testing whether you can reason from a concrete decision — “I need to change this column's type” or “this number is wrong, where did it come from?” — back to what the metadata must support.

Lineage is a directed graph. Nodes are datasets, jobs, columns, dashboards, BI reports, sometimes ML models. Edges carry semantics — produced, read, transformed — and ideally the version and time they were observed. Two queries matter more than any other: upstream root-cause tracing (walk backward from a wrong number to the earliest source of the anomaly) and downstream impact analysis (take the transitive closure from a proposed change and enumerate everything that breaks). If you can state those two traversals crisply, you've shown the interviewer you've operated a lineage system rather than admired one.

Identity and lifecycle come before enrichment

Give every resource a stable identity before adding metadata. Distinguish a logical dataset from a physical table, file, stream, API service, materialized view or versioned distribution. Define namespace rules across accounts, regions and environments, and handle rename, move, clone and recreation without silently conflating resources. Display names change; durable identity, aliases and lifecycle history preserve continuity. A rename that mints a new GUID without an alias fragments history; a clone that reuses a path without a new identity conflates tenants.

What the catalog carries

Useful catalog metadata includes business purpose and definition, grain and coverage, owner and steward, authoritative source, schema and version, refresh/freshness, quality expectations and incidents, classification, access path, permitted use, retention, geographic and temporal scope, distributions or services, lineage and limitations. Technical harvesters populate schemas and locations; accountable humans must maintain business meaning and policy. Completeness is not accuracy — a field that appears populated can still be wrong for every consumer.

Discovery should expose enough safe metadata for a prospective user to decide relevance and request access without revealing protected contents. Search uses titles, synonyms, glossary terms, domains, tags and usage signals, but popularity must not be confused with authority or fitness. Results should show ownership, freshness, quality and known limitations, and should distinguish certified, deprecated, experimental and inaccessible resources.

Granularity, and what it costs

Lineage comes in levels: system-level (services talk to services), table-level or dataset-level, column-level or field-level, and row-level. It also splits into logical lineage (business process X feeds report Y) and physical lineage (this specific job, script or query). Column-level is the differentiator and the expensive one. Table-level lineage is often sufficient for impact triage; field-level evidence matters for sensitive attributes, metric debugging and schema change — “does this hashed email flow into this dashboard?” is a column question that no table-level graph answers. A projection differs from a transformation, aggregation, join, filter, mask or an opaque user-defined function, and an output can depend on an input for row selection without copying its value. Record the supported dependency type and confidence; expose unknown or partial lineage rather than guessing. Choose granularity from the decisions and the risk, not from a promise of universal column-level capture.

How lineage is actually captured

Capture methods fall along a spectrum of coverage and cost:

  • SQL/query parsing — parse queries issued through an engine or collected from logs; precise for what ran, blind to what never ran.
  • ETL and transformation manifests — dbt's manifest.json records nodes, sources, exposures, parent_map/child_map and unique_ids: the compiled graph. This is build-time declaration, not every runtime branch of a run-operation.
  • Runtime/log-based observation — orchestrators, query engines, ingestion services and model pipelines emit events at execution boundaries. This is what actually executed, but it can be incomplete during failures, retries and unsupported paths.
  • Inference — ML-based matching of schemas and values. Useful gap-filler, weakest evidence; sampling values to infer relationships creates privacy leakage risk and false edges, and needs separate authority.
  • Manual annotation — the fallback for everything un-instrumentable.

Parsers break on dynamic SQL, stored procedures, UDFs, hand-rolled Python and in-app transforms. Static code parsing reveals declared dependencies but can miss runtime branches, late-binding sources and external side effects. The practical answer interviewers look for: reconcile declared and observed lineage, and treat neither as total truth.

Static versus dynamic lineage

This distinction is the most common senior-level probe. Static lineage is what the pipeline declares at compile or build time — the DAG, the manifest, the code. Dynamic lineage is what actually executed at run time — which branches fired, which late-bound tables were read, which runs failed before writing. Conditional control flow, loops over source lists, and late-binding sources mean the two can disagree. A concrete case: a job reads yesterday_snapshot or today_snapshot depending on run date; static analysis shows one edge, the run log shows the other. During impact analysis, downstream consumers that only exist on the untaken branch are a blind spot; during incident triage, only the dynamic graph tells you which versions of the data actually flowed.

Lineage as event processing

Lineage ingestion is an event-processing system. Events can be duplicated, delayed, reordered, missing or replayed. Use idempotent event identity, lifecycle semantics, checkpoints, dead-letter handling and reconciliation against scheduler and query history. Define how incomplete or failed runs affect published edges. Never infer that a START event produced a valid dataset, and never overwrite a completed run with a later-arriving older event.

In PROV terms, an entity can be generated by an activity, which can use other entities and be associated with an agent. OpenLineage similarly distinguishes datasets, jobs and runs, and attaches facets for schema, source code, quality and lifecycle. Preserve event time, observed time, producer, spec version and stable run identity — a graph edge without the process, version and evidence time can mislead. In PROV-O's Entity/Activity/Agent triangle, wasGeneratedBy, used, wasAssociatedWith, wasDerivedFrom and wasInvalidatedBy distinguish creation from use from retirement. OpenLineage's object model uses namespace plus name for dataset identity, a job (namespace, name, facets) and a run UUID; its column-lineage facet maps each output field to input fields whose transformations carry type (DIRECT/INDIRECT), subtype (IDENTITY, TRANSFORMATION, AGGREGATION, FILTER and so on), description and masking. Claiming IDENTITY with masking: false on a hashed email is a false edge. Custom facets must be namespaced to avoid colliding with later standard names.

DCAT 3 models a catalog of resources: a Dataset is the conceptual asset; a Distribution is an accessible form (file, API, format). Confusing them produces lineage that points at a CSV download as if it were the logical customer table.

Vendor surfaces exist but only know instrumented paths: Google Cloud's Dataplex/Data Catalog lineage exposes table- and column-level views across services; Apache Atlas's Lineage REST retrieves a graph by entity GUID with direction and depth, where depth limits and GUID loss after recreation are operational failure modes, not UI glitches.

Access control on the metadata itself

Access control applies to metadata and graphs, not just data. Schemas, names, SQL, source-code locations, quality statistics and lineage paths can disclose sensitive business structure or the existence of protected datasets. Provide purpose- and role-aware views, redact or summarize facets, avoid embedding credentials or personal values in metadata, log access and maintain retention. A user may be allowed to see that an asset exists without seeing its columns or upstream systems.

Operational workflows, and what interviews probe

The graph earns its keep in four workflows:

  • Breaking schema change — find downstream consumers, owners and recent actual use (dynamic evidence, not just the declared graph); notify and track migration.
  • Incident — trace affected versions and runs forward to decisions and backward to sources.
  • Privacy and retention — locate controlled copies, but verify completeness outside the graph. GDPR rights of access, rectification and erasure need lineage plus a copy inventory: a graph that omits extracts and notebooks is incomplete evidence for a controller.
  • Retirement — combine lineage, query telemetry, contracts and direct owner confirmation, because shadow and offline consumers may not appear in the graph.

What interviewers probe, and what weak answers sound like:

  • “How would you find everything affected by changing this column?” Weak: “search the catalog.” Strong: take the transitive downstream closure at column granularity, intersect with recent actual use so dormant branches don't dominate the blast radius, and name the known blind spots (un-instrumented consumers).
  • “Your lineage says table X feeds table Y, but the number is wrong. How do you debug?” Weak: trusting the edge. Strong: pull run-level evidence — which run, which code version, which input versions — before walking upstream.
  • “Do you need column-level lineage everywhere?” Weak: yes, for completeness. Strong: only where risk and decisions demand it — classified fields, shared metrics, schema-change surface — and label coverage honestly elsewhere.
  • “Is lineage ever wrong?” Weak: no. Strong: yes, always partially — declare static vs dynamic capture, event lag and reconciliation, and confidence per edge.

Likely follow-ups: how you handle rename without fragmenting history, how failed runs appear in the graph, how you'd audit lineage completeness for a regulator, and who is accountable for each metadata field. A weak answer treats the catalog as a documentation project; a strong one treats it as an event-processing and identity system with humans accountable for business meaning.

Measuring it

Measure harvest lag (source change to catalog observation), event completeness (runs in the scheduler versus COMPLETE events), sampled edge precision, and column-lineage coverage on classified fields. Test failed runs, retried attempts, streaming checkpoints, views and cross-project copies. A catalog succeeds when users can act safely and know what it does not know.