Skip to content
Tech Interview Prep home
Technical interview guide

Master Data Management

Maintaining one authoritative version of core business entities — customers, products — across many systems.

Read
34 min
Practice MCQs
25
Interview QA
25
Edition
v4
Editorial status
Reviewed
Relevant for
Data Architect

Scope: IBM InfoSphere and watsonx.data MDM guidance, Microsoft Purview MDM guidance, W3C PROV-O, EU GDPR, and NIST Privacy Framework current 2026-08-31.

Overview

Curated: · Written: · Reviewed:

Master data management creates governed identity, not infallible truth

Master data management (MDM) establishes durable identities and governed attributes for shared business entities — customers, organizations, products, suppliers, employees, locations. It reconciles member records from multiple systems so authorized consumers get a consistent view. The phrase "golden record" describes an intended trusted composite; it does not mean every selected value is objectively true, current for every purpose, or safe for every user.

In an interview, the question behind the question is almost always: who owns truth, and what happens when sources disagree? If you answer MDM with a dedupe story, you've answered the wrong question.

Master data as a category

Master data is the nouns the business runs on; transactional data is the verbs. An order, a payment, a shipment is transactional — it happens once, is immutable in substance, and ages into history. A customer, a product, a supplier is master data — it has identity that persists across many transactions and gets revised over time. Reference data sits between them: small, slow-changing code lists (country codes, currency codes, status enums) that classify other data but rarely need matching. Metadata describes the data itself — schemas, lineage, ownership tags.

The distinction that trips people up: master data is shared and reused across contexts. A cost center used by one system is that system's reference data; the same cost center, once three systems must agree on it, becomes master data. In practice the entities that earn an MDM program are customer, product, supplier, employee, and location — anything where the same real-world thing is represented independently in multiple systems and where disagreeing representations cost money (duplicate mailings, double-paying a supplier, mismatched inventory).

A weak answer treats MDM as "a database of customers." A strong answer names the entity, names the systems that independently author it, and names the decisions that break when they disagree.

Start with the entity and its grain, not a tool

Define what counts as one customer, product, or legal organization; which relationships and lifecycle states matter; whose authority governs meaning; and which consumers need operational, analytical, or regulatory views. An account is not necessarily a person, a household is not a legal entity, and a product model is not the same grain as a sellable SKU. Ambiguous grain makes sophisticated matching confidently wrong — the matcher will happily merge two accounts belonging to one household, or split one person with two emails.

Interviewers probe grain deliberately: "Is a customer the person, the account, or the household?" There is no universal answer; the right answer names the decision that needs the data (marketing to households, billing to accounts, KYC on legal persons) and picks grain per use case — possibly mastering multiple related grains with explicit relationships between them.

Authoritative source and system-of-record ownership

"One golden record" is a decision about ownership and precedence, not a dedupe job. For each attribute, someone must be able to answer: which system wins, and why? Shipping address might be authored by the order-management system; legal name by the CRM; credit status only by finance. The composite is only as trustworthy as the precedence table behind it.

A system of record owns an attribute end-to-end: it is where corrections are made, and other systems are consumers, not co-authors. An authoritative source is the designated winner for a specific attribute, which may differ attribute by attribute. When interviewers ask "how do you build the golden record," they are usually testing whether you reach for a matching algorithm first (weak) or an ownership model first (strong). The algorithm finds that two records describe the same entity; ownership decides what the merged record says.

Reconciliation: survivorship when sources disagree

Once records are linked, survivorship (composition) decides which value represents the entity. Common rules: prefer the authoritative source for that attribute; prefer verified over unverified; prefer the most complete; prefer the most recent; prefer the value effective on a given date. Most-recent-wins is the default people reach for and the one that fails most often — timestamps mean different things in different systems (last edit vs last sync vs last login), and a low-trust feed can overwrite a steward-verified value. Most-trusted-wins is safer but staler.

The case that separates candidates: two sources of equal authority disagree, both recently updated. There is no algorithmic answer. You need a deterministic tie-breaker (declared source priority, then lexical, then steward review), competing values preserved rather than discarded, provenance on every value (observed, steward-entered, inferred, composed — and when it was effective), and a governed manual override. A master record without lineage is difficult to explain, correct, split, or lawfully delete.

Entity resolution and matching

Entity resolution standardizes comparable fields, blocks or buckets plausible candidates, computes evidence across attributes, and classifies pairs or clusters. Deterministic matching on governed identifiers (tax ID, DUNS, email) is strong where those identifiers are reliable and rare where they aren't. Probabilistic and fuzzy matching on names and addresses handles noise — transliteration, nicknames, typos, apartment-vs-building addresses — but every threshold encodes a cost tradeoff between false merges and missed matches. Neither approach is universally superior; thresholds must be tuned by entity, population, language, channel, and use case.

False merges are usually worse than missed matches downstream: a bridge record can incorrectly join two entities, and one bad merge contaminates many attributes and every system that consumed the composite. But the ratio depends on the use case — a marketing dedupe tolerates more false merges than a sanctions screening list. Evaluate with labeled representative pairs, and remember clustering adds failure modes beyond pairwise precision/recall: sample the auto-match, auto-nonmatch, and clerical-review bands; slices with sparse or changing identifiers; protected or underrepresented populations. Monitor drift after source, normalization, or model changes.

Manual merge and unmerge is the escape hatch, not the design. If stewards are merging thousands of pairs by hand, the thresholds are wrong; if there is no unmerge, a single false merge is unrecoverable.

Stewardship

Stewardship handles ambiguity automation can't safely resolve. A review task should show candidate evidence, source records, consequences, and permitted actions without exposing unnecessary personal data. Merge and split operations need separation of duties for high-risk entities, reason codes, audit trails, reversible or compensating procedures, downstream propagation, and service-level objectives. Stewards need domain authority and capacity, not merely a queue.

Implementation styles: registry, consolidation, coexistence, centralized

The style question is really: where does the golden record physically live, and how much do you couple to the sources?

  • Registry (virtual): sources stay authoritative; the hub holds only identifiers and crosswalks, assembling the view at read time. Lowest source disruption, but every consumer query touches the sources, so latency and availability inherit the worst source, and cross-source survivorship happens per-request.
  • Consolidation: sources stay authoritative for authoring; the hub builds a physical consolidated record used for analytics, not fed back to operations. Clean separation, but operational systems never benefit from the mastered view.
  • Coexistence: the hub maintains a physical golden record and pushes cleaned values back to sources, so sources and hub stay in sync. Best consumer experience, hardest correctness problem — you now have bidirectional flows, loop risk, and conflict rules for the write-back itself.
  • Centralized (hub): the hub is the system of record; authoring happens there and sources become consumers. Strongest consistency and governance, highest migration cost and organizational resistance.

Style can differ by domain or attribute — centralized for product, registry for supplier. Whatever you pick, document exactly where each attribute may be authored and how conflicts, loops, and outages resolve. Interviewers follow up with "what breaks when the hub is down?" — a registry style degrades to source-local truth, a centralized style takes down the golden view with it.

Distribution is part of correctness

Master changes reach consumers by API, events, batch files, or replication. Give changes stable identifiers, entity version, event time, ordering scope, and idempotency semantics. Consumers need schemas, deprecation, replay, reconciliation, and deletion/correction behavior. "Exactly once" cannot be assumed across independent systems; design idempotent application and periodic comparison to repair missed, duplicated, or reordered messages.

Privacy and security

Linkage increases identifiability and inference even when sources were separately permitted. Combining them into a 360-degree view is further processing: it needs its own purpose and lawful basis, not a presumption that each source's original collection purpose covers the composite. Define purpose and lawful authority, minimize attributes, restrict views by consumer purpose, protect match keys, log use, and enforce retention and rights across member records, mastered values, candidates, exports, and backups. Pseudonymization does not automatically make linked records anonymous, and broad 360-degree access is not a legitimate default.

Rollout and operation

Roll out by bounded domain and measurable outcome: profile sources, establish ownership and semantics, create a crosswalk, tune matching, define composition and stewardship, validate a shadow master, then migrate selected consumers with reconciliation and rollback. Avoid big-bang replacement of operational identifiers. Measure duplicate reduction, false merges, missed matches, review aging, attribute provenance and freshness, downstream convergence, correction time, and business outcome — not just records processed.

Operate MDM as a versioned control system. Changes to normalization, blocking, thresholds, models, source priority, or identity rules can reshape many entities. Use simulation, impact reports, approvals proportional to blast radius, canary populations, immutable evidence, rollback or repair plans, and post-change monitoring. Test late events, source replay, identifier reuse, acquisitions, household changes, consent and rights requests, merge-then-split, vendor exit, and disaster recovery.

What interviewers probe, and what a weak answer sounds like

Likely follow-ups, in rough order of frequency:

  • "How do you decide which value wins when two systems disagree?" — they want an ownership/precedence model with a tie-breaker, not "most recent wins."
  • "What's the cost of a false merge versus a missed match here?" — they want the answer tied to the use case, not a generic recall figure.
  • "Registry or centralized hub — how would you choose?" — they want latency, coupling, and authoring-ownership trade-offs, not a vendor name.
  • "How do you undo a bad merge that already propagated to five downstream systems?" — they want unmerge design, versioning, and replay, not "we'd fix it manually."
  • "Does combining these datasets need a new lawful basis?" — increasingly common at staff level; the composite is further processing.

A weak answer sounds like: a tool pitch ("we use tool X"), a dedupe story with no ownership model, "latest wins" stated as a default rather than a risk, or a golden record described as true rather than governed. A strong answer names the entity and its grain, names who owns each attribute, states the false-merge/missed-match tradeoff in terms of downstream consequence, and treats the golden record as a decision with provenance, not an output of an algorithm.