Overview
Curated: · Written: · Reviewed:
Schema evolution is a reader-writer and meaning migration
The one-sentence answer interviewers want
Schema evolution changes a data interface while producers, stored records, processors and consumers run different versions of the code that reads and writes it. Compatibility is directional: backward compatibility asks whether a new reader can consume old data; forward compatibility asks whether an old reader can consume new data; full requires both. Every "is this change safe?" question in an interview resolves to three sub-questions: which direction do you need, in which format's resolution rules, and in what deployment order?
A weak answer stops at "adding a field is backward compatible." A strong answer names the format (Avro, Protobuf, Parquet, Iceberg), the direction, the consumers that were actually verified, and the rollout order that makes the direction true. Interviewers probe exactly that gap: they want to hear "compatible relative to what," not a policy slogan.
Compatibility semantics: the vocabulary that decides everything
Backward compatibility is what lets you upgrade readers first: new code reads old data, so you can deploy consumers before producers. Forward compatibility is what lets you upgrade writers first: old code tolerates the new payload. Which you need is a deployment-order question, not a taste question.
Mapping change types to directions, for the common formats:
- Adding a field with a default (Avro: default in the reader schema; Protobuf: new tag number, old binaries skip it) is backward compatible and, in Protobuf, forward compatible too. In Avro, forward compatibility of an added field requires old readers to tolerate the unknown field — Avro achieves this via writer-schema resolution, but only if the writer schema travels with the data.
- Renaming, retyping, dropping, or narrowing (removing an enum value, tightening a numeric type, making an optional field required) breaks at least one direction. Renaming a Protobuf field keeps the wire format (tags, not names, carry identity) but breaks JSON mappings and generated-code APIs. Renaming an Avro field breaks resolution unless aliases are used where the implementation supports them.
- Reusing a deleted Protobuf field number is the classic silent-corruption bug: old binaries will decode the new bytes as the old field. Reserve deleted numbers and names.
"Compatible" is always relative to a format, a policy, the two versions compared, and the actual consumers. A registry can reject structurally incompatible schemas; it cannot prove business meaning, code behavior, or what a downstream Python job does with an unknown key.
Additive is not safe: the misconception interviewers test
The most common trap: "adding an optional field can't break anything." It can break strict deserializers, exhaustive pattern matches, function signatures, snapshot tests, SELECT * pipelines ordered by position, closed JSON objects, and consumers configured to reject unknown keys. JSON Schema 2020-12's unevaluatedProperties and OpenAPI 3.1.1's Schema Object make "additive object field" a real breaking change for strict validators.
Defaults cut both ways. A truthful default lets old data be read; a fabricated default turns "unknown" into apparently observed information — country = "US" on records that never had the field silently skews every aggregate built on it. Test real old/new binaries against historical fixtures, not just the registry's checker.
Choosing a compatibility mode: why "full" is often the wrong default
Registries (e.g., Confluent Schema Registry) offer BACKWARD, FORWARD, FULL, each optionally transitive (checked against all retained versions, not just the latest). The trade-off: stricter modes protect consumers but freeze the schema. FULL TRANSITIVE on a high-churn subject means every additive change must be forward-compatible forever, which in practice means "no one changes the schema" or "someone disables the check" — and a disabled check is worse than an honest BACKWARD policy.
What a platform team should require per tier, not globally:
- Public event streams and APIs (many unknown consumers): BACKWARD_TRANSITIVE at minimum; breaking changes go through versioned topics/endpoints, not mutation.
- Internal streams with known, co-deployed consumers: BACKWARD against the last N versions actually still alive in replay windows.
- Owned analytics tables (Iceberg/Delta): additive evolution under the engine's rules; identity via field IDs, not names.
Interviewers follow up with: "who is allowed to disable a compatibility check, and what does the audit trail look like?" Disabling checks or registering under a fresh subject to bypass a rejection is the failure mode a strong answer names unprompted.
Breaking changes: expand–migrate–contract
When a change genuinely breaks — new grain, retyped field, renamed concept — you do not mutate the shared schema in place. In-place mutation is the failure mode: some reader is always mid-deployment against the old shape. Instead:
- Expand: add the new field/column/version alongside the old. In PostgreSQL terms:
ADD COLUMNnullable, no volatile default (a volatile default can rewrite the heap; a new constraint can scan every row while holding locks that block writes). - Migrate: dual-write or dual-publish, or backfill from one authority in batches. Deploy readers that tolerate both representations.
- Validate parity: compare old and new paths on real traffic before trusting the new one.
- Contract: cut consumers over, stop old writes, then remove the old representation after a telemetry-confirmed window plus rollback headroom.
Versioned topics, tables, and endpoints are the same idea at the contract level: orders.v2 lets old consumers keep reading orders.v1 on their own schedule.
Registry enforcement in CI, not in someone's head
Human review does not scale to dozens of producers and consumers. The working pattern: producers register schemas through a registry; PR/CI runs compatibility checks against the subject's policy before merge; the registry gates production serializers at runtime. Subject naming determines which records share a compatibility history — get it wrong and two different event types silently share (or dodge) one policy. Protect the registry's configuration, audit changes to it, and test the exact serializer and client versions production uses, because a policy is only as real as the code that enforces it.
Format and engine specifics that change the answer
Avro. Resolution is writer-schema against reader-schema. Reader defaults supply fields absent from the writer schema; they do not make a writer field optional, and binary encoding still writes every writer-schema field (JSON encoding may omit a value equal to its default). Unions, enum symbols, numeric promotion, and record names each have precise resolution rules. The writer schema or a stable schema ID must travel with the data or none of this works.
Protobuf. Wire compatibility is about field numbers and wire types. Unknown fields are preserved by binary runtimes but can be lost through JSON or application transformations — a wire-safe change can still be behaviorally unsafe.
Parquet. Logical annotations (decimal precision/scale, TIMESTAMP_MICROS vs MILLIS, legacy INT96) sit above physical types; a warehouse that "just reads Parquet" still has a type-promotion matrix, and reader-side projection means consumers see only the columns they name.
Iceberg. Stable field IDs make rename an ID-preserving operation, not drop-and-add; schema, partition-spec, and sort-order evolution proceed without rewriting existing files under the spec's promotion rules (e.g., int → long). The table spec versions snapshots and manifests: each commit produces a new snapshot referencing its own manifest list, so files written under the old schema stay readable at the snapshot that wrote them — that's time travel, and it's why evolution must be additive rather than rewriting history. A reader that doesn't understand a required spec feature must refuse the table rather than silently drop nested fields. Mixing engines that identify columns only by name with Iceberg's IDs is a known compatibility hole.
Delta Lake. Reader/writer protocol versions and table features (per PROTOCOL.md) communicate required client capabilities; column mapping changes how Parquet physical names relate to logical columns, so an under-versioned client can corrupt or misread the table. Delta's transaction log preserves history per table version, giving the same time-trivial property: a schema change creates a new table version, and readers pinned to an older version see the older schema over the same files.
Time travel changes the compatibility question in one specific way: your old readers include every query that names a past snapshot or table version. A rename that is safe for current readers can still break a dashboard that replays last quarter's snapshot, so the rollback window for a table-format change is bounded by your snapshot-retention policy, not by your deploy cadence.
Semantic evolution: the change that passes every check
Changing milliseconds to seconds, gross to net, event time to processing time, or customer-grain to account-grain passes every structural compatibility check and silently corrupts every downstream decision. This is the senior-level point: physical compatibility is necessary, not sufficient. Name units and time semantics in the schema, version controlled vocabularies, and when meaning changes, ship a new field or a new version even when the type is unchanged. A discriminator change in an OpenAPI payload is wire-visible JSON that still passes a registry's Avro check — different contract surface, same incident.
Null, absent, and default are three different facts. Define whether missing means unknown, not applicable, not collected, or redacted, and preserve that distinction through serialization, database, and API layers. Tightening requiredness needs population evidence and staged enforcement; relaxing a constraint can expose states old code considered impossible.
Worked trace: an Avro rename that breaks, then the fix
Writer schema v1 (what's stored and still being produced):
{ "name": "Purchase", "type": "record", "fields": [
{ "name": "amt", "type": "double"}
]}
Reader schema v2 renames amt to amount with no alias:
{ "name": "Purchase", "type": "record", "fields": [
{ "name": "amount", "type": "double", "default": 0.0}
]}
Trace the resolution of a v1 record amt=42.5:
- Reader looks for
amountin the writer schema — not found. amounthas a default, so resolution succeeds — withamount = 0.0.amtexists in the writer schema but not the reader schema — skipped.
Result: no error, no log line, and the value 42.5 is silently replaced by 0.0 in every old record. This is why "the registry accepted it" and "the data is correct" are different claims. The fix is an alias ("aliases": ["amt"] on the reader field, where the implementation supports it) or an expand–contract: add amount, dual-write, backfill, then remove amt.
What interviewers probe, and the weak versions of each
- "Walk me through a breaking change you shipped." Weak: "we added a field, it was fine." Strong: names the direction, the rollout order, the parity validation, and what broke anyway.
- "Backward vs forward — which do you need first?" Weak: definitions recited. Strong: ties the answer to deployment order — readers-first needs backward, writers-first needs forward.
- "Why not just require full transitive everywhere?" Weak: "safest." Strong: names the agility cost and the bypass behavior it causes.
- "A change passed all compatibility checks but dashboards went wrong. What happened?" Semantic drift — units, grain, or default fabrication. This follow-up separates senior candidates from policy-reciters.
- "How do you know it's safe?" A test matrix: old writer/new reader and new writer/old reader, each supported language/runtime, oldest stored data replayed, null/unknown/enum-addition/timezone cases, registry outage, and rollback. Safe means consumer decisions and stored history remain correct — not merely that a registry accepted the schema.
Operate schema change as a governed release: classify risk, identify owners, generate a structural and semantic diff, run contract tests against historical fixtures, canary, monitor decode errors and semantic metrics, and keep a rollback or forward-repair path. Emergency changes still need evidence and a bounded exception, not a disabled check.
