Overview
Curated: · Written: · Reviewed:
Make feature values temporally correct, consistent, and operable
A feature store is a system for defining, discovering, computing, storing and serving model inputs with lineage and governance. It can reduce duplicated transformations and training-serving skew, but it is not automatically a source of truth or a guarantee of correctness. The real contract is: for an entity and decision time, return the intended feature definition and version, computed only from information available then, within freshness and latency objectives.
Start with feature semantics. Record name, owner, description, entity/join keys, type/schema, event time, source, transformation code, window, aggregation, default/null behavior, freshness/TTL, version, consumers, sensitivity and quality expectations. A feature called thirty_day_spend is ambiguous without currency, inclusion, time zone, window boundary, late-event and refund policy. Catalog discovery without precise semantics spreads mistakes faster.
Separate offline and online responsibilities. Offline stores retain historical values for exploration, point-in-time training retrieval and batch inference. Online stores provide low-latency lookup and often keep only the latest value per entity. The offline history may be authoritative for reconstruction; online is a materialized serving view. Do not assume online history exists or that the latest record is correct merely because its event timestamp is greatest.
Point-in-time correctness prevents future leakage. For every labeled entity row at prediction time t, retrieve the newest eligible feature value whose event time was available by t, subject to TTL and created/ingestion-time policy. Event time describes when reality occurred; created/ingestion time describes when the platform learned it. Late and corrected data require explicit behavior. Entity-only joins or latest-value joins make historical training unrealistically informed.
Use the same versioned transformation semantics for training and serving. This may mean shared code, a compiler to batch/stream forms, or parity tests—not necessarily one runtime. Compare offline and online values for the same entities and effective times, including nulls, types, rounding, window boundaries, late events and defaults. Skew can arise from code divergence, source delay, different clocks, materialization lag or incompatible schema.
Materialization computes and copies features into serving stores. Define start/end windows, schedule, watermark, overlap, idempotency, retry, backfill and correction semantics. Monitor source freshness, job state, rows/entities, lag, failures, rejected writes and online age. A successful job that processes zero unexpected rows is not healthy. Backfills must not overwrite newer online values accidentally or send stale event-time records to low-latency serving.
Entity keys are security and correctness boundaries. Define canonical identity, namespace/tenant, composite-key encoding and lifecycle. Prevent collisions, cross-tenant retrieval and silent coercion. Manage entity merges/deletes and privacy requests across offline history, online state, caches, logged features and saved datasets. A broadly shared feature server can expose sensitive attributes even when the model endpoint is protected.
Version feature definitions immutably for behavior-changing changes. Models pin required feature-set versions and schemas. Additive compatible schema evolution may be supported by a platform, but semantic changes still require a new version, backfill and consumer migration. Do not edit an old definition so historical model lineage points to new behavior. Deprecate with usage inventory and a replacement plan.
Design retrieval failure explicitly. Missing entity, expired TTL, late source, online outage and authorization failure are different. Defaults can hide incidents and bias groups; choose them only with model training alignment and safe meaning. Otherwise fail closed, degrade to a reviewed fallback or route to human handling. Record a presence/freshness signal and monitor fallback rate by slice.
Protect and govern features as data products. Apply least privilege by producer/consumer and environment, encryption, private connectivity where needed, audit reads/writes/definition changes, retention and purpose limitation. Separate registry administration, transformation execution and raw feature access. Treat transformations and materialization images as supply-chain code. Cross-account sharing needs explicit provenance, residency, licensing and revocation.
Test before adoption: point-in-time fixtures around boundaries, late/backfilled/corrected events, duplicate/out-of-order ingestion, entity collision, offline-online parity, TTL expiration, default behavior, schema evolution, job retry, online outage, regional failure, authorization and deletion. Trace a production prediction to model, feature definition, materialization and source evidence.
Measure feature coverage and ownership, definition/lineage completeness, offline-online parity, freshness/TTL violations, point-in-time test failures, materialization reliability/lag, missing/default/fallback rate, entity-key errors, access anomalies, unused/duplicate features, consumer count, cost and incidents caused by skew or leakage. A successful feature store makes correct reuse easier while keeping time, identity, provenance and failure visible.
Leakage from a naive join is large enough to invalidate a project, and it is invisible in offline metrics because the offline metrics are the thing being corrupted. Take a churn model whose training rows are labelled at a prediction date and joined to a support_tickets_30d feature computed from the latest available value rather than the value as of that date: rows labelled churned pick up the tickets the customer filed while cancelling, which are in the future relative to the decision. The model reads that feature as the dominant predictor and scores 0.94 AUC offline; served on requests where the future does not exist, it lands near 0.71, and the gap is attributed to drift rather than to the join. The test that catches it is cheap — build a fixture with events deliberately straddling the decision timestamp and assert that the retrieved value ignores everything after it.
A feature store is infrastructure with a real fixed cost, and adopting one before the problem exists trades a small duplication for a large operational surface. For a single model with one training pipeline and one serving path, the same transformation code called from both places gives point-in-time correctness and parity with no registry, no materialization schedule and no online store to keep warm. The economics change with the number of independent consumers: once several teams recompute overlapping features, the duplicated definitions diverge quietly and each divergence is a skew bug nobody owns. Adopt when reuse across teams is real rather than anticipated, and judge the result by parity failures caught and definitions retired, since a feature store that mainly adds a second place for the same logic to live has made skew more likely rather than less.
