Overview
Curated: · Written: · Reviewed:
Key takeaways
- A cache is justified only when requests reuse results and the product can define an acceptable freshness window.
- Every cache creates two modes—hit and miss—and production capacity must survive a shift toward misses.
- Freshness is a product requirement, not a TTL guessed by an engineer. Express it as bounded staleness and invalidation behavior.
- Prevent stampedes before launch with request coalescing, jittered expiry, controlled refresh, and downstream backpressure.
- Cache keys are security boundaries. Include every dimension that changes authorization or representation.
- Measure backend load, miss penalty, stale responses, evictions, and hot-key concentration—not hit rate alone.
1. Why cache—and when not to
A cache stores a reusable representation closer to the caller or in a cheaper execution path. A hit can avoid database execution, a network round trip, serialization, remote API cost, or expensive computation. The business case is normally one or more of:
- lower tail latency;
- higher throughput from the protected dependency;
- lower infrastructure or API cost; or
- bounded availability during a dependency brownout by serving an explicitly acceptable stale value.
Caching helps only when the workload has reuse. If nearly every request has a unique key, entries expire before reuse, or correctness requires the latest committed value, the added cache lookup and invalidation machinery can make the system slower and less reliable. Start with a measured bottleneck and a freshness requirement; do not start with Redis.
A small latency and capacity model
Suppose a cache hit costs 2 ms, a miss costs 2 ms plus a 48 ms database call, and the hit rate is 90%:
expected latency = 0.90 × 2 ms + 0.10 × 50 ms = 6.8 ms
That average hides the operational risk. At 10,000 requests/second, the database receives roughly 1,000 requests/second at a 90% hit rate. A cold cache sends almost 10,000 requests/second. If the database was downsized around normal cache behavior, a cache restart can become a database outage. Capacity-test the miss path and cap fallback traffic.
Worked example: a 90% hit rate is still 1,000 source rps
10,000 requests/second, 2 ms hit, 50 ms miss.
| cache state | source load | expected latency |
|---|---|---|
| 90% hit | 1,000 rps | 6.8 ms |
| cold / fail-open | ~10,000 rps | 50 ms |
| coalesced refresh of one hot key | 1 source read | waiters share that fill |
The interview answer is the miss-path capacity, not the hit-rate slide.
2. Where caching can live
| Layer | Best fit | Main tradeoff |
|---|---|---|
| Browser/private HTTP cache | User-specific reusable responses and assets | Harder to invalidate centrally; device-local |
| CDN/shared HTTP cache | Public responses reused across many users | Cache-key and privacy mistakes have large blast radius |
| Gateway/reverse proxy | Whole API responses with protocol-level policy | Becomes part of the request availability path |
| In-process cache | Tiny, extremely hot, slow-changing data | Per-instance copies diverge and multiply cold-start load |
| Distributed cache | Shared dynamic data across an application fleet | Network hop plus another service to scale and protect |
| Database-adjacent/materialized result | Expensive queries or aggregates | Dependency-aware refresh and invalidation complexity |
A two-level design—small local L1 plus shared distributed L2—can reduce network latency without making every miss reach the database. It also creates a coherence problem: invalidating L2 does not automatically remove every L1 copy. Define propagation, maximum local TTL, and behavior when invalidation delivery is missed.
browser/private cache
│
▼
CDN/shared HTTP cache ── public representations only
│
▼
gateway/API ──▶ per-instance L1 ──▶ shared L2 ──▶ source of truth
▲ │
└── invalidation ─┘
Each arrow is a policy boundary: define the cache key, who may reuse the representation, the maximum stale age, and what happens when the next layer is unavailable.
3. Read and write patterns
Cache-aside
The application owns the algorithm: read cache, fetch the source on miss, then populate the cache. It is simple and keeps unused data out of the cache, but the first request pays miss latency and concurrent misses can stampede.
request A ──GET key──▶ cache: MISS
request B ──GET key──▶ cache: MISS
│
without coalescing: two source reads
│
with single-flight: A refreshes; B waits
▼
source of truth
Read-through
The cache or data-access layer loads misses behind a uniform API. Application code is simpler, but the caching layer is now more tightly coupled to source access and can become part of the dependency's availability equation.
Write-through and write-around
Write-through updates the authoritative store and cache in the write path. It is intended to improve read-after-write behavior, but it does not guarantee that the two systems are never inconsistent: partial failures, concurrent writers, ordering, replication lag, and retries still need a defined protocol. Prefer committing the source of truth first, then updating or invalidating the cache with idempotent retry.
Write-around commits only to the source and lets a later read populate the cache. It avoids polluting the cache with write-once data, at the cost of a miss on the first subsequent read.
Write-behind/write-back
The cache acknowledges first and flushes writes asynchronously. This can batch work and reduce caller latency, but the cache or its pending-write log temporarily becomes authoritative. Durability, ordering, replay, idempotency, backpressure, and recovery are mandatory design questions. Redis itself supports multiple persistence modes; choosing a mode is an explicit performance-versus-data-loss decision, not proof that every cached write is durable.
4. Freshness and invalidation
Translate product language into a contract: “prices may be at most 30 seconds old,” “permission revocation must be visible immediately,” or “serve the last known catalog for five minutes during an outage.” Then choose mechanisms:
- TTL: simple upper bound when stale data is acceptable; it does not detect a change immediately.
- Explicit invalidation/update: react to a successful source write; needs retry and missed-event repair.
- Versioned keys: put a content/schema version in the key so new deployments cannot read incompatible entries.
- Soft/hard TTL: refresh after the soft deadline but allow the last good value until the hard deadline during a controlled failure.
- Validators: HTTP ETag or Last-Modified lets a cache revalidate without downloading an unchanged body.
Use TTL as a safety net even with events. Redis keyspace notifications use Pub/Sub's fire-and-forget delivery, so a disconnected subscriber misses events; a durable queue/outbox or reconciliation job is required when loss would violate the freshness contract. An event-driven path reduces normal invalidation lag, but it is not literally instantaneous and still needs monitoring, retries, ordering rules, and idempotent consumers.
The classic cache-aside write race is:
- reader misses and begins loading version 7;
- writer commits version 8 and invalidates the cache;
- slow reader stores version 7 after the invalidation.
Mitigate with version checks, compare-and-set, short bounded TTLs, or a design where stale refreshes cannot overwrite a newer version. “Delete after database write” is necessary but not sufficient under concurrency.
5. Stampedes, locks, hot keys, and cold starts
When a popular entry expires, thousands of callers can miss together. Use a combination of:
- single-flight/request coalescing so one refresh is in flight per key;
- TTL jitter so many keys do not expire at the same instant;
- refresh-ahead or stale-while-revalidate for predictable hot entries;
- negative caching with a shorter TTL for repeated not-found or controlled error results; and
- backpressure/load shedding so cache failure cannot send unbounded traffic to the source.
A distributed lock is not the whole solution. The lease can expire while a slow worker still computes, the lock service can fail, and a paused old holder can overwrite a newer value. Use short ownership scope, safe release, bounded waits, and—where stale writers would be harmful—monotonic versions or fencing tokens.
Hot keys can saturate one cache shard even when total memory and cluster throughput look healthy. Detect per-key or per-partition concentration, replicate or locally memoize safe hot values, split aggregations when possible, and rate-limit pathological callers.
Cold starts occur after restart, resharding, fleet expansion, or mass invalidation. Warm only the demonstrably hot working set, ramp traffic gradually, and test with caching disabled. Blindly preloading the entire dataset delays recovery and may evict the useful subset.
6. Eviction, admission, sizing, and key design
TTL controls freshness; eviction controls what leaves under memory pressure. LRU favors recent reuse, LFU favors sustained frequency, and neither is universally best. Redis implements configurable policies and approximate sampling, so validate policy behavior with the real access distribution.
Admission matters as much as eviction. A one-time scan can flood a cache with entries that will never be reused, evicting the hot working set. Consider minimum-frequency admission, separate pools, or bypassing the cache for scan-shaped traffic.
Keys must include every representation dimension: tenant/user, authorization scope, locale, currency, API or schema version, feature flag/experiment, and normalized query parameters where relevant. Omitting one can leak data or serve the wrong variant. Including unbounded raw input can create cardinality and memory attacks, so canonicalize and bound key space.
7. HTTP caching without privacy incidents
HTTP caching follows RFC 9111:
max-agecontrols freshness generally;s-maxageoverrides it for shared caches.privatepermits private-cache storage but prevents shared-cache reuse;no-storeprohibits storage.no-cachemeans a stored response must be validated before reuse—it does not mean “never store.”Varyadds named request headers to response selection. Varying on a high-cardinality header can destroy hit rate.ETagwithIf-None-Matchcan produce304 Not Modified, saving body transfer while still contacting the origin.stale-while-revalidateandstale-if-errorare extensions defined by RFC 5861.
Shared caching of authenticated or personalized responses requires deliberate directives and a correct cache key. Use private for user-specific responses that a private cache may retain; add no-store when no compliant cache should retain them. These directives govern cache behavior and do not replace transport or application security. Never assume HTTPS alone prevents an intermediary or CDN configuration from caching a response.
8. Observability and economics
Track metrics by endpoint and key class, not only globally:
- request and byte hit rate;
- hit latency, miss latency, and miss penalty;
- source request rate, saturation, errors, and throttling;
- memory utilization, item count, evictions, expirations, and rejected writes;
- hot-key/shard concentration;
- refresh duration, coalesced waiters, lock contention, and refresh failures;
- stale responses, negative hits, invalidation lag, and age at serve time.
A high hit rate can still be poor: tiny objects may hit while large expensive objects miss, or hits may save little work. Compare source capacity/cost and end-to-end latency before and after. Include cache fleet cost and operational burden in the result.
9. Failure and recovery
Choose per operation whether to fail open or fail closed. Public product descriptions may serve stale; authorization and account-balance decisions usually must not. If the cache is unavailable, an unconstrained “just query the database” fallback can amplify the incident. Bound fallback concurrency, shed noncritical load, use circuit breakers, and preserve source headroom.
A production runbook should cover node loss, full-cluster loss, resharding, mass expiry, poisoned serialization, missed invalidations, reconnect storms, and cold recovery. Exercise cache-disabled load tests and confirm observability still works when the cache itself is the failing dependency.
10. Interview decision framework
For any caching design, state:
- Goal: which measured latency, throughput, availability, or cost problem is being solved?
- Reuse: which keys repeat, and what is the working-set distribution?
- Freshness: how stale may each data class be, including failures?
- Placement/pattern: which layer and read/write strategy match the ownership model?
- Correctness/security: what belongs in the key, and how are races and invalidations handled?
- Protection: what prevents stampedes, hot keys, and source overload?
- Operations: which metrics, capacity tests, failure drills, and rollback plan prove the cache helps?
A senior answer does not say “use Redis with a five-minute TTL.” It makes the consistency and failure contract explicit, calculates the miss-path load, and explains what evidence would justify changing the design.
