Skip to content
Tech Interview Prep home
Technical interview guide

Caching Strategies

Production caching for technical interviews: placement, read/write patterns, freshness, stampedes, HTTP caching, observability, failure recovery, and decision tradeoffs.

Read
24 min
Practice MCQs
25
Interview QA
25
Edition
v4
Editorial status
Reviewed

Scope: HTTP caching per RFC 9111 and RFC 5861; Redis behavior referenced from current official documentation accessed 2026-08-30..

Overview

Curated: · Written: · Reviewed:

Key takeaways

  • A cache is justified only when requests reuse results and the product can define an acceptable freshness window.
  • Every cache creates two modes—hit and miss—and production capacity must survive a shift toward misses.
  • Freshness is a product requirement, not a TTL guessed by an engineer. Express it as bounded staleness and invalidation behavior.
  • Prevent stampedes before launch with request coalescing, jittered expiry, controlled refresh, and downstream backpressure.
  • Cache keys are security boundaries. Include every dimension that changes authorization or representation.
  • Measure backend load, miss penalty, stale responses, evictions, and hot-key concentration—not hit rate alone.

1. Why cache—and when not to

A cache stores a reusable representation closer to the caller or in a cheaper execution path. A hit can avoid database execution, a network round trip, serialization, remote API cost, or expensive computation. The business case is normally one or more of:

  • lower tail latency;
  • higher throughput from the protected dependency;
  • lower infrastructure or API cost; or
  • bounded availability during a dependency brownout by serving an explicitly acceptable stale value.

Caching helps only when the workload has reuse. If nearly every request has a unique key, entries expire before reuse, or correctness requires the latest committed value, the added cache lookup and invalidation machinery can make the system slower and less reliable. Start with a measured bottleneck and a freshness requirement; do not start with Redis.

A small latency and capacity model

Suppose a cache hit costs 2 ms, a miss costs 2 ms plus a 48 ms database call, and the hit rate is 90%:

expected latency = 0.90 × 2 ms + 0.10 × 50 ms = 6.8 ms

That average hides the operational risk. At 10,000 requests/second, the database receives roughly 1,000 requests/second at a 90% hit rate. A cold cache sends almost 10,000 requests/second. If the database was downsized around normal cache behavior, a cache restart can become a database outage. Capacity-test the miss path and cap fallback traffic.

Worked example: a 90% hit rate is still 1,000 source rps

10,000 requests/second, 2 ms hit, 50 ms miss.

cache statesource loadexpected latency
90% hit1,000 rps6.8 ms
cold / fail-open~10,000 rps50 ms
coalesced refresh of one hot key1 source readwaiters share that fill

The interview answer is the miss-path capacity, not the hit-rate slide.

2. Where caching can live

LayerBest fitMain tradeoff
Browser/private HTTP cacheUser-specific reusable responses and assetsHarder to invalidate centrally; device-local
CDN/shared HTTP cachePublic responses reused across many usersCache-key and privacy mistakes have large blast radius
Gateway/reverse proxyWhole API responses with protocol-level policyBecomes part of the request availability path
In-process cacheTiny, extremely hot, slow-changing dataPer-instance copies diverge and multiply cold-start load
Distributed cacheShared dynamic data across an application fleetNetwork hop plus another service to scale and protect
Database-adjacent/materialized resultExpensive queries or aggregatesDependency-aware refresh and invalidation complexity

A two-level design—small local L1 plus shared distributed L2—can reduce network latency without making every miss reach the database. It also creates a coherence problem: invalidating L2 does not automatically remove every L1 copy. Define propagation, maximum local TTL, and behavior when invalidation delivery is missed.

browser/private cache
        │
        ▼
CDN/shared HTTP cache ── public representations only
        │
        ▼
gateway/API ──▶ per-instance L1 ──▶ shared L2 ──▶ source of truth
                                      ▲                 │
                                      └── invalidation ─┘

Each arrow is a policy boundary: define the cache key, who may reuse the representation, the maximum stale age, and what happens when the next layer is unavailable.

3. Read and write patterns

Cache-aside

The application owns the algorithm: read cache, fetch the source on miss, then populate the cache. It is simple and keeps unused data out of the cache, but the first request pays miss latency and concurrent misses can stampede.

request A ──GET key──▶ cache: MISS
request B ──GET key──▶ cache: MISS
                     │
          without coalescing: two source reads
                     │
          with single-flight: A refreshes; B waits
                     ▼
                source of truth

Read-through

The cache or data-access layer loads misses behind a uniform API. Application code is simpler, but the caching layer is now more tightly coupled to source access and can become part of the dependency's availability equation.

Write-through and write-around

Write-through updates the authoritative store and cache in the write path. It is intended to improve read-after-write behavior, but it does not guarantee that the two systems are never inconsistent: partial failures, concurrent writers, ordering, replication lag, and retries still need a defined protocol. Prefer committing the source of truth first, then updating or invalidating the cache with idempotent retry.

Write-around commits only to the source and lets a later read populate the cache. It avoids polluting the cache with write-once data, at the cost of a miss on the first subsequent read.

Write-behind/write-back

The cache acknowledges first and flushes writes asynchronously. This can batch work and reduce caller latency, but the cache or its pending-write log temporarily becomes authoritative. Durability, ordering, replay, idempotency, backpressure, and recovery are mandatory design questions. Redis itself supports multiple persistence modes; choosing a mode is an explicit performance-versus-data-loss decision, not proof that every cached write is durable.

4. Freshness and invalidation

Translate product language into a contract: “prices may be at most 30 seconds old,” “permission revocation must be visible immediately,” or “serve the last known catalog for five minutes during an outage.” Then choose mechanisms:

  • TTL: simple upper bound when stale data is acceptable; it does not detect a change immediately.
  • Explicit invalidation/update: react to a successful source write; needs retry and missed-event repair.
  • Versioned keys: put a content/schema version in the key so new deployments cannot read incompatible entries.
  • Soft/hard TTL: refresh after the soft deadline but allow the last good value until the hard deadline during a controlled failure.
  • Validators: HTTP ETag or Last-Modified lets a cache revalidate without downloading an unchanged body.

Use TTL as a safety net even with events. Redis keyspace notifications use Pub/Sub's fire-and-forget delivery, so a disconnected subscriber misses events; a durable queue/outbox or reconciliation job is required when loss would violate the freshness contract. An event-driven path reduces normal invalidation lag, but it is not literally instantaneous and still needs monitoring, retries, ordering rules, and idempotent consumers.

The classic cache-aside write race is:

  1. reader misses and begins loading version 7;
  2. writer commits version 8 and invalidates the cache;
  3. slow reader stores version 7 after the invalidation.

Mitigate with version checks, compare-and-set, short bounded TTLs, or a design where stale refreshes cannot overwrite a newer version. “Delete after database write” is necessary but not sufficient under concurrency.

5. Stampedes, locks, hot keys, and cold starts

When a popular entry expires, thousands of callers can miss together. Use a combination of:

  • single-flight/request coalescing so one refresh is in flight per key;
  • TTL jitter so many keys do not expire at the same instant;
  • refresh-ahead or stale-while-revalidate for predictable hot entries;
  • negative caching with a shorter TTL for repeated not-found or controlled error results; and
  • backpressure/load shedding so cache failure cannot send unbounded traffic to the source.

A distributed lock is not the whole solution. The lease can expire while a slow worker still computes, the lock service can fail, and a paused old holder can overwrite a newer value. Use short ownership scope, safe release, bounded waits, and—where stale writers would be harmful—monotonic versions or fencing tokens.

Hot keys can saturate one cache shard even when total memory and cluster throughput look healthy. Detect per-key or per-partition concentration, replicate or locally memoize safe hot values, split aggregations when possible, and rate-limit pathological callers.

Cold starts occur after restart, resharding, fleet expansion, or mass invalidation. Warm only the demonstrably hot working set, ramp traffic gradually, and test with caching disabled. Blindly preloading the entire dataset delays recovery and may evict the useful subset.

6. Eviction, admission, sizing, and key design

TTL controls freshness; eviction controls what leaves under memory pressure. LRU favors recent reuse, LFU favors sustained frequency, and neither is universally best. Redis implements configurable policies and approximate sampling, so validate policy behavior with the real access distribution.

Admission matters as much as eviction. A one-time scan can flood a cache with entries that will never be reused, evicting the hot working set. Consider minimum-frequency admission, separate pools, or bypassing the cache for scan-shaped traffic.

Keys must include every representation dimension: tenant/user, authorization scope, locale, currency, API or schema version, feature flag/experiment, and normalized query parameters where relevant. Omitting one can leak data or serve the wrong variant. Including unbounded raw input can create cardinality and memory attacks, so canonicalize and bound key space.

7. HTTP caching without privacy incidents

HTTP caching follows RFC 9111:

  • max-age controls freshness generally; s-maxage overrides it for shared caches.
  • private permits private-cache storage but prevents shared-cache reuse; no-store prohibits storage.
  • no-cache means a stored response must be validated before reuse—it does not mean “never store.”
  • Vary adds named request headers to response selection. Varying on a high-cardinality header can destroy hit rate.
  • ETag with If-None-Match can produce 304 Not Modified, saving body transfer while still contacting the origin.
  • stale-while-revalidate and stale-if-error are extensions defined by RFC 5861.

Shared caching of authenticated or personalized responses requires deliberate directives and a correct cache key. Use private for user-specific responses that a private cache may retain; add no-store when no compliant cache should retain them. These directives govern cache behavior and do not replace transport or application security. Never assume HTTPS alone prevents an intermediary or CDN configuration from caching a response.

8. Observability and economics

Track metrics by endpoint and key class, not only globally:

  • request and byte hit rate;
  • hit latency, miss latency, and miss penalty;
  • source request rate, saturation, errors, and throttling;
  • memory utilization, item count, evictions, expirations, and rejected writes;
  • hot-key/shard concentration;
  • refresh duration, coalesced waiters, lock contention, and refresh failures;
  • stale responses, negative hits, invalidation lag, and age at serve time.

A high hit rate can still be poor: tiny objects may hit while large expensive objects miss, or hits may save little work. Compare source capacity/cost and end-to-end latency before and after. Include cache fleet cost and operational burden in the result.

9. Failure and recovery

Choose per operation whether to fail open or fail closed. Public product descriptions may serve stale; authorization and account-balance decisions usually must not. If the cache is unavailable, an unconstrained “just query the database” fallback can amplify the incident. Bound fallback concurrency, shed noncritical load, use circuit breakers, and preserve source headroom.

A production runbook should cover node loss, full-cluster loss, resharding, mass expiry, poisoned serialization, missed invalidations, reconnect storms, and cold recovery. Exercise cache-disabled load tests and confirm observability still works when the cache itself is the failing dependency.

10. Interview decision framework

For any caching design, state:

  1. Goal: which measured latency, throughput, availability, or cost problem is being solved?
  2. Reuse: which keys repeat, and what is the working-set distribution?
  3. Freshness: how stale may each data class be, including failures?
  4. Placement/pattern: which layer and read/write strategy match the ownership model?
  5. Correctness/security: what belongs in the key, and how are races and invalidations handled?
  6. Protection: what prevents stampedes, hot keys, and source overload?
  7. Operations: which metrics, capacity tests, failure drills, and rollback plan prove the cache helps?

A senior answer does not say “use Redis with a five-minute TTL.” It makes the consistency and failure contract explicit, calculates the miss-path load, and explains what evidence would justify changing the design.