Skip to content
Tech Interview Prep home
Technical interview guide

Agent Memory & Context Management

Giving agents state across turns and sessions despite a fixed, finite context window.

Read
45 min
Practice MCQs
25
Interview QA
25
Edition
v4
Editorial status
Reviewed

Scope: MemGPT, Generative Agents, Reflexion, RAG, Lost in the Middle, Sentence-BERT, HNSW, Model Cards, NIST Privacy Framework and GenAI Profile, and OWASP Prompt Injection/Sensitive Disclosure references reviewed 2026-09-04.

Overview

Curated: · Written: · Reviewed:

Memory is governed state, not an infinite transcript

Review status: rewritten after review — coverage and interview framing below bar.

An agent has no durable memory merely because a conversation feels continuous. The model sees a finite serialized context window; persistence exists only when the application writes state externally and retrieves it later. Every interview question about agent memory reduces to one mental model: memory is a retrieval corpus the agent writes into, governed like any other user data store — with schema, ACLs, retention, and a fixed read budget.

The weak answer sounds like: "we use a vector database so the agent remembers everything." That answer has no write policy, no deletion story, no budget arithmetic, and no failure modes. The strong answer treats memory as a system with a write path, a read path, a lifecycle, and an evaluation baseline.

The mental model: a fixed budget, an external store, two paths

Three facts drive everything else:

  1. The context window is a hard token budget. A 128k-token window is not "effectively unlimited" — attention over long contexts is uneven (performance degrades on retrieval-over-long-context tasks as input grows; the classic "needle in a haystack" demo is not a workload), and cost scales linearly with input tokens.
  2. Anything not in the current window is gone unless the application persisted it.
  3. What comes back into the window is chosen by your ranking code, not by the model. If it retrieves the wrong thing, the model confidently uses the wrong thing.

So the architecture is always the same shape:

user turn ──▶ [read path: scope → filter → rank → top-k] ──▶ assemble prompt
                                                              │
                                                              ▼
                                                          model + tools
                                                              │
              [write path: extract → score → dedupe → store] ◀─┘

Write path and read path are separate design problems, and the write path matters more. A perfect ranker cannot rescue a corpus full of raw conversation dumps and speculative inferences; a mediocre ranker over a curated corpus still works. Interviewers probe this: "what triggers a memory write?" is where most candidates hand-wave.

Taxonomy: which tier an agent reaches for

tierstoreslifetimeexampletypical store
workingcurrent objective, active plan, recent turns, validated intermediate resultsthis context window"step 3 of 5 done, tests passing"prompt itself / scratchpad
episodicparticular past events with time and provenanceacross sessions"on 2025-01-12 user cancelled plan X"event log / DB rows
semanticextracted facts, entities, preferencesacross sessions"user's deployment target is us-east-1"structured records + embedding index
proceduralapproved reusable methods, tool sequencesacross sessions"for this repo, always run pnpm test before commit"skill/workflow library

The tiers have different provenance, sensitivity, and verification requirements. A generated reflection ("the user seems anxious about deadlines") is an inference, not a user-stated fact — it belongs in a separate field with lower confidence semantics, or not at all. On a given task the agent reaches for working memory first (it's already in context), then semantic facts relevant to the current entities, then episodic history when the task references the past, then procedural memory when the task resembles a solved one. Retrieving all four every turn wastes budget and displaces current instructions.

The write path: policy determines memory quality

What triggers a write, in order of trustworthiness:

  • Explicit user statements ("remember that I prefer TypeScript") — highest consent, write directly.
  • Authoritative structured state (a completed transaction, a changed setting) — write from the system of record, not from the model's paraphrase.
  • After-turn extraction — a model pass over the turn proposes candidate memories; each candidate passes gates before storage.
  • Batch consolidation — periodic jobs merge/dedupe/decay the corpus.

The gates, in order:

def admit(candidate):
    if not candidate.purpose_allowed:
        return reject("no purpose")
    if candidate.sensitivity == "high" and not candidate.user_confirmed:
        return reject("high-impact fact needs confirmation")
    if dupes := find_duplicates(candidate, threshold=0.92):  # embedding + entity match
        return supersede(dupes, candidate)   # entity-level update, not append
    if candidate.confidence < 0.6:           # extraction model's self-score, discounted
        return reject("low confidence")
    return store(candidate)

Prefer user-confirmed stable preferences and tool-verified state over speculative personality traits. Separate observed, user-stated, tool-verified, and model-inferred fields in the schema — a later reader must be able to tell them apart. High-impact facts (payment details, medical data, anything that could gate an action) require authoritative verification and may be unsuitable for model-controlled writes entirely.

Every record needs immutable identity, subject and tenant, type, content, source event and span, writer and model/prompt version, timestamps, confidence semantics, sensitivity, consent/purpose, ACL, expiration, supersession and deletion status. Store canonical structured facts in a transactional database; embeddings are a retrieval index over governed records, never the system of record. A vector does not erase the sensitivity or deletion obligations of its source.

The read path: scope first, similarity never grants access

Retrieval begins with the current authenticated scope, not with the query. Apply tenant, subject, ACL, purpose, type, time and deletion filters before or within keyword/vector search, reranking and caching. Cache keys must include authorization and memory-index versions, or a stale cache leaks deleted records.

Ranking is a documented blend of relevance, recency, importance, confidence and task eligibility — calibrate on representative tasks, and don't let the model's self-assigned importance dominate. Too many recalled items crowd out current instructions and create confirmation loops. Cap results and tokens, and retain why each item was selected (for audit and for the ablation tests below).

A worked trace of the failure interviewers love (figures hypothetical but the arithmetic is real):

Session 1, user: "sure, auto-pay the $9.99 plan." Session 2, the agent attempts a $500 vendor payout.

what was storedretrieval into session 2$500 POST authorized?
raw quote treated as executable consentranks #1, cosine 0.91yes, if the action layer trusts memory
typed plan.autopay=true, SKU basic, amount $9.99retrieved, amount checked at executionno — $500 ≠ $9.99, destination not the merchant
user deleted the memory; tombstone + index drop0 hits after the 2 s serving SLAno

The lesson: similarity is ranking among eligible records; it does not mint rights. A stored sentence like "the user approved all future transfers" cannot replace current authenticated consent. The server derives identity and scope, validates each action, and checks authoritative postconditions. Personalization may influence presentation and defaults within policy — never permissions.

Context budget: a scheduling problem, not a summarization problem

Given a 128k-token window, a realistic allocation before the user's turn even arrives:

consumertokens (example)
system + developer instructions2,000
tool schemas (12 tools)6,000
retrieved memories (top-k=8)3,000
recent turns (sliding window, last 10)8,000
working state / scratchpad4,000
output reserve8,000
headroom left for tool results~97,000

One large tool result — a 40k-token log dump — eats 40% of that headroom. The decisions that matter:

  • Priority and eviction rules. Current trusted requirements and safety constraints outrank old conversational detail. Write the rules down; don't let whichever component appended last win.
  • Sliding window + compaction. Auto-compact stale turns into a checkpointed summary, but keep lineage to source events so a bad summary can be rebuilt. Summarization is a lossy transformation: distinguish quotes from generated conclusions, retain constraints, decisions and unresolved items, and never upgrade tentative content to fact.
  • Scratchpad and tool-result pruning. Truncate or summarize tool outputs visibly (the model should know it got a truncated result), never silently.
  • Position and length testing. Long context is not uniformly used; test that a constraint placed at token 90,000 still binds. Naive truncation silently drops exactly the constraints the agent still needs — usually the oldest system-level ones.

Deletion must propagate into derived summaries and embeddings, or compaction resurrects deleted data.

Retrieval vs memory vs fine-tuning

Three ways to give an agent persistent knowledge, and when each is right:

approachright whenwrong when
retrieval (docs, RAG)knowledge is shared, current, externalknowledge is user-specific or behavioral
long-term memoryknowledge is per-user, changes with use, must be deletableknowledge is stable and shared (that's retrieval)
fine-tuningstable behavioral patterns across all usersanything user-specific, anything that must be forgotten

Long-term memory is really a curated, agent-written retrieval corpus: hybrid search (keyword + vector), recency decay, entity-level updates — not a database of everything ever said. Fine-tuning cannot delete one user's data and cannot update per-user facts cheaply; retrieval can't learn a preference the user never wrote down in a document. Memory sits between them: per-user, mutable, governed.

Consolidation, forgetting, and conflict

Time changes truth. Preserve valid_from, observed_at, and superseded_at rather than overwriting history. New statements can confirm, refine or conflict with older state:

  • Entity-level update when the fact is about a known entity: the user's plan changed from basic to pro — supersede, don't append.
  • Append when the event itself is the fact: "user cancelled on 2025-03-01" — history matters.
  • Conflict policy when sources disagree: apply source authority (tool-verified beats model-inferred; newer user-stated beats older user-stated), and ask the user or consult the authoritative system when ambiguity matters. Never silently treat a recent model inference as stronger than an older user-confirmed preference.

Decay or expiration reduces retrieval eligibility; legal retention and deletion are independent concerns. Stale or contradictory memories corrupt later turns quietly — the model retrieves "user is in Germany" from March and bills a VAT rate the user's move in May invalidated. That's why stale-use and missing-recall need separate metrics: retrieving harmful stale memory and failing to retrieve relevant memory are different errors with different fixes.

Memory is untrusted at read time

Users, documents, tools and earlier model outputs can all store instructions designed to hijack later sessions. Defenses, in layers:

  • Render memory as delimited data with source and trust labels, not as instructions.
  • Scan and quarantine suspicious content — assuming detection can miss.
  • Prevent memory from changing system policy, permissions, tool schemas or confirmation requirements; limit capabilities so a successful injection has bounded impact.
  • Test cross-tenant identifiers, group changes, shared accounts, stale caches, score and snippet leakage, and policy outages. Fail closed when scope cannot be established.

Privacy, evaluation, and operations

Privacy controls cover collection, use, access, portability, correction, retention and deletion: minimize raw text and sensitive attributes, encrypt records and indexes, separate tenants, restrict operator access, audit reads and writes, prevent automatic training reuse. Deletion must tombstone serving state immediately, remove or rebuild keyword/vector indexes and caches, update summaries, honor backup policy, and verify with reconciliation probes.

Evaluate against a no-memory baseline — the question "is memory actually better than just asking again?" is the one most teams skip. Measure task completion, correct recall, precision, stale/conflicting use, hallucinated memory, over-personalization, context displacement, latency and cost, by task, user and language. Test consent and write policy, multi-session continuity, corrections, expiry, deletion, tenant isolation, injection, corrupted indexes, migrations, outages and recovery. Inspect stored state, not just final behavior — a plausible response can conceal a forbidden write.

Version the memory schema, extraction prompt/model, embedding, index, ranking, compaction and retention policy. Migrate through backfills with counts and checksums, dual reads or shadow evaluation, atomic aliases and rollback. Roll out to a small cohort with stop thresholds.

What interviewers probe, and the weak answers

Likely follow-ups, with the answers that fail:

  • "How do you decide what to write?" Weak: "the model decides." Strong: named gates — purpose, consent, sensitivity, confidence, dedup — and why write policy dominates read policy.
  • "What happens when a fact changes?" Weak: "we update it." Strong: entity-level supersession vs append, conflict policy by source authority, and the VAT example of stale-use as a distinct metric.
  • "How does deletion work with embeddings?" Weak: "delete the vector." Strong: tombstone first, then index rebuild, cache invalidation, summary propagation, backup policy, reconciliation probes.
  • "Why not fine-tune instead?" Weak: no answer. Strong: per-user mutability and the right to be forgotten rule out fine-tuning for user-specific knowledge.
  • "What breaks at scale?" Weak: nothing. Strong: context displacement as corpora grow, ranking drift as the corpus distribution shifts, cross-tenant leakage under ACL changes, and cost growth if every turn retrieves top-k=20.
  • "When should you not build memory at all?" Single-session tasks, high-sensitivity domains where any persistence is a liability, and workloads where retrieval over authoritative documents already answers everything.

A useful memory system is attributable, correctable, scoped, forgettable, and demonstrably better than simply asking again. If you can't show the baseline comparison, you've built a liability with a cosine score attached.