Overview
Curated: · Written: · Reviewed:
Diagnose the earliest broken contract, not the most visible model output
A generic report that "the AI hallucinated" collapses at least six distinct layers — prompt construction, model decoding, retrieval, context assembly, the agent loop, and tool execution — and sends engineers toward random prompt edits. In an interview, this topic is a favorite probe for staff-level judgment: the interviewer wants to hear you classify before you change, name the cheapest experiment that separates two candidate causes, and refuse to accept "works on my prompt" as a diagnosis. A weak answer jumps straight to "I'd tweak the prompt" or recites a checklist with no decision logic; a strong answer walks a symptom down the stack and says which trace field settles each fork.
Start with symptom, impact, and containment
Troubleshooting begins with a precise symptom: which user, task, cohort and time; expected versus observed external state; severity; reproducibility; and what changed. Contain critical disclosure, unsafe action, authorization failure or duplicate side effects before experimenting. Preserve privacy-safe evidence and exact versions.
The layer-isolation method: classify, then run one cheap experiment
Classify the failure before changing anything. Common classes: requirements or expected-answer error; malformed or adversarial input; template variable or instruction conflict; model capability or sampling; truncation and structured-output validation; ingestion, retrieval, ranking or context; tool schema, authorization or provider; agent routing, loop or state; memory provenance; cache/version mixing; capacity, timeout and retry; evaluator or telemetry defects; external dependency change.
The method is a decision tree where each fork has a cheap separating experiment:
| Fork | Cheapest experiment | What settles it |
|---|---|---|
| Prompt vs. model | Replay the exact captured prompt against the same pinned model, then a known-good baseline model | Same failure on baseline → prompt/context; failure only on current model → model/decoding |
| Model vs. retrieval | Answer the question with the retrieved context pasted manually (no RAG pipeline) | Correct answer from pasted context → retrieval/ranking is fine, synthesis is broken; still wrong → retrieval miss or unsupported claim |
| Retrieval miss vs. ranking miss | Check recall@k: is the answer-bearing chunk in the candidate set at all? | Chunk absent from top-100 → recall problem (ingestion/embedding); present at rank 47 but not top-5 → ranking problem |
| Ranking vs. context assembly | Present the correct chunk to the model directly | Model answers correctly → truncation/ordering dropped it; still wrong → synthesis failure |
| Agent loop vs. tool layer | Replay the recorded tool call with the recorded arguments against the tool directly | Tool succeeds with same args → loop/orchestration bug; tool fails → schema/auth/provider issue |
| Tool vs. infra | Compare finish reason, latency percentile, and provider error codes against a pinned baseline call | 429/timeout spikes → capacity; deterministic 400 on the same args → contract bug |
Vary one layer at a time. A successful retry under different retrieval, prompt and model settings reveals nothing — you've changed three variables.
RAG failure signatures and the metrics that tell them apart
For RAG, ask first whether the answer-bearing authorized source was ingested, parsed, chunked and available in the exact serving index. Then inspect query filters and rewrites, candidate recall, lexical/dense contribution, reranking, deduplication, selected spans, context order and truncation. Finally test claim support and citations.
Each failure mode has a distinct signature in a trace:
| Failure | Trace signature | Metric evidence |
|---|---|---|
| Retrieval miss | Answer chunk absent from top-k candidates; query rewrite dropped the key term | recall@k ≈ 0 for gold chunk |
| Ranking miss | Gold chunk present at rank 47 of 100, never selected | recall@5 low, recall@100 fine; MRR collapses |
| Chunking/embedding mismatch | Gold chunk exists but the answer sentence was split across two chunks; or embedding model version differs between index and query | recall@k low only for split-answer queries; inventory shows mixed embedding spaces |
| Context-assembly truncation | Gold chunk retrieved but dropped by token budget; finish_reason length | context precision high at retrieval, low at assembly; token counts show chunk cut |
| Synthesis failure | Perfect retrieved chunk in context; model contradicts or ignores it | faithfulness score low despite context precision high |
Reranking cannot recover absent candidates, and a perfect retrieved chunk can still be ignored by generation. Interviewers probe exactly this: "recall is fine, faithfulness is bad — what do you fix?" The answer is generation/prompt, not retrieval.
Agent-loop failure modes and how they present in traces
For agents, replay the explicit graph: task contract, routing, dependencies, input/output artifacts, state versions, messages, budgets, tool calls, retries, cancellation and terminal criteria.
- Tool-call/schema errors and malformed arguments: trace shows the model emitting a call whose arguments fail server-side validation (wrong type, missing required field). Distinguish model error (arguments wrong against a correct schema) from schema drift (tool changed, model was never told).
- Infinite or thrashing loops: trace shows repeated identical tool calls or alternating plans with no state change. Look for a missing no-progress detector or a terminal criterion the model can never satisfy.
- Planning drift: step N+1's subgoal no longer serves step 1's contract; trace shows the accumulated context growing while the task contract never re-enters the window.
- State/memory corruption: a later step reads a value an earlier step never wrote, or two steps write the same key without versioning. Trace shows state-version regressions.
- Wrong-tool selection: two tools with overlapping schemas; the model picks the read-only one and the task silently no-ops.
- Unhandled tool timeouts: the trace shows a timeout followed by a blind retry. Transport failure does not prove no side effect — verify authoritative state before retrying, and use idempotency keys.
Non-determinism and reproducibility
The same input yields different outputs because of sampling (temperature > 0), tool nondeterminism (a search API returns fresh results), retrieval ties (equal scores resolved by arbitrary order), and concurrency (shared state between parallel steps). "Works on my prompt" is not a diagnosis — it is one sample from a distribution.
To make an incident reproducible: pin temperature to 0 and the seed where the provider exposes one, snapshot the exact inputs and retrieved chunks, capture the full trace with model/prompt versions, and replay against fixtures for external systems. Note that temperature 0 is still not fully deterministic on some providers (batching and floating-point reduction order vary), so treat near-determinism as "stable enough to bisect," not proof. Mutable aliases and missing tool results make failures irreproducible; record uncertainty rather than inventing the absent state.
Observability as the diagnostic instrument
A useful trace must capture: the rendered prompt (post-template, pre-send), retrieved chunks with scores and source IDs, every tool call with arguments and raw result, latency per span, token counts in/out, finish reason, and model/prompt/index versions. Redaction constraints: never log secrets or full PII prompts to broadly accessible stores; log hashes or sampled, access-controlled prompt snapshots instead.
Trace review maps directly onto the layer-isolation method: prompt span → template check; retrieval span → recall/ranking metrics; assembly span → token budget and order; tool spans → schema and postcondition; end-to-end latency percentiles → infra. A trace that only records the final completion is a symptom recorder, not a diagnostic instrument.
Worked example: "the model hallucinated a second $50 refund"
Ticket: customer charged twice. The first suggestion on the ticket is a prompt edit. Trace the actual calls:
| check | observation | diagnosis |
|---|---|---|
finish_reason + output tokens | length, JSON cut at "amount": | truncation — raise the output reserve, reject incomplete structures |
| gold policy chunk in top-20 | rank 47, dense-only retrieval | retrieval/ranking miss — fix hybrid recall before blaming generation |
| refund tool POST after a 2.1s timeout | first call returned 202; retry used a fresh idempotency key | duplicate spend — pin the key (refund:4419) and verify authoritative state before retry |
Three layers produced the same complaint. Classify, then change one thing.
Worked examples for the practice questions
The tables below attach the same checks to the first five questions; each remaining question carries its own table in its explanation and model answer.
Question 1
Question 2
Question 3
Question 4
Question 5
Fix, verify, and close
Fix the narrowest causal layer, rerun the original reproducer plus surrounding regression, safety, security and performance suites, and compare against the pinned baseline. Release through shadow or a small sticky canary with exact attribution, stop thresholds and rollback. During incidents: contain first, scope cohorts and versions, notify owners, preserve evidence, document root cause and detection gap, and add monitoring and tests. Troubleshooting is complete only when the expected invariant is restored and recurrence is observable.
Likely follow-ups to prepare: "How do you debug a failure that only reproduces in production?" (capture-then-replay with fixtures; sampling traces by cohort). "What if the evaluator itself is wrong?" (verify dataset version, labels, judge model/prompt/parser, calibration; compare semantic claims with deterministic checks). "When do you weaken an ACL to fix missing results?" (never — trace permission-revocation lag and index aliases instead).
