Overview
Curated: · Written: · Reviewed:
Retrieval quality is a measurable evidence supply chain
Advanced RAG work begins after vector search returns plausible text. The production question is whether the system consistently finds the authorized, current, answer-bearing evidence under real query language, corpus scale, latency budgets and failure. Define a retrieval contract: target users and tasks, source-of-truth precedence, relevance unit, acceptable no-answer behavior, permissions, freshness, latency/cost, and release thresholds. A larger top-k or stronger generator cannot repair missing or forbidden evidence.
Build a versioned evaluation collection from representative queries. Preserve ordinary, hard, exact-identifier, paraphrase, multilingual, ambiguous, multi-hop, time-sensitive, conflicting-source and unanswerable cases. Have qualified annotators identify relevant documents and spans, graded relevance, required evidence sets and allowed source versions. Record disagreement and incomplete judgments. Group related queries before splits, protect a frozen release set, and inspect public benchmark and production-feedback contamination.
Chunking controls the unit available to rank. Preserve document hierarchy, headings, tables, code, lists, captions, page/section anchors and relationships. Fixed or recursive chunks are simple; semantic boundaries can improve coherence but depend on another model; parent-child and sentence-window retrieval combine precise matching with surrounding context. Measure span and document recall, redundancy, answer support, context tokens and latency by document type instead of choosing one universal size.
Overlap can protect facts at boundaries while multiplying near duplicates. Duplicate chunks crowd out source diversity, distort score distributions and consume context. Use stable content hashes plus semantic duplicate groups, retain canonical provenance, and decide whether copied policies represent one or several authorities. Diversity methods such as maximal marginal relevance can reduce redundancy, but may discard multiple independent evidence pieces; evaluate multi-hop completeness explicitly.
Lexical, learned sparse, dense and late-interaction retrieval have different strengths. Lexical methods retain exact names, IDs and rare terms. Dense dual encoders capture semantic paraphrases. Learned sparse representations add expansion while remaining indexable. Late interaction preserves token-level matching at higher storage or compute. Hybrid systems form union candidates and combine ranks through calibrated scores or reciprocal rank fusion; log each component contribution so failures are diagnosable.
Security filtering is part of retrieval, not post-processing. Derive tenant, subject, groups, attributes and policy version from authenticated context, and apply filters before or within every candidate generator, reranker, cache and snippet path. Unauthorized items must not displace permitted top-k, reveal score or metadata, enter evaluator logs, or reach the model. Test permission removal, shared resources, cross-tenant identifiers and policy-service failure, and measure revocation propagation.
Query transformations must preserve intent. Expansion can add lexical variants; multi-query retrieval broadens recall; decomposition targets multi-hop evidence; HyDE embeds a hypothetical answer. Each generated query can drop negation, version, entity, date or permission constraints and can amplify injected instructions. Retain the original, label every derivative, bound fan-out and recursion, deduplicate, attribute candidates, and fall back safely when transformation fails.
A reranker improves ordering only within its candidate set. Candidate-generation recall is therefore an upper bound. Specify model/revision, truncation, pair construction, score interpretation, batch and hardware. Cross-encoders model query-document interaction; ColBERT-style late interaction supports more precomputation. Calibrate cutoffs on the target domain, measure relevance and end-to-end outcomes, and test timeouts, long passages, multilingual text and distribution shift.
Contextual compression selects or summarizes parts of candidates to save tokens. It can omit qualifiers, provenance or contradictory evidence, so preserve exact source spans and evaluate claim support against originals. Context ordering also matters: models may underuse evidence placed in the middle. Test permutations and length boundaries, reserve output tokens, present source identity and authority, and state conflict and insufficient-evidence behavior.
Citations are a claim-to-evidence contract. Decompose verifiable claims, link them to exact immutable source versions and spans actually provided, and check entailment, contradiction, completeness and viewer authorization. A citation marker or high-ranked source is not proof. Conflicts require transparent attribution and precedence rules, not silent synthesis. Links may move; retain governed checksums or snapshots and propagate deletion obligations.
Freshness spans connector detection, parsing, chunking, embeddings, indexes, caches and serving aliases. Track source-update-to-query availability and permission/deletion revocation lag. Use idempotent events, immutable index versions, tombstones, atomic alias switches and reconciliation between source, chunk and index inventories. An embedding migration must rebuild a compatible parallel index; mixing vector spaces, preprocessing or distance conventions silently corrupts results.
Approximate indexes trade recall for latency, memory and build/update cost. Establish exact nearest-neighbor results on a representative sample, then sweep HNSW, IVF, probes, compression or other relevant parameters. Vector recall is not semantic relevance, so report both. Include filtered queries, concurrent traffic, inserts/deletes, restarts, fragmentation and corrupted artifacts. Preserve configuration, corpus checksum and embedding identity for rollback.
Evaluate layer by layer. Candidate retrieval uses recall@k and permission correctness; ranking uses MRR, nDCG and precision; context uses relevance, redundancy and required-evidence coverage; generation uses correctness, completeness, faithfulness, citation precision/coverage and abstention; end-to-end uses verified task outcomes, latency and cost. Segment by query/source/language/time/access/risk. Calibrate model judges against expert labels and keep critical disclosure or unsupported-action failures outside aggregate scores.
Operate with versioned traces: original and derived queries, filters and policy, candidates and component ranks, reranker, selected spans, context order, generator and prompt, citations, cache, latency, tokens, cost and outcome, under privacy and retention controls. Shadow or canary embedding, ranking, chunking and prompt changes against a pinned baseline, prevent cache/index mixing, define stop thresholds and preserve atomic rollback. Retrieval quality is maintained through fresh evidence, regression cases, incident response and continuous reconciliation.
Worked example: 12 error-code queries vs a dense-only rerank
Gold set: 100 support queries. Twelve of them are exact identifiers (E-4419, SKUs). Dense retrieval paraphrases well and misses codes.
| pipeline | recall@20 overall | recall@20 on 12 code queries | nDCG@10 after cross-encoder | p95 ms |
|---|---|---|---|---|
| dense k=20 | 0.81 | 0.17 (2/12) | 0.74 | 40 |
| BM25 k=20 | 0.62 | 1.00 (12/12) | 0.58 | 12 |
| RRF hybrid k=20 | 0.89 | 1.00 | 0.86 | 48 |
| rerank dense-only top-20 | 0.81 ceiling | 0.17 ceiling | 0.79 | 180 |
The reranker cannot invent E-4419 if candidate generation never retrieved it. Interview the recall ceiling, not the nDCG of a pretty rerank.
