AI · RAG architecture

Production RAG: from chunks to grounded LLM answers

A production design for chunking, embeddings, Qdrant retrieval, reranking, context assembly, citations and evaluation—with clear alternatives to RAG.

TOPIC HUBAI, RAG & Vector Search
Conceptual production RAG pipeline showing documents becoming chunks and embeddings, passing through a vector database, retrieval and reranking gates, then forming cited context for an LLM.
An editorial interpretation of the topic, followed by a practical execution diagram.

A production RAG pipeline is not “put PDFs in a vector database and ask an LLM.” It is two versioned systems joined by an evidence contract: an offline path that parses, chunks, embeds and indexes authorized source material, and an online path that understands a query, retrieves candidates, reranks them, builds a bounded cited context and asks the model to answer only from that evidence. RAG is best when answers depend on large, changing, unstructured knowledge; SQL or ordinary search is usually better for exact structured facts, transactions and simple document lookup.

This guide uses the original RAG paper, official OpenAI, Anthropic, Qdrant and Voyage documentation reviewed on 5 October 2026. Provider examples illustrate interfaces and trade-offs, not universal benchmark promises. Anthropic's published Contextual Retrieval results, for example, are company experiments on its chosen datasets and configurations; teams should reproduce the evaluation on their own corpus.

Start with the decision: RAG, search or SQL

RAG earns its complexity when the user asks natural-language questions whose answers must be synthesized from many unstructured documents, when the knowledge changes faster than model training, and when provenance matters. Policies, manuals, contracts, support histories and technical documentation are common fits. The original RAG paper framed the approach as combining parametric model memory with explicit non-parametric memory, partly to make knowledge updateable and provide provenance.

Do not route every question through embeddings. Use SQL for “what is invoice 1042's paid amount?” because the schema, current value and exact predicate are known. Use a search engine for finding pages containing an error code, product name or quoted phrase. Use a deterministic API for account state and permissions. Use fine-tuning when the problem is behavior, style or a repeatable transformation rather than missing facts. If a small, stable corpus fits comfortably in the model context and latency and privacy allow it, passing the complete corpus can be simpler than retrieval.

A useful router classifies the request before retrieval: transactional or structured requests go to tools and SQL; navigational requests go to lexical search; knowledge synthesis goes to RAG; unsupported or sensitive requests stop. The router itself needs evaluation because a wrong route can be more damaging than a weak similarity score.

Split the architecture into offline and online paths

The offline path owns source ingestion. It accepts a document version, stores the original, extracts structured text, records headings and page locations, creates chunks, produces embeddings, writes searchable points and marks the version searchable only after every required artifact is durable. Each stage should be idempotent and resume-safe. A failed embedding batch must not leave half a document visible beside the previous complete version.

The online path owns the latency budget. It authenticates the requester, derives tenant and ACL scope, normalizes or rewrites the query when justified, runs one or more retrieval branches, fuses and reranks candidates, hydrates canonical text, assembles context, invokes the LLM and validates citations. Do not make the request path parse files or repair an index. Send those jobs to separate queues with independent concurrency because extraction, embedding and indexing have different memory, rate-limit and network profiles.

Join both paths with stable identifiers: `tenant_id`, `document_id`, `document_version`, `chunk_id`, section path, source location, ACL attributes, content checksum, embedding model and index version. The vector point is a retrieval projection, not the source of truth. Hydrate final text and current permissions from the authoritative store when the risk justifies it.

Chunk by meaning and structure, not one global number

Chunking decides what can be retrieved as one unit. Fixed token windows are a reasonable baseline, but they cut tables, procedures and definitions away from their headings. Prefer structure-aware boundaries: headings and paragraphs for prose; function or class boundaries for code; rows plus headers for tables; clauses for contracts; turns for conversations. Preserve the structural path as metadata even when the text stands alone.

A chunk should be small enough to represent one retrieval idea and large enough to answer a useful sub-question. Overlap can protect facts crossing boundaries, but large overlap multiplies embedding cost and returns near-duplicates. Instead of guessing, build a question set from real tasks and sweep chunk size, boundary rule and overlap. Measure whether at least one answer-bearing chunk appears in the candidate set.

Parent-child retrieval is useful when narrow chunks search well but the answer needs surrounding explanation: index the child, then hydrate its parent section or immediate neighbors after retrieval. Contextualization can also restore missing identity. Anthropic's Contextual Retrieval prepends a short chunk-specific description derived from the whole document before embedding and lexical indexing. Anthropic reported lower failed-retrieval rates in its experiments, but this is not a reason to generate context blindly: contextual text can introduce errors, cost and versioning work, so evaluate original and contextualized chunks separately.

A reliable RAG system separates offline indexing from online retrieval, then measures each gate before asking the LLM to answer.
A reliable RAG system separates offline indexing from online retrieval, then measures each gate before asking the LLM to answer. Open for a larger view

Treat embeddings as a versioned data contract

An embedding is produced by a specific model, input mode, preprocessing pipeline and output dimension. Changing any of those changes the vector space. Store `embedding_model`, dimensions and pipeline version with every point, and never query a collection with a vector from an incompatible model. A migration should build a new named vector or collection, backfill it, compare recall and latency, switch reads gradually and retain a rollback window.

The official OpenAI embeddings guide describes embeddings as vector representations useful for search, clustering, recommendations and classification. The production implication is to align the input type with the model contract: distinguish document and query modes when the provider exposes them, normalize text consistently, batch within provider limits, retry with idempotent writes and record failures per chunk. Never mark the document indexed until all intended chunks have the expected embedding version.

Deduplicate before paying for embeddings. A content checksum can reuse unchanged chunks across document versions, but include the contextual prefix and preprocessing version in the hash. Do not use the raw file name as identity. Encrypt originals and extracted text as required, and avoid sending documents to an embedding provider whose data handling is incompatible with the product's security obligations.

Model Qdrant around retrieval and isolation

A Qdrant collection defines vector size and distance behavior, so it should follow the embedding contract rather than the business entity count. Put point payload fields that affect filtering or provenance next to the vector: tenant, document and version IDs, content type, language, timestamps, ACL labels and source location. Create payload indexes for fields used in filters; a field existing in payload does not guarantee an efficient filter.

For multitenancy, never accept a tenant filter directly from the public request. Build it from authenticated scope and inject it into every query branch. Qdrant's multitenancy guidance recommends payload partitioning for many small tenants and documents a tenant-aware keyword index; larger tenants can use dedicated or tiered shards. The right topology depends on tenant size, isolation, noisy-neighbor risk and operations—not on a slogan that one collection or one collection per tenant is always correct.

Filters are security-sensitive query construction. Test dense, sparse, recommendation and scroll endpoints for the same enforced scope. If a query rewrite expands entities or time ranges, it must not expand authorization. Store canonical text outside the vector database when policy must be re-evaluated at read time, then hydrate selected `chunk_id` values through an authorized repository.

Retrieve broadly, then rerank narrowly

Dense retrieval captures semantic similarity but can miss identifiers and exact domain terms. Lexical search catches those strings but can miss paraphrases. Hybrid retrieval runs both and fuses their ranks, producing a candidate set with higher coverage. The previous Qdrant hybrid-search guide explains RRF, DBSF and payload filtering in detail; in this pipeline, the key is to treat fusion as candidate generation rather than the final answer context.

Rerankers score the query against each candidate with a richer interaction than independent embeddings. They are more expensive, so apply them to a bounded candidate set, not the whole corpus. The Voyage reranker documentation exposes query-plus-document relevance scoring; similar products have different limits and score meanings. Select the reranked context count using evaluation and token budget rather than copying a vendor's top-N and top-K values.

Query rewriting can resolve pronouns, expand abbreviations or create lexical alternatives, but preserve the original query and evaluate the rewrite as its own component. A model-generated rewrite can drift from the user's intent. Run exact-ID extraction before semantic rewriting, and combine filters deterministically. If initial candidates already reach the required recall and order, skip reranking to save latency and cost.

Build context as evidence, not a text dump

Context construction is a deterministic product component. Remove duplicate and near-duplicate chunks, prefer the newest authorized version, merge adjacent fragments only when they belong to the same source, and reserve tokens for instructions, user input and the answer. Order evidence by usefulness, not merely vector score; a short definition may precede a long supporting section. Keep stable citation IDs mapped to document version and location.

Separate instructions from retrieved text. Documents are untrusted data and may contain prompt-injection strings. Tell the model that retrieved content cannot override system rules or request tools. Strip active markup where practical, cap contribution per source to avoid one document dominating, and do not pass hidden ACL metadata unless the answer needs it. Context should include enough source metadata to render a human-usable citation without revealing internal storage paths.

An illustrative context builder can enforce the budget before the model call:

const candidates = await retrieve({ query, scope, dense: 40, sparse: 40 });
const ranked = await rerank(query, deduplicate(candidates));
const hydrated = await repository.loadAuthorized(scope, ranked.map(x => x.chunkId));

const context = buildContext(hydrated, {
  maxTokens: 7_000,
  maxPerDocument: 3,
  includeSource: ['documentId', 'version', 'page', 'section'],
});

return generateGroundedAnswer({ query, context, requireCitations: true });

The numbers are example budgets, not recommendations. Measure with the chosen model, corpus and service-level objective.

Make the answer cite and abstain

Ask the model to answer from supplied evidence, attach a citation to each material claim and say when the context is insufficient or conflicting. Citation syntax alone is not grounding: after generation, verify that every citation points to an included chunk and that quoted or factual claims are supported by that chunk. Invalid citation IDs should fail validation, not silently disappear.

Conflicts need policy. Prefer an effective-dated authoritative source when one exists; otherwise present the disagreement and cite both. Do not let similarity score decide which policy is legally current. For high-impact actions, the RAG answer should inform a user or propose a tool call, not directly mutate production state without deterministic validation and authorization.

Keep generation temperature, prompt version, model version and context IDs in the trace. Store enough to reproduce failures without retaining sensitive user content longer than allowed. The model response is a derived view; canonical facts remain in their source systems.

Evaluate each gate, not only the final prose

A fluent answer can hide a retrieval failure. Build a versioned evaluation set with real questions, expected answer-bearing documents, difficult negatives, unanswerable questions and permission boundaries. Split metrics by domain, language, tenant size and query type so an average does not conceal a broken class. OpenAI's evaluation guidance recommends task-specific evals, representative data and continuous evaluation as systems change.

Measure retrieval before generation: recall@k asks whether an answer-bearing chunk was retrieved; MRR or nDCG measures ranking quality; filter tests verify that forbidden points never appear. Then measure the reranker, context builder and answer separately: evidence precision, citation validity, supported-claim rate, abstention quality and task correctness. Track p50/p95 latency, tokens, embedding and reranking cost, queue age and index freshness alongside quality.

Run ablations. Compare lexical only, dense only, hybrid, hybrid plus reranking, neighboring-chunk expansion and contextualized chunks on the same questions. Change one component at a time. An improvement in retrieval recall can still reduce answer quality if it fills the context with distractors. Promote a version only when it meets quality, security, latency and cost gates.

Failure modes to design out

Common failures are silent: a parser drops tables; chunks lose their headings; query and document embeddings use different versions; a document is visible before its full index exists; a deleted source survives in Qdrant; a tenant filter is omitted from one branch; reranking times out and returns an unbounded fallback; context truncation removes citations; or the model confidently answers an unanswerable question. Give every stage an explicit state and observable reason.

Use blue-green index versions and an alias or application-side switch. Reconcile the authoritative document table against vector points and tombstone deletions. Cap candidates and context even during dependency degradation. Define fallbacks: lexical-only may be acceptable for an embedding outage; skipping reranking may be acceptable when candidate ranking is strong; generating without retrieval is usually not acceptable for a knowledge-grounded route.

Security tests must attempt cross-tenant retrieval, prompt injection inside documents, poisoned sources, oversized chunks and citation spoofing. Reliability tests must replay jobs, duplicate events, interrupt migrations and expire provider credentials. A pipeline is production-ready when it fails closed or degrades predictably—not merely when the happy-path demo answers one PDF correctly.

Production rollout checklist

Before launch, define the questions RAG should answer and the ones routed to SQL or search; version parser, chunker and embeddings; keep originals and canonical text; index only complete document versions; derive tenant and ACL filters from trusted identity; create Qdrant payload indexes; evaluate dense, lexical, fusion and reranking stages; build a deterministic token budget; preserve citation provenance; validate claims and abstention; trace component versions and latency; reconcile deletions; and maintain a representative, permission-aware eval set.

For focused follow-ups, read evaluating retrieval quality, Qdrant hybrid search, secure SaaS tenant context and PostgreSQL tenant isolation. These are the deeper contracts around the full pipeline.

Official references

These references document the tools discussed. Examples and design decisions are illustrative and should be adapted to the project and its versions.

Prepared by: Noor Yasser

FROM DECISION TO DELIVERY

Working through a similar engineering challenge?

I help teams turn architecture decisions into a clear scope and dependable, reviewable implementation.

Book a 30-minute callRelated serviceAI integrations & retrieval systemsRelevant projectMurshid