AI · RAG chunking

RAG chunking for complex PDFs: preserve structure, tables and context

A production guide to layout-aware PDF chunking, table handling, parent-child retrieval, overlap, contextual metadata and measurable RAG evaluation.

TOPIC HUBAI, RAG & Vector Search
Conceptual editorial illustration comparing structure-aware PDF chunks that preserve headings, tables and diagrams with arbitrary broken text strips; not a product interface.
An editorial interpretation of the topic, followed by a practical execution diagram.

A reliable PDF chunking strategy starts with document structure, not a universal token number. Parse headings, paragraphs, lists, tables, figures and page coordinates into typed elements; form retrievable child chunks inside those boundaries; retain parent, section and page lineage; then measure whether the required evidence appears for real questions. This matters for policy manuals, contracts, reports and technical specifications where a fixed window can separate an exception from its rule or detach a table value from its header. It is useful for engineers building production RAG, not for data that should be queried live from SQL or an API.

The referenced documentation was verified on 6 October 2026. Documented parser and splitter behavior is separated below from the engineering recommendations for storage, retrieval expansion and evaluation.

Classify the document before choosing a splitter

Do not send every file through one generic recursive splitter. A born-digital PDF with a clean heading hierarchy, a scanned invoice, a two-column research paper and a financial report with cross-page tables are different inputs. First classify whether text is selectable, whether OCR is required, whether reading order is reliable, whether tables carry critical facts and whether the document is primarily narrative, reference material or transactional data.

For clean Markdown or HTML, header-aware splitting can preserve a strong hierarchy. LangChain's documented Markdown splitter groups content by configured headers and keeps those headers in metadata; a second size-based splitter can run inside each group. For PDFs, the hierarchy must first be recovered from layout. If extraction produces the wrong reading order, tuning chunk size cannot repair the lost structure.

Parse layout into typed elements before chunking

The Unstructured chunking documentation distinguishes partitioning from chunking. Partitioning detects semantic elements such as titles, narrative text and tables. Chunking then combines consecutive whole elements and falls back to text splitting only when one element exceeds the hard maximum. This is materially different from flattening the entire PDF to text and cutting every fixed number of characters.

Microsoft's Document Layout skill similarly extracts document structure and can output Markdown sections with header hierarchy or text sections with page-location metadata. It supports common office and image formats, but the documentation warns that formats can produce different image behavior and that long processing can time out. Treat parser selection and OCR quality as observable stages with their own failure states.

Build a semantic element model with stable lineage

Store the parsed element stream before producing embeddings. Each element should have a stable document version, ordinal position, type, page number, bounding coordinates when available, section path and a pointer to the source artifact. Preserve original element IDs so a chunk can be rebuilt after changing size or embedding model without rerunning expensive OCR unnecessarily.

A practical model separates content identity from vector identity:

{
  "document_version": "manual:v17",
  "element_ids": ["e_184", "e_185"],
  "section_path": ["Returns", "Damaged items"],
  "pages": [12, 13],
  "chunk_kind": "narrative",
  "chunk_hash": "sha256:...",
  "embedding_version": "embed-v4"
}

This is an application schema, not a provider contract. It lets citations, reindexing, deletion and evaluation refer to the same evidence even when vectors are regenerated.

A reliable PDF ingestion path preserves layout elements first, creates retrievable child chunks with parent and page metadata, then expands only the evidence selected for context.
A reliable PDF ingestion path preserves layout elements first, creates retrievable child chunks with parent and page metadata, then expands only the evidence selected for context. Open for a larger view

Keep headings and section paths with every child chunk

A paragraph reading "This exception applies for 30 days" is hard to retrieve if the product, policy and subsection exist only several pages earlier. Preserve the heading path as metadata and usually include a compact textual form in the embedded representation. LangChain documents keeping header metadata while further constraining chunk size; Unstructured's `by_title` strategy closes a chunk when a new detected title begins.

Do not blindly trust every detected title. OCR can classify a short bullet or decorative label as a heading. Add validation for impossible hierarchy jumps, repeated running headers and sections containing only a few characters. Keep the raw extracted label, the normalized hierarchy and a confidence signal so low-confidence documents can enter review rather than silently creating hundreds of tiny chunks.

Treat tables as structured evidence, not ordinary paragraphs

Unstructured documents that a `Table` element is isolated rather than combined with surrounding narrative. That is a useful default because a table has its own schema. Preserve its caption, column headers, row headers, footnotes, page range and nearby section path. Store a readable representation for retrieval, but keep the parsed cells separately so the answer layer can verify exact values.

For a small reference table, a self-contained Markdown or HTML representation may be enough. For a wide or multi-page table, create row groups that repeat the required headers and carry a parent table ID. Never embed a cell value such as "12.4%" without the metric, period and entity that give it meaning. If users need aggregation, filtering or the latest balance across many rows, load the data into SQL or an analytical store and use an authorized query instead of asking semantic retrieval to calculate from flattened text.

Use parent-child chunks for long sections

A small child chunk is easier to retrieve precisely, while a larger parent section may be better evidence for generation. Store both identities: retrieve against compact children, then hydrate the approved parent or a bounded neighborhood after ranking. The parent can be a section, clause group or table, not necessarily the whole document.

Do not place both every child and every parent in the final context. That duplicates text, spends tokens and can overweight one source. Deduplicate by source range, prefer the smallest evidence that answers the question and expand only when the child depends on definitions, exceptions or labels outside its boundary. Keep expansion deterministic so the same retrieval result produces an explainable context package.

Add neighboring chunks after retrieval, not everywhere

Adjacency is valuable when a match lands in the middle of a multi-part explanation. Save previous and next chunk IDs inside the same document version and section. After retrieval and tenant filtering, pull one neighbor when the hit begins with a continuation, references a preceding definition or ends before a sentence or list completes. Give neighbors lower retrieval priority and retain the original hit that justified expansion.

Do not make adjacency cross a tenant, document version or section boundary. Do not expand every result automatically: five hits plus two neighbors each can turn into fifteen overlapping chunks and bury the relevant evidence. Measure whether expansion improves answer support and recall on questions that genuinely require multiple local passages.

Use overlap as an exception for oversized elements

Overlap cannot reconstruct a heading hierarchy or table schema. The Unstructured defaults apply overlap when an oversized semantic element itself must be text-split; its documentation warns that applying overlap between normal chunks can pollute otherwise clean boundaries. Start with no overlap between well-formed elements. Add a small overlap only within oversized narrative text or code where a boundary would otherwise cut a sentence, definition or block.

Record the original character or token ranges so duplicated text can be deduplicated after retrieval. Tune overlap with evaluation rather than intuition. Larger overlap increases index size and can fill top-k with near-duplicates, reducing source diversity without adding evidence.

Contextualize only chunks that are ambiguous alone

Anthropic's Contextual Retrieval prepends short, chunk-specific context before building semantic and BM25 indexes. Its example repairs a financial sentence that omits the company and quarter. This can help chunks whose local wording uses pronouns, abbreviations or section-dependent terms. It is not a substitute for recovering layout correctly.

Generate context from an authorized document version, keep it distinct from source text and store the prompt and model version. Never present generated context as a quotation or citation. For many PDFs, the deterministic section path, document title, date and table headers already supply enough context at lower cost and risk. Compare metadata-derived context with generated context on the same held-out questions before adopting another ingestion stage.

Know when search, SQL or a live API is better than RAG

Use semantic or hybrid search for discovery across narrative knowledge. Use lexical search for exact error codes, model numbers and clauses. Use SQL for calculations, grouping, ranges and authoritative tabular facts. Use a live API for inventory, order status or other rapidly changing operational state. RAG can explain the policy around a refund, but the current refund status belongs to the transaction system.

A useful boundary is reproducibility: if the answer must be recomputed exactly from typed rows, use a query engine and let the model explain the returned result. If the task is locating and synthesizing passages whose wording varies, retrieval is appropriate. A production assistant can use both, with application code choosing the authorized tool rather than letting retrieved text redefine permissions.

Evaluate chunking before changing embeddings

Create questions that expose boundaries: a rule plus its exception, a table value plus its header, a definition used on a later page, a cross-page list, a duplicated running header, a scanned paragraph and an unanswerable request. Label the smallest sufficient evidence range and any required structured lookup. Measure child recall at k, parent-evidence recall, duplicate ratio, citation page accuracy and the share of context tokens that support the answer.

Inspect failures by stage. If OCR lost a minus sign, retrieval tuning is irrelevant. If the correct child is retrieved but parent expansion omits the exception, fix context construction. If evidence is present but the answer is wrong, then evaluate generation. Keep the held-out set stable while changing one variable. The RAG evaluation guide covers this separation in more depth.

Production blueprint and rollout checklist

A dependable path is: classify file, parse layout, persist typed elements, validate reading order, form section-bounded children, isolate tables, attach lineage, embed and index, retrieve under tenant scope, rerank, expand approved parents or neighbors, deduplicate source ranges, construct cited context and evaluate the answer. Make every stage versioned and resumable.

Start with one document family instead of enabling every PDF. Compare the structure-aware baseline against a simple fixed-window baseline on the same questions. Track parsing failures, tiny-chunk rate, oversized-element splits, table extraction confidence, index growth, retrieval latency and citation correctness. Reindex from stored elements when changing chunk rules, and retire old chunk versions atomically so stale evidence does not remain searchable.

Connect the result to the full production RAG pipeline and Qdrant hybrid retrieval. The key design rule is simple: preserve the document's meaning before optimizing the vector index. A clean chunk is not merely short; it is independently retrievable, traceable to its source and expandable without crossing the security or semantic boundary.

Official references

These references document the tools discussed. Examples and design decisions are illustrative and should be adapted to the project and its versions.

Prepared by: Noor Yasser

FROM DECISION TO DELIVERY

Working through a similar engineering challenge?

I help teams turn architecture decisions into a clear scope and dependable, reviewable implementation.

Book a 30-minute callRelated serviceAI integrations & retrieval systemsRelevant projectMurshid