RAG Chunking Strategies
Retrieval-augmented generation works by finding the passages most relevant to a question and giving them to a language model. Before anything can be retrieved, documents must be split into chunks: pieces small enough to embed precisely and fit into the model's context window, but large enough to carry meaning on their own. Chunking is one of the most influential, and most overlooked, decisions in a RAG pipeline.
Chunks that are too large dilute embeddings and waste context. Chunks that are too small lose the surrounding meaning ("it increased by 12%": what did?). Splits that ignore document structure cut tables in half and separate headings from their content. Good chunking preserves meaning boundaries, carries useful metadata, and is validated against real queries.
TL;DR
- Chunk to balance retrieval precision (smaller) against self-contained context (larger). Typical starting points are 200–800 tokens with 10–20% overlap.
- Prefer structure-aware splitting (headings, paragraphs, list items, code blocks, table rows) over naive character counts.
- Recursive splitting tries paragraph, then sentence, then word boundaries until chunks fit.
- Semantic chunking splits where embedding similarity between sentences drops, which suits unstructured prose.
- Attach metadata (source, title, section path, dates, permissions) to every chunk for filtering and citations.
- Parent-child / small-to-big retrieval and contextual chunk headers get precise matching and rich context.
Quick Example
Structure-aware chunking of Markdown documentation with section context:
Each chunk stays within one section, carries its heading path for context and citations, and records who may see it.
Core Concepts
Why Chunk at All?
- Embedding precision: an embedding summarizes a whole chunk as one vector. A 10-page chunk averages many topics and matches everything weakly. See embeddings.
- Context budget: you can fit many focused chunks, from several documents, into a prompt, but only a few large ones.
- Citations: smaller units let answers point to the exact passage.
- Cost and latency: fewer irrelevant tokens per request.
Chunking Methods
Size and Overlap
There's no universal best size. It depends on the embedding model, document type, and question style:
- Factoid Q&A ("what's the refund window?") favors smaller chunks (100–300 tokens).
- Explanatory or synthesis questions favor larger chunks (500–1,000 tokens) or parent-document retrieval.
- Overlap (10–20%) prevents information at boundaries from being cut off in both neighbors. More overlap increases index size and duplicate results.
- Measure size in tokens with the embedding model's tokenizer, not characters. See tokenization.
Metadata
Every chunk should carry metadata used for filtering (product, version, language, date range, tenant), access control (only retrieve what the user may see), citations (URL, title, page, section), and freshness (updated timestamps for recency boosts). Metadata filters in the vector database often matter as much as embedding quality.
Parent-Child and Small-to-Big Retrieval
Index small chunks (sentences or short passages) for precise matching, but return their parent (the full section or a window of neighboring chunks) to the model. You get precise retrieval without starving the model of context. Variants include sentence-window retrieval and hierarchical indexes (summaries → sections → passages).
Contextual Retrieval
Chunks often lose meaning out of context ("The company's revenue grew 3% over the previous quarter": which company, which quarter?). Contextual chunk headers prepend the document title and section path. Contextual retrieval goes further: an LLM writes a short, chunk-specific context sentence for each chunk before embedding and keyword indexing. Anthropic's published results showed large reductions in retrieval failures, especially combined with hybrid search and reranking. Prompt caching of the full document keeps the generation cost reasonable.
Special Content Types
- Tables: keep rows with their header row; convert to Markdown or key-value text; consider one chunk per row for large reference tables.
- Code: split by function, class, or module using a syntax-aware splitter, and include file path and signatures.
- PDFs: use layout-aware parsers to avoid mixing columns, headers, and footers. Extract tables and figures separately, and OCR scanned pages.
- Conversations and transcripts: split by speaker turns or time windows, keeping speaker labels.
- FAQs: one question and answer per chunk, which is a naturally ideal unit.
Best Practices
Clean Before You Chunk
Strip navigation, boilerplate, cookie banners, repeated headers and footers, and tracking noise. Garbage text produces garbage embeddings and wastes tokens.
Keep Stable, Deterministic Chunk IDs
Derive IDs from document ID plus section or position, so re-ingestion updates or deletes the right vectors instead of duplicating them. Re-chunk only changed documents.
Evaluate Chunking Empirically
Build a set of real questions with known supporting passages, and compare chunking strategies by retrieval recall@k and answer quality. Intuition about "good chunk size" is often wrong for a given corpus. See RAG evaluation.
Match Chunks to the Embedding Model
Embedding models have maximum input lengths, and often perform best well below them. Chunks exceeding the limit get truncated silently.
Common Mistakes
Fixed Character Splits Through Structure
The exception list is separated from its rule and mixed with another section. Split on structure first, then size.
One Chunk per Document
Embedding whole long documents yields vague vectors that match many queries loosely and flood the context with irrelevant text. Split into focused passages.
Losing Access Control in the Index
Chunking a restricted document without copying its permissions into chunk metadata lets retrieval surface confidential passages to anyone. Enforce ACL filters at query time. See AI guardrails.
FAQ
What's the best chunk size for RAG?
It depends on your documents, questions, and embedding model. A common starting point is 300–600 tokens with 10–20% overlap, split on structural boundaries. Then evaluate a few sizes against real queries and pick what maximizes retrieval recall and answer quality.
Does chunk overlap help?
Usually a little. It reduces the chance that an answer spanning a boundary is split across two chunks that each score poorly. Too much overlap inflates the index and returns near-duplicate chunks. Structure-aware splitting reduces the need for overlap.
Is semantic chunking better than recursive chunking?
Not reliably. Semantic chunking can help with long unstructured text, but it costs more to compute, and studies show mixed results against well-tuned recursive or structure-aware splitting. Evaluate on your data before adopting it.
Do long-context models make chunking unnecessary?
Not entirely. For a few documents, you can often skip retrieval and send them whole. For large or changing corpora, retrieval over chunks remains far cheaper and faster, and chunking quality still determines what gets retrieved. See context windows.
Related Topics
- RAG — Retrieval-augmented generation overview
- Embeddings — Turning chunks into vectors
- Hybrid Search — Combining keyword and vector retrieval
- Reranking — Reordering retrieved chunks by relevance
- RAG Evaluation — Measuring whether chunking works
- Vector Databases — Storing chunks and metadata