Strategies for chunking
For a local RAG system, start with chunks of 300 to 500 tokens and 0 to 10% overlap, split at paragraphs and headings rather than by character, then measure performance on your own questions before adjusting. Semantic chunking has not proven worth its cost: a 2024 study concludes that its gains do not justify the additional computation. The setting that matters most is splitting at the text’s natural boundaries.
How you split documents into chunks determines what search can retrieve: a chunk cut at the wrong point never contains the complete answer, while an overly large chunk buries the answer in noise. This guide compares common strategies, provides starting sizes backed by published studies, highlights the units trap (tokens or characters), and proposes a protocol for validating your choice on your documents.
#Why chunking matters as much as the embedding model
Chunking is the operation that splits a document into passages before indexing them: each passage receives a vector, and it is a passage, never the entire document, that will be retrieved and then passed to the language model. Everything that follows depends on this split. If the sentence answering the question is split between two passages, neither contains it in full and the search misses it; if a passage mixes three topics, its vector is a vague average that does not strongly resemble any question. A more powerful embedding cannot fix faulty chunking, because it can only vectorize what it is given.
A Chroma study on evaluating chunking strategies shows this: heuristic methods such as RecursiveCharacterTextSplitter often produce good results in practice when properly configured, and results vary significantly depending on the selected chunk size and overlap. This guide therefore promises no general quantified improvement: the only reliable approach is to measure on your documents.
#What chunk size should you choose?
Your documents, your AI: a reliable local RAG over your PDFs, notes and mail — nothing leaves your machine.
- Lifetime online access
- PDF + files
- Lifetime updates
There is no universal size, but there are useful guidelines. In Chroma's study, the recursive strategy outperforms token-based splitting for sizes of 400 tokens or less without overlap, while behavior differs for larger sizes and with overlap. OpenAI's file search tool default configuration, 800 tokens with 400 overlap, achieves slightly below-average recall in this benchmark and the lowest scores on the other metrics. The guidelines below are derived from these findings, without claiming exactness for your corpus.
| Size | Suitable for | Risk |
|---|---|---|
| 200 to 300 tokens | Precise factual questions (a date, a clause, a value) in dense documents | Loses the context around the answer; consider including the section title |
| 300 to 500 tokens | A general starting point for documentation, procedures, and articles | Low risk; this is the area to test first |
| 500 to 800 tokens | Argumentative texts whose ideas span multiple paragraphs | The vector dilutes several ideas; the prompt fills up faster |
| More than 1,000 tokens | Rarely suitable for retrieval; reserve it for a parent level in a hierarchical split | Too-general vector, passages that are difficult to sort |
#Tokens or characters: the units trap
Libraries do not all use the same unit. LlamaIndex's SentenceSplitter expresses chunk_size and chunk_overlap in tokens, with 1024 and 200 as the defaults according to its documentation. Other tools, including LangChain's RecursiveCharacterTextSplitter, count characters by default; check the length_function parameter in your version. Setting “500” in either one produces chunks whose size varies by a factor of roughly four. Second constraint: the embedding model has its own maximum input length. The BGE-M3 model handles inputs of up to 8,192 tokens according to its fact sheet; other, older, or smaller embedding models accept significantly fewer, and beyond that limit the end of the passage is ignored. Check your model's limit before choosing the size, and count with the model's tokenizer rather than by eye.
#Overlap: useful, but not always
Overlap repeats the end of one chunk at the beginning of the next, so a sentence cut at a boundary remains readable on at least one side. It has a cost: more chunks, a larger index, and duplicates in the results. The Chroma study notes that reducing overlap improves the IoU score, a metric that penalizes redundant information. Overlap makes sense if you split at a fixed size; it matters less if you already split at paragraphs and headings, since cuts then fall at natural boundaries.
- 0 tokens
- Sufficient with paragraph- or structure-based chunking; the most economical option.
- 10 to 15% of the size
- A good compromise when splitting on sentences or characters.
- More than 25%
- Rarely justified: lots of redundancy, with nearly identical passages in the results.
#Chunking strategies, from simplest to most expensive
#1. By fixed character count
Cut every N characters or tokens, without looking at the text. The simplest—and most destructive—option: it cuts through the middle of words, sentences, and tables. Reserve it for prototypes.
#2. By paragraphs or by sentences
Split on double line breaks or punctuation, then regroup units until reaching the target size. A clear improvement at zero cost, since each chunk starts and ends at a natural boundary.
#3. Recursive
We try the largest separators first (paragraph, line, sentence, space), and only split at the character level as a last resort. This is the default behavior in the main frameworks and a solid baseline: the Chroma study finds that this type of splitting, when properly configured, often performs well.
#4. Depending on the document structure
We preserve headings, lists, tables, and code blocks. LlamaIndex's MarkdownNodeParser, for example, splits by headings and attaches to each node the path of headings leading to it, providing context that the text alone does not have. This is the best choice for technical documentation, wikis, and exported HTML pages.
#5. Semantics
We calculate one embedding per sentence, then split where similarity drops between two neighboring sentences. In LlamaIndex, SemanticSplitterNodeParser takes a buffer_size (the number of sentences compared together, 1 by default) and a breakpoint_percentile_threshold (95 by default; a lower value creates more nodes). The cost is an additional embedding calculation during indexing, and the benefit is unproven: an October 2024 study of three retrieval tasks concluded that the computational cost of semantic splitting was not justified by consistent performance gains. Test it on your corpus before adopting it.
#6. Hierarchical (parent and child)
We index small chunks for search accuracy, then send the larger parent block back to the model for context. LlamaIndex’s HierarchicalNodeParser produces this kind of hierarchy, for example at three levels of 2048, 512, and 128 tokens according to its documentation. This addresses long passages: neither small-only nor large-only.
| Document type | Recommended strategy |
|---|---|
| Documentation, wiki, Markdown, HTML | Structure (headings), then recursive splitting within long sections |
| Contracts, legal texts | Structure by article or clause; 300 to 500 tokens in size; article title repeated in each chunk |
| Unstructured free-form prose (emails, notes, transcripts) | Recursive, 10% overlap, optionally hierarchical |
| PDF with tables | Extract the structure upstream, with one table per chunk or per row depending on the questions |
| Source code | Split by function or class, never in the middle of a block |
#Implementations with LlamaIndex
#Restoring context for each chunk
A passage torn from its document loses its bearings: “The deadline is 30 days” does not say what it refers to. Three simple practices address this. Prefix each chunk with the document title and section path; store this information in metadata for filtering (by date, source, or document type); and, for passages that begin with a pronoun or a reference (“this clause”), consider the parent level of the hierarchical split. A single line of context before the text is often enough, and it costs far less than changing models.
#Evaluate your chunking before locking it in
- 01Write 30 to 50 real questionsQuestions your users would ask, each with the expected source passage; write them before looking at the results.
- 02Index with two or three configurationsFor example, 250, 400, and 700 tokens, with overlap of 0 and 10%. Keep the other parameters identical.
- 03Measure recall in the top 5 resultsFor each question, does the expected passage appear in the first 5 results? The percentage gives recall at 5.
- 04Also consider accuracy and redundancyAn identical prompt with more varied results is a better setting. Count duplicates in the top 5.
- 05Read the failures one by oneFor each failed question, open the chunk that should have answered it: truncated, too broad, broken extraction? The cause determines the fix.
#Common pitfalls
- Broken tables
- A PDF extractor that flattens a table produces meaningless cell rows: extract the structure beforehand, before chunking.
- Truncated code blocks
- A splitter without syntax awareness cuts a block in the middle; use structure-aware splitting.
- Mixed documents
- A chunk that is half French and half English produces an unhelpful average vector: split by language if the corpus is mixed.
- Nearly empty chunks
- A standalone title (“3.2.1 Obligations”) without the following text is noise: filter out chunks that are too short or merge them with the next one.
- Change chunking without reindexing
- Chunking is fixed at indexing time: any change requires recalculating the vectors. Plan for a reindexing script from the start.
#Frequently asked questions about chunking
What chunk size should you use for a local RAG?+
Is semantic chunking worth the cost?+
Do chunks need to overlap?+
Tokens or characters: how should you configure chunk_size?+
How do you split a PDF with tables?+
Do you need to reindex when changing chunk size?+
- Add a reranker to your pipeline
- Hybrid BM25 + vector search
- The best French embedding models
- Local RAG with ChromaDB and Ollama
- LlamaIndex in practice
- Source: Chroma, Evaluating Chunking Strategies for Retrieval
- Source: Is Semantic Chunking Worth the Computational Cost?
- Source: LlamaIndex, node parsers
- Source: BAAI/bge-m3 Hugging Face page
Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.