Advanced 11 minOptimization

Strategies for chunking

Direct response

For a local RAG system, start with chunks of 300 to 500 tokens and 0 to 10% overlap, split at paragraphs and headings rather than by character, then measure performance on your own questions before adjusting. Semantic chunking has not proven worth its cost: a 2024 study concludes that its gains do not justify the additional computation. The setting that matters most is splitting at the text’s natural boundaries.

How you split documents into chunks determines what search can retrieve: a chunk cut at the wrong point never contains the complete answer, while an overly large chunk buries the answer in noise. This guide compares common strategies, provides starting sizes backed by published studies, highlights the units trap (tokens or characters), and proposes a protocol for validating your choice on your documents.

By Mohamed Meguedmi·Update 2026-09-30·Tested on Windows, macOS, and Linux

#Why chunking matters as much as the embedding model

Chunking is the operation that splits a document into passages before indexing them: each passage receives a vector, and it is a passage, never the entire document, that will be retrieved and then passed to the language model. Everything that follows depends on this split. If the sentence answering the question is split between two passages, neither contains it in full and the search misses it; if a passage mixes three topics, its vector is a vague average that does not strongly resemble any question. A more powerful embedding cannot fix faulty chunking, because it can only vectorize what it is given.

A Chroma study on evaluating chunking strategies shows this: heuristic methods such as RecursiveCharacterTextSplitter often produce good results in practice when properly configured, and results vary significantly depending on the selected chunk size and overlap. This guide therefore promises no general quantified improvement: the only reliable approach is to measure on your documents.

i
A good chunk
A good chunk contains one complete idea that can be understood on its own. If it's too small, the idea gets cut off and the vector is vague; if it's too large, several ideas get mixed together and the search returns noise.

#What chunk size should you choose?

The Local RAG Kit

Your documents, your AI: a reliable local RAG over your PDFs, notes and mail — nothing leaves your machine.

  • Lifetime online access
  • PDF + files
  • Lifetime updates

There is no universal size, but there are useful guidelines. In Chroma's study, the recursive strategy outperforms token-based splitting for sizes of 400 tokens or less without overlap, while behavior differs for larger sizes and with overlap. OpenAI's file search tool default configuration, 800 tokens with 400 overlap, achieves slightly below-average recall in this benchmark and the lowest scores on the other metrics. The guidelines below are derived from these findings, without claiming exactness for your corpus.

Starting sizes by document and question type
SizeSuitable forRisk
200 to 300 tokensPrecise factual questions (a date, a clause, a value) in dense documentsLoses the context around the answer; consider including the section title
300 to 500 tokensA general starting point for documentation, procedures, and articlesLow risk; this is the area to test first
500 to 800 tokensArgumentative texts whose ideas span multiple paragraphsThe vector dilutes several ideas; the prompt fills up faster
More than 1,000 tokensRarely suitable for retrieval; reserve it for a parent level in a hierarchical splitToo-general vector, passages that are difficult to sort
→
Starting point
Start at 400 tokens with an overlap of 0 to 50 tokens, test 250 and 700, and don’t change anything until you’ve measured recall on your questions.

#Tokens or characters: the units trap

Libraries do not all use the same unit. LlamaIndex's SentenceSplitter expresses chunk_size and chunk_overlap in tokens, with 1024 and 200 as the defaults according to its documentation. Other tools, including LangChain's RecursiveCharacterTextSplitter, count characters by default; check the length_function parameter in your version. Setting “500” in either one produces chunks whose size varies by a factor of roughly four. Second constraint: the embedding model has its own maximum input length. The BGE-M3 model handles inputs of up to 8,192 tokens according to its fact sheet; other, older, or smaller embedding models accept significantly fewer, and beyond that limit the end of the passage is ignored. Check your model's limit before choosing the size, and count with the model's tokenizer rather than by eye.

#Overlap: useful, but not always

Overlap repeats the end of one chunk at the beginning of the next, so a sentence cut at a boundary remains readable on at least one side. It has a cost: more chunks, a larger index, and duplicates in the results. The Chroma study notes that reducing overlap improves the IoU score, a metric that penalizes redundant information. Overlap makes sense if you split at a fixed size; it matters less if you already split at paragraphs and headings, since cuts then fall at natural boundaries.

0 tokens
Sufficient with paragraph- or structure-based chunking; the most economical option.
10 to 15% of the size
A good compromise when splitting on sentences or characters.
More than 25%
Rarely justified: lots of redundancy, with nearly identical passages in the results.

#Chunking strategies, from simplest to most expensive

#1. By fixed character count

Cut every N characters or tokens, without looking at the text. The simplest—and most destructive—option: it cuts through the middle of words, sentences, and tables. Reserve it for prototypes.

#2. By paragraphs or by sentences

Split on double line breaks or punctuation, then regroup units until reaching the target size. A clear improvement at zero cost, since each chunk starts and ends at a natural boundary.

#3. Recursive

We try the largest separators first (paragraph, line, sentence, space), and only split at the character level as a last resort. This is the default behavior in the main frameworks and a solid baseline: the Chroma study finds that this type of splitting, when properly configured, often performs well.

#4. Depending on the document structure

We preserve headings, lists, tables, and code blocks. LlamaIndex's MarkdownNodeParser, for example, splits by headings and attaches to each node the path of headings leading to it, providing context that the text alone does not have. This is the best choice for technical documentation, wikis, and exported HTML pages.

#5. Semantics

We calculate one embedding per sentence, then split where similarity drops between two neighboring sentences. In LlamaIndex, SemanticSplitterNodeParser takes a buffer_size (the number of sentences compared together, 1 by default) and a breakpoint_percentile_threshold (95 by default; a lower value creates more nodes). The cost is an additional embedding calculation during indexing, and the benefit is unproven: an October 2024 study of three retrieval tasks concluded that the computational cost of semantic splitting was not justified by consistent performance gains. Test it on your corpus before adopting it.

#6. Hierarchical (parent and child)

We index small chunks for search accuracy, then send the larger parent block back to the model for context. LlamaIndex’s HierarchicalNodeParser produces this kind of hierarchy, for example at three levels of 2048, 512, and 128 tokens according to its documentation. This addresses long passages: neither small-only nor large-only.

Which strategy for which document
Document typeRecommended strategy
Documentation, wiki, Markdown, HTMLStructure (headings), then recursive splitting within long sections
Contracts, legal textsStructure by article or clause; 300 to 500 tokens in size; article title repeated in each chunk
Unstructured free-form prose (emails, notes, transcripts)Recursive, 10% overlap, optionally hierarchical
PDF with tablesExtract the structure upstream, with one table per chunk or per row depending on the questions
Source codeSplit by function or class, never in the middle of a block

#Implementations with LlamaIndex

Sentence-by-sentence splitting with overlap
from llama_index.core.node_parser import SentenceSplitter

splitter = SentenceSplitter(chunk_size=400, chunk_overlap=40)  # en tokens
nodes = splitter.get_nodes_from_documents(docs)
Depending on the Markdown structure
from llama_index.core.node_parser import MarkdownNodeParser

parser = MarkdownNodeParser()
nodes = parser.get_nodes_from_documents(docs)
# le chemin des titres est stocké dans les métadonnées de chaque nœud
Hierarchical
from llama_index.core.node_parser import HierarchicalNodeParser, get_leaf_nodes

parser = HierarchicalNodeParser.from_defaults(chunk_sizes=[2048, 512, 128])
nodes = parser.get_nodes_from_documents(docs)
leaves = get_leaf_nodes(nodes)  # ce sont les feuilles qu'on vectorise
Semantics (test before adopting)
from llama_index.core.node_parser import SemanticSplitterNodeParser
from llama_index.embeddings.huggingface import HuggingFaceEmbedding

embed = HuggingFaceEmbedding(model_name="BAAI/bge-m3")
splitter = SemanticSplitterNodeParser(embed_model=embed, buffer_size=1, breakpoint_percentile_threshold=95)
nodes = splitter.get_nodes_from_documents(docs)

#Restoring context for each chunk

A passage torn from its document loses its bearings: “The deadline is 30 days” does not say what it refers to. Three simple practices address this. Prefix each chunk with the document title and section path; store this information in metadata for filtering (by date, source, or document type); and, for passages that begin with a pronoun or a reference (“this clause”), consider the parent level of the hierarchical split. A single line of context before the text is often enough, and it costs far less than changing models.

#Evaluate your chunking before locking it in

  1. 01
    Write 30 to 50 real questions
    Questions your users would ask, each with the expected source passage; write them before looking at the results.
  2. 02
    Index with two or three configurations
    For example, 250, 400, and 700 tokens, with overlap of 0 and 10%. Keep the other parameters identical.
  3. 03
    Measure recall in the top 5 results
    For each question, does the expected passage appear in the first 5 results? The percentage gives recall at 5.
  4. 04
    Also consider accuracy and redundancy
    An identical prompt with more varied results is a better setting. Count duplicates in the top 5.
  5. 05
    Read the failures one by one
    For each failed question, open the chunk that should have answered it: truncated, too broad, broken extraction? The cause determines the fix.

#Common pitfalls

Broken tables
A PDF extractor that flattens a table produces meaningless cell rows: extract the structure beforehand, before chunking.
Truncated code blocks
A splitter without syntax awareness cuts a block in the middle; use structure-aware splitting.
Mixed documents
A chunk that is half French and half English produces an unhelpful average vector: split by language if the corpus is mixed.
Nearly empty chunks
A standalone title (“3.2.1 Obligations”) without the following text is noise: filter out chunks that are too short or merge them with the next one.
Change chunking without reindexing
Chunking is fixed at indexing time: any change requires recalculating the vectors. Plan for a reindexing script from the start.
!
Don't optimize blindly
Without a question set, you can’t tell whether a setting improves or degrades performance. A size that “seems reasonable” has a measurable effect, in either direction, on recall.

#Frequently asked questions about chunking

FAQ
What chunk size should you use for a local RAG?+
Start between 300 and 500 tokens, with an overlap of 0 to 10%, and split on paragraphs or headings. Then test 250 and 700 on 30 to 50 of your real questions, comparing recall in the 5 top results. There is no universal value: it depends on your documents and questions.
Is semantic chunking worth the cost?+
Not systematically. It requires an additional embedding computation during indexing, and an October 2024 study on three retrieval tasks concluded that this cost was not justified by consistent gains over fixed-size chunking. Test it on your corpus: if it doesn't win, keep the recursive approach.
Do chunks need to overlap?+
Not always. If you already split on paragraphs and headings, zero overlap is often sufficient. If you split by fixed size or sentences, 10 to 15% avoids losing a sentence at the boundary. High overlap multiplies duplicates in the results and inflates the index.
Tokens or characters: how should you configure chunk_size?+
Check your library's unit. LlamaIndex's SentenceSplitter counts tokens, while other tools count characters by default, changing the actual size by a factor of about four. Also check your embedding model's maximum input length: beyond that, the passage is truncated during indexing.
How do you split a PDF with tables?+
First extract the structure with a tool that recognizes tables, then split it: one table per chunk if it is small, or one table row per chunk accompanied by the header if it is large. An extractor that flattens the table into plain text produces unusable passages, regardless of how they are split afterward.
Do you need to reindex when changing chunk size?+
Yes, always. Vectors are calculated from passages as they were at indexing time: changing the chunk size, overlap, or splitter requires recalculating all vectors. Keep a script that rebuilds the index from the source documents, and version the chunking parameters used.
Did this guide help you?

Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.