Add a reranker to your pipeline
A reranker is a cross-encoder that rereads each question-passage pair and reorders the candidates returned by vector search: retrieve broadly (20 to 100 passages), then keep the 3 to 5 best ones for the model. In French, BAAI/bge-reranker-v2-m3 (Apache 2.0 license, approximately 568 million parameters) is the simplest starting point, locally, on a GPU, or even on a CPU if the volume remains modest.
Embeddings retrieve passages that are close to a question, but not necessarily passages that answer it. The reranker fixes this weakness by judging each question-passage pair in a single computation. This guide explains when it is worth the cost, which model to choose for French, how to connect it in thirty lines, and how to verify on your own documents that it genuinely improves the results.
#Reranker: what it does in a RAG pipeline
A reranker takes the question and a passage, reads them together, and returns a relevance score; you then sort the candidates by descending score and pass only the best ones to the language model. In a local RAG pipeline, it sits between the vector database (ChromaDB, Qdrant, Weaviate) and the LLM: vector search retrieves, for example, 30 passages, and the reranker keeps 5. BAAI's official documentation describes it this way: unlike an embedding model, the reranker receives the question and document as input and directly produces a similarity score instead of a vector. The benefit is greatest when the correct answer is among the candidates but not at the top: passages that discuss the right general subject without answering the precise question occupy the top positions, and the model then generates an off-target or fabricated answer. If the correct answer isn't among the candidates, a reranker can't help: it doesn't search; it sorts.
An embedding produces a vector for each document, independently of the question asked, and the distance measures overall topical similarity. So two passages on the same subject can receive similar scores even though only one contains the answer. The cross-encoder looks at the pair as a whole: it can see whether the requested date, name, or condition appears in the passage. It's more precise for this local judgment and more expensive because it requires one passage through the transformer per pair instead of a single computation per document.
#Bi-encoder and cross-encoder: why combine them
Your documents, your AI: a reliable local RAG over your PDFs, notes and mail — nothing leaves your machine.
- Lifetime online access
- PDF + files
- Lifetime updates
- Bi-encoder (embedding)
- Encode the question and each document separately; document vectors are computed once during indexing, and search is reduced to a vector comparison. Fast and suitable for millions of passages.
- Cross-encoder (reranker)
- It encodes the question and passage together and outputs a score. No computation can be reused: it must be redone for every question and candidate. Precise, but impossible to apply to the entire corpus.
The Sentence Transformers documentation explains the reasoning: evaluating thousands or millions of pairs would be rather slow, so the retriever is used to produce a set of candidates, such as a hundred, which the cross-encoder re-ranks. This two-stage design is the same as that used by traditional search engines. It also explains the main tuning parameter: the number of candidates retrieved controls both recall—the wider the selection, the more likely it is to contain the right answer—and latency, since each additional candidate requires a pass through the cross-encoder.
#Which reranking model should you choose for French?
The choice comes down to three criteria: language, license, and memory. The size figures below come from the Hugging Face model cards; note that file size depends on weight precision: F32 for bge-reranker-v2-m3 (about 4 bytes per parameter), and F16 for mxbai.
| Model | Languages | Advertised size | License | Verdict for French |
|---|---|---|---|---|
| BAAI/bge-reranker-v2-m3 | Multilingual | 0.6 billion parameters, weights in F32 (about 2.3 GB) | Apache 2.0 | Default choice: multilingual, lightweight, with good tooling |
| BAAI/bge-reranker-v2-gemma | Multilingual | 3 billion parameters, F32 weights (about 10 GB) | Apache 2.0 | For demanding workloads with a dedicated GPU; Gemma-2B-based reranker |
| mixedbread-ai/mxbai-rerank-large-v1 | English | 0.4 billion parameters, F16 weights | Apache 2.0 | Avoid for a French corpus: the specifications list English |
| Cohere Rerank | Multilingual | Hosted service | Commercial | Irrelevant for a 100% local pipeline: the passages leave your machine |
The bge-reranker-v2-m3 card presents it as a lightweight reranker with strong multilingual capabilities that is easy to deploy and fast at inference; the bge-reranker-v2-gemma card targets multilingual contexts, with good results in both English and multilingual settings. Two useful corrections to what is often reported: the bge-reranker-v2-m3 file weighs more than 2 GB, not 560 MB (568 million parameters in F32), and mxbai-rerank-large-v1 is an English-language model and should not be chosen for French documents. Once loaded in half precision, the m3 model uses approximately 1.1 GB for the weights (568 million parameters × 2 bytes, calculated in this guide), plus the activations for the processed batch.
#The pipeline before and after the reranker
Two parameters govern the whole system: k_retrieve, the number of candidates returned by the vector database, and k_final, the number of passages sent to the LLM. A k_final of 3 to 5 suits most local models with 7 to 14 billion parameters; beyond that, you fill the context window without a clear benefit, and prompt-processing time increases. k_retrieve depends on corpus difficulty: 20 is a good first try, 50 to 100 if questions are vague or the corpus contains many similar passages. The context-window guide explains the cost of adding more passages to the prompt.
#Implementation: sentence-transformers, FlagEmbedding, llama.cpp
#With sentence-transformers
The CrossEncoder class loads the model and evaluates pairs. Its rank method directly accepts the question and document list and returns the best ones; the top_k parameter limits the number of results (without it, all documents are returned).
#With FlagEmbedding, the model authors' library
The model card uses the FlagEmbedding library. It states that the raw score can be mapped to a value between 0 and 1 using a sigmoid function with normalize=True, and that use_fp16=True speeds up computation at the cost of a slight quality reduction. Remember that the raw score has no absolute scale and is often negative for irrelevant passages; only the ranking matters unless you set a threshold.
#With llama.cpp, without Python
The llama.cpp server provides a reranking endpoint, disabled by default. The documentation says it requires a reranking model, cites bge-reranker-v2-m3 as an example, and runs with the --embedding and --pooling rank options. You need a GGUF version of the model. This is the option to choose if your stack is already built around llama.cpp and you want to avoid installing PyTorch. Check the exact option in the installed version: the documentation warns that this endpoint may change.
#With LlamaIndex
#Cost: latency, memory, passage length
No latency figure is guaranteed: it depends on the card, precision (FP16 or FP32), number of candidates, and passage length. Focus on the proportions. Reranking time grows linearly with the number of pairs: going from 20 to 100 candidates multiplies the workload by five. It also grows with passage length, since each pair is encoded in full. Measure on your machine with your passages instead of relying on a number from a blog: time one hundred real queries and look at the median and worst-case times.
- Number of candidates
- First lever. Start at 20, measure recall, and increase to 50 only if good answers are still being left out.
- Accuracy
- use_fp16=True (FlagEmbedding) or loading in half precision reduces memory usage and speeds up computation, with a slight performance drop depending on the model card.
- Batch size
- The sentence-transformers rank method processes 32 pairs per batch by default. Reduce this to 8 or 16 if memory is tight; increase it if the GPU is underutilized.
- Maximum length
- 512 tokens per pair (the max_length value in the official examples). A longer passage is truncated: if your chunks exceed this size, the end of the passage isn't read. Shorten the chunks before increasing the limit.
- CPU or GPU
- On CPU, the reranker works, but each request with 20 to 50 candidates takes seconds; that's acceptable for an internal document assistant, less so for interactive chat.
#Reranking and chunking: the two settings interact
A cross-encoder judges an entire passage. If the passage mixes three topics, its score will be average for the three corresponding questions; if it is too short, it loses the context that would allow it to be recognized as an answer. Splitting text into medium-sized passages with slight overlap gives the reranker enough material without exceeding its maximum length. If you change the chunk size, rerun the recall test: the best k_retrieve and the reranker gain change with it. The guide to chunking strategies explains these choices in detail.
Another interaction: hybrid search. When keyword search (BM25) and vector search are combined, the candidate pool is more diverse, giving the reranker more chances to find a good answer than either method alone would have had. The reranker is the final layer, and hybrid search is the second: they complement rather than replace each other.
#Measure the improvement on your documents
The improvement reported in blogs ranges from twofold to fivefold depending on the corpus, and no general figure applies to yours. On a highly structured corpus (clean technical documentation), vector search alone is already good; on a noisy corpus (emails, notes, poorly extracted PDFs), the gap is more pronounced. The only way to know is to evaluate it.
- 01Prepare 30 to 50 questionsUse real user questions, each with the passage containing the answer (an ID is enough).
- 02Measure recall at 5 without a rerankerFor each question, check whether the relevant passage appears in the vector database's top 5 results.
- 03Measure recall at 5 with a rerankerRetrieve 20 candidates, rerank them, and count the relevant passages again among the top 5.
- 04Compare latency as wellRecord the end-to-end median time. A gain of a few recall points does not necessarily justify an additional second of waiting.
- 05Reviewing failuresFor each missed question, check whether the correct answer was among the 20 candidates. If not, the problem is upstream: chunking, embeddings, or text extraction.
#Do you need a reranker? Decision guide
| Situation | Decision |
|---|---|
| A corpus of a few dozen clean documents, with good vector top-5 results already | No, measure first |
| Good answer frequently between 6th and 30th place | Yes: that's the typical use case |
| Vague questions, noisy or highly heterogeneous corpus | Yes, with 30 to 50 candidates |
| Interactive chat on a machine without a GPU | Caution: measure latency on the CPU and reduce to 10 to 20 candidates |
| The correct answer is missing even from the first 50 candidates | No: fix the chunking, embeddings, or extraction first |
#Frequently asked questions about reranking
Does a reranker replace vector search?+
Which reranker should you choose for documents in French?+
How many candidates should you rerank?+
Do you need a GPU for a reranker?+
Can you use a reranker with Ollama?+
How can I tell whether the reranker is really improving my results?+
- Hybrid search: BM25 + vector search
- Chunking strategies
- The best French embedding models
- Understanding the context window
- Local RAG with ChromaDB and Ollama
- Source: BAAI/bge-reranker-v2-m3 Hugging Face page
- Source: BAAI/bge-reranker-v2-gemma specifications
- Source: Sentence Transformers, Retrieve & Re-Rank
- Source: llama.cpp server documentation
Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.