Advanced 12 minOptimization

Add a reranker to your pipeline

Direct response

A reranker is a cross-encoder that rereads each question-passage pair and reorders the candidates returned by vector search: retrieve broadly (20 to 100 passages), then keep the 3 to 5 best ones for the model. In French, BAAI/bge-reranker-v2-m3 (Apache 2.0 license, approximately 568 million parameters) is the simplest starting point, locally, on a GPU, or even on a CPU if the volume remains modest.

Embeddings retrieve passages that are close to a question, but not necessarily passages that answer it. The reranker fixes this weakness by judging each question-passage pair in a single computation. This guide explains when it is worth the cost, which model to choose for French, how to connect it in thirty lines, and how to verify on your own documents that it genuinely improves the results.

By Mohamed Meguedmi·Update 2026-09-30·Tested on Windows, macOS, and Linux

#Reranker: what it does in a RAG pipeline

A reranker takes the question and a passage, reads them together, and returns a relevance score; you then sort the candidates by descending score and pass only the best ones to the language model. In a local RAG pipeline, it sits between the vector database (ChromaDB, Qdrant, Weaviate) and the LLM: vector search retrieves, for example, 30 passages, and the reranker keeps 5. BAAI's official documentation describes it this way: unlike an embedding model, the reranker receives the question and document as input and directly produces a similarity score instead of a vector. The benefit is greatest when the correct answer is among the candidates but not at the top: passages that discuss the right general subject without answering the precise question occupy the top positions, and the model then generates an off-target or fabricated answer. If the correct answer isn't among the candidates, a reranker can't help: it doesn't search; it sorts.

An embedding produces a vector for each document, independently of the question asked, and the distance measures overall topical similarity. So two passages on the same subject can receive similar scores even though only one contains the answer. The cross-encoder looks at the pair as a whole: it can see whether the requested date, name, or condition appears in the passage. It's more precise for this local judgment and more expensive because it requires one passage through the transformer per pair instead of a single computation per document.

i
The metaphor
The embedding model is the librarian who leads you to the right section and hands you twenty books. The reranker is the expert who skims those twenty books with your question in mind and puts the three that answer it on top.

#Bi-encoder and cross-encoder: why combine them

The Local RAG Kit

Your documents, your AI: a reliable local RAG over your PDFs, notes and mail — nothing leaves your machine.

  • Lifetime online access
  • PDF + files
  • Lifetime updates
Bi-encoder (embedding)
Encode the question and each document separately; document vectors are computed once during indexing, and search is reduced to a vector comparison. Fast and suitable for millions of passages.
Cross-encoder (reranker)
It encodes the question and passage together and outputs a score. No computation can be reused: it must be redone for every question and candidate. Precise, but impossible to apply to the entire corpus.

The Sentence Transformers documentation explains the reasoning: evaluating thousands or millions of pairs would be rather slow, so the retriever is used to produce a set of candidates, such as a hundred, which the cross-encoder re-ranks. This two-stage design is the same as that used by traditional search engines. It also explains the main tuning parameter: the number of candidates retrieved controls both recall—the wider the selection, the more likely it is to contain the right answer—and latency, since each additional candidate requires a pass through the cross-encoder.

#Which reranking model should you choose for French?

The choice comes down to three criteria: language, license, and memory. The size figures below come from the Hugging Face model cards; note that file size depends on weight precision: F32 for bge-reranker-v2-m3 (about 4 bytes per parameter), and F16 for mxbai.

Rerankers usable locally (Hugging Face cards, September 2026)
ModelLanguagesAdvertised sizeLicenseVerdict for French
BAAI/bge-reranker-v2-m3Multilingual0.6 billion parameters, weights in F32 (about 2.3 GB)Apache 2.0Default choice: multilingual, lightweight, with good tooling
BAAI/bge-reranker-v2-gemmaMultilingual3 billion parameters, F32 weights (about 10 GB)Apache 2.0For demanding workloads with a dedicated GPU; Gemma-2B-based reranker
mixedbread-ai/mxbai-rerank-large-v1English0.4 billion parameters, F16 weightsApache 2.0Avoid for a French corpus: the specifications list English
Cohere RerankMultilingualHosted serviceCommercialIrrelevant for a 100% local pipeline: the passages leave your machine

The bge-reranker-v2-m3 card presents it as a lightweight reranker with strong multilingual capabilities that is easy to deploy and fast at inference; the bge-reranker-v2-gemma card targets multilingual contexts, with good results in both English and multilingual settings. Two useful corrections to what is often reported: the bge-reranker-v2-m3 file weighs more than 2 GB, not 560 MB (568 million parameters in F32), and mxbai-rerank-large-v1 is an English-language model and should not be chosen for French documents. Once loaded in half precision, the m3 model uses approximately 1.1 GB for the weights (568 million parameters × 2 bytes, calculated in this guide), plus the activations for the processed batch.

→
The embedding model and reranker are independent
You can change one without reindexing the other: the reranker reads only the passage text, never their vectors. To choose an embedding model, see the dedicated guide to French embedding models.

#The pipeline before and after the reranker

Before / after
AVANT :
  Question → Embedding → Base vectorielle (top-5) → LLM

APRÈS :
  Question → Embedding → Base vectorielle (top-20 à top-50)
                       → Reranker (top-5) → LLM

On récupère large, puis on ordonne finement.

Two parameters govern the whole system: k_retrieve, the number of candidates returned by the vector database, and k_final, the number of passages sent to the LLM. A k_final of 3 to 5 suits most local models with 7 to 14 billion parameters; beyond that, you fill the context window without a clear benefit, and prompt-processing time increases. k_retrieve depends on corpus difficulty: 20 is a good first try, 50 to 100 if questions are vague or the corpus contains many similar passages. The context-window guide explains the cost of adding more passages to the prompt.

#Implementation: sentence-transformers, FlagEmbedding, llama.cpp

#With sentence-transformers

The CrossEncoder class loads the model and evaluates pairs. Its rank method directly accepts the question and document list and returns the best ones; the top_k parameter limits the number of results (without it, all documents are returned).

CrossEncoder.rank
from sentence_transformers import CrossEncoder

reranker = CrossEncoder("BAAI/bge-reranker-v2-m3", max_length=512)

def retrieve_and_rerank(question, k_retrieve=20, k_final=5):
    # 1. Récupération par embedding (ChromaDB, Qdrant, etc.)
    candidats = embedding_search(question, top_k=k_retrieve)  # liste de textes

    # 2. Scoring par le cross-encoder, tri et coupe en une seule étape
    resultats = reranker.rank(question, candidats, top_k=k_final, batch_size=16)
    # resultats = [{'corpus_id': 3, 'score': 0.91}, ...]
    return [candidats[r['corpus_id']] for r in resultats]

#With FlagEmbedding, the model authors' library

The model card uses the FlagEmbedding library. It states that the raw score can be mapped to a value between 0 and 1 using a sigmoid function with normalize=True, and that use_fp16=True speeds up computation at the cost of a slight quality reduction. Remember that the raw score has no absolute scale and is often negative for irrelevant passages; only the ranking matters unless you set a threshold.

FlagReranker
from FlagEmbedding import FlagReranker

reranker = FlagReranker('BAAI/bge-reranker-v2-m3', use_fp16=True)
score = reranker.compute_score(['ma question', 'un passage'], normalize=True)  # entre 0 et 1

#With llama.cpp, without Python

The llama.cpp server provides a reranking endpoint, disabled by default. The documentation says it requires a reranking model, cites bge-reranker-v2-m3 as an example, and runs with the --embedding and --pooling rank options. You need a GGUF version of the model. This is the option to choose if your stack is already built around llama.cpp and you want to avoid installing PyTorch. Check the exact option in the installed version: the documentation warns that this endpoint may change.

llama-server (adapt to your version)
llama-server -m bge-reranker-v2-m3-Q8_0.gguf --embedding --pooling rank --reranking --port 8081

#With LlamaIndex

SentenceTransformerRerank
from llama_index.core.postprocessor import SentenceTransformerRerank

reranker = SentenceTransformerRerank(model="BAAI/bge-reranker-v2-m3", top_n=5)

query_engine = index.as_query_engine(
    similarity_top_k=20,
    node_postprocessors=[reranker],
)

#Cost: latency, memory, passage length

No latency figure is guaranteed: it depends on the card, precision (FP16 or FP32), number of candidates, and passage length. Focus on the proportions. Reranking time grows linearly with the number of pairs: going from 20 to 100 candidates multiplies the workload by five. It also grows with passage length, since each pair is encoded in full. Measure on your machine with your passages instead of relying on a number from a blog: time one hundred real queries and look at the median and worst-case times.

Number of candidates
First lever. Start at 20, measure recall, and increase to 50 only if good answers are still being left out.
Accuracy
use_fp16=True (FlagEmbedding) or loading in half precision reduces memory usage and speeds up computation, with a slight performance drop depending on the model card.
Batch size
The sentence-transformers rank method processes 32 pairs per batch by default. Reduce this to 8 or 16 if memory is tight; increase it if the GPU is underutilized.
Maximum length
512 tokens per pair (the max_length value in the official examples). A longer passage is truncated: if your chunks exceed this size, the end of the passage isn't read. Shorten the chunks before increasing the limit.
CPU or GPU
On CPU, the reranker works, but each request with 20 to 50 candidates takes seconds; that's acceptable for an internal document assistant, less so for interactive chat.
→
Skip the reranker when it is unnecessary
On a small, clean corpus (a few dozen well-structured pages), the vector top-5 often already contains the answer. Measure first, and keep the reranker only if recall improves.

#Reranking and chunking: the two settings interact

A cross-encoder judges an entire passage. If the passage mixes three topics, its score will be average for the three corresponding questions; if it is too short, it loses the context that would allow it to be recognized as an answer. Splitting text into medium-sized passages with slight overlap gives the reranker enough material without exceeding its maximum length. If you change the chunk size, rerun the recall test: the best k_retrieve and the reranker gain change with it. The guide to chunking strategies explains these choices in detail.

Another interaction: hybrid search. When keyword search (BM25) and vector search are combined, the candidate pool is more diverse, giving the reranker more chances to find a good answer than either method alone would have had. The reranker is the final layer, and hybrid search is the second: they complement rather than replace each other.

#Measure the improvement on your documents

The improvement reported in blogs ranges from twofold to fivefold depending on the corpus, and no general figure applies to yours. On a highly structured corpus (clean technical documentation), vector search alone is already good; on a noisy corpus (emails, notes, poorly extracted PDFs), the gap is more pronounced. The only way to know is to evaluate it.

  1. 01
    Prepare 30 to 50 questions
    Use real user questions, each with the passage containing the answer (an ID is enough).
  2. 02
    Measure recall at 5 without a reranker
    For each question, check whether the relevant passage appears in the vector database's top 5 results.
  3. 03
    Measure recall at 5 with a reranker
    Retrieve 20 candidates, rerank them, and count the relevant passages again among the top 5.
  4. 04
    Compare latency as well
    Record the end-to-end median time. A gain of a few recall points does not necessarily justify an additional second of waiting.
  5. 05
    Reviewing failures
    For each missed question, check whether the correct answer was among the 20 candidates. If not, the problem is upstream: chunking, embeddings, or text extraction.
Evaluate recall
def rappel_a_k(pipeline, questions, cibles, k=5):
    ok = 0
    for q, cible in zip(questions, cibles):
        ok += cible in [p.id for p in pipeline(q)[:k]]
    return ok / len(questions)

sans = rappel_a_k(pipeline_sans_reranker, questions, cibles)
avec = rappel_a_k(pipeline_avec_reranker, questions, cibles)
print(f"Sans : {sans:.0%}   Avec : {avec:.0%}")
!
Reranker limitations
It won’t fix text extracted incorrectly from a PDF, a split that cuts the answer in two, or an ambiguous question. It may also penalize short passages with a low score even when they contain the correct number: check failures instead of relying on the average.

#Do you need a reranker? Decision guide

When to add a reranker
SituationDecision
A corpus of a few dozen clean documents, with good vector top-5 results alreadyNo, measure first
Good answer frequently between 6th and 30th placeYes: that's the typical use case
Vague questions, noisy or highly heterogeneous corpusYes, with 30 to 50 candidates
Interactive chat on a machine without a GPUCaution: measure latency on the CPU and reduce to 10 to 20 candidates
The correct answer is missing even from the first 50 candidatesNo: fix the chunking, embeddings, or extraction first

#Frequently asked questions about reranking

FAQ
Does a reranker replace vector search?+
No. It does not scan the corpus: it reranks the few dozen candidates provided by the vector search. Without a fast first stage, it would have to run on every passage in the database, which would be far too slow. The two stages complement each other: the vector search ensures recall, while the reranker improves precision at the top of the ranking.
Which reranker should you choose for documents in French?+
BAAI/bge-reranker-v2-m3 is a reasonable starting point: multilingual, licensed under Apache 2.0, and approximately 0.6 billion parameters, so it can run on a modest graphics card. Avoid mxbai-rerank-large-v1, whose model card lists English. The larger bge-reranker-v2-gemma is justified only if the first model is insufficient for your question set.
How many candidates should you rerank?+
Start with 20 candidates and keep 5 passages. If good answers remain outside the top 20 during evaluation, increase it to 50. Each additional candidate adds a computation, so latency grows roughly linearly. Beyond 100, the gains become rare: this often indicates a chunking or embedding problem.
Do you need a GPU for a reranker?+
No, but it helps. On CPU, the m3 model works, with a delay of around one second or more for a batch of candidates depending on the machine—measure it on your system. That remains acceptable for internal document use. For a smooth chat experience, even a modest GPU or a lighter model changes how it feels to use.
Can you use a reranker with Ollama?+
Ollama serves the generation model and embeddings, but reranking is handled separately: with sentence-transformers or FlagEmbedding in Python, or with the llama.cpp server, which exposes a reranking endpoint. The reranker is a separate model, loaded in its own process, and coexists with the generation model if memory allows.
How can I tell whether the reranker is really improving my results?+
Build 30 to 50 real questions with the expected passage, then compare recall in the top 5 results with and without a reranker, along with median latency. On a clean corpus, the difference may be small; on a noisy corpus, it is more pronounced. Then analyze the failures to determine whether they originate upstream.
Did this guide help you?

Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.