Advanced 11 minOptimization

Hybrid search: BM25 + vectoriel

Direct response

Hybrid search launches a keyword search (BM25) and a vector search for the same query in parallel, then merges the two rankings, most often using Reciprocal Rank Fusion (RRF). It handles cases where the embedding fails (identifiers, proper names, rare terms) without losing paraphrase understanding. Qdrant, Weaviate, and Elasticsearch offer it natively; with ChromaDB, you can build it in thirty lines of Python.

A vector database retrieves meaning, not literal text: it misses a reference such as “RG/2024-117” or a ticket number. Conversely, keyword search does not understand that “automobile” and “car” refer to the same thing. This guide shows how to combine both, which fusion method to choose, what Qdrant and Weaviate do, and a French tokenization pitfall that silently ruins the BM25 component.

By Mohamed Meguedmi·Update 2026-09-30·Tested on Windows, macOS, and Linux

#Why purely vector-based search misses some questions

Vector search turns each passage into a vector that summarizes its general meaning, then returns the passages whose vectors are closest to the question's vector. This compression works very well for paraphrases and poorly for anything that must be found verbatim: an identifier, an error code, a person's name, or an internal acronym. For the embedding model, « PROD-4817 » and « PROD-4871 » look similar, and an unrelated ticket may appear before the right one. Hybrid search addresses this weakness: it adds a second ranking based on exact word matches, then combines the two so each method compensates for the other's blind spots.

Identifiers and references
Ticket, case, and contract numbers, SKUs, error codes: these are strings that an embedding does not reliably preserve and that BM25 finds as soon as they appear in the passage.
Rare specialized vocabulary
Infrequent medical, legal, or technical terms: a rare word has a high weight in BM25, while its embedding may be vague.
Very short queries
Two words such as “2024 invoice” provide little material for an embedding; keywords, on the other hand, can be compared as-is.
Paraphrases and rewordings
The converse case: “fixed-term contract” and “CDD,” or “car” and “automobile,” share no words; only vector search brings them closer together.
i
A typical case
In a ticket database, the query “PROD-4817” using pure vector search returns tickets with similar content, while keyword search immediately isolates the only ticket bearing that number.

#BM25: how keyword search works

The Local RAG Kit

Your documents, your AI: a reliable local RAG over your PDFs, notes and mail — nothing leaves your machine.

  • Lifetime online access
  • PDF + files
  • Lifetime updates

BM25 is the lexical ranking function used by Lucene, Elasticsearch, and OpenSearch, and offered by SQLite in its FTS5 module (the documentation describes its bm25() function as returning a value that indicates how well a row matches the query). It is based on three ideas: a rare word weighs more than a frequent word, a repeated word matters less and less beyond a certain number of occurrences, and a long passage is slightly penalized compared with a short one. The k1 parameter controls this saturation: according to Elastic's description, it limits the influence a single query term can have on a document's score.

Strengths
Rare terms, identifiers, short queries, specialized vocabulary, no model to load or train, compact index, explainable results (you know which word surfaced the passage).
Weaknesses
Synonyms, reformulations, paraphrases, typos; without lemmatization, “signed” and “signature” are two different words.
Prerequisites
Careful tokenization: lowercasing, removing accents, and possibly stemming. This is where most homegrown implementations go wrong; see below.

#Hybrid search: the principle

We run both searches on the same question; each returns its own candidate list (20 to 50 passages), then we merge the two lists into a single ranking. The improvement comes from a simple observation: passages that appear in both lists are almost always relevant, and each method also brings back passages the other missed. Weaviate defines hybrid search as combining the results of vector search and keyword search by merging the two result sets, with a configurable fusion method and relative weights. The measured improvement depends on the corpus and the questions: no general figure is reliable, and you need to measure it on your own documents (see below).

#Merge without normalization: Reciprocal Rank Fusion

The classic trap is adding the raw scores together. A BM25 score is a positive, unbounded number; a vector similarity score is a distance or cosine value within a bounded range: the two scales are incomparable, and even a minor corpus change shifts them. Reciprocal Rank Fusion sidesteps the problem by using only the ranks. Elasticsearch's documentation describes it as a method that requires no tuning and whose relevance indicators do not need to be correlated.

RRF formula
score(doc) = somme, sur chaque liste i où le doc apparaît, de 1 / (k + rang_i(doc))

k = 60 par défaut (valeur par défaut d'Elasticsearch)
rang_i = 1 pour le premier de la liste, 2 pour le deuxième, etc.
Un doc absent d'une liste n'ajoute rien pour cette liste.

A passage ranked first by BM25 and fifth by the vector search gets 1/61 + 1/65, or about 0.0318; a passage ranked fifteenth in both lists gets 2/75, or about 0.0267. The constant k reduces the advantage of the very top rank: the larger it is, the more distant ranks count. Elasticsearch documents this constant as rank_constant, with a default of 60, and a window size (rank_window_size) that sets the length of each list before merging. A larger window improves relevance at the cost of performance, the documentation notes.

#What about score fusion?

Some engines offer an alternative: normalize the scores from each list, then combine them using a weight. Weaviate documents both methods: rank-based fusion and relative-score fusion, the latter being the default since version 1.24; it is required to use autocut with the hybrid operator. Qdrant offers RRF and DBSF, with the latter preserving raw scores while normalizing their distribution (mean and standard deviation) before combining them. Rank fusion is more robust when the score distribution is unknown; score fusion allows finer tuning when it can be measured.

#Which tool: engines with native hybrid support

Hybrid search depending on the tool (official documentation, September 2026)
ToolNative hybridMergeWeight setting
QdrantYes, via the Query API (available since version 1.10)RRF or DBSFAdjustable weight per query and k constant in recent versions
WeaviateYes, hybrid operatorRelative ranks or scores (default since 1.24)Alpha parameter: 1 = pure vector, 0 = pure keywords
ElasticsearchYes, RRF retrieverRRFrank_constant (60 by default) and rank_window_size
SQLite FTS5 + vector extensionTo assembleTo writeOver to you
ChromaDB + rank_bm25To assemble in PythonTo write (RRF in 6 lines)Over to you

The fundamental difference isn't the merge itself, which takes only a few lines, but the index: a native engine keeps both indexes updated together, whereas a homegrown setup keeps the BM25 index in memory and has to rebuild it every time a document is added. For a corpus of a few thousand passages that rarely changes, the homegrown setup is more than sufficient. Beyond that, or as soon as documents change every day, a native engine avoids inconsistencies between the two indexes. The Weaviate guide covers this tool in detail.

#In-house implementation: ChromaDB, rank_bm25, and RRF

The following code assembles the three components. It fixes a common issue in examples: the normalization function must remove diacritics before filtering; otherwise, each accented letter splits a word in two.

Hybrid BM25 + ChromaDB with RRF
import re, unicodedata
import chromadb
from rank_bm25 import BM25Okapi

def normalize(txt):
    txt = unicodedata.normalize("NFKD", txt.lower())
    txt = "".join(c for c in txt if not unicodedata.combining(c))  # retire les accents
    return re.findall(r"[a-z0-9]+", txt)

# Indexation BM25 (en mémoire) : l'indice de la liste = l'identifiant du passage
all_docs = [d["text"] for d in load_docs()]
bm25 = BM25Okapi([normalize(d) for d in all_docs])

# ChromaDB : les ids doivent être les mêmes, sous forme de chaînes "0", "1", ...
coll = chromadb.PersistentClient("./chroma_db").get_collection("docs")

def rrf_fuse(rankings, weights=None, k=60):
    weights = weights or [1.0] * len(rankings)
    scores = {}
    for ranking, w in zip(rankings, weights):
        for rank, doc_id in enumerate(ranking, start=1):
            scores[doc_id] = scores.get(doc_id, 0) + w / (k + rank)
    return sorted(scores, key=scores.get, reverse=True)

def hybrid_search(question, top_k=5, n_candidates=20):
    s = bm25.get_scores(normalize(question))
    bm25_top = sorted(range(len(all_docs)), key=lambda i: -s[i])[:n_candidates]
    bm25_ranking = [str(i) for i in bm25_top]
    vec_ranking = coll.query(query_texts=[question], n_results=n_candidates)["ids"][0]
    fused = rrf_fuse([bm25_ranking, vec_ranking])
    return [all_docs[int(i)] for i in fused[:top_k]]
!
French tokenization pitfall
With the common normalization “NFKD, then replace everything that is not a-z or a digit with a space,” “référence” becomes “re fe rence”: each accent leaves a combining mark that is replaced with a space. The BM25 component continues to work on identifiers, but misses all accented words. Test normalize on three sentences before indexing.

Two practical precautions. First, the identifiers must be identical in both indexes: here, the list index, converted to a string, serves as the Chroma identifier. Second, the in-memory BM25 index disappears when the program stops: rebuild it at startup, which takes a few seconds for tens of thousands of passages, or save the document list alongside the database.

#Tune the balance between BM25 and vector search

By default, RRF treats both lists equally. If your corpus is rich in references (case law, tickets, catalogs), give BM25 more weight; if the questions are conversational, keep the balance or favor the vector search. In the previous function, simply pass weights=[0.6, 0.4] to assign BM25 a 60% weight. Qdrant offers the same setting: its documentation states that each query's weight is 1 by default, restoring the original RRF formula, and that the constant k can be adjusted in recent versions. In Weaviate, set alpha: 1 means pure vector search, while 0 means pure keyword search.

Starting point by corpus type
Corpus and questionsStarting BM25 / vector weightsWhat you monitor
References, numbers, proper names (legal, tickets, catalogs)60 / 40Do identifier-based questions surface first?
Written documentation, natural-language questions40 / 60Do paraphrases retrieve the right passage?
Mixed or unknown corpus50 / 50Recall@5 on 30 to 50 real questions
Acronyms and industry jargon55 / 45, with a synonym dictionary on the BM25 sideDo the acronym and its expanded form retrieve the same passage?
→
Measure before tuning
These values are starting points, not measured results. Assemble 30 to 50 real questions with the expected passage, calculate recall in the first 5 results for BM25 alone, vector search alone, and then hybrid search, and keep the hybrid approach only if it wins on your documents.

#After the merge: add a reranker

Fusion produces a more varied candidate set than either method alone. A reranker can then sort this set by reading each passage alongside the question: the two techniques complement each other. The usual order is hybrid search, then fusion, then reranking the top 20 to 50 results, followed by placing the 3 to 5 best passages in the prompt. The reranker guide covers this final step in detail, and the chunking guide explains why passage size affects both BM25 and embeddings so much.

#Frequently asked questions about hybrid search

FAQ
What is hybrid search in RAG?+
This combines two searches for the same question: a vector search, which retrieves passages with similar meaning, and a keyword search (BM25), which retrieves passages containing the same terms. The two rankings are merged into one, most often using Reciprocal Rank Fusion, before the best passages are passed to the model.
Should BM25 and vector scores be normalized before adding them together?+
That is not necessary with RRF, which uses only ranks—and that is the point: the two scales are not comparable. If you prefer to combine scores, normalize them first, as Weaviate’s relative-score fusion or Qdrant’s DBSF does, and verify that the result remains stable when the corpus changes.
What value of k should you choose in RRF?+
Keep 60, Elasticsearch's default, unless you have a reason to do otherwise. A smaller value favors the very top ranks more strongly, while a larger one gives distant ranks more influence. This setting matters little compared with tokenization quality and the number of candidates in each list.
Can you do hybrid search with ChromaDB?+
Yes, by combining Chroma's vector search, a BM25 index in Python (the rank_bm25 library), and an RRF fusion consisting of a few lines. Make sure to use the same identifiers in both indexes and strip accents during tokenization. For a corpus that changes frequently, a native hybrid engine such as Qdrant or Weaviate avoids maintaining two indexes.
Does hybrid search slow queries down significantly?+
Very little, generally: BM25 on an in-memory index is very fast, and the vector search has already taken place. The real cost comes from any reranker that follows and from the lexical index's memory usage. Measure end-to-end time on your queries, not the time for each isolated step.
How can you tell whether hybrid is worthwhile for your documents?+
On 30 to 50 real questions where you know the expected passage, compare recall in the top 5 results for BM25 alone, vector search alone, and the hybrid approach. If the hybrid approach gains nothing, your corpus may be purely conversational and vector search is sufficient. If it performs best mainly on identifiers, increase BM25's weight.
Did this guide help you?

Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.