Hybrid search: BM25 + vectoriel
Hybrid search launches a keyword search (BM25) and a vector search for the same query in parallel, then merges the two rankings, most often using Reciprocal Rank Fusion (RRF). It handles cases where the embedding fails (identifiers, proper names, rare terms) without losing paraphrase understanding. Qdrant, Weaviate, and Elasticsearch offer it natively; with ChromaDB, you can build it in thirty lines of Python.
A vector database retrieves meaning, not literal text: it misses a reference such as “RG/2024-117” or a ticket number. Conversely, keyword search does not understand that “automobile” and “car” refer to the same thing. This guide shows how to combine both, which fusion method to choose, what Qdrant and Weaviate do, and a French tokenization pitfall that silently ruins the BM25 component.
#Why purely vector-based search misses some questions
Vector search turns each passage into a vector that summarizes its general meaning, then returns the passages whose vectors are closest to the question's vector. This compression works very well for paraphrases and poorly for anything that must be found verbatim: an identifier, an error code, a person's name, or an internal acronym. For the embedding model, « PROD-4817 » and « PROD-4871 » look similar, and an unrelated ticket may appear before the right one. Hybrid search addresses this weakness: it adds a second ranking based on exact word matches, then combines the two so each method compensates for the other's blind spots.
- Identifiers and references
- Ticket, case, and contract numbers, SKUs, error codes: these are strings that an embedding does not reliably preserve and that BM25 finds as soon as they appear in the passage.
- Rare specialized vocabulary
- Infrequent medical, legal, or technical terms: a rare word has a high weight in BM25, while its embedding may be vague.
- Very short queries
- Two words such as “2024 invoice” provide little material for an embedding; keywords, on the other hand, can be compared as-is.
- Paraphrases and rewordings
- The converse case: “fixed-term contract” and “CDD,” or “car” and “automobile,” share no words; only vector search brings them closer together.
#BM25: how keyword search works
Your documents, your AI: a reliable local RAG over your PDFs, notes and mail — nothing leaves your machine.
- Lifetime online access
- PDF + files
- Lifetime updates
BM25 is the lexical ranking function used by Lucene, Elasticsearch, and OpenSearch, and offered by SQLite in its FTS5 module (the documentation describes its bm25() function as returning a value that indicates how well a row matches the query). It is based on three ideas: a rare word weighs more than a frequent word, a repeated word matters less and less beyond a certain number of occurrences, and a long passage is slightly penalized compared with a short one. The k1 parameter controls this saturation: according to Elastic's description, it limits the influence a single query term can have on a document's score.
- Strengths
- Rare terms, identifiers, short queries, specialized vocabulary, no model to load or train, compact index, explainable results (you know which word surfaced the passage).
- Weaknesses
- Synonyms, reformulations, paraphrases, typos; without lemmatization, “signed” and “signature” are two different words.
- Prerequisites
- Careful tokenization: lowercasing, removing accents, and possibly stemming. This is where most homegrown implementations go wrong; see below.
#Hybrid search: the principle
We run both searches on the same question; each returns its own candidate list (20 to 50 passages), then we merge the two lists into a single ranking. The improvement comes from a simple observation: passages that appear in both lists are almost always relevant, and each method also brings back passages the other missed. Weaviate defines hybrid search as combining the results of vector search and keyword search by merging the two result sets, with a configurable fusion method and relative weights. The measured improvement depends on the corpus and the questions: no general figure is reliable, and you need to measure it on your own documents (see below).
#Merge without normalization: Reciprocal Rank Fusion
The classic trap is adding the raw scores together. A BM25 score is a positive, unbounded number; a vector similarity score is a distance or cosine value within a bounded range: the two scales are incomparable, and even a minor corpus change shifts them. Reciprocal Rank Fusion sidesteps the problem by using only the ranks. Elasticsearch's documentation describes it as a method that requires no tuning and whose relevance indicators do not need to be correlated.
A passage ranked first by BM25 and fifth by the vector search gets 1/61 + 1/65, or about 0.0318; a passage ranked fifteenth in both lists gets 2/75, or about 0.0267. The constant k reduces the advantage of the very top rank: the larger it is, the more distant ranks count. Elasticsearch documents this constant as rank_constant, with a default of 60, and a window size (rank_window_size) that sets the length of each list before merging. A larger window improves relevance at the cost of performance, the documentation notes.
#What about score fusion?
Some engines offer an alternative: normalize the scores from each list, then combine them using a weight. Weaviate documents both methods: rank-based fusion and relative-score fusion, the latter being the default since version 1.24; it is required to use autocut with the hybrid operator. Qdrant offers RRF and DBSF, with the latter preserving raw scores while normalizing their distribution (mean and standard deviation) before combining them. Rank fusion is more robust when the score distribution is unknown; score fusion allows finer tuning when it can be measured.
#Which tool: engines with native hybrid support
| Tool | Native hybrid | Merge | Weight setting |
|---|---|---|---|
| Qdrant | Yes, via the Query API (available since version 1.10) | RRF or DBSF | Adjustable weight per query and k constant in recent versions |
| Weaviate | Yes, hybrid operator | Relative ranks or scores (default since 1.24) | Alpha parameter: 1 = pure vector, 0 = pure keywords |
| Elasticsearch | Yes, RRF retriever | RRF | rank_constant (60 by default) and rank_window_size |
| SQLite FTS5 + vector extension | To assemble | To write | Over to you |
| ChromaDB + rank_bm25 | To assemble in Python | To write (RRF in 6 lines) | Over to you |
The fundamental difference isn't the merge itself, which takes only a few lines, but the index: a native engine keeps both indexes updated together, whereas a homegrown setup keeps the BM25 index in memory and has to rebuild it every time a document is added. For a corpus of a few thousand passages that rarely changes, the homegrown setup is more than sufficient. Beyond that, or as soon as documents change every day, a native engine avoids inconsistencies between the two indexes. The Weaviate guide covers this tool in detail.
#In-house implementation: ChromaDB, rank_bm25, and RRF
The following code assembles the three components. It fixes a common issue in examples: the normalization function must remove diacritics before filtering; otherwise, each accented letter splits a word in two.
Two practical precautions. First, the identifiers must be identical in both indexes: here, the list index, converted to a string, serves as the Chroma identifier. Second, the in-memory BM25 index disappears when the program stops: rebuild it at startup, which takes a few seconds for tens of thousands of passages, or save the document list alongside the database.
#Tune the balance between BM25 and vector search
By default, RRF treats both lists equally. If your corpus is rich in references (case law, tickets, catalogs), give BM25 more weight; if the questions are conversational, keep the balance or favor the vector search. In the previous function, simply pass weights=[0.6, 0.4] to assign BM25 a 60% weight. Qdrant offers the same setting: its documentation states that each query's weight is 1 by default, restoring the original RRF formula, and that the constant k can be adjusted in recent versions. In Weaviate, set alpha: 1 means pure vector search, while 0 means pure keyword search.
| Corpus and questions | Starting BM25 / vector weights | What you monitor |
|---|---|---|
| References, numbers, proper names (legal, tickets, catalogs) | 60 / 40 | Do identifier-based questions surface first? |
| Written documentation, natural-language questions | 40 / 60 | Do paraphrases retrieve the right passage? |
| Mixed or unknown corpus | 50 / 50 | Recall@5 on 30 to 50 real questions |
| Acronyms and industry jargon | 55 / 45, with a synonym dictionary on the BM25 side | Do the acronym and its expanded form retrieve the same passage? |
#After the merge: add a reranker
Fusion produces a more varied candidate set than either method alone. A reranker can then sort this set by reading each passage alongside the question: the two techniques complement each other. The usual order is hybrid search, then fusion, then reranking the top 20 to 50 results, followed by placing the 3 to 5 best passages in the prompt. The reranker guide covers this final step in detail, and the chunking guide explains why passage size affects both BM25 and embeddings so much.
#Frequently asked questions about hybrid search
What is hybrid search in RAG?+
Should BM25 and vector scores be normalized before adding them together?+
What value of k should you choose in RRF?+
Can you do hybrid search with ChromaDB?+
Does hybrid search slow queries down significantly?+
How can you tell whether hybrid is worthwhile for your documents?+
- Add a reranker to your pipeline
- Chunking strategies
- Weaviate: hybrid search and multi-tenancy
- Local RAG with ChromaDB and Ollama
- The best French embedding models
- Source: Qdrant, hybrid queries
- Source: Weaviate, hybrid search
- Source: Elasticsearch, Reciprocal Rank Fusion
- Source: SQLite FTS5, bm25() function
Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.