Intermediate 11 minEmbeddings

BGE-M3: embeddings that truly speak French

Direct response

BGE-M3 (BAAI, MIT license) is an embedding model that handles more than 100 languages, accepts up to 8,192 tokens per passage, and produces three representations in one pass: dense (overall meaning), lexical (exact terms, BM25-style), and multi-vector (ColBERT-style reranking). On a French corpus, it is a safer choice than the English-language models in public rankings. It fits in 1.2 GB via Ollama (bge-m3:567m) and runs without a graphics card.

Most of the highest-ranked embedding models are trained primarily on English. Used on a French corpus, they work without errors and retrieve less effectively—the most insidious failure because nothing signals it. BGE-M3, published by the Chinese laboratory BAAI, is one of the few truly multilingual models, with an additional useful feature: it produces three types of representations at once, at no additional cost.

By Mohamed Meguedmi·Update 2026-09-30·Tested on Windows, macOS, and Linux

#The French problem

An embedding model learns its notion of similarity from its training corpus. If that corpus is 90% English, the model correctly positions English texts relative to one another, but handles French texts much less precisely. In practice, on a French document corpus, this means relevant passages fail to surface while off-topic passages do—with no error message, no warning in the logs, and often without anyone noticing until a user reports an irrelevant answer.

That's why the answer to “which embedding model?” differs by language. Public benchmarks (such as MTEB) are heavily dominated by English-language tasks; they don't predict what will happen on your documents or specialized vocabulary. That's exactly the use case for which the Chinese lab BAAI (Beijing Academy of Artificial Intelligence) designed BGE-M3: the name stands for Multi-Functionality, Multi-Linguality, Multi-Granularity—three search functions, more than 100 languages, and inputs ranging from short sentences to long documents, combined in a single model rather than spread across multiple tools.

#What BGE-M3 brings

The Local RAG Kit

Your documents, your AI: a reliable local RAG over your PDFs, notes and mail — nothing leaves your machine.

  • Lifetime online access
  • PDF + files
  • Lifetime updates
Multilingual by design
According to its official profile, the model supports more than 100 working languages; it handles French correctly and also lets you query an English corpus in French.
Long context: 8 192 tokens
It accepts significantly longer passages than most embedding models (often limited to 512 tokens), giving you more flexibility in the size of indexed chunks.
Three representations at once
Dense, lexical, and multi-vector, produced in a single pass with no additional compute cost, according to BAAI.
MIT License
Usable in business without particular restrictions; always check the model card when integrating it, as the license may change from one version to another.
No instruction to add
Unlike older BGE models, BGE-M3 no longer requires adding an instruction to queries: use it directly, which simplifies integration.

#Dense, lexical, multi-vector: what it is for

Three outputs, three use cases (definitions taken from the model's official documentation)
RepresentationWhat it capturesWhat it offers
DenseThe text reduced to a single vector that captures its overall meaningTraditional semantic search
Lexical (sparse)One weight per vocabulary term, zero except for words present in the textFinds references, acronyms, and proper names that semantic search misses, like BM25
Multi-vecteurMultiple vectors per text, in the style of ColBERTA more precise re-ranking of the top candidates

The benefit is being able to run hybrid search without running two separate models: the official page explicitly recommends the hybrid search + reranking combination, using lexical weights obtained “at no additional cost” alongside the dense embedding. The lexical component precisely addresses the dense component's failures, while the multivector component reranks the twenty or so selected candidates. Not all vector databases can use all three outputs, but dense-only is already enough for typical use—it's also the simplest configuration to implement when getting started, before adding the lexical layer if accuracy is still lacking.

!
The choice is made before the first import
Changing the embedding model means re-encoding everything: vectors from two models are not comparable. On a substantial corpus, this takes several hours. Testing two or three models on a sample of your own questions takes half a day and avoids having to redo the work.

#Use it locally

The fastest way to test BGE-M3 locally is Ollama, which distributes a quantized version of the model under the name bge-m3:567m — 567 million parameters, a download of approximately 1.2 GB, and an advertised context window of 8,000 tokens on the model page. It is built on XLM-RoBERTa-large, whose context length was first extended to 8,192 tokens through dedicated pretraining (RetroMAE), followed by unified fine-tuning for the three retrieval tasks in a single training stage, rather than three separate models to maintain.

Terminal — download and test the model
ollama pull bge-m3
ollama run bge-m3 "Quel est le délai de rétractation légal en France ?"

For a complete RAG pipeline, only dense output is directly usable by most standard vector databases (Qdrant, Milvus, pgvector) without specific configuration; lexical and multi-vector outputs require an engine that knows how to combine them, as documented, for example, by the Milvus and Vespa integrations cited by BAAI.

In a standard RAG pipeline with a local LLM, BGE-M3 simply replaces the default embedding model: each document chunk is encoded once during indexing, the user’s question is encoded for each query, and the best retrieved passages are then sent to the LLM (Qwen, Mistral, Llama…) loaded under Ollama or LM Studio. Choosing the embedding model and choosing the generation model are two independent decisions: nothing requires using models from the same family for both steps, and in practice that is rarely the case.

i
A detail that often trips people up
The embedding model used during indexing must remain strictly identical to the one used for the query. Changing the version, even a minor one, or the normalization parameter between the two steps degrades relevance without causing a visible error—the same silent problem described in the introduction, but self-inflicted this time.

#Check against your own documents

A product page's “more than 100 languages” is not a guarantee of uniform quality: it is coverage, not a per-language score. BAAI evaluates BGE-M3 on MIRACL (multilingual retrieval) and MLDR, a long-document retrieval dataset covering 13 languages that the lab published specifically for this model, with the corresponding evaluation pipeline openly available. These evaluations provide a general trend measured on standardized research corpora; they do not replace testing on your own corpus, in your own domain, with your own technical vocabulary and your own question phrasing.

  1. 01
    Create a small representative sample
    Twenty to thirty real documents from the target corpus, with their usual diversity of length and vocabulary—not a handpicked selection.
  2. 02
    Write questions the way your users would ask them
    Ten to twenty real questions, including questions that use synonyms or wording different from the source document.
  3. 03
    Compare two or three models on the same sample
    BGE-M3 against a well-regarded English-language model and, if possible, another multilingual model. Note how many times the correct passage appears in the top three results.
  4. 04
    Decide before the bulk import
    The test takes half a day; changing models after indexing an entire corpus means re-encoding everything, as noted above.
→
The site catalog as a reference point
To put the size of an embedding model in perspective relative to the LLMs you already run locally, the QuelLLM catalog (quelllm.fr/api/models.json) lists the sizes and memory requirements of the models referenced on the site.

#What it costs

This is a heavier embedding model than the small English-language models currently in fashion: its 567 million parameters and 1.2 GB download make it mid-weight for the category, far from the models with a few tens of millions of parameters that dominate pure speed rankings. It encodes more slowly and uses more memory. You’ll notice this during the initial indexing of a large corpus; for a single question, the difference is imperceptible compared with the generation time of the language model that follows.

Another consequence: its dense vectors have a dimensionality of 1,024, so the index takes up more space than models with 384 or 768 dimensions. With one million passages, the arithmetic is no longer negligible, and it is time to look at your vector database’s compression options.

BGE-M3 in the model family published by BAAI
ModelSizeMax lengthScope
BAAI/bge-m31 0248,192 tokensMultilingual, three unified retrieval methods
BAAI/bge-large-en-v1.51 024512 tokensEnglish only
BAAI/bge-base-en-v1.5768512 tokensEnglish only, lighter
BAAI/bge-small-en-v1.5384512 tokensEnglish only, the lightest

The contrast is clear: at the same dimension (1,024), BGE-M3 accepts sixteen times more tokens per passage than bge-large-en-v1.5 and covers more than 100 languages, whereas BAAI's entire “en” family is limited to English. That's the price, in download size and encoding time, of not having to choose a different model for each corpus language or maintain multiple separate indexes based on the language of incoming documents.

#When to choose it and when to skip it

Decide honestly
SituationChoice
French or multilingual corpusBGE-M3 or another genuinely multilingual model
Questions in French about English documentsA multilingual model, mandatory
English-only corpus, speed firstA small specialized English-language model
Long passages (over 512 tokens) that we do not want to split furtherBGE-M3, for its 8,192-token input length
Highly constrained machineA lighter model, even if it means losing relevance in French

#FAQ

Is BGE-M3 good in French?+
Yes, that is precisely its advantage: its official page states that it supports more than 100 working languages, whereas most of the highest-ranked models are predominantly English-language. Still, verify it with your own questions, since public rankings remain dominated by English.
Can you query English documents in French?+
This is one of the benefits of a multilingual model such as BGE-M3: the question and the document do not need to be in the same language to end up close in vector space, unlike a model trained on a single language. This is useful for a mixed document repository, such as technical documentation in English queried by a French-speaking team.
Do you need a graphics card?+
No, it runs on the processor: the Ollama version weighs about 1.2 GB for 567 million parameters, a reasonable weight for CPU use. A GPU significantly reduces indexing time for a large corpus; for an isolated query, the difference is negligible compared with the language model that processes the response afterward.
What are its three outputs for?+
Dense retrieval handles classic semantic search, lexical (sparse) retrieval catches exact terms such as references and acronyms in the manner of BM25, and multi-vector retrieval, ColBERT-style, finely reranks the best candidates. Using only the dense output is already a valid approach and the easiest to integrate into a standard vector database.
How many tokens does BGE-M3 accept per passage?+
Up to 8 192 tokens according to its official specifications, versus 512 for the same publisher's bge-large-en-v1.5, bge-base-en-v1.5, and bge-small-en-v1.5 family of English-language models. This is useful for indexing long passages such as contracts or reports without splitting them further, at the cost of slower encoding for large volumes of documents.
Can I change it later?+
Yes, at the cost of fully re-encoding the corpus: vectors from two models aren’t comparable, even when they have the same dimensionality. It’s better to test two or three models on a sample before the first large-scale import than to discover the problem afterward on a corpus that’s already indexed.
Is BGE-M3 better than OpenAI models for French?+
An independent comparison cited by BAAI ranks it first both in English and in other languages against OpenAI models on this specific test, but a single comparison does not replace testing on your corpus. The best practice remains to test several models on your own questions before choosing.
Did this guide help you?

Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.