The best embedding models FR
For French, BGE-M3 is the safest starting point: multilingual, 8,192-token context, MIT license, available in Ollama. Qwen3-Embedding and EmbeddingGemma are the recent alternatives to test. English-focused models (nomic-embed-text, mxbai-embed-large, all-MiniLM) should be avoided for French. Validate the final choice on your own documents.
An embedding model transforms each piece of text into a vector: it determines that “CDD” and “contrat à durée déterminée” are close neighbors. For French, the MTEB-French benchmark measures a 22-point retrieval gap between BGE-M3 and all-MiniLM-L12-v2. This page ranks open models usable locally, cites published measurements and their limitations, then covers what costs money in production: prefixes, truncation, dimensions, storage, and reindexing.
#What an embedding model is, and why French makes it more difficult
An embedding model is a neural network that converts text (a sentence, paragraph, or chunk) into a vector of numbers, from 384 to 4,096 dimensions depending on the model. Two texts with similar meanings produce nearby vectors, most often measured using cosine similarity. This is what allows a RAG system to retrieve a passage about a “fixed-term contract” when the user writes “CDD,” with no words in common. The same model must encode the questions and documents, as the Ollama documentation points out; changing models therefore requires reindexing the entire corpus. To choose, three questions are enough: was the model trained on French, does its context cover your chunks, and does its license allow your use?
French adds its own difficulties: accents, elisions (“l'employeur”), legal acronyms, and technical Anglicisms mixed into the text. The gap between models is clear. In the MTEB-French benchmark, the retrieval score ranges from 0.43 for all-MiniLM-L12-v2 to 0.65 for BGE-M3, a 22-point difference on the same question set.
#The five criteria that truly distinguish the models
Your documents, your AI: a reliable local RAG over your PDFs, notes and mail — nothing leaves your machine.
- Lifetime online access
- PDF + files
- Lifetime updates
- Training languages
- BGE-M3 and Qwen3-Embedding claim more than 100 languages, while multilingual-e5-large supports 94. The model cards for nomic-embed-text-v1.5 and mxbai-embed-large-v1, which are heavily downloaded on Ollama, are labeled English.
- Maximum context
- 512 tokens for multilingual-e5-large, Solon, and mxbai-embed-large; 2 048 for EmbeddingGemma; 8 192 for BGE-M3; 32 000 for Qwen3-Embedding and Granite R2. A longer chunk is truncated, not rejected.
- Dimensions and storage
- From 384 to 4,096 dimensions, or 4 bytes per dimension and per vector in float32 (calculation below).
- License
- MIT for BGE-M3, multilingual-e5-large, and Solon; Apache 2.0 for Qwen3-Embedding, Granite R2, nomic, and mxbai; Gemma terms for EmbeddingGemma; CC BY-NC 4.0 (non-commercial) for jina-clip-v2.
- Prefixes and instructions
- Some models require a prefix before each text ("query:" and "passage:" for E5, even in French). Forgetting it degrades retrieval without producing an error.
#The 2026 ranking for French: nine models compared
This ranking is editorial, not based on in-house testing: it combines French-language coverage, published benchmarks, and local availability. The weights come from the Ollama library when the model is listed there; otherwise, from the float32 estimate in the MTEB-French study.
| Model | Dimensions / context | License | Weights | What the sources say about French |
|---|---|---|---|---|
| BGE-M3 (BAAI) | 1 024 / 8 192 | MIT | 1.2 GB (Ollama) | Retrieval MTEB-French 0.65. Dense, sparse, and multi-vector. Recommended starting point. |
| Qwen3-Embedding 0.6B / 4B / 8B | 1 024 / 2 560 / 4 096 ; 32 000 | Apache 2.0 | 639 MB / 2.5 GB / 4.7 GB (Ollama) | More than 100 languages. Multilingual MTEB according to Qwen: 64.33 / 69.45 / 70.58. |
| EmbeddingGemma 300M (Google) | 768 (512, 256, 128) / 2 048 | Gemma | 622 MB (Ollama) | More than 100 languages. Multilingual MTEB v2: 61.15 according to Google. |
| multilingual-e5-large | 1 024 / 512 | MIT | 2.24 GB (float32) | Retrieval MTEB-French 0.59. Prefixes required. |
| Solon-embeddings-large-0.1 | 1 024 / 512 | MIT | 2.24 GB (float32) | French model published by Ordalie Technologies. Retrieval MTEB-French 0.63. |
| Granite Embedding 311M Multilingual R2 (IBM) | 768 / 32 768 | Apache 2.0 | 311M parameters | French among the 52 strengthened languages. 65.2 on multilingual MTEB retrieval according to IBM; no independent French measurement consulted. |
| nomic-embed-text v1.5 | 768 / 8,192 (2,048 under Ollama) | Apache 2.0 | 274 MB (Ollama) | English-labeled card. Reserve it for English corpora. |
| mxbai-embed-large-v1 | 512 context | Apache 2.0 | 670 MB (Ollama) | English-labeled card. Reserve it for English corpora. |
| all-MiniLM-L12-v2 | 384 | Apache 2.0 | 33M parameters | Retrieval MTEB-French 0.43. CPU-only prototype. |
BGE-M3 remains the best default: its French retrieval has been published, its 8,192-token context avoids most forced chunking, and it also provides the lexical weights useful for hybrid search. Qwen3-Embedding is the candidate when a graphics card is available: its 8B version topped the multilingual MTEB leaderboard when it launched, on June 5, 2025, according to Qwen, but its 4,096 dimensions weigh on storage. EmbeddingGemma targets modest machines.
Solon, often presented as a specialist in French legal and administrative language, does not outperform any of the three retrieval datasets in the MTEB-French study: it matches BGE-M3 on the Syntec collective bargaining agreement and trails it on legal articles (BSARD) and school questions (Alloprof).
#What French benchmarks say—and what they don’t
The published reference is MTEB-French (Ciancone et al., May 2024): 18 datasets, 8 task categories, and 51 models compared. The authors conclude that large multilingual models pretrained on sentence similarity perform exceptionally well. Here is the retrieval (NDCG@10) on three datasets: French legal articles (BSARD), the Syntec collective agreement, and Alloprof school questions.
| Model | Average retrieval | Law (BSARD) | Syntec agreement | Alloprof |
|---|---|---|---|---|
| text-embedding-3-large (OpenAI API) | 0,73 | 0,73 | 0,87 | 0,60 |
| mistral-embed (API Mistral) | 0,68 | 0,68 | 0,79 | 0,57 |
| BGE-M3 | 0,65 | 0,60 | 0,85 | 0,49 |
| Solon-embeddings-large-0.1 | 0,63 | 0,58 | 0,85 | 0,47 |
| multilingual-e5-large | 0,59 | 0,59 | 0,81 | 0,38 |
| multilingual-e5-small | 0,52 | 0,52 | 0,76 | 0,27 |
| paraphrase-multilingual-MiniLM-L12-v2 | 0,44 | 0,38 | 0,66 | 0,27 |
| all-MiniLM-L12-v2 | 0,43 | 0,34 | 0,61 | 0,33 |
Three limitations prevent this from being a definitive ranking. The study dates from 2024: it includes neither Qwen3-Embedding (June 2025), nor EmbeddingGemma (September 2025), nor Granite R2, and no comparable French measurements of these models were consulted for this page. The corpora (legal articles, collective bargaining agreements, school questions) may not resemble yours. Finally, the multilingual scores published by vendors are self-reported and combine all languages.
#Which model for which use case: decision table
| Your situation | Model to test first | Why | Key consideration |
|---|---|---|---|
| French or French + English corpus, capable machine | BGE-M3 | French retrieval score of 0.65, on par with the study's best open models | 1.2 GB in Ollama; no prefix |
| Legal, administration, HR | BGE-M3, then Solon for comparison | BGE-M3 matches or exceeds Solon on all three French datasets in the study | A model without an internal legal test set remains a gamble |
| Modest machine or CPU only | EmbeddingGemma 300M, or multilingual-e5-small | 622 MB in Ollama; designed by Google for laptops and mobile devices | 2,048-token context: short chunks |
| Available graphics card, maximum quality | Qwen3-Embedding-4B or 8B | 32,000 tokens, adjustable dimensions, more than 100 languages | 2,560 and 4,096 dimensions: storage and index limits |
| Long chunks, structured documents | BGE-M3, Qwen3-Embedding, or Granite R2 | 8 192 to 32 768 context tokens | A very long chunk dilutes the meaning |
| Commercial use, license audit | BGE-M3, multilingual-e5-large, Solon, Qwen3-Embedding, Granite R2 | MIT or Apache 2.0 | Review the Gemma terms; jina-clip-v2 is non-commercial |
#Custom enterprise embeddings: open model, fine-tuning, or API
“Custom” does not mean training your own model from scratch. The order that avoids wasting time: a generic open model, careful chunking, hybrid search, and a reranker; then measure recall on your real questions; then, only if failures persist and are tied to business vocabulary (internal references, in-house abbreviations), fine-tune.
In June 2024, Philipp Schmid fine-tuned the bge-base-en-v1.5 model on 6,300 question-passage pairs drawn from financial documents: the retrieval score improved by approximately 7.4% on its test set, with three minutes of training on a consumer graphics card. This is an English case, on a single corpus, with pairs generated by an LLM: an order of magnitude, not a promise. BAAI also documents fine-tuning BGE-M3.
| Option | When to choose it | Limitations |
|---|---|---|
| Generic open model locally (BGE-M3, Qwen3-Embedding) | Default starting point; text stays on your network; MIT or Apache license | Domain-specific vocabulary not learned; evaluate it on your documents |
| Fine-tuning an open model | Recall stuck despite hybrid search and reranking; several thousand question-passage pairs | Corpus to re-encode; model to version like software; risk of overfitting |
| Embeddings API (Mistral, OpenAI) | No GPU, moderate volume; text-embedding-3-large and mistral-embed lead the 2024 MTEB-French study | Text leaves your network; the model may change on the provider's side; recurring cost |
#Multimodal embedding: searching images and document pages
A multimodal embedding model places text and images in the same vector space. You can then retrieve a photo from a sentence, or a scanned PDF page without using OCR. The models from the Ollama library consulted for this page (BGE-M3, Qwen3-Embedding, EmbeddingGemma, nomic-embed-text, mxbai-embed-large) accept text only: the multimodal models below are used through Hugging Face (Sentence-Transformers or Transformers).
| Model | Inputs | License | Good to know |
|---|---|---|---|
| Qwen3-VL-Embedding 2B / 8B | Text, images, screenshots, video | Apache 2.0 | 32,000 tokens; 2,048 and 4,096 dimensions; adjustable dimensions |
| jina-clip-v2 | Text and images; 94 languages | CC BY-NC 4.0 | Non-commercial: unsuitable for a paid product or service without the publisher’s permission |
| ColPali v1.3 | Image-based document pages, multi-vector | MIT (spec sheet) | English-labeled sheet; multiple vectors per page change storage requirements |
A multimodal model is not better for pure text. In the tables published by Qwen, Qwen3-VL-Embedding-2B scores 63.87 on multilingual MTEB, compared with 64.33 for Qwen3-Embedding-0.6B, a text model more than three times smaller. For PDFs whose useful content is text, convert them to clean text (Docling or OCR) and keep a text embedding model. Reserve multimodal models for visual corpora: diagrams, plans, screenshots, and slides.
#Managing your embeddings: prefixes, truncation, storage, and reindexing
A RAG system often degrades more because of these details than because of the model choice. First detail: each model family expects a different input format for queries and documents.
| Model | When faced with the question | In front of the document |
|---|---|---|
| BGE-M3 | None | None |
| multilingual-e5-large | query: | passage: (required, even in French) |
| Solon-embeddings-large-0.1 | query: | None |
| nomic-embed-text v1.5 | search_query: | search_document: |
| mxbai-embed-large-v1 | Represent this sentence for searching relevant passages: | None |
| Qwen3-Embedding | Instruct: {task in one sentence}, line break, Query: {question} | None |
| EmbeddingGemma | task: search result | query: {question} | title: {title or none} | text: {content} |
With Ollama, the official mxbai-embed-large example places the prefix directly in the text sent: add it yourself in your code. Qwen indicates that instructions generally provide a 1 to 5% improvement, and recommends writing them in English even for a multilingual corpus.
#The Ollama trap: silent truncation
The /api/embed endpoint of Ollama has a truncate parameter whose default value is true: input that exceeds the model's context window is cut off without an error. With mxbai-embed-large (512 tokens in Ollama), a 700-token chunk is indexed without its ending, and no one notices. The context window of Ollama may also differ from the original specifications: 2K for nomic-embed-text versus the 8,192 claimed by Nomic. Set truncate to false during testing: an error is better than truncated text.
#Dimensions, storage, and indexing limits
Storage is calculated simply: number of chunks × dimensions × 4 bytes in float32, excluding the index. For one million chunks, the vectors alone occupy:
| Dimensions | Model example | Storage |
|---|---|---|
| 256 | Reduced EmbeddingGemma or Qwen3 models (Matryoshka) | 1.0 GB |
| 384 | all-MiniLM-L12-v2 | 1.5 GB |
| 768 | EmbeddingGemma, Granite R2, nomic | 3.1 GB |
| 1 024 | BGE-M3, multilingual-e5-large, Qwen3-0.6B | 4.1 GB |
| 2 560 | Qwen3-Embedding-4B | 10.2 GB |
| 4 096 | Qwen3-Embedding-8B | 16.4 GB |
Databases impose limits too. With pgvector, an index supports vectors with at most 2,000 dimensions, or 4,000 in half precision (halfvec): Qwen3-Embedding-8B’s 4,096 dimensions exceed even that limit. The workaround is to truncate the vectors, which models trained with Matryoshka (Qwen3-Embedding, EmbeddingGemma, nomic v1.5) support, and then renormalize them, as Google specifies for EmbeddingGemma. The /api/embed API of Ollama accepts a dimensions parameter, intended for these models.
#Version and reindex
- Metadata to retain
- Exact model name, revision, dimension, prefix template, chunking parameters, indexing date. Without them, no one knows why retrieval changed.
- Model change
- Re-encode the entire corpus into a new collection, evaluate it, then switch over. Never mix two models in the same collection.
- Normalization
- Ollama returns L2-normalized vectors: use the same metric (cosine or dot product) everywhere.
#Use a model in practice and test it on your documents
Keep only the ordering of the two scores in mind: the first must clearly exceed the second. To seriously compare two or three candidates, the following procedure supersedes all published rankings.
- 01Build a real-world test setCollect 50 to 100 questions asked by real users (support, tickets, frequently asked questions) and note for each the chunk containing the answer. Questions invented by an LLM make the results look better than they are.
- 02Encode with each candidate modelEncode the chunks and then the questions according to the prefix table, using the same chunk length and storage dimension for all candidates.
- 03Measure recall@5For each question, check whether the correct chunk appears among the top five results. The script below performs this calculation.
- 04Decide based on costKeep the smallest model whose recall@5 remains within one or two points of the best: a model eight times larger costs more in storage and time on every reindexing.
- Sentence Transformers: local embeddings
- Chunking strategies: chunk size versus the model’s context
- Hybrid BM25 + vector search
- Add a reranker before fine-tuning
- pgvector: dimension limits and indexes
- Ragas: evaluate your RAG with numbers
- Source: official BGE-M3 specification sheet (BAAI)
- Source: MTEB-French, study by Ciancone et al. (2024)
- Source: Ollama embedding documentation
- Source: Qwen3-Embedding specifications
- Source: EmbeddingGemma fact sheet (Google)
What is the best embedding model for French?+
Can you use nomic-embed-text or mxbai-embed-large in French?+
Is there a bge-m4 model?+
Do you need to reindex when changing the embedding model?+
Can a multimodal embedding model replace a text model?+
Do you need a custom enterprise embedding model?+
Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.