Intermediate 14 minEmbeddings

The best embedding models FR

Direct response

For French, BGE-M3 is the safest starting point: multilingual, 8,192-token context, MIT license, available in Ollama. Qwen3-Embedding and EmbeddingGemma are the recent alternatives to test. English-focused models (nomic-embed-text, mxbai-embed-large, all-MiniLM) should be avoided for French. Validate the final choice on your own documents.

An embedding model transforms each piece of text into a vector: it determines that “CDD” and “contrat à durée déterminée” are close neighbors. For French, the MTEB-French benchmark measures a 22-point retrieval gap between BGE-M3 and all-MiniLM-L12-v2. This page ranks open models usable locally, cites published measurements and their limitations, then covers what costs money in production: prefixes, truncation, dimensions, storage, and reindexing.

By Mohamed Meguedmi·Update 2026-09-29·Tested on Windows, macOS, and Linux

#What an embedding model is, and why French makes it more difficult

An embedding model is a neural network that converts text (a sentence, paragraph, or chunk) into a vector of numbers, from 384 to 4,096 dimensions depending on the model. Two texts with similar meanings produce nearby vectors, most often measured using cosine similarity. This is what allows a RAG system to retrieve a passage about a “fixed-term contract” when the user writes “CDD,” with no words in common. The same model must encode the questions and documents, as the Ollama documentation points out; changing models therefore requires reindexing the entire corpus. To choose, three questions are enough: was the model trained on French, does its context cover your chunks, and does its license allow your use?

French adds its own difficulties: accents, elisions (“l'employeur”), legal acronyms, and technical Anglicisms mixed into the text. The gap between models is clear. In the MTEB-French benchmark, the retrieval score ranges from 0.43 for all-MiniLM-L12-v2 to 0.65 for BGE-M3, a 22-point difference on the same question set.

i
Size ≠ quality
A longer vector is not inherently more accurate. On the multilingual MTEB benchmark, EmbeddingGemma goes from 61.15 to 60.71 when reducing its vectors from 768 to 512 dimensions, and Google also lists sizes of 256 and 128 dimensions. The storage savings are proportional; the quality loss is not.

#The five criteria that truly distinguish the models

The Local RAG Kit

Your documents, your AI: a reliable local RAG over your PDFs, notes and mail — nothing leaves your machine.

  • Lifetime online access
  • PDF + files
  • Lifetime updates
Training languages
BGE-M3 and Qwen3-Embedding claim more than 100 languages, while multilingual-e5-large supports 94. The model cards for nomic-embed-text-v1.5 and mxbai-embed-large-v1, which are heavily downloaded on Ollama, are labeled English.
Maximum context
512 tokens for multilingual-e5-large, Solon, and mxbai-embed-large; 2 048 for EmbeddingGemma; 8 192 for BGE-M3; 32 000 for Qwen3-Embedding and Granite R2. A longer chunk is truncated, not rejected.
Dimensions and storage
From 384 to 4,096 dimensions, or 4 bytes per dimension and per vector in float32 (calculation below).
License
MIT for BGE-M3, multilingual-e5-large, and Solon; Apache 2.0 for Qwen3-Embedding, Granite R2, nomic, and mxbai; Gemma terms for EmbeddingGemma; CC BY-NC 4.0 (non-commercial) for jina-clip-v2.
Prefixes and instructions
Some models require a prefix before each text ("query:" and "passage:" for E5, even in French). Forgetting it degrades retrieval without producing an error.

#The 2026 ranking for French: nine models compared

This ranking is editorial, not based on in-house testing: it combines French-language coverage, published benchmarks, and local availability. The weights come from the Ollama library when the model is listed there; otherwise, from the float32 estimate in the MTEB-French study.

Embedding models for French: specifications published by vendors
ModelDimensions / contextLicenseWeightsWhat the sources say about French
BGE-M3 (BAAI)1 024 / 8 192MIT1.2 GB (Ollama)Retrieval MTEB-French 0.65. Dense, sparse, and multi-vector. Recommended starting point.
Qwen3-Embedding 0.6B / 4B / 8B1 024 / 2 560 / 4 096 ; 32 000Apache 2.0639 MB / 2.5 GB / 4.7 GB (Ollama)More than 100 languages. Multilingual MTEB according to Qwen: 64.33 / 69.45 / 70.58.
EmbeddingGemma 300M (Google)768 (512, 256, 128) / 2 048Gemma622 MB (Ollama)More than 100 languages. Multilingual MTEB v2: 61.15 according to Google.
multilingual-e5-large1 024 / 512MIT2.24 GB (float32)Retrieval MTEB-French 0.59. Prefixes required.
Solon-embeddings-large-0.11 024 / 512MIT2.24 GB (float32)French model published by Ordalie Technologies. Retrieval MTEB-French 0.63.
Granite Embedding 311M Multilingual R2 (IBM)768 / 32 768Apache 2.0311M parametersFrench among the 52 strengthened languages. 65.2 on multilingual MTEB retrieval according to IBM; no independent French measurement consulted.
nomic-embed-text v1.5768 / 8,192 (2,048 under Ollama)Apache 2.0274 MB (Ollama)English-labeled card. Reserve it for English corpora.
mxbai-embed-large-v1512 contextApache 2.0670 MB (Ollama)English-labeled card. Reserve it for English corpora.
all-MiniLM-L12-v2384Apache 2.033M parametersRetrieval MTEB-French 0.43. CPU-only prototype.

BGE-M3 remains the best default: its French retrieval has been published, its 8,192-token context avoids most forced chunking, and it also provides the lexical weights useful for hybrid search. Qwen3-Embedding is the candidate when a graphics card is available: its 8B version topped the multilingual MTEB leaderboard when it launched, on June 5, 2025, according to Qwen, but its 4,096 dimensions weigh on storage. EmbeddingGemma targets modest machines.

Solon, often presented as a specialist in French legal and administrative language, does not outperform any of the three retrieval datasets in the MTEB-French study: it matches BGE-M3 on the Syntec collective bargaining agreement and trails it on legal articles (BSARD) and school questions (Alloprof).

#What French benchmarks say—and what they don’t

The published reference is MTEB-French (Ciancone et al., May 2024): 18 datasets, 8 task categories, and 51 models compared. The authors conclude that large multilingual models pretrained on sentence similarity perform exceptionally well. Here is the retrieval (NDCG@10) on three datasets: French legal articles (BSARD), the Syntec collective agreement, and Alloprof school questions.

MTEB-French: retrieval (NDCG@10) by dataset, 2024 study
ModelAverage retrievalLaw (BSARD)Syntec agreementAlloprof
text-embedding-3-large (OpenAI API)0,730,730,870,60
mistral-embed (API Mistral)0,680,680,790,57
BGE-M30,650,600,850,49
Solon-embeddings-large-0.10,630,580,850,47
multilingual-e5-large0,590,590,810,38
multilingual-e5-small0,520,520,760,27
paraphrase-multilingual-MiniLM-L12-v20,440,380,660,27
all-MiniLM-L12-v20,430,340,610,33

Three limitations prevent this from being a definitive ranking. The study dates from 2024: it includes neither Qwen3-Embedding (June 2025), nor EmbeddingGemma (September 2025), nor Granite R2, and no comparable French measurements of these models were consulted for this page. The corpora (legal articles, collective bargaining agreements, school questions) may not resemble yours. Finally, the multilingual scores published by vendors are self-reported and combine all languages.

!
None of these figures are yours
These tables come from a published study and vendor datasheets, not an in-house test. They indicate where to look, not which model will win for you. To decide between two similar models, the method in the final section uses your own data.

#Which model for which use case: decision table

Choosing an embedding model for French
Your situationModel to test firstWhyKey consideration
French or French + English corpus, capable machineBGE-M3French retrieval score of 0.65, on par with the study's best open models1.2 GB in Ollama; no prefix
Legal, administration, HRBGE-M3, then Solon for comparisonBGE-M3 matches or exceeds Solon on all three French datasets in the studyA model without an internal legal test set remains a gamble
Modest machine or CPU onlyEmbeddingGemma 300M, or multilingual-e5-small622 MB in Ollama; designed by Google for laptops and mobile devices2,048-token context: short chunks
Available graphics card, maximum qualityQwen3-Embedding-4B or 8B32,000 tokens, adjustable dimensions, more than 100 languages2,560 and 4,096 dimensions: storage and index limits
Long chunks, structured documentsBGE-M3, Qwen3-Embedding, or Granite R28 192 to 32 768 context tokensA very long chunk dilutes the meaning
Commercial use, license auditBGE-M3, multilingual-e5-large, Solon, Qwen3-Embedding, Granite R2MIT or Apache 2.0Review the Gemma terms; jina-clip-v2 is non-commercial

#Custom enterprise embeddings: open model, fine-tuning, or API

“Custom” does not mean training your own model from scratch. The order that avoids wasting time: a generic open model, careful chunking, hybrid search, and a reranker; then measure recall on your real questions; then, only if failures persist and are tied to business vocabulary (internal references, in-house abbreviations), fine-tune.

In June 2024, Philipp Schmid fine-tuned the bge-base-en-v1.5 model on 6,300 question-passage pairs drawn from financial documents: the retrieval score improved by approximately 7.4% on its test set, with three minutes of training on a consumer graphics card. This is an English case, on a single corpus, with pairs generated by an LLM: an order of magnitude, not a promise. BAAI also documents fine-tuning BGE-M3.

Three ways to get an embedding suited to a business
OptionWhen to choose itLimitations
Generic open model locally (BGE-M3, Qwen3-Embedding)Default starting point; text stays on your network; MIT or Apache licenseDomain-specific vocabulary not learned; evaluate it on your documents
Fine-tuning an open modelRecall stuck despite hybrid search and reranking; several thousand question-passage pairsCorpus to re-encode; model to version like software; risk of overfitting
Embeddings API (Mistral, OpenAI)No GPU, moderate volume; text-embedding-3-large and mistral-embed lead the 2024 MTEB-French studyText leaves your network; the model may change on the provider's side; recurring cost
→
The right test before fine-tuning
If the relevant passage appears in the top twenty results but not the top five, that is a reranker use case: BAAI also recommends hybrid search followed by reranking. Embedding fine-tuning targets passages that never appear in the results.

#Multimodal embedding: searching images and document pages

A multimodal embedding model places text and images in the same vector space. You can then retrieve a photo from a sentence, or a scanned PDF page without using OCR. The models from the Ollama library consulted for this page (BGE-M3, Qwen3-Embedding, EmbeddingGemma, nomic-embed-text, mxbai-embed-large) accept text only: the multimodal models below are used through Hugging Face (Sentence-Transformers or Transformers).

Open multimodal embedding models
ModelInputsLicenseGood to know
Qwen3-VL-Embedding 2B / 8BText, images, screenshots, videoApache 2.032,000 tokens; 2,048 and 4,096 dimensions; adjustable dimensions
jina-clip-v2Text and images; 94 languagesCC BY-NC 4.0Non-commercial: unsuitable for a paid product or service without the publisher’s permission
ColPali v1.3Image-based document pages, multi-vectorMIT (spec sheet)English-labeled sheet; multiple vectors per page change storage requirements

A multimodal model is not better for pure text. In the tables published by Qwen, Qwen3-VL-Embedding-2B scores 63.87 on multilingual MTEB, compared with 64.33 for Qwen3-Embedding-0.6B, a text model more than three times smaller. For PDFs whose useful content is text, convert them to clean text (Docling or OCR) and keep a text embedding model. Reserve multimodal models for visual corpora: diagrams, plans, screenshots, and slides.

#Managing your embeddings: prefixes, truncation, storage, and reindexing

A RAG system often degrades more because of these details than because of the model choice. First detail: each model family expects a different input format for queries and documents.

Prefixes expected by the model (questions and documents)
ModelWhen faced with the questionIn front of the document
BGE-M3NoneNone
multilingual-e5-largequery: passage: (required, even in French)
Solon-embeddings-large-0.1query: None
nomic-embed-text v1.5search_query: search_document:
mxbai-embed-large-v1Represent this sentence for searching relevant passages: None
Qwen3-EmbeddingInstruct: {task in one sentence}, line break, Query: {question}None
EmbeddingGemmatask: search result | query: {question}title: {title or none} | text: {content}

With Ollama, the official mxbai-embed-large example places the prefix directly in the text sent: add it yourself in your code. Qwen indicates that instructions generally provide a 1 to 5% improvement, and recommends writing them in English even for a multilingual corpus.

#The Ollama trap: silent truncation

The /api/embed endpoint of Ollama has a truncate parameter whose default value is true: input that exceeds the model's context window is cut off without an error. With mxbai-embed-large (512 tokens in Ollama), a 700-token chunk is indexed without its ending, and no one notices. The context window of Ollama may also differ from the original specifications: 2K for nomic-embed-text versus the 8,192 claimed by Nomic. Set truncate to false during testing: an error is better than truncated text.

#Dimensions, storage, and indexing limits

Storage is calculated simply: number of chunks × dimensions × 4 bytes in float32, excluding the index. For one million chunks, the vectors alone occupy:

One million float32 chunks, vectors only (calculation)
DimensionsModel exampleStorage
256Reduced EmbeddingGemma or Qwen3 models (Matryoshka)1.0 GB
384all-MiniLM-L12-v21.5 GB
768EmbeddingGemma, Granite R2, nomic3.1 GB
1 024BGE-M3, multilingual-e5-large, Qwen3-0.6B4.1 GB
2 560Qwen3-Embedding-4B10.2 GB
4 096Qwen3-Embedding-8B16.4 GB

Databases impose limits too. With pgvector, an index supports vectors with at most 2,000 dimensions, or 4,000 in half precision (halfvec): Qwen3-Embedding-8B’s 4,096 dimensions exceed even that limit. The workaround is to truncate the vectors, which models trained with Matryoshka (Qwen3-Embedding, EmbeddingGemma, nomic v1.5) support, and then renormalize them, as Google specifies for EmbeddingGemma. The /api/embed API of Ollama accepts a dimensions parameter, intended for these models.

#Version and reindex

Metadata to retain
Exact model name, revision, dimension, prefix template, chunking parameters, indexing date. Without them, no one knows why retrieval changed.
Model change
Re-encode the entire corpus into a new collection, evaluate it, then switch over. Never mix two models in the same collection.
Normalization
Ollama returns L2-normalized vectors: use the same metric (cosine or dot product) everywhere.

#Use a model in practice and test it on your documents

Ollama: download BGE-M3 and generate an embedding
ollama pull bge-m3

curl http://localhost:11434/api/embed -d '{
  "model": "bge-m3",
  "input": "Le contrat débute le 1er mai."
}'
Ollama in Python: two similar texts, one distant text
import numpy as np
import ollama

textes = [
    "Le contrat débute le 1er mai.",
    "Date de début : 01/05.",
    "La voiture est rouge.",
]
vecteurs = np.array(ollama.embed(model="bge-m3", input=textes)["embeddings"])

# Les vecteurs d'Ollama sont normalisés : le produit scalaire vaut le cosinus
print(vecteurs[0] @ vecteurs[1])
print(vecteurs[0] @ vecteurs[2])

Keep only the ordering of the two scores in mind: the first must clearly exceed the second. To seriously compare two or three candidates, the following procedure supersedes all published rankings.

  1. 01
    Build a real-world test set
    Collect 50 to 100 questions asked by real users (support, tickets, frequently asked questions) and note for each the chunk containing the answer. Questions invented by an LLM make the results look better than they are.
  2. 02
    Encode with each candidate model
    Encode the chunks and then the questions according to the prefix table, using the same chunk length and storage dimension for all candidates.
  3. 03
    Measure recall@5
    For each question, check whether the correct chunk appears among the top five results. The script below performs this calculation.
  4. 04
    Decide based on cost
    Keep the smallest model whose recall@5 remains within one or two points of the best: a model eight times larger costs more in storage and time on every reindexing.
Sentence-Transformers: recall@5 on your data
import numpy as np
from sentence_transformers import SentenceTransformer

modele = SentenceTransformer("BAAI/bge-m3")

chunks = [...]      # vos chunks (liste de textes)
questions = [...]   # 50 à 100 vraies questions
cible = [...]       # pour chaque question, l'index du bon chunk

C = modele.encode(chunks, normalize_embeddings=True)
Q = modele.encode(questions, normalize_embeddings=True)

top5 = np.argsort(-(Q @ C.T), axis=1)[:, :5]
recall5 = np.mean([cible[i] in top5[i] for i in range(len(questions))])
print(f"recall@5 = {recall5:.0%}")
→
Persist the vectors
Re-encoding the entire corpus at every launch is the slowest part of a beginner RAG setup. Store the vectors in a database (Chroma, Qdrant, pgvector, FAISS) and encode only new chunks.
FAQ
What is the best embedding model for French?+
BGE-M3 is the best starting point: its French retrieval performance is published (0.65 on MTEB-French), it supports 8,192 tokens, its license is MIT, and Ollama offers it. Qwen3-Embedding and EmbeddingGemma are newer and absent from this study. Test them on fifty questions from your corpus before deciding.
Can you use nomic-embed-text or mxbai-embed-large in French?+
Their Hugging Face cards are labeled English, and they are absent from the MTEB-French study. They technically work with French, but quality is not guaranteed. In Ollama, nomic-embed-text exposes only 2,048 context tokens, versus the 8,192 claimed by Nomic. Between the two, nomic offers more context (2,048 tokens in Ollama versus 512); mxbai requires a prefix only for queries. Prefer BGE-M3.
Is there a bge-m4 model?+
The official FlagEmbedding repository mentions no bge-m4. The multilingual model in the family is called BGE-M3, and M3 refers to three qualities: multi-function (dense, sparse, multi-vector), multilingual (more than 100 languages), and multi-granularity (up to 8,192 tokens). If a site offers bge-m4, verify its origin before downloading.
Do you need to reindex when changing the embedding model?+
Yes, completely. The vectors from two models live in different spaces and often don't even have the same dimensions. Create a new collection, re-encode all chunks with the new model and its prefix template, measure recall, then switch over. Never mix two models in the same collection.
Can a multimodal embedding model replace a text model?+
No, unless the content is visual. In the Qwen tables, Qwen3-VL-Embedding-2B (63,87) remains below Qwen3-Embedding-0.6B (64,33) for text alone. For text-based PDFs, convert them to text and keep a text embedding. Multimodal models are useful for diagrams, screenshots, and slides; note that jina-clip-v2 is non-commercial.
Do you need a custom enterprise embedding model?+
Rarely at first. A generic open model, careful chunking, hybrid search, and a reranker handle most cases. Fine-tuning makes sense when recall remains limited by domain-specific vocabulary and you have several thousand question-passage pairs, followed by a complete reindexing.
Did this guide help you?

Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.