Intermediate 11 minEmbeddings

Sentence Transformers: embeddings in local

Direct response

Sentence Transformers is a Python library that loads an embedding model from Hugging Face and turns text into vectors with model.encode, on the CPU or GPU; model.similarity then compares those vectors. For French, choose a multilingual model: an English-language model such as all-MiniLM-L6-v2 performs poorly. Keep your passages below the model's maximum length, because anything beyond it is truncated without warning.

This guide shows how to install the library, encode a corpus, choose a model that understands French, avoid silent errors (length limits, query prefixes, two mixed models), and speed up indexing on your machine. It also compares Sentence Transformers with the embeddings API from Ollama, to help you decide which one to use.

By Mohamed Meguedmi·Update 2026-09-30·Tested on Windows, macOS, and Linux

#What an embedding is and what the library does

An embedding model transforms text into a list of numbers—a vector—so that two texts with similar meanings have similar vectors. The Sentence Transformers documentation describes these models, known as bi-encoders, as computing a fixed-size representation for a text, with embedding computation that is often efficient and similarity computation that is very fast. This is the foundation of RAG: encode the question, find the passages whose vectors are closest, and give them to the language model. Two consequences to remember: the model that encodes the documents must be the one that encodes the questions, and its notion of “close” comes from its training. A model trained primarily on English is an unreliable judge of French.

The first example in the official documentation
from sentence_transformers import SentenceTransformer

model = SentenceTransformer("sentence-transformers/all-MiniLM-L6-v2")
sentences = [
    "The weather is lovely today.",
    "It's so sunny outside!",
    "He drove to the stadium.",
]
embeddings = model.encode(sentences)
print(embeddings.shape)  # [3, 384]
similarities = model.similarity(embeddings, embeddings)

This model produces 384-dimensional vectors and serves as a starting example, but it is designed for English: for a French corpus, replace it as indicated below. The library automatically places the model on the best available device (cuda, mps, or cpu), and you can force the choice with the device parameter.

#Install and encode a French corpus

The Local RAG Kit

Your documents, your AI: a reliable local RAG over your PDFs, notes and mail — nothing leaves your machine.

  • Lifetime online access
  • PDF + files
  • Lifetime updates
  1. 01
    Install the library
    Use pip install -U sentence-transformers in a Python environment. Version 6.1.0 was released on September 18, 2026, and requires Python 3.10 or later, according to PyPI.
  2. 02
    Choose a multilingual model
    Choose a model advertised as multilingual (see the table below), not an all-* model trained for English.
  3. 03
    Batch-encode documents
    Pass a list of texts to encode: the library processes them in batches, making much better use of the hardware than one call per text.
  4. 04
    Encode the question with the same model
    Use the same model and prefixes as for the documents, then compare with similarity.
  5. 05
    Keep the model name with the index
    Note it in the metadata: if you change it, you'll need to re-encode the entire corpus.
Semantic search with a multilingual model
from sentence_transformers import SentenceTransformer

model = SentenceTransformer("intfloat/multilingual-e5-large")

docs = [
    "Le contrat peut être résilié avec un préavis de trois mois.",
    "La facture est payable à trente jours fin de mois.",
]
question = "Quel est le délai pour mettre fin au contrat ?"

doc_emb = model.encode(docs, prompt="passage: ", normalize_embeddings=True)
q_emb = model.encode(question, prompt="query: ", normalize_embeddings=True)
print(model.similarity(q_emb, doc_emb))

The query: prefix for questions and passage: prefix for documents is the one cited in the Sentence Transformers documentation for this model. Without it, retrieval degrades without an error message. The encode prompt parameter applies the prefix to each text.

#Choose a model that understands French

The documentation suggests original models and recommends consulting the MTEB leaderboard for inspiration, with two caveats: filter out models that are too large for your hardware, and experiment, because highly ranked models may not perform well on your own tasks. The following table summarizes what the documentation says about the cited models.

Models cited in the Sentence Transformers documentation
ModelWhat the documentation saysFor French
all-MiniLM-L6-v2About 5 times faster than all-mpnet-base-v2, good quality; 384 dimensions, 256 tokens maximumNo: English-oriented
all-mpnet-base-v2Best quality in the all-* family, general-purpose modelNo: English-oriented
multi-qa-mpnet-base-cos-v1Trained for semantic search on 215 million question-answer pairsNo: English-oriented
paraphrase-multilingual-MiniLM-L12-v2Trained on parallel data for more than 50 languagesYes, designed for sentence similarity
paraphrase-multilingual-mpnet-base-v2Same family, more than 50 languagesYes, heavier
distiluse-base-multilingual-cased-v1Supports 15 languages, including FrenchYes, limited languages
multilingual-e5-largeRequires the query: and passage: prefixes (as specified in the documentation)Yes, multilingual retrieval
i
Sentence similarity and document search are not the same task
A model trained to recognize paraphrases is not necessarily the best at retrieving a passage that answers a question. For RAG, prefer a model trained for retrieval. Our guides on French embeddings and BGE-M3 compare the candidates.

#Errors that ruin a corpus without an error message

An English-language model on French
Everything works, yet retrieval is still skewed. This is the most common error.
Passages longer than the model limit
The documentation specifies that longer texts are truncated to the first tokens up to max_seq_length: the end of each long passage becomes invisible.
A different model for documents and questions
The vectors are no longer talking about the same thing: bizarre results, generally introduced by updating only part of the pipeline.
Forgotten prefixes
Some models require a different prefix for questions and documents; omitting it degrades retrieval.
Unnormalized vectors with a poor score
If the model expects cosine similarity, normalize the vectors or use the corresponding metric; otherwise, ranking quality degrades.

#Measuring similarity: normalization and cosine similarity

Two vectors are compared using a similarity measure, most often cosine similarity. The Ollama documentation states that its endpoint returns normalized vectors and recommends cosine for most semantic searches. With normalized vectors, cosine and dot product produce the same ranking, allowing you to use an index optimized for dot product. The point to watch is consistency: the vector database's metric must match the one the model expects, and it is listed in its documentation.

For length, the documentation gives an order of magnitude: 512 tokens for many BERT-type models, or 300 to 400 words in English, and 256 tokens for all-MiniLM-L6-v2. French uses more tokens per word: split your passages below the limit with some headroom, and check model.max_seq_length. You can reduce it, but not increase it beyond what the model supports. The documentation adds that a model trained on short texts represents long texts less effectively.

#Index quickly on your machine

CPU or GPU
The model runs on the best available device. A GPU speeds up bulk indexing but makes little difference for a single query.
Reduced precision
On a GPU, switching to float16 or bfloat16 speeds up inference with minimal precision loss, according to the documentation.
ONNX and OpenVINO backends
The documentation claims possible speedups of up to 2 to 3 times depending on the hardware and backend; test them on your machine.
Multiple GPUs or processes
encode accepts a list of devices: useful for large corpora, less so for small ones because of startup overhead.
More compact vectors
Binary or integer quantization of vectors and models with truncatable dimensions reduce storage and search costs.
Caching and indexing
Don't re-encode unchanged documents: the cache key is based on the content, not the filename. On a machine that also serves a language model, index when no one is using it.

#How much a vector index weighs

The size of an index is calculated by multiplying the number of vectors by their dimensionality and by four bytes if the numbers are float32. A 384-dimensional vector takes 1,536 bytes: one million passages represent about 1.5 GB. At 1,024 dimensions, the same million takes about 4 GB. These figures exclude metadata and the vector database's index structures. The model choice therefore directly affects the RAM required for search, and vector quantization mentioned above can reduce this footprint.

#Sentence Transformers or Ollama's embeddings API

Ollama also exposes embeddings: the documentation states that the ollama run peut command produces vectors and that the api/embed endpoint returns L2-normalized vectors. It recommends the embeddinggemma, qwen3-embedding, and all-minilm models, with vectors typically ranging from 384 to 1 024 dimensions, and reiterates two rules: cosine similarity for most searches, and the same model for indexing and querying.

Which tool should you use for your embeddings?
CriterionSentence TransformersOllama embeddings API
InstallationPython library to installAlready present if you use Ollama
Model selectionAny Hugging Face-compatible modelModels from the Ollama library
Prefix and length controlEnd: prompt, max_seq_length, named promptsDepending on the model and the call
Training or fine-tuningYes, the library supports itNo
Typical useBulk indexing, experiments, rerankingSimple integration into a pipeline already built on Ollama

#Embedding is only a first-pass filter

Sentence Transformers documentation describes the bi-encoder as the first step in a two-stage search, where a cross-encoder, or reranker, re-ranks the best results. The cross-encoder reads the question and each passage together: it is slower but more accurate on a handful of candidates. Retrieving around twenty passages by similarity, then re-ranking them to keep only a few, is one of the least expensive improvements to an RAG pipeline. The dedicated reranker guide explains how to set it up; measure the gain on your twenty questions before keeping it, because it adds latency to every query, a second model to maintain, and additional memory usage.

A good first step, even before comparing models, is to look at what your corpus contains: short, self-contained sentences or long technical paragraphs? Questions phrased as keywords or complete sentences? The answer guides your choice of length limit, chunking strategy, and model. A dense legal corpus calls for different settings than a short question-and-answer database, and no public ranking can make that decision for you.

#Evaluate the model choice before indexing

  1. 01
    Write twenty real-world questions
    Use questions your users will ask, phrased in their words, not the document's.
  2. 02
    Record the expected transition
    For each question, identify the passage or passages that contain the answer.
  3. 03
    Compare two or three models
    Encode the same corpus with each model and check whether the correct passage appears among the top five results.
  4. 04
    Replay on every change
    Model, chunking, or prefix changed: rerun this short test before reindexing.

#FAQ

FAQ
Do you need a GPU for Sentence Transformers?+
No. Embedding models are small enough to run on the CPU, and the library automatically selects the best available device. A GPU is mainly useful for indexing a large corpus faster; for encoding a single question, the difference is small in most cases.
Which Sentence Transformers model should you choose for French?+
Choose a multilingual model: the documentation cites paraphrase-multilingual-MiniLM-L12-v2 for more than 50 languages, and multilingual-e5-large with the query: and passage: prefixes. Test two or three candidates on your own questions instead of relying solely on the MTEB ranking, whose documentation notes that it does not predict your results.
Can you use its conversational model for encoding?+
No. An embedding model is trained to place texts with similar meanings close together; a conversational model is not, and retrieval becomes poor even though the code appears to work without errors. Use a dedicated embedding model, separate from the model that writes the answers.
What happens if I change the embedding model?+
You need to re-encode the entire corpus because vectors from two models are not comparable: their dimensions and coordinates differ. Record the model name with the index, and allow time to rebuild the vector database before switching, keeping the old index active during the operation.
What happens if my text exceeds the maximum length?+
It is truncated to the first max_seq_length tokens without an error: the end of the passage is ignored by the search. With all-MiniLM-L6-v2, the limit is 256 tokens. Split your texts into passages shorter than the model’s limit, with some headroom.
Sentence Transformers or Ollama for embeddings?+
Sentence Transformers offers more control (prefixes, length, Hugging Face models, training). The Ollama API is simpler if your pipeline already relies on Ollama. In both cases, use the same model for indexing and querying, with the similarity measure recommended for that model.

#Go further

Did this guide help you?

Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.