Sentence Transformers: embeddings in local
Sentence Transformers is a Python library that loads an embedding model from Hugging Face and turns text into vectors with model.encode, on the CPU or GPU; model.similarity then compares those vectors. For French, choose a multilingual model: an English-language model such as all-MiniLM-L6-v2 performs poorly. Keep your passages below the model's maximum length, because anything beyond it is truncated without warning.
This guide shows how to install the library, encode a corpus, choose a model that understands French, avoid silent errors (length limits, query prefixes, two mixed models), and speed up indexing on your machine. It also compares Sentence Transformers with the embeddings API from Ollama, to help you decide which one to use.
#What an embedding is and what the library does
An embedding model transforms text into a list of numbers—a vector—so that two texts with similar meanings have similar vectors. The Sentence Transformers documentation describes these models, known as bi-encoders, as computing a fixed-size representation for a text, with embedding computation that is often efficient and similarity computation that is very fast. This is the foundation of RAG: encode the question, find the passages whose vectors are closest, and give them to the language model. Two consequences to remember: the model that encodes the documents must be the one that encodes the questions, and its notion of “close” comes from its training. A model trained primarily on English is an unreliable judge of French.
This model produces 384-dimensional vectors and serves as a starting example, but it is designed for English: for a French corpus, replace it as indicated below. The library automatically places the model on the best available device (cuda, mps, or cpu), and you can force the choice with the device parameter.
#Install and encode a French corpus
Your documents, your AI: a reliable local RAG over your PDFs, notes and mail — nothing leaves your machine.
- Lifetime online access
- PDF + files
- Lifetime updates
- 01Install the libraryUse pip install -U sentence-transformers in a Python environment. Version 6.1.0 was released on September 18, 2026, and requires Python 3.10 or later, according to PyPI.
- 02Choose a multilingual modelChoose a model advertised as multilingual (see the table below), not an all-* model trained for English.
- 03Batch-encode documentsPass a list of texts to encode: the library processes them in batches, making much better use of the hardware than one call per text.
- 04Encode the question with the same modelUse the same model and prefixes as for the documents, then compare with similarity.
- 05Keep the model name with the indexNote it in the metadata: if you change it, you'll need to re-encode the entire corpus.
The query: prefix for questions and passage: prefix for documents is the one cited in the Sentence Transformers documentation for this model. Without it, retrieval degrades without an error message. The encode prompt parameter applies the prefix to each text.
#Choose a model that understands French
The documentation suggests original models and recommends consulting the MTEB leaderboard for inspiration, with two caveats: filter out models that are too large for your hardware, and experiment, because highly ranked models may not perform well on your own tasks. The following table summarizes what the documentation says about the cited models.
| Model | What the documentation says | For French |
|---|---|---|
| all-MiniLM-L6-v2 | About 5 times faster than all-mpnet-base-v2, good quality; 384 dimensions, 256 tokens maximum | No: English-oriented |
| all-mpnet-base-v2 | Best quality in the all-* family, general-purpose model | No: English-oriented |
| multi-qa-mpnet-base-cos-v1 | Trained for semantic search on 215 million question-answer pairs | No: English-oriented |
| paraphrase-multilingual-MiniLM-L12-v2 | Trained on parallel data for more than 50 languages | Yes, designed for sentence similarity |
| paraphrase-multilingual-mpnet-base-v2 | Same family, more than 50 languages | Yes, heavier |
| distiluse-base-multilingual-cased-v1 | Supports 15 languages, including French | Yes, limited languages |
| multilingual-e5-large | Requires the query: and passage: prefixes (as specified in the documentation) | Yes, multilingual retrieval |
#Errors that ruin a corpus without an error message
- An English-language model on French
- Everything works, yet retrieval is still skewed. This is the most common error.
- Passages longer than the model limit
- The documentation specifies that longer texts are truncated to the first tokens up to max_seq_length: the end of each long passage becomes invisible.
- A different model for documents and questions
- The vectors are no longer talking about the same thing: bizarre results, generally introduced by updating only part of the pipeline.
- Forgotten prefixes
- Some models require a different prefix for questions and documents; omitting it degrades retrieval.
- Unnormalized vectors with a poor score
- If the model expects cosine similarity, normalize the vectors or use the corresponding metric; otherwise, ranking quality degrades.
#Measuring similarity: normalization and cosine similarity
Two vectors are compared using a similarity measure, most often cosine similarity. The Ollama documentation states that its endpoint returns normalized vectors and recommends cosine for most semantic searches. With normalized vectors, cosine and dot product produce the same ranking, allowing you to use an index optimized for dot product. The point to watch is consistency: the vector database's metric must match the one the model expects, and it is listed in its documentation.
For length, the documentation gives an order of magnitude: 512 tokens for many BERT-type models, or 300 to 400 words in English, and 256 tokens for all-MiniLM-L6-v2. French uses more tokens per word: split your passages below the limit with some headroom, and check model.max_seq_length. You can reduce it, but not increase it beyond what the model supports. The documentation adds that a model trained on short texts represents long texts less effectively.
#Index quickly on your machine
- CPU or GPU
- The model runs on the best available device. A GPU speeds up bulk indexing but makes little difference for a single query.
- Reduced precision
- On a GPU, switching to float16 or bfloat16 speeds up inference with minimal precision loss, according to the documentation.
- ONNX and OpenVINO backends
- The documentation claims possible speedups of up to 2 to 3 times depending on the hardware and backend; test them on your machine.
- Multiple GPUs or processes
- encode accepts a list of devices: useful for large corpora, less so for small ones because of startup overhead.
- More compact vectors
- Binary or integer quantization of vectors and models with truncatable dimensions reduce storage and search costs.
- Caching and indexing
- Don't re-encode unchanged documents: the cache key is based on the content, not the filename. On a machine that also serves a language model, index when no one is using it.
#How much a vector index weighs
The size of an index is calculated by multiplying the number of vectors by their dimensionality and by four bytes if the numbers are float32. A 384-dimensional vector takes 1,536 bytes: one million passages represent about 1.5 GB. At 1,024 dimensions, the same million takes about 4 GB. These figures exclude metadata and the vector database's index structures. The model choice therefore directly affects the RAM required for search, and vector quantization mentioned above can reduce this footprint.
#Sentence Transformers or Ollama's embeddings API
Ollama also exposes embeddings: the documentation states that the ollama run peut command produces vectors and that the api/embed endpoint returns L2-normalized vectors. It recommends the embeddinggemma, qwen3-embedding, and all-minilm models, with vectors typically ranging from 384 to 1 024 dimensions, and reiterates two rules: cosine similarity for most searches, and the same model for indexing and querying.
| Criterion | Sentence Transformers | Ollama embeddings API |
|---|---|---|
| Installation | Python library to install | Already present if you use Ollama |
| Model selection | Any Hugging Face-compatible model | Models from the Ollama library |
| Prefix and length control | End: prompt, max_seq_length, named prompts | Depending on the model and the call |
| Training or fine-tuning | Yes, the library supports it | No |
| Typical use | Bulk indexing, experiments, reranking | Simple integration into a pipeline already built on Ollama |
#Embedding is only a first-pass filter
Sentence Transformers documentation describes the bi-encoder as the first step in a two-stage search, where a cross-encoder, or reranker, re-ranks the best results. The cross-encoder reads the question and each passage together: it is slower but more accurate on a handful of candidates. Retrieving around twenty passages by similarity, then re-ranking them to keep only a few, is one of the least expensive improvements to an RAG pipeline. The dedicated reranker guide explains how to set it up; measure the gain on your twenty questions before keeping it, because it adds latency to every query, a second model to maintain, and additional memory usage.
A good first step, even before comparing models, is to look at what your corpus contains: short, self-contained sentences or long technical paragraphs? Questions phrased as keywords or complete sentences? The answer guides your choice of length limit, chunking strategy, and model. A dense legal corpus calls for different settings than a short question-and-answer database, and no public ranking can make that decision for you.
#Evaluate the model choice before indexing
- 01Write twenty real-world questionsUse questions your users will ask, phrased in their words, not the document's.
- 02Record the expected transitionFor each question, identify the passage or passages that contain the answer.
- 03Compare two or three modelsEncode the same corpus with each model and check whether the correct passage appears among the top five results.
- 04Replay on every changeModel, chunking, or prefix changed: rerun this short test before reindexing.
#FAQ
Do you need a GPU for Sentence Transformers?+
Which Sentence Transformers model should you choose for French?+
Can you use its conversational model for encoding?+
What happens if I change the embedding model?+
What happens if my text exceeds the maximum length?+
Sentence Transformers or Ollama for embeddings?+
#Go further
- The best embedding models for French
- BGE-M3: embeddings that really speak French
- Add a reranker to your pipeline
- Chunking strategies
- Qdrant: the vector database for local RAG
- Local RAG: introduction
- Source: Sentence Transformers, quickstart
- Source: Sentence Transformers, embedding computation
- Source: Sentence Transformers, pretrained models
- Source: Ollama, embeddings
Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.