Intermediate 11 minStack

RAG with ChromaDB and Mistral

Direct response

For a local RAG setup with Mistral, the simplest stack is Ollama (generation and embeddings with bge-m3) plus ChromaDB in file mode, with no server or PyTorch. For the model, ministral-3:8b (6.0 GB, Apache 2.0 license, advertised 256K context) is a good default for an 8 to 12 GB card, ministral-3:14b for 16 GB, and mistral-small3.2:24b (15 GB) beyond that. The setting you must not forget: Ollama's context window, which you need to increase so the passages fit in the prompt.

This guide builds a complete document assistant in two Python scripts, with a Mistral model running on your machine: your PDFs and text files are split up, indexed in ChromaDB, and then the retrieved passages are provided to the model, which answers with citations. It also explains which Mistral model to choose based on your graphics memory and the pitfalls that cause an RAG system to answer beside the point.

By Mohamed Meguedmi·Update 2026-09-30·Tested on Windows, macOS, and Linux

#What we're building: a fully local Mistral RAG

RAG (retrieval-augmented generation) consists of finding the passages in your documents relevant to the question, then inserting them into the model’s prompt so it can answer from them. The intended result here is a small command-line tool: one script indexes a document folder; a second reads a question, finds the five closest passages in ChromaDB, sends them to a Mistral model via Ollama along with the question, and displays the answer followed by the files consulted. Nothing leaves the machine: Ollama serves the generation model and the embedding model, while ChromaDB stores the vectors in a local folder.

Two meanings of “Mistral” are in circulation: Mistral AI's open-weight models, which you download and run yourself (the subject of this guide), and the company's hosted APIs, which send your passages to its servers. For confidential documents, only the former meets the “100% local” requirement.

#Which Mistral model to choose for RAG

The Local RAG Kit

Your documents, your AI: a reliable local RAG over your PDFs, notes and mail — nothing leaves your machine.

  • Lifetime online access
  • PDF + files
  • Lifetime updates

RAG has specific requirements: the model must follow a strict instruction (“answer only from the passages”), read multiple passages without getting lost, and respond in French. Size matters less than for open-ended conversation; available context memory matters more. The sizes below are those displayed by the Ollama library, using the default quantization.

Mistral models in the Ollama library (September 2026 snapshot)
ModelSize OllamaAdvertised contextWho it's for
ministral-3:3b3.0 GB256KMachine without a dedicated GPU; simple responses, low tolerance for complex instructions
ministral-3:8b6.0 GB256KReasonable default for an 8 to 12 GB card or a laptop with 16 GB of memory
ministral-3:14b9.1 GB256K16 GB card, or 12 GB with a moderate context
mistral-nemo (12B)view the Ollama page128KOlder alternative, still widely used
mistral-small3.2:24b15 GB128K24 GB card or 32 GB or more of unified memory; strongest on format instructions
mistral (7B, version 0.3)4.4 GB32KOlder model: reserve it for very limited machines

The Ministral 3 family (3B, 8B, and 14B) is released under the Apache 2.0 license, as stated in the Mistral 3 announcement, and the Ollama page describes it as designed for edge deployment and capable of running on a wide range of hardware. Mistral Small 4, published in 2026 with 119 billion parameters in total according to the name of its Hugging Face listing, targets server hardware: it is not a candidate for a personal machine. For an estimate of the required memory, the site’s VRAM calculator gives the model’s size plus the context cache.

i
Advertised context and useful context
A 128K or 256K context is the model's maximum capacity, not a setting: Ollama uses far less by default, and a large context consumes additional memory. For a RAG with five passages, 8,000 tokens are enough.

#The technical stack

Generation
A Mistral model served by Ollama through the local HTTP API on port 11434.
Embeddings
bge-m3 served by Ollama: the library page describes it as a versatile, multilingual, multi-granularity model from BAAI with 567 million parameters. It avoids installing PyTorch and sentence-transformers.
Vector database
ChromaDB in local mode (PersistentClient): one folder, no server. Chroma provides a wrapper, OllamaEmbeddingFunction, that calls the embeddings API of Ollama.
Reading files
pypdf for PDFs containing text, direct reading for Markdown and plain text. A scanned PDF is an image: it first requires optical character recognition.

#Prepare the environment

  1. 01
    Install Ollama and fetch the models
    Install Ollama, then download the generation model and the embedding model using the two commands below.
  2. 02
    Create the Python environment
    Python 3.10 or later. A virtual environment keeps the project's dependencies separate.
  3. 03
    Place the documents
    Copy your PDFs, Markdown files, and text into a docs/ folder alongside the scripts.
Models and dependencies
ollama pull ministral-3:8b
ollama pull bge-m3

mkdir mon-rag && cd mon-rag
python3 -m venv venv
source venv/bin/activate   # .\venv\Scripts\activate sous Windows
pip install chromadb pypdf requests

#2. Index the documents in ChromaDB

The script reads each file, splits the text into passages of about 1,800 characters at paragraph breaks, then passes them to Chroma, which calls bge-m3 through Ollama to calculate the vectors. Two details matter: each passage keeps the filename as metadata (to cite the source), and additions are made in batches rather than one passage at a time.

index.py
from pathlib import Path
import chromadb
from chromadb.utils.embedding_functions.ollama_embedding_function import OllamaEmbeddingFunction
from pypdf import PdfReader

ef = OllamaEmbeddingFunction(url="http://localhost:11434", model_name="bge-m3")
coll = chromadb.PersistentClient(path="./chroma_db").get_or_create_collection("mes_docs", embedding_function=ef)

def lire(path: Path) -> str:
    if path.suffix.lower() == ".pdf":
        return "\n\n".join(p.extract_text() or "" for p in PdfReader(str(path)).pages)
    return path.read_text(encoding="utf-8", errors="ignore")

def decouper(texte: str, max_chars=1800):
    """Regroupe des paragraphes entiers jusqu'à max_chars ; un paragraphe trop long est coupé."""
    chunks, courant = [], ""
    for para in (p.strip() for p in texte.split("\n\n")):
        if not para:
            continue
        while len(para) > max_chars:
            if courant:
                chunks.append(courant); courant = ""
            chunks.append(para[:max_chars]); para = para[max_chars:]
        if len(courant) + len(para) + 2 > max_chars and courant:
            chunks.append(courant); courant = ""
        courant = (courant + "\n\n" + para).strip()
    if courant:
        chunks.append(courant)
    return chunks

n = 0
for path in sorted(Path("docs").rglob("*")):
    if path.suffix.lower() not in {".pdf", ".md", ".txt"}:
        continue
    chunks = decouper(lire(path))
    if not chunks:
        print(f"  ! {path.name} : aucun texte extrait (PDF scanné ?)")
        continue
    for i in range(0, len(chunks), 32):  # par lots de 32
        lot = chunks[i:i + 32]
        coll.upsert(
            ids=[f"{path.name}-{i + j}" for j in range(len(lot))],
            documents=lot,
            metadatas=[{"source": path.name}] * len(lot),
        )
    n += len(chunks)
    print(f"  + {path.name} : {len(chunks)} passages")
print(f"Terminé : {n} passages indexés")

Using upsert with identifiers built from the filename and pass number makes the script rerunnable: reindexing the same folder updates the passages instead of duplicating them. However, note that if a document gets shorter, the old surplus passages remain in the database; for a major change, delete the chroma_db folder and reindex. The choice of passage size is detailed in the guide to chunking strategies.

#3. Query: search, then generation

The second script embeds the question, retrieves the five closest passages, and builds the prompt. The instruction is crucial: it asks the model to answer only from the passages, admit when information is missing, and cite the file. The num_ctx parameter increases the context window: Ollama’s documentation states that the default window is 4,096 tokens and that the OLLAMA_CONTEXT_LENGTH variable or the num_ctx parameter changes it. With five passages of 400 to 500 tokens, the instruction, and the answer, 4,096 tokens is borderline: an overly short context is truncated silently, and the model answers without reading the end of your passages.

ask.py
import sys, requests
import chromadb
from chromadb.utils.embedding_functions.ollama_embedding_function import OllamaEmbeddingFunction

MODELE = "ministral-3:8b"
ef = OllamaEmbeddingFunction(url="http://localhost:11434", model_name="bge-m3")
coll = chromadb.PersistentClient(path="./chroma_db").get_collection("mes_docs", embedding_function=ef)

SYSTEME = (
    "Tu réponds en français, uniquement à partir des passages fournis. "
    "Si la réponse n'y figure pas, dis-le clairement au lieu de deviner. "
    "Termine chaque affirmation par le nom du fichier source entre crochets."
)

def repondre(question: str, k: int = 5):
    res = coll.query(query_texts=[question], n_results=k)
    passages = list(zip(res["documents"][0], res["metadatas"][0]))
    contexte = "\n\n---\n\n".join(f"[{m['source']}]\n{p}" for p, m in passages)
    r = requests.post("http://localhost:11434/api/chat", json={
        "model": MODELE,
        "stream": False,
        "options": {"temperature": 0.2, "num_ctx": 8192},
        "messages": [
            {"role": "system", "content": SYSTEME},
            {"role": "user", "content": f"PASSAGES :\n{contexte}\n\nQUESTION : {question}"},
        ],
    }, timeout=300)
    r.raise_for_status()
    return r.json()["message"]["content"], sorted({m["source"] for _, m in passages})

if __name__ == "__main__":
    q = " ".join(sys.argv[1:]) or input("Question : ")
    reponse, sources = repondre(q)
    print("\n" + reponse)
    print("\nSources consultées :", ", ".join(sources))
Run
python index.py
python ask.py "Quel est le délai de préavis prévu au contrat ?"

#Check what ChromaDB returns before blaming the model

When an answer is bad, the cause is in one of two places: retrieval failed to return the right passage, or the model used it incorrectly. You can distinguish the two by displaying the retrieved passages with their distance, without calling the model. If the right passage is missing from the top five, change the chunking, add keyword search, or use a reranker. If it is present and the answer is still wrong, the problem comes from the prompt, truncated context, or the model: try the next larger model before drawing a conclusion.

debug.py: display the passages and their distance
import sys
import chromadb
from chromadb.utils.embedding_functions.ollama_embedding_function import OllamaEmbeddingFunction

ef = OllamaEmbeddingFunction(url="http://localhost:11434", model_name="bge-m3")
coll = chromadb.PersistentClient(path="./chroma_db").get_collection("mes_docs", embedding_function=ef)
res = coll.query(query_texts=[" ".join(sys.argv[1:])], n_results=8)
for doc, meta, dist in zip(res["documents"][0], res["metadatas"][0], res["distances"][0]):
    print(f"{dist:.3f}  {meta['source']}  {doc[:120]!r}")

#Memory budget: what must fit at the same time

RAG runs two models side by side: the one that generates and the one that computes vectors, plus the first model's context cache. Ollama loads each model on demand and can unload one to make room for the other, adding a delay at every switch when memory is tight. The table provides a rough estimate for three configurations; the model weight comes from the Ollama library, while the rest is a calculation to refine with the site's VRAM calculator.

Memory required (weights Ollama, context of 8,192 tokens)
ConfigurationGeneration model weightTo addTarget card
ministral-3:8b + bge-m36.0 GBContext cache, embedding model (567 million parameters, just over one GB in half precision), system headroom8 to 12 GB
ministral-3:14b + bge-m39.1 GBSame here; long context becomes the limiting factor at 12 GB12 to 16 GB
mistral-small3.2:24b + bge-m315 GBSame; allow plenty of headroom24 GB or more
→
If memory is insufficient
First reduce num_ctx (8,192 is already generous for five passages), then switch to the smaller model. Also avoid increasing the number of passages: too many passages dilute the response as much as they fill memory.

#The pitfalls that make it answer the wrong question

The default context is too short
See above: without increasing num_ctx, the final passages are truncated. Typical symptom: ChromaDB retrieves the correct answer, but the model says it can’t find it.
Scanned PDFs
pypdf only reads text that is already present. A scan returns nothing: the script displays it. Run the document through OCR first, as described in the guide on Tesseract.
Passages without context
A passage ripped from its document (« the deadline is 30 days ») does not say what it refers to. Prefix each passage with the document or section title.
Question with no answer in the documents
Without the instruction “say it clearly,” a model fills the gap with what it knows. Always test a question whose answer is not in your files.
Exact identifiers and terms
A contract or case number is not retrieved reliably by embeddings: add keyword search, as described in the guide to hybrid search.
!
Verify before trusting
A RAG cites its sources, but that does not prove the answer is correct: open the cited file for decisions that matter (contracts, figures, deadlines).

#Go further

Improvements ranked by effort and impact
ImprovementEffortTo do when
Increase the number of passages (k) from 5 to 8One lineThe answer is spread across several passages
Chunking by headings instead of paragraphsMediumStructured documents (documentation, contracts with numbered articles)
Hybrid BM25 + vector searchMediumQuestions by identifier, acronym, or proper name
Reranker (bge-reranker-v2-m3)MediumThe correct answer is retrieved but ranked below the 5th position
Chat interface (Open WebUI, FastAPI API)VariableOther people need to use the tool
Planned backup and reindexingLowThe document folder changes every week

Each improvement has its own guide: measure recall on 30 to 50 real questions before and after, rather than piling on techniques. If you prefer a ready-made interface without writing code, the no-code RAG guide covers Open WebUI and AnythingLLM.

#Frequently asked questions about RAG with Mistral

FAQ
Which Mistral model for a local RAG?+
For an 8 to 12 GB card, ministral-3:8b (6.0 GB in Ollama, Apache 2.0 license) is a good starting point; ministral-3:14b (9.1 GB) for 16 GB; mistral-small3.2:24b (15 GB) for 24 GB and above. Choose based on the memory left after the model weights: you need room for the context.
Can Ollama calculate embeddings instead of sentence-transformers?+
Yes: Ollama exposes an embeddings API and offers bge-m3, a multilingual model with 567 million parameters. ChromaDB provides the OllamaEmbeddingFunction wrapper to call it. The advantage is having only one engine to install, without PyTorch; the drawback is that you need to keep Ollama active during indexing.
Why does the model say it can't find the answer when it's in my documents?+
Two common causes. Either Ollama’s context window is too short and the passages are truncated: raise num_ctx to 8 192. Or the retrieved passages don’t contain the answer: check what ChromaDB returns before generation, then adjust the chunking or search.
Can you use the Mistral API instead of Ollama?+
Technically yes, but your passages would then be sent to the company's servers, which conflicts with the confidentiality requirement for sensitive documents. To stay local, keep the open-weight models running through Ollama. For non-sensitive documents, the API is possible; read the provider's data-processing terms.
How do you add new documents without reindexing everything?+
Copy them into docs/ and rerun index.py: thanks to upsert and stable identifiers, existing passages are updated and new ones are added. If you change the passage size or embedding model, delete the chroma_db folder and reindex everything: the old vectors are no longer comparable.
Do you need a GPU for this RAG?+
Not necessarily: ministral-3:3b runs on a recent processor with 8 GB of memory, with slow responses. Indexing runs only once. A GPU mainly improves response time: greater speed and the ability to use a larger model. Measure response time on your machine before deciding to invest.
Did this guide help you?

Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.