RAG with ChromaDB and Mistral
For a local RAG setup with Mistral, the simplest stack is Ollama (generation and embeddings with bge-m3) plus ChromaDB in file mode, with no server or PyTorch. For the model, ministral-3:8b (6.0 GB, Apache 2.0 license, advertised 256K context) is a good default for an 8 to 12 GB card, ministral-3:14b for 16 GB, and mistral-small3.2:24b (15 GB) beyond that. The setting you must not forget: Ollama's context window, which you need to increase so the passages fit in the prompt.
This guide builds a complete document assistant in two Python scripts, with a Mistral model running on your machine: your PDFs and text files are split up, indexed in ChromaDB, and then the retrieved passages are provided to the model, which answers with citations. It also explains which Mistral model to choose based on your graphics memory and the pitfalls that cause an RAG system to answer beside the point.
#What we're building: a fully local Mistral RAG
RAG (retrieval-augmented generation) consists of finding the passages in your documents relevant to the question, then inserting them into the model’s prompt so it can answer from them. The intended result here is a small command-line tool: one script indexes a document folder; a second reads a question, finds the five closest passages in ChromaDB, sends them to a Mistral model via Ollama along with the question, and displays the answer followed by the files consulted. Nothing leaves the machine: Ollama serves the generation model and the embedding model, while ChromaDB stores the vectors in a local folder.
Two meanings of “Mistral” are in circulation: Mistral AI's open-weight models, which you download and run yourself (the subject of this guide), and the company's hosted APIs, which send your passages to its servers. For confidential documents, only the former meets the “100% local” requirement.
#Which Mistral model to choose for RAG
Your documents, your AI: a reliable local RAG over your PDFs, notes and mail — nothing leaves your machine.
- Lifetime online access
- PDF + files
- Lifetime updates
RAG has specific requirements: the model must follow a strict instruction (“answer only from the passages”), read multiple passages without getting lost, and respond in French. Size matters less than for open-ended conversation; available context memory matters more. The sizes below are those displayed by the Ollama library, using the default quantization.
| Model | Size Ollama | Advertised context | Who it's for |
|---|---|---|---|
| ministral-3:3b | 3.0 GB | 256K | Machine without a dedicated GPU; simple responses, low tolerance for complex instructions |
| ministral-3:8b | 6.0 GB | 256K | Reasonable default for an 8 to 12 GB card or a laptop with 16 GB of memory |
| ministral-3:14b | 9.1 GB | 256K | 16 GB card, or 12 GB with a moderate context |
| mistral-nemo (12B) | view the Ollama page | 128K | Older alternative, still widely used |
| mistral-small3.2:24b | 15 GB | 128K | 24 GB card or 32 GB or more of unified memory; strongest on format instructions |
| mistral (7B, version 0.3) | 4.4 GB | 32K | Older model: reserve it for very limited machines |
The Ministral 3 family (3B, 8B, and 14B) is released under the Apache 2.0 license, as stated in the Mistral 3 announcement, and the Ollama page describes it as designed for edge deployment and capable of running on a wide range of hardware. Mistral Small 4, published in 2026 with 119 billion parameters in total according to the name of its Hugging Face listing, targets server hardware: it is not a candidate for a personal machine. For an estimate of the required memory, the site’s VRAM calculator gives the model’s size plus the context cache.
#The technical stack
- Generation
- A Mistral model served by Ollama through the local HTTP API on port 11434.
- Embeddings
- bge-m3 served by Ollama: the library page describes it as a versatile, multilingual, multi-granularity model from BAAI with 567 million parameters. It avoids installing PyTorch and sentence-transformers.
- Vector database
- ChromaDB in local mode (PersistentClient): one folder, no server. Chroma provides a wrapper, OllamaEmbeddingFunction, that calls the embeddings API of Ollama.
- Reading files
- pypdf for PDFs containing text, direct reading for Markdown and plain text. A scanned PDF is an image: it first requires optical character recognition.
#Prepare the environment
- 01Install Ollama and fetch the modelsInstall Ollama, then download the generation model and the embedding model using the two commands below.
- 02Create the Python environmentPython 3.10 or later. A virtual environment keeps the project's dependencies separate.
- 03Place the documentsCopy your PDFs, Markdown files, and text into a docs/ folder alongside the scripts.
#2. Index the documents in ChromaDB
The script reads each file, splits the text into passages of about 1,800 characters at paragraph breaks, then passes them to Chroma, which calls bge-m3 through Ollama to calculate the vectors. Two details matter: each passage keeps the filename as metadata (to cite the source), and additions are made in batches rather than one passage at a time.
Using upsert with identifiers built from the filename and pass number makes the script rerunnable: reindexing the same folder updates the passages instead of duplicating them. However, note that if a document gets shorter, the old surplus passages remain in the database; for a major change, delete the chroma_db folder and reindex. The choice of passage size is detailed in the guide to chunking strategies.
#3. Query: search, then generation
The second script embeds the question, retrieves the five closest passages, and builds the prompt. The instruction is crucial: it asks the model to answer only from the passages, admit when information is missing, and cite the file. The num_ctx parameter increases the context window: Ollama’s documentation states that the default window is 4,096 tokens and that the OLLAMA_CONTEXT_LENGTH variable or the num_ctx parameter changes it. With five passages of 400 to 500 tokens, the instruction, and the answer, 4,096 tokens is borderline: an overly short context is truncated silently, and the model answers without reading the end of your passages.
#Check what ChromaDB returns before blaming the model
When an answer is bad, the cause is in one of two places: retrieval failed to return the right passage, or the model used it incorrectly. You can distinguish the two by displaying the retrieved passages with their distance, without calling the model. If the right passage is missing from the top five, change the chunking, add keyword search, or use a reranker. If it is present and the answer is still wrong, the problem comes from the prompt, truncated context, or the model: try the next larger model before drawing a conclusion.
#Memory budget: what must fit at the same time
RAG runs two models side by side: the one that generates and the one that computes vectors, plus the first model's context cache. Ollama loads each model on demand and can unload one to make room for the other, adding a delay at every switch when memory is tight. The table provides a rough estimate for three configurations; the model weight comes from the Ollama library, while the rest is a calculation to refine with the site's VRAM calculator.
| Configuration | Generation model weight | To add | Target card |
|---|---|---|---|
| ministral-3:8b + bge-m3 | 6.0 GB | Context cache, embedding model (567 million parameters, just over one GB in half precision), system headroom | 8 to 12 GB |
| ministral-3:14b + bge-m3 | 9.1 GB | Same here; long context becomes the limiting factor at 12 GB | 12 to 16 GB |
| mistral-small3.2:24b + bge-m3 | 15 GB | Same; allow plenty of headroom | 24 GB or more |
#The pitfalls that make it answer the wrong question
- The default context is too short
- See above: without increasing num_ctx, the final passages are truncated. Typical symptom: ChromaDB retrieves the correct answer, but the model says it can’t find it.
- Scanned PDFs
- pypdf only reads text that is already present. A scan returns nothing: the script displays it. Run the document through OCR first, as described in the guide on Tesseract.
- Passages without context
- A passage ripped from its document (« the deadline is 30 days ») does not say what it refers to. Prefix each passage with the document or section title.
- Question with no answer in the documents
- Without the instruction “say it clearly,” a model fills the gap with what it knows. Always test a question whose answer is not in your files.
- Exact identifiers and terms
- A contract or case number is not retrieved reliably by embeddings: add keyword search, as described in the guide to hybrid search.
#Go further
| Improvement | Effort | To do when |
|---|---|---|
| Increase the number of passages (k) from 5 to 8 | One line | The answer is spread across several passages |
| Chunking by headings instead of paragraphs | Medium | Structured documents (documentation, contracts with numbered articles) |
| Hybrid BM25 + vector search | Medium | Questions by identifier, acronym, or proper name |
| Reranker (bge-reranker-v2-m3) | Medium | The correct answer is retrieved but ranked below the 5th position |
| Chat interface (Open WebUI, FastAPI API) | Variable | Other people need to use the tool |
| Planned backup and reindexing | Low | The document folder changes every week |
Each improvement has its own guide: measure recall on 30 to 50 real questions before and after, rather than piling on techniques. If you prefer a ready-made interface without writing code, the no-code RAG guide covers Open WebUI and AnythingLLM.
#Frequently asked questions about RAG with Mistral
Which Mistral model for a local RAG?+
Can Ollama calculate embeddings instead of sentence-transformers?+
Why does the model say it can't find the answer when it's in my documents?+
Can you use the Mistral API instead of Ollama?+
How do you add new documents without reindexing everything?+
Do you need a GPU for this RAG?+
- Local RAG with ChromaDB and Ollama: Python tutorial
- Chunking strategies
- Hybrid BM25 + vector search
- Add a reranker to your pipeline
- VRAM Calculator
- Local RAG with Ollama without coding
- Source: Ollama, ministral-3
- Source: Ollama, mistral-small3.2
- Source: Ollama, bge-m3
- Source: Chroma, embeddings Ollama
- Source: Ollama FAQ, context window
Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.