Local RAG with ChromaDB and Ollama: tutorial Python
Building a local ChromaDB/Ollama/Python RAG involves three interlocking pieces: a vector store that persists to disk (ChromaDB), an embeddings model that turns your chunks into vectors (nomic-embed-text via Ollama), and a chat LLM that answers based on the retrieved passages. No API key, no data leakage. This guide takes you from a raw PDF to a chatbot that cites its sources in 22 minutes.
#Why this stack for a local RAG
Many RAG tutorials start with LangChain or LlamaIndex. These frameworks are powerful but hide what's happening under the hood. Here, we write the pipeline by hand with only three dependencies. You'll understand every step and know what to optimize later.
- ChromaDB
- Open-source vector store, pure Python, with built-in persistent mode (SQLite + HNSW index). No server to start.
- Ollama
- Serves both the embeddings model (nomic-embed-text) and the chat LLM (Qwen 3.5, Granite 4.2, Gemma 4). Single HTTP endpoint on localhost:11434.
- Native Python
- A few functions, no framework. You can add LangChain later if needed, but it is not necessary to get started.
#Prerequisites
This guide gets you to the model. The kit gets you to the coding copilot in your editor.
- Lifetime online access
- PDF + files
- Lifetime updates
- Python 3.10+
- ChromaDB requires 3.10 minimum. Check with python --version.
- Ollama installed and running
- The daemon listens on http://localhost:11434 by default. If you’re starting from scratch, follow the Ollama installation guide first.
- 8 GB of RAM
- 16 GB is comfortable. The 9B chat model in Q4 uses ~6 GB, and the embeddings model uses ~300 MB.
- A GPU isn't required
- CPU inference works, just more slowly. For ingesting a large corpus, a 6 GB+ GPU greatly speeds up embeddings.
#1. Install ChromaDB and prepare Ollama
We create a clean virtual environment, install the three required libraries, and download the models on the Ollama side.
Three packages: chromadb for the vector store, ollama for the official Python client, and pypdf for reading PDFs. That's it.
nomic-embed-text is a 137M-parameter multilingual embedding model that produces 768-dimensional vectors. Lightweight, fast, and good at French. Qwen 3.5 9B (6.6 GB, 256k context, multilingual, Apache 2.0) is used for the final chat: it is the default 8 GB choice in 2026. You can replace it with granite4.2:8b (more frugal) or gemma4:12b without changing the code.
#2. Configure the embeddings model
An embedding is a vector that represents the meaning of a piece of text. Two semantically similar texts have similar vectors. This is what powers RAG: we search for the chunks whose embedding most closely resembles the question’s.
You should see Vector dimension: 768. If it fails with model not found, ollama pull nomic-embed-text was not done.
#3. French PDF ingestion
Ingestion does three things: reads the pages of a PDF, splits the text into reasonably sized chunks, and stores each chunk with its embedding in ChromaDB in persistent mode.
The chunker splits content into 800-character blocks with 100 characters of overlap. This is a starting point: not too small (lack of context) and not too large (diluted signal). For very dense legal content, reduce it to 500. For well-spaced technical manuals, increase it to 1200.
Start ingestion on a ./pdfs/ folder containing your documents:
#4. Top-k search in ChromaDB
Once the chunks are indexed, retrieval consists of embedding the question and then asking Chroma for the k closest vectors by cosine distance. It's instantaneous, even with 100,000 chunks.
k=4 is a good default. Too small, and you miss relevant context; too large, and you drown the LLM in noise and blow up the context window. For very precise questions, k=2 is enough. For cross-cutting questions, increase it to 6.
#5. Chat loop with citations
Now let’s put it together: we find the relevant chunks, build a prompt with the context, send it to Qwen 3.5 via Ollama, and ask the model to cite its sources.
Three details matter. First, temperature=0.2: we want a factual response, not a creative one. Second, num_ctx=8192: the default window of Ollama (2048) is too short once we inject 4 chunks of 800 characters. Third, the system prompt forces the model to say “I don't know” rather than hallucinate—this is RAG's primary anti-hallucination safeguard.
#6. Concrete example: legal chatbot for contracts
Imagine a firm that wants to query 200 PDF service contracts. With the stack above, in under an hour we have an assistant that can answer questions such as:
- Typical question
- “Which contracts include a post-termination non-compete clause longer than 12 months?”
- What happens
- The question embedding retrieves the chunks containing semantically related keywords (non-compete, post-termination, duration). Qwen 3.5 reads these 4 passages and responds with the names of the relevant files.
- Confidentiality guarantee
- No data leaves the workstation. No API key. No telemetry. That's what distinguishes a local RAG system from an OpenAI wrapper.
#Troubleshooting
- ChromaDB is slow during ingestion
- The bottleneck is almost always the embeddings call to Ollama. Check that nomic-embed-text is running on the GPU with ollama ps. On CPU, expect ~50 chunks/second; on GPU, ~500.
- “model not found”
- Ollama can’t find nomic-embed-text. Restart ollama pull nomic-embed-text and check with ollama list.
- Responses that fabricate sources
- A 9B model still hallucinates sometimes. Move to mistral-small (24B, ~14 GB, very good in French) or qwen3.8:27b if you have the VRAM. Or add a reranker (cross-encoder) after ChromaDB to filter false positives.
- Poor-quality embeddings in French
- nomic-embed-text is multilingual but not optimal for French-only content. For legal or medical content, test Solon-embeddings-large-0.1 or bge-m3 (load via sentence-transformers, outside Ollama).
- ChromaDB grows without limit
- Each reindexing adds duplicates. Before reingesting a PDF, run collection.delete(where={"source": name}) to purge the old chunks.
#Go further
You have a working RAG system. Here are the natural next steps to take it further:
- Compare French embedding models
- Our guide “The Best French Embedding Models” compares BGE, E5, Solon, and nomic on French-language content.
- Improve chunking
- “Chunking strategies” covers semantic chunking, chunking by Markdown headings, and chunking by paragraphs—often the approaches that unlock the biggest gains in precision.
- Add a reranker
- “Add a reranker to your pipeline”: +15% relevance by placing a cross-encoder after Chroma. The logical next step.
- Hybrid search
- “Hybrid BM25 + vector search” combines lexical and semantic search, which is essential when there's a lot of jargon or many proper names.
Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.