Intermediate 22 minRAG

Local RAG with ChromaDB and Ollama: tutorial Python

Building a local ChromaDB/Ollama/Python RAG involves three interlocking pieces: a vector store that persists to disk (ChromaDB), an embeddings model that turns your chunks into vectors (nomic-embed-text via Ollama), and a chat LLM that answers based on the retrieved passages. No API key, no data leakage. This guide takes you from a raw PDF to a chatbot that cites its sources in 22 minutes.

By Mohamed Meguedmi·Update 2026-08-27·Tested on Windows, macOS, and Linux

#Why this stack for a local RAG

Many RAG tutorials start with LangChain or LlamaIndex. These frameworks are powerful but hide what's happening under the hood. Here, we write the pipeline by hand with only three dependencies. You'll understand every step and know what to optimize later.

ChromaDB
Open-source vector store, pure Python, with built-in persistent mode (SQLite + HNSW index). No server to start.
Ollama
Serves both the embeddings model (nomic-embed-text) and the chat LLM (Qwen 3.5, Granite 4.2, Gemma 4). Single HTTP endpoint on localhost:11434.
Native Python
A few functions, no framework. You can add LangChain later if needed, but it is not necessary to get started.
i
What you get
A Python script of about 150 lines that ingests a folder of PDFs, chunks them, indexes them in ChromaDB, and answers questions in French with citations. Fully local, with zero outbound requests.

#Prerequisites

The Local Copilot Kit

This guide gets you to the model. The kit gets you to the coding copilot in your editor.

  • Lifetime online access
  • PDF + files
  • Lifetime updates
Python 3.10+
ChromaDB requires 3.10 minimum. Check with python --version.
Ollama installed and running
The daemon listens on http://localhost:11434 by default. If you’re starting from scratch, follow the Ollama installation guide first.
8 GB of RAM
16 GB is comfortable. The 9B chat model in Q4 uses ~6 GB, and the embeddings model uses ~300 MB.
A GPU isn't required
CPU inference works, just more slowly. For ingesting a large corpus, a 6 GB+ GPU greatly speeds up embeddings.

#1. Install ChromaDB and prepare Ollama

We create a clean virtual environment, install the three required libraries, and download the models on the Ollama side.

Python environment
python -m venv .venv
source .venv/bin/activate  # sous Windows : .venv\Scripts\activate
pip install chromadb ollama pypdf

Three packages: chromadb for the vector store, ollama for the official Python client, and pypdf for reading PDFs. That's it.

Ollama models
ollama pull nomic-embed-text
ollama pull qwen3.5:9b

nomic-embed-text is a 137M-parameter multilingual embedding model that produces 768-dimensional vectors. Lightweight, fast, and good at French. Qwen 3.5 9B (6.6 GB, 256k context, multilingual, Apache 2.0) is used for the final chat: it is the default 8 GB choice in 2026. You can replace it with granite4.2:8b (more frugal) or gemma4:12b without changing the code.

→
Verify that Ollama responds
A simple curl http://localhost:11434/api/tags should list your models. If nothing appears, the daemon is not running: run ollama serve in another terminal.

#2. Configure the embeddings model

An embedding is a vector that represents the meaning of a piece of text. Two semantically similar texts have similar vectors. This is what powers RAG: we search for the chunks whose embedding most closely resembles the question’s.

embed.py — quick test
import ollama

resp = ollama.embeddings(
    model="nomic-embed-text",
    prompt="Le contrat est résilié de plein droit en cas de manquement grave."
)

vec = resp["embedding"]
print(f"Dimension du vecteur : {len(vec)}")
print(f"5 premières valeurs : {vec[:5]}")

You should see Vector dimension: 768. If it fails with model not found, ollama pull nomic-embed-text was not done.

i
Why nomic-embed-text
On French benchmarks (MTEB-fr), nomic-embed-text ranks among the top 5 models with fewer than 200M parameters. For pure French, mxbai-embed-large often performs better but weighs 670M. nomic is an excellent quality/speed compromise to get started.

#3. French PDF ingestion

Ingestion does three things: reads the pages of a PDF, splits the text into reasonably sized chunks, and stores each chunk with its embedding in ChromaDB in persistent mode.

ingest.py
import os
import chromadb
import ollama
from pypdf import PdfReader

client = chromadb.PersistentClient(path="./chroma_db")
collection = client.get_or_create_collection(name="docs")

def chunk_text(text, size=800, overlap=100):
    chunks = []
    start = 0
    while start < len(text):
        end = min(start + size, len(text))
        chunks.append(text[start:end])
        start += size - overlap
    return chunks

def ingest_pdf(path):
    reader = PdfReader(path)
    name = os.path.basename(path)
    for page_num, page in enumerate(reader.pages):
        text = page.extract_text() or ""
        for i, chunk in enumerate(chunk_text(text)):
            emb = ollama.embeddings(
                model="nomic-embed-text",
                prompt=chunk
            )["embedding"]
            collection.add(
                ids=[f"{name}-p{page_num}-c{i}"],
                embeddings=[emb],
                documents=[chunk],
                metadatas=[{"source": name, "page": page_num + 1}],
            )
    print(f"OK : {name} ingéré ({len(reader.pages)} pages)")

if __name__ == "__main__":
    for f in os.listdir("./pdfs"):
        if f.endswith(".pdf"):
            ingest_pdf(f"./pdfs/{f}")

The chunker splits content into 800-character blocks with 100 characters of overlap. This is a starting point: not too small (lack of context) and not too large (diluted signal). For very dense legal content, reduce it to 500. For well-spaced technical manuals, increase it to 1200.

→
ChromaDB's persistent mode
PersistentClient(path="./chroma_db") creates a directory that survives restarts. SQLite stores the metadata, and an HNSW index stores the vectors. No server to start, no Docker. To switch to client/server mode later, simply replace it with HttpClient.

Start ingestion on a ./pdfs/ folder containing your documents:

Start ingestion
mkdir -p pdfs
# placez vos PDF dans ./pdfs/
python ingest.py
!
Scanned PDFs = no text
pypdf extracts only native text. If your PDFs are image scans, extract_text() will return nothing. You then need to use OCR (Tesseract, or a vision model such as Qwen 3.5 9B, multimodal, via Ollama) before ingestion.

#4. Top-k search in ChromaDB

Once the chunks are indexed, retrieval consists of embedding the question and then asking Chroma for the k closest vectors by cosine distance. It's instantaneous, even with 100,000 chunks.

search.py
import chromadb
import ollama

client = chromadb.PersistentClient(path="./chroma_db")
collection = client.get_collection(name="docs")

def search(question, k=4):
    q_emb = ollama.embeddings(
        model="nomic-embed-text",
        prompt=question
    )["embedding"]
    results = collection.query(
        query_embeddings=[q_emb],
        n_results=k,
    )
    chunks = results["documents"][0]
    metas = results["metadatas"][0]
    return list(zip(chunks, metas))

if __name__ == "__main__":
    hits = search("Quelles sont les conditions de résiliation ?")
    for chunk, meta in hits:
        print(f"[{meta['source']} p.{meta['page']}]")
        print(chunk[:200], "...\n")

k=4 is a good default. Too small, and you miss relevant context; too large, and you drown the LLM in noise and blow up the context window. For very precise questions, k=2 is enough. For cross-cutting questions, increase it to 6.

#5. Chat loop with citations

Now let’s put it together: we find the relevant chunks, build a prompt with the context, send it to Qwen 3.5 via Ollama, and ask the model to cite its sources.

chat.py
import ollama
from search import search

SYSTEM = """Tu es un assistant qui répond uniquement à partir du CONTEXTE fourni.
Si la réponse n'est pas dans le contexte, dis-le clairement.
Cite tes sources entre crochets sous la forme [source.pdf p.X]."""

def ask(question):
    hits = search(question, k=4)
    context = "\n\n".join(
        f"[{m['source']} p.{m['page']}]\n{c}" for c, m in hits
    )
    prompt = f"CONTEXTE :\n{context}\n\nQUESTION : {question}"
    resp = ollama.chat(
        model="qwen3.5:9b",
        messages=[
            {"role": "system", "content": SYSTEM},
            {"role": "user", "content": prompt},
        ],
        options={"temperature": 0.2, "num_ctx": 8192},
    )
    return resp["message"]["content"]

if __name__ == "__main__":
    while True:
        q = input("\nQuestion (vide pour quitter) > ").strip()
        if not q:
            break
        print("\n" + ask(q))

Three details matter. First, temperature=0.2: we want a factual response, not a creative one. Second, num_ctx=8192: the default window of Ollama (2048) is too short once we inject 4 chunks of 800 characters. Third, the system prompt forces the model to say “I don't know” rather than hallucinate—this is RAG's primary anti-hallucination safeguard.

→
Streaming for a better UX
Replace ollama.chat with ollama.chat(..., stream=True) and iterate over the response to display tokens as they arrive. This is crucial as soon as you integrate this code into a real interface (FastAPI + WebSocket, or Streamlit).

#6. Concrete example: legal chatbot for contracts

Imagine a firm that wants to query 200 PDF service contracts. With the stack above, in under an hour we have an assistant that can answer questions such as:

Typical question
“Which contracts include a post-termination non-compete clause longer than 12 months?”
What happens
The question embedding retrieves the chunks containing semantically related keywords (non-compete, post-termination, duration). Qwen 3.5 reads these 4 passages and responds with the names of the relevant files.
Confidentiality guarantee
No data leaves the workstation. No API key. No telemetry. That's what distinguishes a local RAG system from an OpenAI wrapper.
!
Limitations to know
A basic RAG handles targeted questions well (“what is clause X?”) and aggregate questions poorly (“how many contracts contain X?”). For the latter, you need either an agent that queries the database in multiple steps or a GraphRAG. That is another story.

#Troubleshooting

ChromaDB is slow during ingestion
The bottleneck is almost always the embeddings call to Ollama. Check that nomic-embed-text is running on the GPU with ollama ps. On CPU, expect ~50 chunks/second; on GPU, ~500.
“model not found”
Ollama can’t find nomic-embed-text. Restart ollama pull nomic-embed-text and check with ollama list.
Responses that fabricate sources
A 9B model still hallucinates sometimes. Move to mistral-small (24B, ~14 GB, very good in French) or qwen3.8:27b if you have the VRAM. Or add a reranker (cross-encoder) after ChromaDB to filter false positives.
Poor-quality embeddings in French
nomic-embed-text is multilingual but not optimal for French-only content. For legal or medical content, test Solon-embeddings-large-0.1 or bge-m3 (load via sentence-transformers, outside Ollama).
ChromaDB grows without limit
Each reindexing adds duplicates. Before reingesting a PDF, run collection.delete(where={"source": name}) to purge the old chunks.

#Go further

You have a working RAG system. Here are the natural next steps to take it further:

Compare French embedding models
Our guide “The Best French Embedding Models” compares BGE, E5, Solon, and nomic on French-language content.
Improve chunking
“Chunking strategies” covers semantic chunking, chunking by Markdown headings, and chunking by paragraphs—often the approaches that unlock the biggest gains in precision.
Add a reranker
“Add a reranker to your pipeline”: +15% relevance by placing a cross-encoder after Chroma. The logical next step.
Hybrid search
“Hybrid BM25 + vector search” combines lexical and semantic search, which is essential when there's a lot of jargon or many proper names.
Did this guide help you?

Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.