Intermediate 18 minRAG

PrivateGPT v2: 100% document chatbot private

PrivateGPT v2 is a 100% private document chatbot for discussing your PDFs, Word files, and Markdown files without a single byte leaving the machine. v2 dropped its built-in LLM engine in favor of Ollama: you retain the quality of a well-built RAG system, gain access to all models available through Ollama, and radically simplify the stack. This guide covers installation, connecting to Ollama, and three concrete business cases (HR, legal, accounting).

By Mohamed Meguedmi·Update 2026-08-27·Tested on Windows, macOS, and Linux

#Why PrivateGPT v2 for a local-document chatbot

The need is straightforward: ask natural-language questions about a corpus of internal documents (contracts, pay stubs, invoices, meeting notes) without sending that data to OpenAI, Anthropic, or Google. That's exactly what PrivateGPT has targeted since its first release in 2023.

Version 2 of the project made a clean break: no more embedded LLM engine, no more internal quantization management. Instead, PrivateGPT v2 delegates generation to an external backend — Ollama in most cases. This makes the project much easier to maintain and provides access to the entire Ollama model library (Qwen, Gemma, Granite, Mistral, etc.) without any special configuration.

100% local
Everything runs on your machine. No telemetry and no outbound calls once the models are downloaded.
Gradio UI included
A web interface opens on localhost, ready to chat with your documents—no need to hack together a frontend.
OpenAI-compatible API
The PrivateGPT server also exposes REST endpoints close to the OpenAI format, which are useful if you want to connect it to a custom app.
Supported formats
PDF, DOCX, PPTX, MD, TXT, HTML, EPUB, CSV, and several others through the underlying LlamaIndex loaders.
i
Two-block architecture
PrivateGPT v2 = RAG pipeline (chunking, embeddings, vector store, retrieval, prompt assembly) + Gradio UI. Ollama = LLM inference engine. The two communicate over local HTTP on port 11434.

#Prerequisites

The Local RAG Kit

Your documents, your AI: a reliable local RAG over your PDFs, notes and mail — nothing leaves your machine.

  • Lifetime online access
  • PDF + files
  • Lifetime updates
Python 3.11
PrivateGPT v2 officially supports only Python 3.11. With 3.12+, some dependencies still break.
Poetry
The project uses Poetry to manage its dependencies with extras (ui, llms-ollama, embeddings-ollama, vector-stores-qdrant).
Ollama installed and active
Ollama daemon reachable at http://localhost:11434. If you don’t have it, install it first—the installer takes 3 minutes.
8 GB of RAM minimum
16 GB is comfortable. If you load an 8B–9B Q4 model into VRAM through Ollama, expect an additional 5 to 7 GB on the GPU side.
Recommended GPU (optional)
A RTX 3060 12 GB or larger lets you use 8B–14B models with a reasonable response time. Without a GPU, stick to 3B (Granite 4.2 3B or Qwen 3.5 2B).
→
No Ollama installed?
First follow our Ollama installation guide for your OS, then come back here. PrivateGPT v2 assumes that the "ollama run" command already works.

#1. Install PrivateGPT v2

The project lives on GitHub at zylon-ai/private-gpt. Clone it, install Poetry if it is not already installed, then install the dependencies with the extras matching your chosen backend.

Clone the repo
git clone https://github.com/zylon-ai/private-gpt.git
cd private-gpt
Install Poetry (if absent)
curl -sSL https://install.python-poetry.org | python3 -
# Ajoutez ~/.local/bin au PATH si ce n'est pas déjà fait
Install PrivateGPT with the Ollama extras
poetry install --extras "ui llms-ollama embeddings-ollama vector-stores-qdrant"

This command installs four groups of extras: the Gradio UI, the Ollama LLM client, the Ollama embeddings client, and Qdrant as an embedded local vector database. Installation takes 5 to 15 minutes depending on your connection — there are many scientific dependencies (numpy, scipy, transformers).

!
Python 3.12 and 3.13: proceed with caution
As of this writing, builds for some native dependencies (notably llama-index and the tokenizers components) have not all been published for 3.12+. If Poetry returns compilation errors, create a dedicated Python 3.11 venv: "pyenv install 3.11.9 && pyenv local 3.11.9".

#2. Configure Ollama as the LLM and embeddings backend

PrivateGPT v2 uses a YAML profile system in settings/. Enable the ollama profile through the PGPT_PROFILES environment variable.

First, pull the two models we'll need: a generation model and an embedding model. For French, Qwen 3.5 9B is the right 2026 default on the LLM side (6.6 GB, 256k context, Apache 2.0 license), and nomic-embed-text remains an excellent lightweight multilingual embedding choice.

Prepare the models in Ollama
# LLM de génération
ollama pull qwen3.5:9b

# Modèle d'embeddings (137M, ~280 Mo)
ollama pull nomic-embed-text

We then check the settings/settings-ollama.yaml file provided in the repo. It must point to the correct model names and the correct Ollama URL:

settings/settings-ollama.yaml
llm:
  mode: ollama
  max_new_tokens: 512
  context_window: 8192

embedding:
  mode: ollama

ollama:
  llm_model: qwen3.5:9b
  embedding_model: nomic-embed-text
  api_base: http://localhost:11434
  embedding_api_base: http://localhost:11434
  request_timeout: 120.0

vectorstore:
  database: qdrant

qdrant:
  path: local_data/private_gpt/qdrant
→
Context window
context_window: 8192 is a reasonable compromise for getting started with Qwen 3.5 9B (which goes up to 256k). On a GPU with more VRAM, move to 16384 or 32768 to handle longer documents without aggressive chunking. Warning: the VRAM used by the KV cache climbs quickly.

#3. Start the server and UI

  1. 01
    Verify that Ollama is running
    Type "ollama list" in a terminal. If the command returns the list of models, the daemon is active. Otherwise, run "ollama serve" in a separate terminal.
  2. 02
    Enable the Ollama profile
    In the terminal where you will launch PrivateGPT, export the PGPT_PROFILES=ollama variable. This tells the project to load settings-ollama.yaml on top of settings.yaml.
  3. 03
    Start the server
    From the repo root, run "PGPT_PROFILES=ollama make run". The server starts on port 8001, and Gradio opens the UI at the same address.
  4. 04
    Open the UI
    Go to http://localhost:8001 in your browser. You will see a chat interface with a sidebar for uploading documents.
Launch command
PGPT_PROFILES=ollama make run

On first launch, you’ll see in the logs that PrivateGPT contacts Ollama to verify that the declared models are available. If a model is missing, the server stops with an explicit message—pull the missing model with ollama pull, then restart.

#4. Index your documents

The Gradio UI provides an “Ingest” tab where you can drag and drop your files. Behind the scenes, PrivateGPT splits each document into chunks (by default ~1024 tokens with an overlap of 200), computes embeddings via Ollama, and stores everything in Qdrant.

PDF
Text extracted with pypdf. Scanned PDFs (image-only) are not OCR'd—run them through a tool such as ocrmypdf before ingestion.
DOCX / PPTX
Supported natively by LlamaIndex loaders. Tables and lists are preserved as plain text.
Markdown / TXT / HTML
Immediate indexing is the format that works best for RAG in practice.
CSV
Each line becomes a chunk. Useful for FAQ databases or business extracts.
!
Indexing time to plan for
Ingestion calls Ollama to calculate embeddings, one document at a time. On a pure CPU, expect about 5 seconds per PDF page. On a GPU, it’s nearly instantaneous. To ingest 500 PDFs, allow plenty of time and run them in batches through the API rather than using drag and drop.
Batch ingestion via the CLI script
# Depuis la racine du repo private-gpt
python scripts/ingest_folder.py /chemin/vers/mes/documents \
  --watch  # surveille en continu les nouveaux fichiers

#Concrete business use cases

#HR: query payslips and collective bargaining agreements

Typical case: an HR service with 200 monthly payslips archived as PDFs, plus the industry's collective agreement (often 100+ pages). PrivateGPT v2 lets an HR employee ask questions like "What is Ms. Dupont's coefficient in 2025?" or "What does the collective agreement say about leave for caring for a sick child?"

Recommended model
Qwen 3.5 9B Q4 is enough for factual extraction. For legal interpretation of the convention, move up to Mistral Small 24B (good in French, 14 GB) or Qwen 3.8 27B (18 GB, 262k context) if your VRAM can handle it.
Cutting
For payroll slips (1 page), one chunk = one document. For the collective agreement, section-based chunking (H2/H3 headings) works better than token-based chunking.
Privacy
This is precisely where PrivateGPT v2 makes a difference: pay slips should never pass through a cloud service, including in “enterprise” mode.

#Legal: analyzing a contract portfolio

Typical case: a legal department with a few hundred customer and supplier contracts in PDF, sometimes scanned. Common questions: "Which contracts expire in the next 6 months?", "Which termination clause applies to the ACME contract?", "Which ones contain an exclusivity clause?"

Preprocessing
Run ocrmypdf on scans before ingestion. Without OCR, these PDFs are invisible to retrieval.
Recommended model
Mistral Small 24B or Qwen 3.8 27B for French legal-reasoning quality. Qwen 3.5 9B is enough for pure research.
System prompt
Define a system prompt that requires citing the contract name and clause number in every response — otherwise the model tends to summarize without citing sources.
i
Always verify sources
PrivateGPT v2's Gradio UI displays the retrieved chunks below the answer. For serious legal use, treat answers as leads and always verify the source passage — RAG is assisted research, not an automatic legal opinion.

#Accountant: query a folder of invoices and statements

Typical case: an accounting firm wants to query a corpus of a client's supplier invoices and bank statements (PDF + CSV). Target questions: "What is the total of Free Mobile invoices in 2025?", "Is there an unpaid invoice with supplier X?".

A limitation to know
RAG does not aggregate naturally: it retrieves relevant passages but does not sum them perfectly across large volumes. For strict analytics, export the invoice CSV to a dedicated tool and use PrivateGPT for qualitative context.
Preferred format
Accounting CSVs are indexed line by line—good for individual lookups, bad for calculations. Attach a summary PDF for each month if you want the LLM to have an overall view.
Model
Qwen 3.5 9B (or even Q8, 11 GB) is very strong on numbers and table reading compared with a small 3B model. If you have a RTX 4070 with 12 GB or more, it is the default choice here.

#Common troubleshooting

"Connection refused" on startup
The Ollama daemon is not running. Check with "curl http://localhost:11434" — you should see "Ollama is running".
Very slow responses
Either Ollama is running on the CPU (check with "ollama ps" — PROCESSOR column), or your context_window is too high. Reduce it to 4096 for testing.
"Model not found"
The model name in settings-ollama.yaml must exactly match a model listed by "ollama list". Watch out for quantization tags: qwen3.5:9b ≠ qwen3.5:9b-q8_0.
Different embeddings between ingestion and querying
If you change the embeddings model after ingestion, you must re-index. Delete local_data/private_gpt/qdrant/ and restart ingestion.
Inaccessible Gradio UI
If you’re on a remote machine, launch with "PGPT_PROFILES=ollama python -m private_gpt" and reconfigure the host in settings.yaml (server.host: 0.0.0.0).

#Go further

PrivateGPT v2 is an excellent entry point for local document RAG, but it is only one building block. A few ideas for going further:

No-code RAG with Open WebUI or AnythingLLM
If PrivateGPT seems cumbersome to install to you (Poetry, Python 3.11), Open WebUI offers a comparable RAG experience that is easier to deploy in Docker.
High-performing French embeddings
nomic-embed-text gets the job done, but specialized French models such as Solon or BGE-M3 deliver better results on French legal or administrative content.
Advanced chunking strategies
Section-based chunking (Markdown headers, DOCX structure) greatly outperforms fixed-token chunking on structured corpora—a substantial relevance boost that is easy to achieve.
Did this guide help you?

Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.