PrivateGPT v2: 100% document chatbot private
PrivateGPT v2 is a 100% private document chatbot for discussing your PDFs, Word files, and Markdown files without a single byte leaving the machine. v2 dropped its built-in LLM engine in favor of Ollama: you retain the quality of a well-built RAG system, gain access to all models available through Ollama, and radically simplify the stack. This guide covers installation, connecting to Ollama, and three concrete business cases (HR, legal, accounting).
#Why PrivateGPT v2 for a local-document chatbot
The need is straightforward: ask natural-language questions about a corpus of internal documents (contracts, pay stubs, invoices, meeting notes) without sending that data to OpenAI, Anthropic, or Google. That's exactly what PrivateGPT has targeted since its first release in 2023.
Version 2 of the project made a clean break: no more embedded LLM engine, no more internal quantization management. Instead, PrivateGPT v2 delegates generation to an external backend — Ollama in most cases. This makes the project much easier to maintain and provides access to the entire Ollama model library (Qwen, Gemma, Granite, Mistral, etc.) without any special configuration.
- 100% local
- Everything runs on your machine. No telemetry and no outbound calls once the models are downloaded.
- Gradio UI included
- A web interface opens on localhost, ready to chat with your documents—no need to hack together a frontend.
- OpenAI-compatible API
- The PrivateGPT server also exposes REST endpoints close to the OpenAI format, which are useful if you want to connect it to a custom app.
- Supported formats
- PDF, DOCX, PPTX, MD, TXT, HTML, EPUB, CSV, and several others through the underlying LlamaIndex loaders.
#Prerequisites
Your documents, your AI: a reliable local RAG over your PDFs, notes and mail — nothing leaves your machine.
- Lifetime online access
- PDF + files
- Lifetime updates
- Python 3.11
- PrivateGPT v2 officially supports only Python 3.11. With 3.12+, some dependencies still break.
- Poetry
- The project uses Poetry to manage its dependencies with extras (ui, llms-ollama, embeddings-ollama, vector-stores-qdrant).
- Ollama installed and active
- Ollama daemon reachable at http://localhost:11434. If you don’t have it, install it first—the installer takes 3 minutes.
- 8 GB of RAM minimum
- 16 GB is comfortable. If you load an 8B–9B Q4 model into VRAM through Ollama, expect an additional 5 to 7 GB on the GPU side.
- Recommended GPU (optional)
- A RTX 3060 12 GB or larger lets you use 8B–14B models with a reasonable response time. Without a GPU, stick to 3B (Granite 4.2 3B or Qwen 3.5 2B).
#1. Install PrivateGPT v2
The project lives on GitHub at zylon-ai/private-gpt. Clone it, install Poetry if it is not already installed, then install the dependencies with the extras matching your chosen backend.
This command installs four groups of extras: the Gradio UI, the Ollama LLM client, the Ollama embeddings client, and Qdrant as an embedded local vector database. Installation takes 5 to 15 minutes depending on your connection — there are many scientific dependencies (numpy, scipy, transformers).
#2. Configure Ollama as the LLM and embeddings backend
PrivateGPT v2 uses a YAML profile system in settings/. Enable the ollama profile through the PGPT_PROFILES environment variable.
First, pull the two models we'll need: a generation model and an embedding model. For French, Qwen 3.5 9B is the right 2026 default on the LLM side (6.6 GB, 256k context, Apache 2.0 license), and nomic-embed-text remains an excellent lightweight multilingual embedding choice.
We then check the settings/settings-ollama.yaml file provided in the repo. It must point to the correct model names and the correct Ollama URL:
#3. Start the server and UI
- 01Verify that Ollama is runningType "ollama list" in a terminal. If the command returns the list of models, the daemon is active. Otherwise, run "ollama serve" in a separate terminal.
- 02Enable the Ollama profileIn the terminal where you will launch PrivateGPT, export the PGPT_PROFILES=ollama variable. This tells the project to load settings-ollama.yaml on top of settings.yaml.
- 03Start the serverFrom the repo root, run "PGPT_PROFILES=ollama make run". The server starts on port 8001, and Gradio opens the UI at the same address.
- 04Open the UIGo to http://localhost:8001 in your browser. You will see a chat interface with a sidebar for uploading documents.
On first launch, you’ll see in the logs that PrivateGPT contacts Ollama to verify that the declared models are available. If a model is missing, the server stops with an explicit message—pull the missing model with ollama pull, then restart.
#4. Index your documents
The Gradio UI provides an “Ingest” tab where you can drag and drop your files. Behind the scenes, PrivateGPT splits each document into chunks (by default ~1024 tokens with an overlap of 200), computes embeddings via Ollama, and stores everything in Qdrant.
- Text extracted with pypdf. Scanned PDFs (image-only) are not OCR'd—run them through a tool such as ocrmypdf before ingestion.
- DOCX / PPTX
- Supported natively by LlamaIndex loaders. Tables and lists are preserved as plain text.
- Markdown / TXT / HTML
- Immediate indexing is the format that works best for RAG in practice.
- CSV
- Each line becomes a chunk. Useful for FAQ databases or business extracts.
#Concrete business use cases
#HR: query payslips and collective bargaining agreements
Typical case: an HR service with 200 monthly payslips archived as PDFs, plus the industry's collective agreement (often 100+ pages). PrivateGPT v2 lets an HR employee ask questions like "What is Ms. Dupont's coefficient in 2025?" or "What does the collective agreement say about leave for caring for a sick child?"
- Recommended model
- Qwen 3.5 9B Q4 is enough for factual extraction. For legal interpretation of the convention, move up to Mistral Small 24B (good in French, 14 GB) or Qwen 3.8 27B (18 GB, 262k context) if your VRAM can handle it.
- Cutting
- For payroll slips (1 page), one chunk = one document. For the collective agreement, section-based chunking (H2/H3 headings) works better than token-based chunking.
- Privacy
- This is precisely where PrivateGPT v2 makes a difference: pay slips should never pass through a cloud service, including in “enterprise” mode.
#Legal: analyzing a contract portfolio
Typical case: a legal department with a few hundred customer and supplier contracts in PDF, sometimes scanned. Common questions: "Which contracts expire in the next 6 months?", "Which termination clause applies to the ACME contract?", "Which ones contain an exclusivity clause?"
- Preprocessing
- Run ocrmypdf on scans before ingestion. Without OCR, these PDFs are invisible to retrieval.
- Recommended model
- Mistral Small 24B or Qwen 3.8 27B for French legal-reasoning quality. Qwen 3.5 9B is enough for pure research.
- System prompt
- Define a system prompt that requires citing the contract name and clause number in every response — otherwise the model tends to summarize without citing sources.
#Accountant: query a folder of invoices and statements
Typical case: an accounting firm wants to query a corpus of a client's supplier invoices and bank statements (PDF + CSV). Target questions: "What is the total of Free Mobile invoices in 2025?", "Is there an unpaid invoice with supplier X?".
- A limitation to know
- RAG does not aggregate naturally: it retrieves relevant passages but does not sum them perfectly across large volumes. For strict analytics, export the invoice CSV to a dedicated tool and use PrivateGPT for qualitative context.
- Preferred format
- Accounting CSVs are indexed line by line—good for individual lookups, bad for calculations. Attach a summary PDF for each month if you want the LLM to have an overall view.
- Model
- Qwen 3.5 9B (or even Q8, 11 GB) is very strong on numbers and table reading compared with a small 3B model. If you have a RTX 4070 with 12 GB or more, it is the default choice here.
#Common troubleshooting
- "Connection refused" on startup
- The Ollama daemon is not running. Check with "curl http://localhost:11434" — you should see "Ollama is running".
- Very slow responses
- Either Ollama is running on the CPU (check with "ollama ps" — PROCESSOR column), or your context_window is too high. Reduce it to 4096 for testing.
- "Model not found"
- The model name in settings-ollama.yaml must exactly match a model listed by "ollama list". Watch out for quantization tags: qwen3.5:9b ≠ qwen3.5:9b-q8_0.
- Different embeddings between ingestion and querying
- If you change the embeddings model after ingestion, you must re-index. Delete local_data/private_gpt/qdrant/ and restart ingestion.
- Inaccessible Gradio UI
- If you’re on a remote machine, launch with "PGPT_PROFILES=ollama python -m private_gpt" and reconfigure the host in settings.yaml (server.host: 0.0.0.0).
#Go further
PrivateGPT v2 is an excellent entry point for local document RAG, but it is only one building block. A few ideas for going further:
- No-code RAG with Open WebUI or AnythingLLM
- If PrivateGPT seems cumbersome to install to you (Poetry, Python 3.11), Open WebUI offers a comparable RAG experience that is easier to deploy in Docker.
- High-performing French embeddings
- nomic-embed-text gets the job done, but specialized French models such as Solon or BGE-M3 deliver better results on French legal or administrative content.
- Advanced chunking strategies
- Section-based chunking (Markdown headers, DOCX structure) greatly outperforms fixed-token chunking on structured corpora—a substantial relevance boost that is easy to achieve.
Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.