Local RAG: introduction
A local RAG (retrieval-augmented generation) lets you query your own documents with a model running on your machine: your files are split into chunks, converted into vectors by an embedding model, and then, for each question, only the closest passages are added to the model's prompt. Nothing needs to leave the computer, either for indexing or for the response.
This introduction explains how local RAG works, the minimal stack for building it on Ollama, no-code options, an example in Python with LlamaIndex, and especially where most first attempts fail: chunking, language, retrieval, and PDFs. It also explains when RAG is not the right solution.
#Local RAG: the definition and what to expect from it
RAG stands for retrieval-augmented generation: relevant passages are first retrieved from a document database, then the model is asked to draft a response based on them. The word local specifies that every step (reading the documents, computing embeddings, storage, and generation) runs on your hardware. This is the most useful application of a private model: questions about contracts, notes, or an internal knowledge base, with responses that cite their sources.
| Component | Role | Examples |
|---|---|---|
| Document reader | Extract clean text from PDFs, Word, Markdown, and HTML | LlamaIndex and Docling readers |
| Chunker | Split the text into useful-sized passages | Splitting by tokens or by structure |
| Embedding model | Turn each passage into a vector | embeddinggemma, qwen3-embedding, all-minilm (recommended by Ollama) |
| Vector database | Store vectors and search for the nearest ones | ChromaDB, Qdrant, FAISS, pgvector |
| Generative model | Write the answer from the passages | Model served by Ollama or LM Studio |
#Why use RAG instead of sending everything to the model
Your documents, your AI: a reliable local RAG over your PDFs, notes and mail — nothing leaves your machine.
- Lifetime online access
- PDF + files
- Lifetime updates
First idea: copy all your documents into the prompt. Two limitations stand in the way. The context window is bounded: 300 PDFs represent millions of tokens, far beyond what a local model can accept. And even when a document fits, quality declines: a landmark study (Liu et al., Stanford) shows that performance can drop significantly when the relevant information is located in the middle of a long context, even for models designed for long contexts. This phenomenon is called lost in the middle.
RAG avoids these two problems by sending the model only a few passages selected for the question. The model reads a few hundred well-targeted tokens instead of tens of thousands of barely useful tokens, with lower compute cost and latency. The whole art lies in retrieving the right information.
As an illustrative calculation, 300 20-page PDFs at about 600 tokens per page represent 3.6 million tokens—hundreds of times the 8,000-token window commonly set for a local model. Split into 400-token passages, they produce about 9,000 passages. A question retrieves five, or 2,000 tokens: the model reads 0.06% of the corpus, but the right 0.06% if retrieval succeeds.
#The concept in two minutes: embeddings and distance
An embedding is a list of numbers that represents the meaning of a text. Two texts about the same thing have nearby vectors, even if they don't use the same words. The documentation for Ollama describes embeddings as numerical vectors that can be stored in a vector database, searched by cosine similarity, or used in a RAG system, with a length that depends on the model, typically between 384 and 1024 dimensions.
A question is transformed into a vector by the same model, then compared with all indexed passages: the closest ones are returned. The point that trips up beginners: the embedding model used for the index and the one used for questions must be identical. LlamaIndex documentation reiterates this in its index-reloading example: it is important to use the same embed_model that was used to build it.
#Anatomy of a RAG: two phases
#Phase 1: indexing, performed once per document
- 01IngestionRead the files (PDF, Word, Markdown, HTML) and extract clean text from them. This is the most underestimated step.
- 02BreakdownSplit into chunks small enough to be precise, but large enough to preserve meaning. Chunks of 200 to 500 tokens are a common starting point.
- 03EmbeddingRun each passage through the embedding model, for example with the ollama run embeddinggemma command or the /api/embed API.
- 04StorageStore the vectors with the original text and metadata (file name, page, date).
#Phase 2: the request, with each question
- 01Question embeddingWith the same model used for indexing.
- 02SearchFind the N passages whose vectors are closest, measured by cosine similarity.
- 03Prompt assemblyBuild a message containing the excerpts and instructions to answer with citations and say when the answer isn't found in them.
- 04GenerationSend this prompt to the local model, which writes the response.
#The minimum stack, plus no-code options
For a local RAG system that works, four components are enough: a generation model served by Ollama, an embedding model, a vector database, and a layer that connects them. For French, choose a multilingual embedding model; the Ollama page recommends three (embeddinggemma, qwen3-embedding, all-minilm), and the vector sizes remain modest, so they can run on a laptop.
| Path | Effort | Control | Suitable for |
|---|---|---|---|
| LM Studio, Chat with Documents | Drag .pdf, .docx, or .txt into a conversation | Low: automatic switching between the full document and RAG | Quick test on a few files |
| AnythingLLM | Application with embedder and vector database selection | Medium | Small team without a developer |
| Open WebUI | Web interface connected to Ollama, knowledge bases | Medium | Daily use of a notes database |
| LlamaIndex or Haystack in Python | Code to write | High: chunking, retrieval, evaluation | Custom-built or ready to deploy |
If you do not want to code, the guide to RAG in LM Studio and the guide to AnythingLLM explain each approach in detail; the guide to NotebookLM and local alternatives covers the query “notebook lm rag”.
#The Python pipeline, step by step
The example below follows the official LlamaIndex tutorial pattern for local models: a directory reader, an embedding model, a model served by Ollama, and then a query engine. First install the llama-index-llms-ollama and llama-index-embeddings-huggingface packages. Adapt the model names to match what you downloaded.
The first run computes all embeddings, which takes more or less time depending on the volume and the machine. To avoid recalculating everything, save the index with index.storage_context.persist, then reload it with load_index_from_storage while reusing the same embedding model. The context_window parameter limits memory usage, as the tutorial explains.
#The pitfalls that derail your first attempts
- Naive chunking breaks structures
- A cut through the middle of a table produces unreadable passages. Use a splitter that preserves headings and tables, or convert the documents first with Docling.
- The embedding language matters
- An embedding model trained primarily on English retrieves French passages less accurately. Choose a multilingual model and test it on 10 real questions before indexing the entire corpus.
- A single top-k rarely works
- Too few passes impoverish the response; too many drown the model. Adjust based on the nature of the documents and verify using questions whose answers you know.
- A bad search produces a confident, incorrect answer
- A model given irrelevant excerpts can make up an answer that sounds confident. Measure excerpt relevance first.
- PDFs are tricky
- Two-column layouts, footers, tables, scans: a significant part of the work is cleaning up ingestion. A scanned PDF requires OCR before any other processing.
- Mixing embedding models
- Indexing with one model and querying with another produces absurd results without an error message.
#When not to use RAG
| Situation | Best approach | Why |
|---|---|---|
| Fewer than twenty pages | Put everything in the context | Simpler, with no loss of passage; that's also what LM Studio does when the document fits |
| Factual question about structured data (2024 revenue) | SQL query or extraction script | RAG retrieves text, not exact calculations |
| Summary question covering the entire corpus | Hierarchical summaries followed by a question about the summaries | RAG retrieves only a few passages, not an overview |
| Documents that are constantly modified | Incremental indexing or keyword search | Reindexing after every change costs more than querying |
| Learn a style or format | Fine-tuning | RAG provides facts, not a writing style |
The guide comparing fine-tuning and RAG covers the latter case in detail.
#Hardware and privacy: what stays on your machine
In a fully local RAG, documents, vectors, and queries stay on the machine, provided every component is local: embeddings served by Ollama or loaded from disk, a vector database in a file or local service, and a local generation model. Make sure no component calls an online service: an embedder or cloud model selected by mistake would send your passages outside the machine.
On the hardware side, the heaviest component is the generation model: in Q4, an 8- to 9-billion-parameter model fits in about 5 GB of memory, plus the context. The context grows here because it contains the excerpts, so leave headroom if you increase the number of passages. The embedding model itself is lightweight; initial indexing of a large corpus remains the longest operation and is performed once. The guide to LLMs without a GPU explains what each amount of RAM allows.
#What next? Improvements in order
A basic RAG system often answers simple questions well. To go further, the most cost-effective order is: first, chunking that respects the structure; then hybrid search (vector and BM25 keyword search, so you do not miss an exact term such as a contract number); then a reranker that reorders the top results; and finally, evaluation on a set of about fifty questions whose answers you know. Without measurement, every change remains an impression.
What is a local RAG?+
What is the difference between RAG and fine-tuning?+
Which Embedding Model Should You Choose for French?+
Do you need a GPU for local RAG?+
What chunk size should you use for indexing?+
How can you tell whether your RAG is responding correctly?+
- What is RAG? A beginner’s guide
- Document splitting strategies
- Embeddings for French
- RAG in LM Studio
- AnythingLLM: RAG tutorial
- Evaluate a RAG with Ragas
- Source: embeddings in Ollama
- Source: LlamaIndex tutorial with local models
- Source: Lost in the Middle (Liu et al.)
- Source: LM Studio, Chat with Documents
Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.