Beginner 11 minConcepts

Local RAG: introduction

Direct response

A local RAG (retrieval-augmented generation) lets you query your own documents with a model running on your machine: your files are split into chunks, converted into vectors by an embedding model, and then, for each question, only the closest passages are added to the model's prompt. Nothing needs to leave the computer, either for indexing or for the response.

This introduction explains how local RAG works, the minimal stack for building it on Ollama, no-code options, an example in Python with LlamaIndex, and especially where most first attempts fail: chunking, language, retrieval, and PDFs. It also explains when RAG is not the right solution.

By Mohamed Meguedmi·Update 2026-09-29·Tested on Windows, macOS, and Linux

#Local RAG: the definition and what to expect from it

RAG stands for retrieval-augmented generation: relevant passages are first retrieved from a document database, then the model is asked to draft a response based on them. The word local specifies that every step (reading the documents, computing embeddings, storage, and generation) runs on your hardware. This is the most useful application of a private model: questions about contracts, notes, or an internal knowledge base, with responses that cite their sources.

What each component of a local RAG does
ComponentRoleExamples
Document readerExtract clean text from PDFs, Word, Markdown, and HTMLLlamaIndex and Docling readers
ChunkerSplit the text into useful-sized passagesSplitting by tokens or by structure
Embedding modelTurn each passage into a vectorembeddinggemma, qwen3-embedding, all-minilm (recommended by Ollama)
Vector databaseStore vectors and search for the nearest onesChromaDB, Qdrant, FAISS, pgvector
Generative modelWrite the answer from the passagesModel served by Ollama or LM Studio

#Why use RAG instead of sending everything to the model

The Local RAG Kit

Your documents, your AI: a reliable local RAG over your PDFs, notes and mail — nothing leaves your machine.

  • Lifetime online access
  • PDF + files
  • Lifetime updates

First idea: copy all your documents into the prompt. Two limitations stand in the way. The context window is bounded: 300 PDFs represent millions of tokens, far beyond what a local model can accept. And even when a document fits, quality declines: a landmark study (Liu et al., Stanford) shows that performance can drop significantly when the relevant information is located in the middle of a long context, even for models designed for long contexts. This phenomenon is called lost in the middle.

RAG avoids these two problems by sending the model only a few passages selected for the question. The model reads a few hundred well-targeted tokens instead of tens of thousands of barely useful tokens, with lower compute cost and latency. The whole art lies in retrieving the right information.

As an illustrative calculation, 300 20-page PDFs at about 600 tokens per page represent 3.6 million tokens—hundreds of times the 8,000-token window commonly set for a local model. Split into 400-token passages, they produce about 9,000 passages. A question retrieves five, or 2,000 tokens: the model reads 0.06% of the corpus, but the right 0.06% if retrieval succeeds.

i
A RAG is not always necessary
LM Studio, in its documentation, illustrates the choice well: if the document is short enough to fit within the model’s context, it adds the entire document to the conversation. It switches to RAG only if the document is very long. This is the right rule of thumb to start with.

#The concept in two minutes: embeddings and distance

An embedding is a list of numbers that represents the meaning of a text. Two texts about the same thing have nearby vectors, even if they don't use the same words. The documentation for Ollama describes embeddings as numerical vectors that can be stored in a vector database, searched by cosine similarity, or used in a RAG system, with a length that depends on the model, typically between 384 and 1024 dimensions.

A question is transformed into a vector by the same model, then compared with all indexed passages: the closest ones are returned. The point that trips up beginners: the embedding model used for the index and the one used for questions must be identical. LlamaIndex documentation reiterates this in its index-reloading example: it is important to use the same embed_model that was used to build it.

#Anatomy of a RAG: two phases

#Phase 1: indexing, performed once per document

  1. 01
    Ingestion
    Read the files (PDF, Word, Markdown, HTML) and extract clean text from them. This is the most underestimated step.
  2. 02
    Breakdown
    Split into chunks small enough to be precise, but large enough to preserve meaning. Chunks of 200 to 500 tokens are a common starting point.
  3. 03
    Embedding
    Run each passage through the embedding model, for example with the ollama run embeddinggemma command or the /api/embed API.
  4. 04
    Storage
    Store the vectors with the original text and metadata (file name, page, date).

#Phase 2: the request, with each question

  1. 01
    Question embedding
    With the same model used for indexing.
  2. 02
    Search
    Find the N passages whose vectors are closest, measured by cosine similarity.
  3. 03
    Prompt assembly
    Build a message containing the excerpts and instructions to answer with citations and say when the answer isn't found in them.
  4. 04
    Generation
    Send this prompt to the local model, which writes the response.

#The minimum stack, plus no-code options

For a local RAG system that works, four components are enough: a generation model served by Ollama, an embedding model, a vector database, and a layer that connects them. For French, choose a multilingual embedding model; the Ollama page recommends three (embeddinggemma, qwen3-embedding, all-minilm), and the vector sizes remain modest, so they can run on a laptop.

Choose your effort level
PathEffortControlSuitable for
LM Studio, Chat with DocumentsDrag .pdf, .docx, or .txt into a conversationLow: automatic switching between the full document and RAGQuick test on a few files
AnythingLLMApplication with embedder and vector database selectionMediumSmall team without a developer
Open WebUIWeb interface connected to Ollama, knowledge basesMediumDaily use of a notes database
LlamaIndex or Haystack in PythonCode to writeHigh: chunking, retrieval, evaluationCustom-built or ready to deploy

If you do not want to code, the guide to RAG in LM Studio and the guide to AnythingLLM explain each approach in detail; the guide to NotebookLM and local alternatives covers the query “notebook lm rag”.

#The Python pipeline, step by step

The example below follows the official LlamaIndex tutorial pattern for local models: a directory reader, an embedding model, a model served by Ollama, and then a query engine. First install the llama-index-llms-ollama and llama-index-embeddings-huggingface packages. Adapt the model names to match what you downloaded.

Minimal RAG with LlamaIndex and Ollama
from llama_index.core import VectorStoreIndex, SimpleDirectoryReader, Settings
from llama_index.llms.ollama import Ollama
from llama_index.embeddings.huggingface import HuggingFaceEmbedding

Settings.embed_model = HuggingFaceEmbedding(model_name="intfloat/multilingual-e5-large")
Settings.llm = Ollama(model="qwen3.5:9b", request_timeout=360.0, context_window=8000)

documents = SimpleDirectoryReader("mes_documents").load_data()
index = VectorStoreIndex.from_documents(documents)
query_engine = index.as_query_engine(similarity_top_k=5)

response = query_engine.query("Quelles sont les échéances du contrat Dupont ?")
print(response)
print(response.source_nodes)

The first run computes all embeddings, which takes more or less time depending on the volume and the machine. To avoid recalculating everything, save the index with index.storage_context.persist, then reload it with load_index_from_storage while reusing the same embedding model. The context_window parameter limits memory usage, as the tutorial explains.

→
Always display sources
Display the retrieved passages (response.source_nodes) in every test. If the excerpts are off-topic, the problem is in retrieval, not the model: there's no point in changing the generation model.

#The pitfalls that derail your first attempts

Naive chunking breaks structures
A cut through the middle of a table produces unreadable passages. Use a splitter that preserves headings and tables, or convert the documents first with Docling.
The embedding language matters
An embedding model trained primarily on English retrieves French passages less accurately. Choose a multilingual model and test it on 10 real questions before indexing the entire corpus.
A single top-k rarely works
Too few passes impoverish the response; too many drown the model. Adjust based on the nature of the documents and verify using questions whose answers you know.
A bad search produces a confident, incorrect answer
A model given irrelevant excerpts can make up an answer that sounds confident. Measure excerpt relevance first.
PDFs are tricky
Two-column layouts, footers, tables, scans: a significant part of the work is cleaning up ingestion. A scanned PDF requires OCR before any other processing.
Mixing embedding models
Indexing with one model and querying with another produces absurd results without an error message.

#When not to use RAG

RAG, long context, or another approach
SituationBest approachWhy
Fewer than twenty pagesPut everything in the contextSimpler, with no loss of passage; that's also what LM Studio does when the document fits
Factual question about structured data (2024 revenue)SQL query or extraction scriptRAG retrieves text, not exact calculations
Summary question covering the entire corpusHierarchical summaries followed by a question about the summariesRAG retrieves only a few passages, not an overview
Documents that are constantly modifiedIncremental indexing or keyword searchReindexing after every change costs more than querying
Learn a style or formatFine-tuningRAG provides facts, not a writing style

The guide comparing fine-tuning and RAG covers the latter case in detail.

#Hardware and privacy: what stays on your machine

In a fully local RAG, documents, vectors, and queries stay on the machine, provided every component is local: embeddings served by Ollama or loaded from disk, a vector database in a file or local service, and a local generation model. Make sure no component calls an online service: an embedder or cloud model selected by mistake would send your passages outside the machine.

On the hardware side, the heaviest component is the generation model: in Q4, an 8- to 9-billion-parameter model fits in about 5 GB of memory, plus the context. The context grows here because it contains the excerpts, so leave headroom if you increase the number of passages. The embedding model itself is lightweight; initial indexing of a large corpus remains the longest operation and is performed once. The guide to LLMs without a GPU explains what each amount of RAM allows.

#What next? Improvements in order

A basic RAG system often answers simple questions well. To go further, the most cost-effective order is: first, chunking that respects the structure; then hybrid search (vector and BM25 keyword search, so you do not miss an exact term such as a contract number); then a reranker that reorders the top results; and finally, evaluation on a set of about fifty questions whose answers you know. Without measurement, every change remains an impression.

Frequently asked questions about local RAG
What is a local RAG?+
It is a system that answers questions from your documents by running reading, indexing, searching, and generation on your machine. Relevant passages are retrieved by vector similarity and then provided to a local model. Your files never leave your computer.
What is the difference between RAG and fine-tuning?+
RAG gives the model excerpts from your documents at query time, providing updatable facts. Fine-tuning modifies the model’s weights: it changes its style or response format, but retains precise facts poorly. To query documents, start with RAG.
Which Embedding Model Should You Choose for French?+
Choose a multilingual model. Ollama recommends embeddinggemma, qwen3-embedding, and all-minilm; the first has 300 million parameters. Test it on ten real questions from your corpus before adopting it: quality depends on your documents. Then keep the same model for indexing and querying, or the results will become inconsistent.
Do you need a GPU for local RAG?+
Not necessarily. Embedding models are small and run on the CPU; indexing is slower without a GPU, but it happens once. The generation model is the main memory consumer: an 8- to 9-billion-parameter model in Q4 requires about 5 GB of memory.
What chunk size should you use for indexing?+
A common starting point is 200 to 500 tokens per passage, with slight overlap. There is no universal value: test two or three sizes on questions whose answers you know. Splitting that preserves headings and tables matters more than the exact size.
How can you tell whether your RAG is responding correctly?+
Prepare around fifty questions whose answers and source passages you know. After each change, check whether the right passage appears in the returned excerpts, then whether the answer cites it correctly. Tools such as Ragas automate this measurement, which is explained in a dedicated guide.
Did this guide help you?

Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.