Beginner 8 minRAG

What is RAG and how does it work (guide beginner)

What is RAG? The short answer: a setup that connects an LLM to your documents so it answers with real facts instead of making things up. The long answer is this guide. No math, no required framework—just the building blocks (embeddings, vector database, LLM) and how they fit together. By the end, you will know why a well-built RAG hallucinates much less and where to start locally.

By Mohamed Meguedmi·Update 2026-08-27·Tested on Windows, macOS, and Linux

#RAG in 30 seconds

RAG stands for Retrieval-Augmented Generation: text generation augmented by retrieval. Instead of asking the LLM directly, “answer this question,” you first search a document database for the most relevant passages, then paste them into the prompt and say: “here are the sources; answer based on them.”

The useful analogy: an LLM on its own is a brilliant student answering an exam from memory. RAG is the same student being allowed to open the textbook on the table. It makes fewer things up, cites the right page, and if you give it a textbook it has never seen (your PDFs, emails, or internal wiki), it can still answer questions about it.

i
In one sentence
RAG = retrieve the right passages from your documents and inject them into the LLM's context before it responds. That's it. Everything else is engineering around these two ideas.

#Why (and when) you need it

The Local RAG Kit

Your documents, your AI: a reliable local RAG over your PDFs, notes and mail — nothing leaves your machine.

  • Lifetime online access
  • PDF + files
  • Lifetime updates

An LLM has two major shortcomings that become apparent as soon as you use it seriously: it makes things up when it does not know (the famous “hallucinations”), and it knows only the data it saw during training. Qwen 3.5 9B has never read your contract, your Notion wiki, or your incident database. Asking it to answer directly from them is like asking someone to imagine the contents of a book they have never opened.

RAG solves both at once: you inject the right excerpts into the prompt, the LLM uses them as a factual basis, and the answers become traceable—you can display the sources.

Chat with your PDFs
Technical notices, contracts, scientific papers, manuals—anything too large to fit in a context window.
Internal team assistant
Wiki, support knowledge base, product documentation. Instead of an approximate Ctrl+F search, get a French-language answer that cites the right pages.
Monitoring and summarization
Index hundreds of articles or reports, ask cross-cutting questions, and compare sources.
Recent or private data
Everything the LLM could not see: your code, your emails, and publications released after its cutoff date.
→
When RAG adds nothing
For purely generative tasks (free-form writing, brainstorming, translation, short code refactoring), an LLM without RAG is more than sufficient. Don't add RAG just because it's trendy—it helps when you have a corpus the model doesn't know, not before.

#The 4-step pipeline

A RAG has two phases: indexing (once, up front) and querying (for every question). Here are the four building blocks in sequence.

  1. 01
    1. Chunking — splitting documents
    Your PDFs, Markdown files, or web pages are first split into chunks of approximately 200 to 800 words. You can't embed an entire book at once, and in any case, you want to retrieve the specific passage that answers the question, not the entire document.
  2. 02
    2. Embeddings — turning text into vectors
    Each chunk passes through an embedding model that transforms it into a vector of numbers, typically 384 to 1,024 dimensions. Two passages discussing the same thing will produce nearby vectors in this space—that's the magic that enables semantic search.
  3. 03
    3. Storage in a vector database
    The vectors and original text are stored in a specialized database (Chroma, Qdrant, FAISS…) that can quickly answer the question, “Which vectors are closest to mine?”
  4. 04
    4. Retrieval + generation
    For the user's question, we calculate its embedding, retrieve the 3 to 10 closest chunks, add them to the LLM prompt with an instruction such as “answer using these excerpts,” and the LLM generates the response.
Pseudocode for the complete pipeline
# Phase 1 : INDEXATION (une fois)
for doc in documents:
    chunks = split(doc, taille=500)              # 1. chunking
    for chunk in chunks:
        vector = embedder.embed(chunk)            # 2. embedding
        vector_db.add(vector, chunk)              # 3. stockage

# Phase 2 : INTERROGATION (à chaque question)
question = "Quel est le délai de résiliation ?"
q_vector = embedder.embed(question)               # même modèle qu'à l'indexation
top_chunks = vector_db.search(q_vector, k=5)      # 4a. retrieval

prompt = f"""Réponds en t'appuyant uniquement sur les extraits ci-dessous.

Extraits :
{top_chunks}

Question : {question}"""
reponse = llm.generate(prompt)                    # 4b. génération
i
The LLM doesn’t “search”—it reads
Many illustrations give the impression that the LLM will “query a database.” False. Retrieval happens before the LLM call. When the LLM steps in, it receives a standard prompt with the passages already inserted. From the model's perspective, it is simply an enriched conversation.

#Embeddings: the heart of retrieval

An embedding model is a mini-LLM specialized for a single task: turning a piece of text into a vector of numbers that captures its “meaning.” Two sentences about the same subject will produce nearby vectors, even if they have no words in common. That is what distinguishes RAG from a basic Ctrl+F.

The final quality of RAG depends as much—often more—on the embedding model as on the LLM behind it. A poor embedding retrieves the wrong chunks, and even the best LLM in the world cannot answer correctly from irrelevant text fragments.

nomic-embed-text
137M parameters, 768 dimensions, 8192-token context. The sensible default offered by Ollama. Good in English, decent in French.
mxbai-embed-large
335M parameters, 1024 dimensions. More precise, 3× slower. Relevant when retrieval quality is the bottleneck.
multilingual-e5-large
560M, 1024 dimensions. The best choice if your documents are in French or multilingual.
bge-m3
Excellent in French and supports long contexts. Heavier to run but a benchmark for multilingual content.
!
The pitfall of English embeddings for French
An English nomic or bge model on French PDFs reduces retrieval relevance by a factor of 1.5 to 2. You'll get passages that are vaguely related to the topic rather than ones that actually answer the question. For a French corpus, choose multilingual-e5-large or bge-m3 from the start—the additional latency is negligible compared with the gain.

#The vector database: where vectors live

A vector database is a database specialized for one operation: “find me the N vectors closest to this one.” Behind the scenes, it uses algorithms (HNSW, IVF…) that make this search fast even across millions of vectors. To get started, you don't need to understand any of these algorithms—just know which one to choose.

Chroma
Open-source vector database, embedded in your Python project or running as a server. Ideal for getting started: zero configuration, disk persistence.
Qdrant
More robust for production: dedicated server, filters, multi-tenant. Runs in a Docker container with one command.
FAISS
Facebook (Meta) library. Very fast, but it is just an index—no metadata management. Good for performance-critical use cases.
Stored in the tool
Open WebUI, AnythingLLM, and LM Studio include their own vector database. It's invisible—you upload a PDF, and it gets indexed. Perfect for getting started without coding.

#The LLM: guided generation

The LLM is the final link in the chain. It receives a prompt that looks like: « here are 5 excerpts from the documentation. Answer the question using only them. If the information is not in the excerpts, say so. »

This framing changes everything. Without injected context, the LLM answers from its training memory—and makes things up when that memory is incomplete. With the right excerpts in the prompt, it has a factual basis in front of it and simply reformulates or summarizes.

What LLM size?
For simple RAG, a small 2026 model that fits in 8 GB (Qwen 3.5 9B, Granite 4.2 8B) is more than enough. Retrieval quality matters more than LLM size.
What context window?
At least 4096 tokens. Retrieved chunks + the question + the system instruction quickly consume 2,000–3,000 tokens. With 8192 or more, you have plenty of room.
Which system prompt?
Something like: “respond in French using only the supplied excerpts. If the information isn't there, say so clearly.”

#Local RAG vs cloud API

You can build a RAG system with the OpenAI or Claude API (quick to set up, maximum performance), or run everything locally with Ollama + a vector database + an embedding model (zero data leakage, zero usage costs). The choice depends on your documents' sensitivity and your budget.

RAG via cloud API
Your documents are sent to the provider (OpenAI, Anthropic, Mistral…) during indexing and with every question. Top-tier performance and quality, but incompatible with confidential data (GDPR, medical confidentiality, client contracts).
100% local RAG
Ollama for the LLM, nomic or bge for embeddings, and Chroma or Qdrant for the database. No data leaves the machine. Ideal for professionals (legal, medical, HR), companies subject to the GDPR, and anyone who wants to stay in control.
Hybrid
Local embeddings, LLM via API: limits exposure (full documents stay local, and only the relevant chunks are sent to the cloud for the query). A pragmatic compromise, but not recommended if the chunks themselves are sensitive.
→
Local RAG works on modest hardware
You don't need a RTX 4090. A PC with 16 GB of RAM and a recent CPU runs a complete RAG (Qwen 3.5 9B + nomic-embed-text + Chroma) at 5-10 tokens/sec. With an 8 GB GPU (RTX 3060, 4060), you reach 30-50 tokens/sec. That's more than enough to chat with your documents.

#Where to start in practice

Three paths depending on your profile. All three run 100% locally on your machine.

  1. 01
    Without coding, with an interface (Open WebUI or AnythingLLM)
    You install Ollama, launch Open WebUI or AnythingLLM in Docker, upload your PDFs to a “Knowledge Base,” and start chatting. Everything—chunking, embeddings, retrieval—is handled for you. 30 minutes from start to finish.
  2. 02
    No coding, all-in-one mode (LM Studio)
    Since version 0.3, LM Studio has a “Chat with Documents” feature: attach a file to the chat, and it processes it. Limit: 5 files per chat and restricted formats (text PDF, DOCX, TXT, MD). Perfect for occasional questions about a document.
  3. 03
    In Python with LangChain or LlamaIndex
    Full control: chunker, embedding model, database, LLM, and prompt selection. Set aside half a day for a clean first prototype, more if you want to optimize (reranker, hybrid search, etc.).
Minimal local RAG stack with Ollama
# 1. Installer Ollama (https://ollama.com)
# 2. Récupérer un LLM et un modèle d'embeddings
ollama pull qwen3.5:9b
ollama pull nomic-embed-text

# 3. Vérifier que le daemon répond
curl http://localhost:11434/api/tags

# 4. Lancer Open WebUI (interface ChatGPT-like avec RAG intégré)
docker run -d -p 3000:8080 \
  --add-host=host.docker.internal:host-gateway \
  -v open-webui:/app/backend/data \
  --name open-webui --restart always \
  ghcr.io/open-webui/open-webui:main

# Ouvrir http://localhost:3000
# Settings -> Documents -> uploader des PDF
# Dans un chat : taper # pour attacher une Knowledge Base

#Common pitfalls when getting started

English embedder, French documents
The number 1 cause of disappointing RAG in France. Make sure your embedding model supports French (multilingual-e5, bge-m3).
Chunks that are too large or too small
500 words is a good starting point. Too small (< 100 words), chunks lose their context; too large (> 1500), the embedding averages everything out and loses precision.
Change the embedder without reindexing
The vectors from one model are not compatible with those from another. If you switch from nomic to bge, you must reindex everything—otherwise retrieval returns nonsense.
Too many chunks in the context
Beyond 8-10 chunks, the LLM starts to lose track. It’s better to use 5 highly relevant chunks than 20 moderately relevant ones (that’s where a reranker comes in, at a more advanced level).
No citations displayed
To verify that a RAG is really working, display the sources used for each answer. Without this visibility, you will not know how to distinguish a good answer from a convincing hallucination.

#Go further

Now that you know what RAG is, here are the natural guides for putting it into practice.

Local RAG with Ollama without coding
The step-by-step Open WebUI + AnythingLLM tutorial for building your first RAG in 30 minutes, even without a GPU.
The best French embedding models
Detailed comparison to choose between nomic, multilingual-e5, bge-m3 and Solon for your corpus.
Local RAG: introduction
The in-depth conceptual guide: chunking, retrieval, reranking, evaluation metrics.
What is Ollama and how does it work
If you haven't installed Ollama yet, go here.
Did this guide help you?

Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.