What is RAG and how does it work (guide beginner)
What is RAG? The short answer: a setup that connects an LLM to your documents so it answers with real facts instead of making things up. The long answer is this guide. No math, no required framework—just the building blocks (embeddings, vector database, LLM) and how they fit together. By the end, you will know why a well-built RAG hallucinates much less and where to start locally.
#RAG in 30 seconds
RAG stands for Retrieval-Augmented Generation: text generation augmented by retrieval. Instead of asking the LLM directly, “answer this question,” you first search a document database for the most relevant passages, then paste them into the prompt and say: “here are the sources; answer based on them.”
The useful analogy: an LLM on its own is a brilliant student answering an exam from memory. RAG is the same student being allowed to open the textbook on the table. It makes fewer things up, cites the right page, and if you give it a textbook it has never seen (your PDFs, emails, or internal wiki), it can still answer questions about it.
#Why (and when) you need it
Your documents, your AI: a reliable local RAG over your PDFs, notes and mail — nothing leaves your machine.
- Lifetime online access
- PDF + files
- Lifetime updates
An LLM has two major shortcomings that become apparent as soon as you use it seriously: it makes things up when it does not know (the famous “hallucinations”), and it knows only the data it saw during training. Qwen 3.5 9B has never read your contract, your Notion wiki, or your incident database. Asking it to answer directly from them is like asking someone to imagine the contents of a book they have never opened.
RAG solves both at once: you inject the right excerpts into the prompt, the LLM uses them as a factual basis, and the answers become traceable—you can display the sources.
- Chat with your PDFs
- Technical notices, contracts, scientific papers, manuals—anything too large to fit in a context window.
- Internal team assistant
- Wiki, support knowledge base, product documentation. Instead of an approximate Ctrl+F search, get a French-language answer that cites the right pages.
- Monitoring and summarization
- Index hundreds of articles or reports, ask cross-cutting questions, and compare sources.
- Recent or private data
- Everything the LLM could not see: your code, your emails, and publications released after its cutoff date.
#The 4-step pipeline
A RAG has two phases: indexing (once, up front) and querying (for every question). Here are the four building blocks in sequence.
- 011. Chunking — splitting documentsYour PDFs, Markdown files, or web pages are first split into chunks of approximately 200 to 800 words. You can't embed an entire book at once, and in any case, you want to retrieve the specific passage that answers the question, not the entire document.
- 022. Embeddings — turning text into vectorsEach chunk passes through an embedding model that transforms it into a vector of numbers, typically 384 to 1,024 dimensions. Two passages discussing the same thing will produce nearby vectors in this space—that's the magic that enables semantic search.
- 033. Storage in a vector databaseThe vectors and original text are stored in a specialized database (Chroma, Qdrant, FAISS…) that can quickly answer the question, “Which vectors are closest to mine?”
- 044. Retrieval + generationFor the user's question, we calculate its embedding, retrieve the 3 to 10 closest chunks, add them to the LLM prompt with an instruction such as “answer using these excerpts,” and the LLM generates the response.
#Embeddings: the heart of retrieval
An embedding model is a mini-LLM specialized for a single task: turning a piece of text into a vector of numbers that captures its “meaning.” Two sentences about the same subject will produce nearby vectors, even if they have no words in common. That is what distinguishes RAG from a basic Ctrl+F.
The final quality of RAG depends as much—often more—on the embedding model as on the LLM behind it. A poor embedding retrieves the wrong chunks, and even the best LLM in the world cannot answer correctly from irrelevant text fragments.
- nomic-embed-text
- 137M parameters, 768 dimensions, 8192-token context. The sensible default offered by Ollama. Good in English, decent in French.
- mxbai-embed-large
- 335M parameters, 1024 dimensions. More precise, 3× slower. Relevant when retrieval quality is the bottleneck.
- multilingual-e5-large
- 560M, 1024 dimensions. The best choice if your documents are in French or multilingual.
- bge-m3
- Excellent in French and supports long contexts. Heavier to run but a benchmark for multilingual content.
#The vector database: where vectors live
A vector database is a database specialized for one operation: “find me the N vectors closest to this one.” Behind the scenes, it uses algorithms (HNSW, IVF…) that make this search fast even across millions of vectors. To get started, you don't need to understand any of these algorithms—just know which one to choose.
- Chroma
- Open-source vector database, embedded in your Python project or running as a server. Ideal for getting started: zero configuration, disk persistence.
- Qdrant
- More robust for production: dedicated server, filters, multi-tenant. Runs in a Docker container with one command.
- FAISS
- Facebook (Meta) library. Very fast, but it is just an index—no metadata management. Good for performance-critical use cases.
- Stored in the tool
- Open WebUI, AnythingLLM, and LM Studio include their own vector database. It's invisible—you upload a PDF, and it gets indexed. Perfect for getting started without coding.
#The LLM: guided generation
The LLM is the final link in the chain. It receives a prompt that looks like: « here are 5 excerpts from the documentation. Answer the question using only them. If the information is not in the excerpts, say so. »
This framing changes everything. Without injected context, the LLM answers from its training memory—and makes things up when that memory is incomplete. With the right excerpts in the prompt, it has a factual basis in front of it and simply reformulates or summarizes.
- What LLM size?
- For simple RAG, a small 2026 model that fits in 8 GB (Qwen 3.5 9B, Granite 4.2 8B) is more than enough. Retrieval quality matters more than LLM size.
- What context window?
- At least 4096 tokens. Retrieved chunks + the question + the system instruction quickly consume 2,000–3,000 tokens. With 8192 or more, you have plenty of room.
- Which system prompt?
- Something like: “respond in French using only the supplied excerpts. If the information isn't there, say so clearly.”
#Local RAG vs cloud API
You can build a RAG system with the OpenAI or Claude API (quick to set up, maximum performance), or run everything locally with Ollama + a vector database + an embedding model (zero data leakage, zero usage costs). The choice depends on your documents' sensitivity and your budget.
- RAG via cloud API
- Your documents are sent to the provider (OpenAI, Anthropic, Mistral…) during indexing and with every question. Top-tier performance and quality, but incompatible with confidential data (GDPR, medical confidentiality, client contracts).
- 100% local RAG
- Ollama for the LLM, nomic or bge for embeddings, and Chroma or Qdrant for the database. No data leaves the machine. Ideal for professionals (legal, medical, HR), companies subject to the GDPR, and anyone who wants to stay in control.
- Hybrid
- Local embeddings, LLM via API: limits exposure (full documents stay local, and only the relevant chunks are sent to the cloud for the query). A pragmatic compromise, but not recommended if the chunks themselves are sensitive.
#Where to start in practice
Three paths depending on your profile. All three run 100% locally on your machine.
- 01Without coding, with an interface (Open WebUI or AnythingLLM)You install Ollama, launch Open WebUI or AnythingLLM in Docker, upload your PDFs to a “Knowledge Base,” and start chatting. Everything—chunking, embeddings, retrieval—is handled for you. 30 minutes from start to finish.
- 02No coding, all-in-one mode (LM Studio)Since version 0.3, LM Studio has a “Chat with Documents” feature: attach a file to the chat, and it processes it. Limit: 5 files per chat and restricted formats (text PDF, DOCX, TXT, MD). Perfect for occasional questions about a document.
- 03In Python with LangChain or LlamaIndexFull control: chunker, embedding model, database, LLM, and prompt selection. Set aside half a day for a clean first prototype, more if you want to optimize (reranker, hybrid search, etc.).
#Common pitfalls when getting started
- English embedder, French documents
- The number 1 cause of disappointing RAG in France. Make sure your embedding model supports French (multilingual-e5, bge-m3).
- Chunks that are too large or too small
- 500 words is a good starting point. Too small (< 100 words), chunks lose their context; too large (> 1500), the embedding averages everything out and loses precision.
- Change the embedder without reindexing
- The vectors from one model are not compatible with those from another. If you switch from nomic to bge, you must reindex everything—otherwise retrieval returns nonsense.
- Too many chunks in the context
- Beyond 8-10 chunks, the LLM starts to lose track. It’s better to use 5 highly relevant chunks than 20 moderately relevant ones (that’s where a reranker comes in, at a more advanced level).
- No citations displayed
- To verify that a RAG is really working, display the sources used for each answer. Without this visibility, you will not know how to distinguish a good answer from a convincing hallucination.
#Go further
Now that you know what RAG is, here are the natural guides for putting it into practice.
- Local RAG with Ollama without coding
- The step-by-step Open WebUI + AnythingLLM tutorial for building your first RAG in 30 minutes, even without a GPU.
- The best French embedding models
- Detailed comparison to choose between nomic, multilingual-e5, bge-m3 and Solon for your corpus.
- Local RAG: introduction
- The in-depth conceptual guide: chunking, retrieval, reranking, evaluation metrics.
- What is Ollama and how does it work
- If you haven't installed Ollama yet, go here.
Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.