Local RAG with Ollama without coding (Open WebUI, AnythingLLM)
Doing local RAG with Ollama means asking a self-hosted LLM questions while having it read your own documents—without sending a single line to the cloud. Two free tools let you set it up without coding: Open WebUI (an ChatGPT-like interface with an integrated knowledge base) and AnythingLLM (workspaces + citations). This guide covers the whole process in 10 minutes, including on a small machine without a GPU.
#Why local RAG with Ollama
A RAG (Retrieval-Augmented Generation) addresses a simple limitation of LLMs: they only know their training data. RAG connects them to your documents: internal procedure PDFs, Markdown notes, email exports, contracts, manuals — anything that is neither public nor recent. The model reads the relevant passages before answering and cites its sources if the tool supports it.
The cloud version (ChatGPT, Claude, Gemini) requires you to upload your documents to third-party servers. For a law firm, doctor, accountant, or even an individual who does not want personal notes sent to the United States, that is a deal-breaker. A local RAG Ollama solves this problem: everything stays on your machine, and you retain control over both costs (zero) and privacy.
- Total privacy
- Your documents never leave the machine. No token is sent to a provider.
- Zero marginal cost
- No subscription, no API quota. You pay for electricity, period.
- Offline
- Once the model is downloaded, RAG works without an internet connection.
- Customizable
- You choose the response model, the embeddings model, and the sources.
#Does Ollama support RAG natively?
Your AI reads your documents without a line of code. The real question remains: does it answer correctly? The Local RAG Kit starts there (ch. 1), takes Open WebUI and AnythingLLM into advanced use (ch. 13), and teaches you to measure false answers and fabricated citations (ch. 14).
- Lifetime online access
- PDF + files
- Lifetime updates
Clear answer: no. Ollama is an inference engine—it downloads and runs LLMs and exposes an OpenAI-compatible API at http://localhost:11434, but it cannot index your documents by itself. It has no vector database, PDF splitter, or interface for dragging and dropping files.
However, it can serve a generation model and an embeddings model in parallel—that's all a third-party tool (Open WebUI, AnythingLLM, LangChain…) needs to build a RAG layer on top. So when we talk about local RAG Ollama, we're always talking about a combination of Ollama + RAG interface.
#Prerequisites
- Ollama installed
- On Windows, macOS, or Linux. It listens on http://localhost:11434 by default.
- A response model
- Granite 4.2 8B or Qwen 3.5 9B in Q4_K_M. Allow for ~5 to 7 GB of VRAM or RAM.
- An embeddings model
- nomic-embed-text or mxbai-embed-large. ~300 MB. Essential for vectorizing documents.
- Docker (option Open WebUI)
- Docker Desktop on Windows/macOS, or docker.io on Linux. This is the recommended installation method.
- Disk space
- 10 GB minimum for the LLM model + embeddings + your document vector index.
#1. Open WebUI: integrated knowledge base
Open WebUI (formerly Ollama WebUI) is the most popular ChatGPT-like interface for Ollama. Its Knowledge feature lets you upload documents and query them with RAG, without any code-side configuration.
- 01Run Open WebUI in DockerIn a terminal: docker run -d -p 3000:8080 --add-host=host.docker.internal:host-gateway -v open-webui:/app/backend/data --name open-webui --restart always ghcr.io/open-webui/open-webui:main. On native Linux, replace host.docker.internal with localhost if you run Ollama outside Docker.
- 02Open the interfaceGo to http://localhost:3000. Create an admin account (the first account is admin by default). Open WebUI automatically detects Ollama on localhost:11434.
- 03Create a knowledge baseGo to Workspace → Knowledge → + Create a Knowledge Base. Give it a name (e.g., 'Client contracts'). This is your vector container.
- 04Import your documentsClick the + at the top right of the database, then upload your PDFs, .docx, .md, and .txt files. Open WebUI splits them and calculates the embeddings in the background.
- 05Configure the embedding modelSettings (admin) → Documents → Embedding Model Engine = Ollama, Embedding Model = nomic-embed-text. Save. Without this, Open WebUI uses a default model that is less suitable for French.
- 06Query the databaseIn a new chat, type # followed by the knowledge base name to attach it. Ask your question—the answer is generated from the relevant chunks, with clickable citation numbers.
#2. AnythingLLM: workspaces and citations
AnythingLLM is the most fully developed alternative to Open WebUI for RAG. While Open WebUI is general-purpose, AnythingLLM is built around the concept of workspaces: each project has its own document database, model, and history. Ideal when you want to separate “Client Contracts” from “Personal Notes” from “Technical Documentation.”
- 01Install AnythingLLMDownload the Desktop installer from useanything.com for Windows, macOS, or Linux. It's an Electron application—Docker is not required for the desktop version.
- 02Choose Ollama as your LLM providerOn first launch, the assistant asks which LLM to use. Select Ollama, enter the URL http://localhost:11434, and choose qwen3.5:9b from the list that loads automatically.
- 03Choose the embedding engineNext step: Embedder. Select Ollama, nomic-embed-text model. AnythingLLM can also use its built-in local embedder if you haven’t pulled nomic-embed-text—less capable, but zero configuration.
- 04Create a workspaceIn the sidebar: + New Workspace. Give it a name. Click it, then click the upload icon to drag in your PDFs, .docx, .csv, .epub, URL, or even YouTube videos (automatic transcription).
- 05Move documents → workspaceAnythingLLM separates ingestion (documents listed on the left) from usage (workspace on the right). Select the documents, click 'Move to Workspace,' then 'Save and Embed.' This is where embeddings are calculated via Ollama.
- 06Discuss with citationsIn the workspace, ask your question. Each answer contains an expandable Citations panel showing exactly which chunks were used. Very useful for verifying that the model did not hallucinate.
- Open WebUI
- Best if you want a general-purpose, multi-user (team) ChatGPT-like interface with web access. RAG is one feature among others.
- AnythingLLM
- Best if RAG is the primary use case: strict workspaces, well-formed citations, support for more formats (URL, YouTube, GitHub), built-in agents.
#3. Which embedding model to choose
The embedding model is the part many people overlook. It transforms your documents into vectors, and its quality directly determines the relevance of the retrieved passages. For French documents, some models are significantly better than others.
- nomic-embed-text (137M, 274 MB)
- The recommended default. Fast, reasonably good multilingual performance, 8192-token context. A very good starting compromise. Available directly: ollama pull nomic-embed-text.
- mxbai-embed-large (335M, 670 MB)
- More accurate than nomic, slightly slower. Primarily English, but handles French for standard documents. Prefer it when quality matters more than speed. ollama pull mxbai-embed-large.
- bge-m3 (567M, 1.2 GB)
- Excellent for multilingual use (100+ languages, including native French), with an 8192-token context. Heavier. Available on Hugging Face, and can be imported into Ollama via a Modelfile.
- snowflake-arctic-embed
- Good multilingual support, lighter than bge-m3. A good alternative if bge-m3 is too heavy.
#4. Local RAG on 4 GB of RAM, without a GPU
Yes, it's feasible. RAG is less resource-intensive than you might think: most of the work (embeddings) happens once during ingestion, and at runtime, it's just the LLM answering with a few thousand additional context tokens. On a small machine, the main trade-off is the answering LLM.
- 4 GB of RAM, slow CPU (old laptop)
- LLM: qwen3.5:2b or granite4.2:3b in Q4_K_M (~2 GB). Embeddings: nomic-embed-text. Slow but readable responses (3–5 tok/s). The LLM context is set to 2048 to stay below the RAM threshold.
- 8 GB of RAM, recent CPU (i5/Ryzen 5)
- LLM: qwen3.5:4b Q4 (~3.4 GB) or gemma4:e2b-it-qat. Embeddings: nomic-embed-text. ~6-10 tok/s. This is the 'pretty good' setup without a GPU.
- 16 GB of RAM, no GPU
- LLM: granite4.2:8b or qwen3.5:9b Q4 (~5-7 GB). Embeddings: nomic or mxbai. ~4-8 tok/s on pure CPU on a recent Ryzen 7 or i7.
#Troubleshooting
- The model “can’t see” my documents
- Verify that the knowledge base is attached to the chat (Open WebUI: # command, AnythingLLM: workspace icon). Also verify that the documents are actually 'embedded', not merely uploaded.
- Very slow responses on small models
- RAG adds 1000 to 3000 tokens to the context. Reduce top-k (3–5 chunks instead of 10) in the RAG settings, and lower the LLM's num_ctx to 2048 or 4096.
- Hallucinations despite the documents
- Increase top-k, and verify that the embedding model is the one that indexed the data (changing models invalidates the database). Enable citations to see what the LLM actually read.
- Open WebUI cannot see Ollama
- On Docker, use --add-host=host.docker.internal:host-gateway, then under Settings → Connections, set http://host.docker.internal:11434 as the Ollama URL.
- AnythingLLM freezes during embedding
- Often a lack of RAM. Close other apps, split the import, or switch to a lighter embeddings model (nomic instead of bge-m3).
#Go further
Once your local RAG Ollama is in place, several logical guides for going deeper:
- Local RAG: introduction
- The conceptual guide explaining what happens under the hood (chunking, embeddings, retrieval)—useful for understanding why it works or doesn't.
- The best French embedding models
- Detailed comparison of BGE, E5, and Solon for French content—when to switch from nomic-embed-text.
- Open WebUI with Ollama: complete guide
- Going beyond RAG on Open WebUI: multi-user support, profiles, custom prompts, pipelines.
- Run an LLM without a GPU
- Detailed CPU-only guide with tokens/sec benchmarks by CPU—if you want to optimize your RAG setup without a graphics card.
Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.