Beginner 10 minRAG

Local RAG with Ollama without coding (Open WebUI, AnythingLLM)

Doing local RAG with Ollama means asking a self-hosted LLM questions while having it read your own documents—without sending a single line to the cloud. Two free tools let you set it up without coding: Open WebUI (an ChatGPT-like interface with an integrated knowledge base) and AnythingLLM (workspaces + citations). This guide covers the whole process in 10 minutes, including on a small machine without a GPU.

By Mohamed Meguedmi·Update 2026-08-27·Tested on Windows, macOS, and Linux

#Why local RAG with Ollama

A RAG (Retrieval-Augmented Generation) addresses a simple limitation of LLMs: they only know their training data. RAG connects them to your documents: internal procedure PDFs, Markdown notes, email exports, contracts, manuals — anything that is neither public nor recent. The model reads the relevant passages before answering and cites its sources if the tool supports it.

The cloud version (ChatGPT, Claude, Gemini) requires you to upload your documents to third-party servers. For a law firm, doctor, accountant, or even an individual who does not want personal notes sent to the United States, that is a deal-breaker. A local RAG Ollama solves this problem: everything stays on your machine, and you retain control over both costs (zero) and privacy.

Total privacy
Your documents never leave the machine. No token is sent to a provider.
Zero marginal cost
No subscription, no API quota. You pay for electricity, period.
Offline
Once the model is downloaded, RAG works without an internet connection.
Customizable
You choose the response model, the embeddings model, and the sources.

#Does Ollama support RAG natively?

The Local RAG Kit

Your AI reads your documents without a line of code. The real question remains: does it answer correctly? The Local RAG Kit starts there (ch. 1), takes Open WebUI and AnythingLLM into advanced use (ch. 13), and teaches you to measure false answers and fabricated citations (ch. 14).

  • Lifetime online access
  • PDF + files
  • Lifetime updates

Clear answer: no. Ollama is an inference engine—it downloads and runs LLMs and exposes an OpenAI-compatible API at http://localhost:11434, but it cannot index your documents by itself. It has no vector database, PDF splitter, or interface for dragging and dropping files.

However, it can serve a generation model and an embeddings model in parallel—that's all a third-party tool (Open WebUI, AnythingLLM, LangChain…) needs to build a RAG layer on top. So when we talk about local RAG Ollama, we're always talking about a combination of Ollama + RAG interface.

i
The right mental model
Ollama = the engine (two services: response LLM + embeddings LLM). Open WebUI or AnythingLLM = the pipeline: document ingestion, chunking, embedding computation, vector storage, search, and prompt injection. No code to write.

#Prerequisites

Ollama installed
On Windows, macOS, or Linux. It listens on http://localhost:11434 by default.
A response model
Granite 4.2 8B or Qwen 3.5 9B in Q4_K_M. Allow for ~5 to 7 GB of VRAM or RAM.
An embeddings model
nomic-embed-text or mxbai-embed-large. ~300 MB. Essential for vectorizing documents.
Docker (option Open WebUI)
Docker Desktop on Windows/macOS, or docker.io on Linux. This is the recommended installation method.
Disk space
10 GB minimum for the LLM model + embeddings + your document vector index.
Prepare the two Ollama models
# Le LLM qui va répondre
ollama pull qwen3.5:9b

# Le modèle d'embeddings (obligatoire pour le RAG)
ollama pull nomic-embed-text

# Vérifier que les deux sont bien là
ollama list

#1. Open WebUI: integrated knowledge base

Open WebUI (formerly Ollama WebUI) is the most popular ChatGPT-like interface for Ollama. Its Knowledge feature lets you upload documents and query them with RAG, without any code-side configuration.

  1. 01
    Run Open WebUI in Docker
    In a terminal: docker run -d -p 3000:8080 --add-host=host.docker.internal:host-gateway -v open-webui:/app/backend/data --name open-webui --restart always ghcr.io/open-webui/open-webui:main. On native Linux, replace host.docker.internal with localhost if you run Ollama outside Docker.
  2. 02
    Open the interface
    Go to http://localhost:3000. Create an admin account (the first account is admin by default). Open WebUI automatically detects Ollama on localhost:11434.
  3. 03
    Create a knowledge base
    Go to Workspace → Knowledge → + Create a Knowledge Base. Give it a name (e.g., 'Client contracts'). This is your vector container.
  4. 04
    Import your documents
    Click the + at the top right of the database, then upload your PDFs, .docx, .md, and .txt files. Open WebUI splits them and calculates the embeddings in the background.
  5. 05
    Configure the embedding model
    Settings (admin) → Documents → Embedding Model Engine = Ollama, Embedding Model = nomic-embed-text. Save. Without this, Open WebUI uses a default model that is less suitable for French.
  6. 06
    Query the database
    In a new chat, type # followed by the knowledge base name to attach it. Ask your question—the answer is generated from the relevant chunks, with clickable citation numbers.
→
Quick attachment tip
If you only have one or two PDFs to query without creating a database, you can also drag and drop a file directly into the chat area. Open WebUI switches to RAG mode for that document only, for the duration of the conversation.

#2. AnythingLLM: workspaces and citations

AnythingLLM is the most fully developed alternative to Open WebUI for RAG. While Open WebUI is general-purpose, AnythingLLM is built around the concept of workspaces: each project has its own document database, model, and history. Ideal when you want to separate “Client Contracts” from “Personal Notes” from “Technical Documentation.”

  1. 01
    Install AnythingLLM
    Download the Desktop installer from useanything.com for Windows, macOS, or Linux. It's an Electron application—Docker is not required for the desktop version.
  2. 02
    Choose Ollama as your LLM provider
    On first launch, the assistant asks which LLM to use. Select Ollama, enter the URL http://localhost:11434, and choose qwen3.5:9b from the list that loads automatically.
  3. 03
    Choose the embedding engine
    Next step: Embedder. Select Ollama, nomic-embed-text model. AnythingLLM can also use its built-in local embedder if you haven’t pulled nomic-embed-text—less capable, but zero configuration.
  4. 04
    Create a workspace
    In the sidebar: + New Workspace. Give it a name. Click it, then click the upload icon to drag in your PDFs, .docx, .csv, .epub, URL, or even YouTube videos (automatic transcription).
  5. 05
    Move documents → workspace
    AnythingLLM separates ingestion (documents listed on the left) from usage (workspace on the right). Select the documents, click 'Move to Workspace,' then 'Save and Embed.' This is where embeddings are calculated via Ollama.
  6. 06
    Discuss with citations
    In the workspace, ask your question. Each answer contains an expandable Citations panel showing exactly which chunks were used. Very useful for verifying that the model did not hallucinate.
Open WebUI
Best if you want a general-purpose, multi-user (team) ChatGPT-like interface with web access. RAG is one feature among others.
AnythingLLM
Best if RAG is the primary use case: strict workspaces, well-formed citations, support for more formats (URL, YouTube, GitHub), built-in agents.

#3. Which embedding model to choose

The embedding model is the part many people overlook. It transforms your documents into vectors, and its quality directly determines the relevance of the retrieved passages. For French documents, some models are significantly better than others.

nomic-embed-text (137M, 274 MB)
The recommended default. Fast, reasonably good multilingual performance, 8192-token context. A very good starting compromise. Available directly: ollama pull nomic-embed-text.
mxbai-embed-large (335M, 670 MB)
More accurate than nomic, slightly slower. Primarily English, but handles French for standard documents. Prefer it when quality matters more than speed. ollama pull mxbai-embed-large.
bge-m3 (567M, 1.2 GB)
Excellent for multilingual use (100+ languages, including native French), with an 8192-token context. Heavier. Available on Hugging Face, and can be imported into Ollama via a Modelfile.
snowflake-arctic-embed
Good multilingual support, lighter than bge-m3. A good alternative if bge-m3 is too heavy.
→
Practical rule for French
Start with nomic-embed-text. If your RAG results are poor on technical French PDFs (legal, medical), move to bge-m3 or a dedicated model such as Solon. Changing the embedding model requires reindexing the entire database—keep that in mind before filling it with 5,000 documents.

#4. Local RAG on 4 GB of RAM, without a GPU

Yes, it's feasible. RAG is less resource-intensive than you might think: most of the work (embeddings) happens once during ingestion, and at runtime, it's just the LLM answering with a few thousand additional context tokens. On a small machine, the main trade-off is the answering LLM.

4 GB of RAM, slow CPU (old laptop)
LLM: qwen3.5:2b or granite4.2:3b in Q4_K_M (~2 GB). Embeddings: nomic-embed-text. Slow but readable responses (3–5 tok/s). The LLM context is set to 2048 to stay below the RAM threshold.
8 GB of RAM, recent CPU (i5/Ryzen 5)
LLM: qwen3.5:4b Q4 (~3.4 GB) or gemma4:e2b-it-qat. Embeddings: nomic-embed-text. ~6-10 tok/s. This is the 'pretty good' setup without a GPU.
16 GB of RAM, no GPU
LLM: granite4.2:8b or qwen3.5:9b Q4 (~5-7 GB). Embeddings: nomic or mxbai. ~4-8 tok/s on pure CPU on a recent Ryzen 7 or i7.
Minimal RAM-friendly setup
# Petit LLM rapide, embeddings standards
ollama pull qwen3.5:2b
ollama pull nomic-embed-text

# Réduire le contexte pour économiser la RAM
# Dans Open WebUI : Settings → Models → qwen3.5 → num_ctx = 2048
# Dans AnythingLLM : Workspace settings → Max Tokens = 2048

# Vérifier l'usage RAM en pleine question :
ollama ps
!
Indexing = peak load
Ingesting large documents (hundreds of pages) is the heaviest step—each chunk has to pass through the embedding model. On a small machine, split the workload: import documents in batches of 10–20, and let indexing finish before continuing. Otherwise, Ollama may saturate RAM and crash halfway through.

#Troubleshooting

The model “can’t see” my documents
Verify that the knowledge base is attached to the chat (Open WebUI: # command, AnythingLLM: workspace icon). Also verify that the documents are actually 'embedded', not merely uploaded.
Very slow responses on small models
RAG adds 1000 to 3000 tokens to the context. Reduce top-k (3–5 chunks instead of 10) in the RAG settings, and lower the LLM's num_ctx to 2048 or 4096.
Hallucinations despite the documents
Increase top-k, and verify that the embedding model is the one that indexed the data (changing models invalidates the database). Enable citations to see what the LLM actually read.
Open WebUI cannot see Ollama
On Docker, use --add-host=host.docker.internal:host-gateway, then under Settings → Connections, set http://host.docker.internal:11434 as the Ollama URL.
AnythingLLM freezes during embedding
Often a lack of RAM. Close other apps, split the import, or switch to a lighter embeddings model (nomic instead of bge-m3).

#Go further

Once your local RAG Ollama is in place, several logical guides for going deeper:

Local RAG: introduction
The conceptual guide explaining what happens under the hood (chunking, embeddings, retrieval)—useful for understanding why it works or doesn't.
The best French embedding models
Detailed comparison of BGE, E5, and Solon for French content—when to switch from nomic-embed-text.
Open WebUI with Ollama: complete guide
Going beyond RAG on Open WebUI: multi-user support, profiles, custom prompts, pipelines.
Run an LLM without a GPU
Detailed CPU-only guide with tokens/sec benchmarks by CPU—if you want to optimize your RAG setup without a graphics card.
Did this guide help you?

Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.