Intermediate 17 minRAG

AnythingLLM: production-ready RAG in local

AnythingLLM (Mintplex Labs) is an open-source RAG platform that deploys in minutes via Docker and turns a local Ollama backend into an enterprise document assistant. Whereas a “homegrown” RAG requires assembling LlamaIndex + Chroma + a UI, AnythingLLM delivers the entire stack: isolated workspaces, multi-user support, built-in agents, and a REST API. This AnythingLLM RAG tutorial covers the complete Docker installation, connecting to Ollama, creating workspaces, using agents, and exposing the API to your applications.

By Mohamed Meguedmi·Update 2026-08-27·Tested on Windows, macOS, and Linux

#Why AnythingLLM?

AnythingLLM occupies a specific niche in the local RAG ecosystem: more opinionated than Open WebUI for document handling, simpler than a custom LlamaIndex script, and more production-ready than a Streamlit demo. The project is open source (MIT) and maintained by Mintplex Labs, a team that has been committing continuously since 2023.

Isolated workspaces
Each workspace has its own document corpus, its own LLM, its own embeddings, and its own system prompt. We never mix legal RAG with customer-support RAG.
Native multi-user support
Authentication, roles (admin / manager / user), and workspace-based permissions. No need for a reverse proxy + basic auth as with a bare Python RAG.
Interchangeable LLM backends
Ollama, LM Studio, the llama.cpp server, vLLM, as well as cloud APIs (OpenAI, Anthropic, etc.). You can switch engines without touching the indexed documents.
Built-in agents
Web scraping, SQL execution, calculations, web search, document storage — callable with @agent in the chat. No need to layer LangChain on top.
Native REST API
An endpoint /api/v1/workspace/{slug}/chat lets you connect any application to a given workspace. Stable, documented response format.
i
When AnythingLLM is the right choice
You want a team RAG system with authentication for 100 to 10,000 documents, without writing a Python pipeline. For an ultra-simple single-user RAG, Msty or Open WebUI are enough. For tens of thousands of documents with custom hybrid search, a Qdrant + LlamaIndex stack remains more flexible.

#Prerequisites

The Local RAG Kit

Your documents, your AI: a reliable local RAG over your PDFs, notes and mail — nothing leaves your machine.

  • Lifetime online access
  • PDF + files
  • Lifetime updates
Docker
Docker Desktop on Mac/Windows, or Docker Engine on Linux. Compose is not required—a docker run is enough to get started.
Ollama installed and working
The daemon must respond on http://localhost:11434. Check with ollama list. If Ollama is not installed yet, install it first—it is this tutorial’s default backend.
An LLM model Ollama
At minimum, an 8-9B model such as qwen3.5:9b (256k ctx, vision) or granite4.2:8b. For French RAG quality, target mistral-small (24B) or qwen3.8:27b if the VRAM allows it.
An embeddings model
nomic-embed-text (default) or bge-m3 for higher-quality multilingual use. Pull with ollama pull nomic-embed-text.
RAM / VRAM
8 GB RAM minimum for the container; 16 GB is comfortable. For the GPU, it depends on the selected Ollama model (≈6–7 GB for a 9B Q4, ≈14 GB for a 24B Q4).
4 GB of disk space
For the container, the internal SQLite database, and the vector database (LanceDB by default). Increase as needed based on the document volume.
→
Test Ollama first
Before launching AnythingLLM, confirm that Ollama responds: curl http://localhost:11434/api/tags should list your models. If the command fails, AnythingLLM won't be able to connect to it, and 80% of "AnythingLLM doesn't work" problems come from this.

#1. Docker installation

The official image is published on Docker Hub under mintplexlabs/anythingllm. Mintplex maintains stable tags (latest, render), and the image includes everything: Node.js, the API server, the frontend, LanceDB as the vector database, and the collection worker.

  1. 01
    Create a storage folder
    AnythingLLM persists everything (configuration, vectors, documents) in a volume. Create a dedicated folder on the host so nothing is lost when the image is updated.
  2. 02
    Start the container
    The command below mounts the storage directory, exposes port 3001, and enables SYS_ADMIN (required by some internal scrapers for PDF/web rendering).
  3. 03
    Open the UI
    Once the container starts, the interface is available at http://localhost:3001. On first launch, a setup wizard guides you through the admin password and LLM backend.
Docker launch (Linux/Mac)
mkdir -p $HOME/anythingllm
touch $HOME/anythingllm/.env

docker run -d -p 3001:3001 \
  --cap-add SYS_ADMIN \
  -v $HOME/anythingllm:/app/server/storage \
  -v $HOME/anythingllm/.env:/app/server/.env \
  -e STORAGE_DIR="/app/server/storage" \
  --name anythingllm \
  --restart unless-stopped \
  mintplexlabs/anythingllm:latest
!
Docker network and Ollama
On Mac and Windows, Ollama runs on the host, but the container is isolated. AnythingLLM must use http://host.docker.internal:11434 to talk to Ollama (not localhost). On Linux, add --add-host=host.docker.internal:host-gateway to docker run, or use the IP address of the docker0 bridge (often 172.17.0.1).
Container verification
docker logs -f anythingllm

# À l'écran : "Primary server in HTTP mode listening on port 3001"
# puis : "Collector hot directory found and ready"

#2. Connect Ollama as the backend

The first time you access http://localhost:3001, AnythingLLM launches an onboarding flow that asks for the LLM Provider, embeddings model, vector database, and admin account creation. You can also configure it later in Settings.

  1. 01
    LLM Provider → Ollama
    Select Ollama from the list. Enter the base URL: http://host.docker.internal:11434 (Mac/Windows) or http://172.17.0.1:11434 (Linux default).
  2. 02
    Choose the chat model
    The dropdown lists your Ollama models. Choose the primary LLM (e.g., mistral-small (24B) for a good quality/VRAM compromise in French, or qwen3.5:9b if VRAM is more limited). Set the context window to 8192 or 16384 if the model supports it.
  3. 03
    Embedding Provider → Ollama
    Use the same backend for embeddings, or choose Native (built-in local model) if you want to avoid loading an embedder into Ollama. For AnythingLLM in French, nomic-embed-text gets the job done; bge-m3 (via Ollama) does better.
  4. 04
    Vector database
    Leave LanceDB as the default. Embedded, with no external dependencies, and performant up to several hundred thousand chunks. If you already manage Qdrant or Chroma elsewhere, you can configure them here.
URL Ollama by OS
Mac / Windows : http://host.docker.internal:11434
Linux (bridge)  : http://172.17.0.1:11434
Linux (--network host) : http://localhost:11434
→
Verify that the connection works
In Settings → LLM Preference, the "Save changes" button triggers a test call to Ollama. A "Could not reach" error almost always points to the host.docker.internal vs localhost URL. Fix it and test before going any further.

#3. Workspaces, documents, and embeddings

A workspace is the fundamental unit of AnythingLLM. It contains a document corpus, an LLM, chat settings, and the associated conversations. You typically create one workspace per domain: Legal, Support, HR, Tech Watch.

  1. 01
    Create a workspace
    Left sidebar → New Workspace. Give it an explicit name (e.g., "contrats-2026"). The slug is generated automatically and will be used in the API URL.
  2. 02
    Upload documents
    Click the upload icon in the workspace. AnythingLLM accepts PDF, DOCX, TXT, MD, CSV, EPUB, and much more. You can also point it to a web URL or a GitHub repo—an internal scraper retrieves the content.
  3. 03
    Move to Workspace + Embed
    Uploaded files first go into the Document Picker (staging area). Select the ones to index, then click Move to Workspace. AnythingLLM chunks them, computes embeddings via Ollama, and stores them in LanceDB.
  4. 04
    Configure the system prompt
    Workspace settings → Chat Settings → Prompt. This is where you define the role ("You are a legal assistant. Always cite the exact contract article"). The retrieval top-K (Document Similarity Threshold) is also configured here.
Chunk size
By default, 1000 characters with 20 characters of overlap. For legal contracts where every clause matters, reduce it to 500. For technical documentation with code blocks, increase it to 1500.
Embedding model
nomic-embed-text (768 dims) is fast but average for French. bge-m3 (1024 dims, multilingual) delivers 10–15% better accuracy on French content. mxbai-embed-large is a good middle ground.
Chat mode vs. query
Chat uses conversation history + RAG. Query is strict RAG: if nothing matches in the documents, the LLM refuses to answer. Query is the right setting for use cases where hallucinations are forbidden.
i
Pin Document
A document can be pinned in the workspace (pin icon). Its full contents are then injected into every prompt in addition to standard RAG retrieval. Ideal for a business glossary or a policy that must always be in context.
Pull embedding models via Ollama
# Modèle par défaut, multilingue correct
ollama pull nomic-embed-text

# Meilleur pour le français, 1024 dimensions
ollama pull bge-m3

# Vérifier qu'ils tournent
ollama list | grep embed

#4. Built-in agents

Beyond strict RAG, AnythingLLM includes an agent system: invoke @agent in the chat, and the LLM can then use skills (tools) to retrieve information outside the document repository. No need for LangChain or to write tool calling: it’s built in.

web-browsing
The agent opens a URL and reads the page (DOM rendering, not just raw HTML). Useful for making the assistant answer questions about information that is not in the RAG.
web-scraping
Variant: scrape a page and add it to the workspace as a document. Useful for enriching the corpus on the fly.
save-document
The agent generates a document (summary, synthesis) and saves it in the workspace. Useful for workflows like “read 10 articles → produce a brief.”
sql-connector
Connect a PostgreSQL/MySQL database and the agent can write and execute SQL queries to answer analytical questions. Obviously, pair this with a read-only SQL account.
rag-memory
Long-term memory across conversations. The agent can save facts it can retrieve in future sessions.
Invoking an agent in the chat
@agent va sur https://blog.example.com/rapport-2026 et fais-moi
un résumé en 5 points des chiffres-clés.

@agent connecte-toi à la base postgres-prod et donne-moi le top 10
des clients par chiffre d'affaires sur le trimestre.

@agent enregistre la conversation précédente sous forme de note
dans ce workspace, titre : "Synthèse veille IA juin 2026".
!
Model powerful enough for tool calling
Agents require an LLM capable of clean function calling. Locally: glm-4.7-flash (MoE 30B-A3B, very good for agents) or qwen3.8:27b, with mistral-small as a solid alternative and qwen3.5:9b as the minimum. The smallest models (2-3B) without tool-use fine-tuning hallucinate tool calls. If the agent loops or misses its calls, this is almost always the problem.

#5. Expose the API

To connect AnythingLLM to your applications (internal chatbot, Slack plugin, business integration), the REST API is the canonical interface. Each workspace becomes an endpoint scoped to its corpus.

  1. 01
    Generate an API key
    Settings → API Keys → Generate New API Key. Record the key; it is displayed only once. You can create several, for example one per client application, and revoke them individually.
  2. 02
    Identify the workspace slug
    It is visible in the URL when you are in the workspace: .../workspace/contrats-2026 → slug = contrats-2026.
  3. 03
    Test with curl
    The main endpoint is POST /api/v1/workspace/{slug}/chat. Authorization header: Bearer YOUR_KEY, JSON body with message and mode (chat or query).
API call with curl
curl -X POST http://localhost:3001/api/v1/workspace/contrats-2026/chat \
  -H "Authorization: Bearer VOTRE_CLE_API" \
  -H "Content-Type: application/json" \
  -d '{
    "message": "Quelle est la durée de préavis dans le contrat ACME ?",
    "mode": "query"
  }'
Python client
import requests

API_KEY   = "votre-cle-api"
WORKSPACE = "contrats-2026"
BASE_URL  = "http://localhost:3001"

def ask(question: str, mode: str = "chat") -> dict:
    response = requests.post(
        f"{BASE_URL}/api/v1/workspace/{WORKSPACE}/chat",
        headers={
            "Authorization": f"Bearer {API_KEY}",
            "Content-Type": "application/json",
        },
        json={"message": question, "mode": mode},
        timeout=120,
    )
    response.raise_for_status()
    return response.json()

result = ask("Résume la clause 4 du contrat ACME signé en mars.")
print(result["textResponse"])
for source in result.get("sources", []):
    print(" -", source["title"])
→
Streaming and advanced endpoints
The API also supports /chat/stream (SSE) for token-by-token streaming, /thread/new for managing multi-turn conversations server-side, and /documents for automating indexing. The full documentation is in Settings → API → Open API Docs (embedded Swagger UI).

#Troubleshooting

"Could not reach Ollama at ..."
Most common error. Check the URL: from the Docker container, localhost does not point to the host. Use host.docker.internal on Mac/Win, the bridge IP, or --network host on Linux.
Very slow embedding
The embedder runs on the CPU by default if you haven't pulled the model into Ollama. Force Ollama as the embedding provider and check ollama ps during indexing to see the GPU working.
RAG fails to retrieve an obvious passage
Three classic causes: chunks that are too large (reduce them from 1000 to 500 characters), an embeddings model that is weak in FR (switch to bge-m3), or a Document Similarity Threshold that is too strict in the workspace settings.
Agent loops on tool call
Model not capable enough. Move to glm-4.7-flash, qwen3.8:27b if possible, or mistral-small. Avoid very small models (2–3B) without tool-use fine-tuning for agents.
Container killed after a few hours
OOM kernel: Docker does not have enough allocated memory. In Docker Desktop, increase the RAM limit to 8–16 GB (Settings → Resources).
Image update
docker pull mintplexlabs/anythingllm:latest puis docker rm -f anythingllm et relancer le run avec les mêmes volumes. Les données dans $HOME/anythingllm sont conservées.

#Go further

Depending on the direction you want to take:

Comparing AnythingLLM with no-code alternatives
“Local RAG with Ollama without coding (Open WebUI, AnythingLLM)” goes head-to-head with Open WebUI on the same Ollama backend.
Optimize French embeddings
“The best French embedding models” compares bge-m3, Solon, and E5, and provides the right settings for AnythingLLM.
Take the production stack further
“Deploying an LLM in production with Docker Compose” shows how to stack AnythingLLM with a Traefik reverse proxy, an external Qdrant instance, and automated backups.
Did this guide help you?

Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.