AnythingLLM: production-ready RAG in local
AnythingLLM (Mintplex Labs) is an open-source RAG platform that deploys in minutes via Docker and turns a local Ollama backend into an enterprise document assistant. Whereas a “homegrown” RAG requires assembling LlamaIndex + Chroma + a UI, AnythingLLM delivers the entire stack: isolated workspaces, multi-user support, built-in agents, and a REST API. This AnythingLLM RAG tutorial covers the complete Docker installation, connecting to Ollama, creating workspaces, using agents, and exposing the API to your applications.
#Why AnythingLLM?
AnythingLLM occupies a specific niche in the local RAG ecosystem: more opinionated than Open WebUI for document handling, simpler than a custom LlamaIndex script, and more production-ready than a Streamlit demo. The project is open source (MIT) and maintained by Mintplex Labs, a team that has been committing continuously since 2023.
- Isolated workspaces
- Each workspace has its own document corpus, its own LLM, its own embeddings, and its own system prompt. We never mix legal RAG with customer-support RAG.
- Native multi-user support
- Authentication, roles (admin / manager / user), and workspace-based permissions. No need for a reverse proxy + basic auth as with a bare Python RAG.
- Interchangeable LLM backends
- Ollama, LM Studio, the llama.cpp server, vLLM, as well as cloud APIs (OpenAI, Anthropic, etc.). You can switch engines without touching the indexed documents.
- Built-in agents
- Web scraping, SQL execution, calculations, web search, document storage — callable with @agent in the chat. No need to layer LangChain on top.
- Native REST API
- An endpoint /api/v1/workspace/{slug}/chat lets you connect any application to a given workspace. Stable, documented response format.
#Prerequisites
Your documents, your AI: a reliable local RAG over your PDFs, notes and mail — nothing leaves your machine.
- Lifetime online access
- PDF + files
- Lifetime updates
- Docker
- Docker Desktop on Mac/Windows, or Docker Engine on Linux. Compose is not required—a docker run is enough to get started.
- Ollama installed and working
- The daemon must respond on http://localhost:11434. Check with ollama list. If Ollama is not installed yet, install it first—it is this tutorial’s default backend.
- An LLM model Ollama
- At minimum, an 8-9B model such as qwen3.5:9b (256k ctx, vision) or granite4.2:8b. For French RAG quality, target mistral-small (24B) or qwen3.8:27b if the VRAM allows it.
- An embeddings model
- nomic-embed-text (default) or bge-m3 for higher-quality multilingual use. Pull with ollama pull nomic-embed-text.
- RAM / VRAM
- 8 GB RAM minimum for the container; 16 GB is comfortable. For the GPU, it depends on the selected Ollama model (≈6–7 GB for a 9B Q4, ≈14 GB for a 24B Q4).
- 4 GB of disk space
- For the container, the internal SQLite database, and the vector database (LanceDB by default). Increase as needed based on the document volume.
#1. Docker installation
The official image is published on Docker Hub under mintplexlabs/anythingllm. Mintplex maintains stable tags (latest, render), and the image includes everything: Node.js, the API server, the frontend, LanceDB as the vector database, and the collection worker.
- 01Create a storage folderAnythingLLM persists everything (configuration, vectors, documents) in a volume. Create a dedicated folder on the host so nothing is lost when the image is updated.
- 02Start the containerThe command below mounts the storage directory, exposes port 3001, and enables SYS_ADMIN (required by some internal scrapers for PDF/web rendering).
- 03Open the UIOnce the container starts, the interface is available at http://localhost:3001. On first launch, a setup wizard guides you through the admin password and LLM backend.
#2. Connect Ollama as the backend
The first time you access http://localhost:3001, AnythingLLM launches an onboarding flow that asks for the LLM Provider, embeddings model, vector database, and admin account creation. You can also configure it later in Settings.
- 01LLM Provider → OllamaSelect Ollama from the list. Enter the base URL: http://host.docker.internal:11434 (Mac/Windows) or http://172.17.0.1:11434 (Linux default).
- 02Choose the chat modelThe dropdown lists your Ollama models. Choose the primary LLM (e.g., mistral-small (24B) for a good quality/VRAM compromise in French, or qwen3.5:9b if VRAM is more limited). Set the context window to 8192 or 16384 if the model supports it.
- 03Embedding Provider → OllamaUse the same backend for embeddings, or choose Native (built-in local model) if you want to avoid loading an embedder into Ollama. For AnythingLLM in French, nomic-embed-text gets the job done; bge-m3 (via Ollama) does better.
- 04Vector databaseLeave LanceDB as the default. Embedded, with no external dependencies, and performant up to several hundred thousand chunks. If you already manage Qdrant or Chroma elsewhere, you can configure them here.
#3. Workspaces, documents, and embeddings
A workspace is the fundamental unit of AnythingLLM. It contains a document corpus, an LLM, chat settings, and the associated conversations. You typically create one workspace per domain: Legal, Support, HR, Tech Watch.
- 01Create a workspaceLeft sidebar → New Workspace. Give it an explicit name (e.g., "contrats-2026"). The slug is generated automatically and will be used in the API URL.
- 02Upload documentsClick the upload icon in the workspace. AnythingLLM accepts PDF, DOCX, TXT, MD, CSV, EPUB, and much more. You can also point it to a web URL or a GitHub repo—an internal scraper retrieves the content.
- 03Move to Workspace + EmbedUploaded files first go into the Document Picker (staging area). Select the ones to index, then click Move to Workspace. AnythingLLM chunks them, computes embeddings via Ollama, and stores them in LanceDB.
- 04Configure the system promptWorkspace settings → Chat Settings → Prompt. This is where you define the role ("You are a legal assistant. Always cite the exact contract article"). The retrieval top-K (Document Similarity Threshold) is also configured here.
- Chunk size
- By default, 1000 characters with 20 characters of overlap. For legal contracts where every clause matters, reduce it to 500. For technical documentation with code blocks, increase it to 1500.
- Embedding model
- nomic-embed-text (768 dims) is fast but average for French. bge-m3 (1024 dims, multilingual) delivers 10–15% better accuracy on French content. mxbai-embed-large is a good middle ground.
- Chat mode vs. query
- Chat uses conversation history + RAG. Query is strict RAG: if nothing matches in the documents, the LLM refuses to answer. Query is the right setting for use cases where hallucinations are forbidden.
#4. Built-in agents
Beyond strict RAG, AnythingLLM includes an agent system: invoke @agent in the chat, and the LLM can then use skills (tools) to retrieve information outside the document repository. No need for LangChain or to write tool calling: it’s built in.
- web-browsing
- The agent opens a URL and reads the page (DOM rendering, not just raw HTML). Useful for making the assistant answer questions about information that is not in the RAG.
- web-scraping
- Variant: scrape a page and add it to the workspace as a document. Useful for enriching the corpus on the fly.
- save-document
- The agent generates a document (summary, synthesis) and saves it in the workspace. Useful for workflows like “read 10 articles → produce a brief.”
- sql-connector
- Connect a PostgreSQL/MySQL database and the agent can write and execute SQL queries to answer analytical questions. Obviously, pair this with a read-only SQL account.
- rag-memory
- Long-term memory across conversations. The agent can save facts it can retrieve in future sessions.
#5. Expose the API
To connect AnythingLLM to your applications (internal chatbot, Slack plugin, business integration), the REST API is the canonical interface. Each workspace becomes an endpoint scoped to its corpus.
- 01Generate an API keySettings → API Keys → Generate New API Key. Record the key; it is displayed only once. You can create several, for example one per client application, and revoke them individually.
- 02Identify the workspace slugIt is visible in the URL when you are in the workspace: .../workspace/contrats-2026 → slug = contrats-2026.
- 03Test with curlThe main endpoint is POST /api/v1/workspace/{slug}/chat. Authorization header: Bearer YOUR_KEY, JSON body with message and mode (chat or query).
#Troubleshooting
- "Could not reach Ollama at ..."
- Most common error. Check the URL: from the Docker container, localhost does not point to the host. Use host.docker.internal on Mac/Win, the bridge IP, or --network host on Linux.
- Very slow embedding
- The embedder runs on the CPU by default if you haven't pulled the model into Ollama. Force Ollama as the embedding provider and check ollama ps during indexing to see the GPU working.
- RAG fails to retrieve an obvious passage
- Three classic causes: chunks that are too large (reduce them from 1000 to 500 characters), an embeddings model that is weak in FR (switch to bge-m3), or a Document Similarity Threshold that is too strict in the workspace settings.
- Agent loops on tool call
- Model not capable enough. Move to glm-4.7-flash, qwen3.8:27b if possible, or mistral-small. Avoid very small models (2–3B) without tool-use fine-tuning for agents.
- Container killed after a few hours
- OOM kernel: Docker does not have enough allocated memory. In Docker Desktop, increase the RAM limit to 8–16 GB (Settings → Resources).
- Image update
- docker pull mintplexlabs/anythingllm:latest puis docker rm -f anythingllm et relancer le run avec les mêmes volumes. Les données dans $HOME/anythingllm sont conservées.
#Go further
Depending on the direction you want to take:
- Comparing AnythingLLM with no-code alternatives
- “Local RAG with Ollama without coding (Open WebUI, AnythingLLM)” goes head-to-head with Open WebUI on the same Ollama backend.
- Optimize French embeddings
- “The best French embedding models” compares bge-m3, Solon, and E5, and provides the right settings for AnythingLLM.
- Take the production stack further
- “Deploying an LLM in production with Docker Compose” shows how to stack AnythingLLM with a Traefik reverse proxy, an external Qdrant instance, and automated backups.
Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.