Local RAG with LM Studio: chat with your documents
LM Studio has included a “Chat with Documents” feature since version 0.3 that runs a complete RAG pipeline locally: indexing, embeddings, retrieval, and generation, without a single byte leaving your machine. It’s probably the fastest way to move from a ChatGPT-like chat to an assistant that answers from your PDFs, contracts, or notes. This guide shows how to enable it, what it can do, where it hits its limits, and when to switch to AnythingLLM or a Python stack.
#Why LM Studio for RAG
Local RAG typically means stacking four components: a file parser, an embedding model, a vector database, and a generative LLM. Most tutorials use Python + LlamaIndex + ChromaDB + Ollama. It's powerful, but that's also four dependencies, an environment to manage, and a script to maintain.
LM Studio hides all of that behind a paperclip. You load a generative model (Qwen 3.5 9B, Granite 4.2 8B, Gemma 4 12B…), load an embeddings model, drop a PDF into the conversation, and ask your question. The interface handles chunking, in-memory indexing, and injecting the relevant passages into the prompt.
- Not a line of code
- Everything goes through the GUI. It's the shortest path between “I have a folder of PDFs” and “I can chat with it.”
- 100% offline
- Embeddings, retrieval, generation: everything runs on your GPU or CPU. No telemetry on document content.
- OpenAI-compatible
- The local server on port 1234 remains accessible. You can use RAG through the GUI and connect a script to the same model in parallel.
#Prerequisites
LM Studio queries your documents. To move from a trial to a reliable tool, the Local RAG toolkit covers what makes the difference: document splitting (ch. 7), choosing a strong French embedding model (ch. 8), and evaluating responses (ch. 14).
- Lifetime online access
- PDF + files
- Lifetime updates
- LM Studio 0.3 or newer
- The Chat with Documents feature appeared in 0.3 and has been enhanced since. Check your version under Settings → About. On an older LM Studio, update before going any further.
- A loaded generative LLM
- Qwen 3.5 9B Q4_K_M (6.6 GB, 256k context), Granite 4.2 8B Q4_K_M (5.3 GB, 128k), or Gemma 4 12B Q4_K_M (7.6 GB, multimodal) are good defaults, all under Apache 2.0. These recent models have large context windows, but still set the Context Length on the LM Studio side (see below)—otherwise you will be limited in how many passages LM Studio can inject.
- An embeddings model
- Required and separate from the LLM. nomic-embed-text-v1.5 (137M, ~80 MB in Q4) is the default recommended by LM Studio. mxbai-embed-large (335M) if you have room to spare. For exclusively French content, multilingual-e5-large is more relevant.
- VRAM or RAM
- Allow for the LLM size + ~200 MB for the embeddings model + 1-2 GB for the in-memory index of an average folder. On a 12 GB VRAM GPU, a 7B Q4 + nomic-embed leaves about 5 GB for the index and context.
- Context window ≥ 8192
- Set the LLM's Context Length to at least 8192, ideally 16384, from the chat's right panel. Below that, LM Studio truncates the retrieved passages and the response loses quality.
#1. Enable Chat with Documents
There is nothing to enable in the strict sense—it is a native chat feature. Simply provide a document in a conversation and LM Studio launches the RAG pipeline in the background.
- 01Open a new chatOpen the Chat tab in the sidebar, then click the + button at the top to create a conversation. Make sure an LLM is loaded in the top dropdown menu. If nothing is loaded, select your generation model and wait for the VRAM to fill.
- 02Drag your file into the input areaDrag and drop a PDF, DOCX, TXT, or MD directly into the message field at the bottom. A thumbnail appears above the field with the file name and size. You can attach several in succession.
- 03Ask your questionType your question normally, as you would in a standard chat. LM Studio detects the document, chunks it, encodes it, searches for relevant passages, and injects them into the prompt — all in 1–5 seconds depending on the size.
- 04Read the answer and sourcesThe model answers based on the passages it retrieves. Depending on the LLM, it may or may not cite the excerpts. To force citation, add “Cite the exact passages from the document” to your question.
#2. The embedding model
The embedding model is what turns a piece of text into a vector of numbers. Retrieval quality—and therefore the final quality of RAG—depends directly on this model, much more than on the generation LLM, contrary to what people initially assume.
- 01Download an embeddings modelIn the Discover tab, filter by “Text Embedding” in the left column. nomic-embed-text-v1.5 is prominently listed—it’s the sensible default. Download the Q4_K_M variant (~80 MB) or Q8_0 (~140 MB) if you insist on maximum quality.
- 02Verify that it is detectedSettings → My Models → Embeddings tab. The model should appear there. If LM Studio can’t see it in this tab even though it has been downloaded, it wasn’t recognized as an embeddings model—check the Hugging Face tags or download it again through Discover.
- 03Select it in the chatWhen a document is attached to a chat, LM Studio displays the name of the embeddings model in use at the bottom of the conversation. Click it to change it. The choice is persisted per chat—handy for comparing nomic and multilingual-e5 on the same document.
- nomic-embed-text-v1.5
- 137M parameters, 768 dimensions, 8192-token context. Excellent English default, adequate French. Recommended by LM Studio.
- mxbai-embed-large-v1
- 335M, 1024 dimensions. Best score on the English MTEB, but 3× slower and 3× heavier. Relevant if retrieval quality is the bottleneck.
- multilingual-e5-large
- 560M, 1024 dimensions. The best choice if your documents are in French or multilingual. Requires prefixing queries with « query: » and passages with « passage: », which LM Studio handles automatically.
- bge-large-en-v1.5
- 335M, 1024 dimensions. Excellent in pure English, avoid for French.
#3. Supported file formats and sizes
LM Studio handles a pragmatic subset of common formats. No OCR for scanned PDFs, no parsing of complex tables, and no direct import from a URL—you need to know these limitations before expecting too much.
- PDF (native text)
- Supported. PDFs generated from Word, Google Docs, or LaTeX work well. Scanned PDFs without an OCR layer can’t be extracted—run them through ocrmypdf or Tesseract first.
- DOCX
- Supported. Formatting is ignored, and the text is extracted as plain text. Tables and images are lost.
- TXT and MD
- Natively supported, it is the ideal format. Markdown preserves its structure (headings, lists), which helps with chunking.
- Source code (.py, .js, .ts…)
- Treated as text. Works, but to chat with an entire repo, prefer Continue.dev or a tool designed for coding.
- CSV, XLSX, JSON, HTML, EPUB
- No official support at this time. Convert to TXT or MD with pandoc, csvkit, or a script before importing.
- Maximum file size
- No strict limit has been announced. In practice, allow 50-100 MB per PDF maximum on the GUI side—beyond that, chunking becomes slow and RAM usage rises quickly.
#4. Test with a real document
The best test is a document you know. Take a PDF of about twenty pages—a user manual, report, or contract—and first ask questions whose answers you know to calibrate your confidence in the system.
- Low temperature
- For RAG, lower the temperature to 0.1–0.3 in the right-hand panel. The goal is to stick to the passages, not create.
- Explicit system prompt
- « You answer only from the provided excerpts. If the information is not present, say so clearly. Quote the exact passages in quotation marks. » This simple addition cuts hallucinations by 3.
- Structured question
- “What is the notice period? Quote the exact clause.” is more useful than “tell me about the notice period.” The quotation forces the model to stay grounded.
#Limitations vs AnythingLLM
LM Studio provides “throwaway” RAG — per conversation, with no persistence between chats and no fine-tuning. That’s intentional: the feature is designed for occasional lookups, not a knowledge base. As soon as you want more, AnythingLLM (free, open source, with a similar GUI approach) takes over.
- Persistent workspaces
- AnythingLLM stores your documents in reusable workspaces. You index 200 PDFs once, and all your workspace chats can access them. LM Studio reindexes with every new chat.
- Clickable citations
- AnythingLLM displays the sources with a link to the exact passage. LM Studio merely injects the passage into the prompt—the model may or may not cite it, as it sees fit.
- Exposed RAG settings
- Top K, chunk size, similarity threshold, embedding model: everything is configurable in AnythingLLM. LM Studio handles it as a black box with reasonable defaults.
- External vector databases
- AnythingLLM connects to ChromaDB, Qdrant, Pinecone, and Weaviate. LM Studio keeps everything in memory within the process, limiting it to modest corpora.
- Multi-utilisateurs
- AnythingLLM supports users and workspace sharing. LM Studio is single-user by design.
- Source connectors
- AnythingLLM can ingest from URLs, Confluence, GitHub, and YouTube transcripts. LM Studio is limited to dragging and dropping files.
#When to move to a Python stack
At some point, the GUI hits its ceiling and coding becomes simpler than working around it. Signs that you should move to Python + LlamaIndex (or Haystack):
- Structure-aware chunking
- Your documents have a strong structure (legal sections, code, scientific articles) that naïve chunking breaks. A Markdown-aware splitter or a docling parser makes a difference.
- Hybrid search
- You want to combine vector search (semantic) and BM25 (lexical) so you don't miss exact terms (clause numbers, variable names, identifiers). No GUI does this natively today.
- Reranking
- You add a cross-encoder (BAAI/bge-reranker-v2-m3) to rerank the top 20 results. +15% relevance, but it requires coding.
- Metadata and filtering
- “Search only within contracts signed in 2024.” Requires metadata per chunk and a filter at query time.
- Automated ingestion pipeline
- A watched folder, a cron job, a webhook that reindexes whenever a new file arrives. At that point, you’ve definitively left the GUI behind.
- Evaluation
- Measure retrieval quality on 50 reference questions. Without Python, you're shooting blind.
Good news: LM Studio can stay in front while Python runs the pipeline. The OpenAI-compatible server on localhost:1234 can be consumed by LlamaIndex in two lines—you keep LM Studio for exploratory chat and code the pipeline behind it.
#Go further
Chat with Documents covers 80% of the needs of personal use or a small team. The logical next steps after this first RAG:
- Understanding the mechanisms
- The local RAG guide: the introduction covers chunking, embeddings, retrieval, and the pitfalls no GUI reveals. Useful reading before changing the settings.
- Choose a better French embedding model
- The guide The Best Embedding Models in French compares Solon, multilingual-e5, and BGE on French-language content. The difference is measurable.
- Switch to API server
- The Transformer LM Studio as an API server guide shows how to expose the OpenAI-compatible endpoint and connect LlamaIndex, Continue.dev, or a custom script without losing your current setup.
Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.