Intermediate 11 minRAG

Local RAG with LM Studio: chat with your documents

LM Studio has included a “Chat with Documents” feature since version 0.3 that runs a complete RAG pipeline locally: indexing, embeddings, retrieval, and generation, without a single byte leaving your machine. It’s probably the fastest way to move from a ChatGPT-like chat to an assistant that answers from your PDFs, contracts, or notes. This guide shows how to enable it, what it can do, where it hits its limits, and when to switch to AnythingLLM or a Python stack.

By Mohamed Meguedmi·Update 2026-08-27·Tested on Windows, macOS, and Linux

#Why LM Studio for RAG

Local RAG typically means stacking four components: a file parser, an embedding model, a vector database, and a generative LLM. Most tutorials use Python + LlamaIndex + ChromaDB + Ollama. It's powerful, but that's also four dependencies, an environment to manage, and a script to maintain.

LM Studio hides all of that behind a paperclip. You load a generative model (Qwen 3.5 9B, Granite 4.2 8B, Gemma 4 12B…), load an embeddings model, drop a PDF into the conversation, and ask your question. The interface handles chunking, in-memory indexing, and injecting the relevant passages into the prompt.

Not a line of code
Everything goes through the GUI. It's the shortest path between “I have a folder of PDFs” and “I can chat with it.”
100% offline
Embeddings, retrieval, generation: everything runs on your GPU or CPU. No telemetry on document content.
OpenAI-compatible
The local server on port 1234 remains accessible. You can use RAG through the GUI and connect a script to the same model in parallel.
i
It is not a replacement for AnythingLLM
LM Studio provides “conversation-based” RAG—you drop documents into a given chat, and the index lives with that chat. There is no concept of a persistent workspace, shared collection, or incremental reindexing. For a stable knowledge base that you query every day, AnythingLLM or a Python stack remain better suited. See the comparison section below.

#Prerequisites

The Local RAG Kit

LM Studio queries your documents. To move from a trial to a reliable tool, the Local RAG toolkit covers what makes the difference: document splitting (ch. 7), choosing a strong French embedding model (ch. 8), and evaluating responses (ch. 14).

  • Lifetime online access
  • PDF + files
  • Lifetime updates
LM Studio 0.3 or newer
The Chat with Documents feature appeared in 0.3 and has been enhanced since. Check your version under Settings → About. On an older LM Studio, update before going any further.
A loaded generative LLM
Qwen 3.5 9B Q4_K_M (6.6 GB, 256k context), Granite 4.2 8B Q4_K_M (5.3 GB, 128k), or Gemma 4 12B Q4_K_M (7.6 GB, multimodal) are good defaults, all under Apache 2.0. These recent models have large context windows, but still set the Context Length on the LM Studio side (see below)—otherwise you will be limited in how many passages LM Studio can inject.
An embeddings model
Required and separate from the LLM. nomic-embed-text-v1.5 (137M, ~80 MB in Q4) is the default recommended by LM Studio. mxbai-embed-large (335M) if you have room to spare. For exclusively French content, multilingual-e5-large is more relevant.
VRAM or RAM
Allow for the LLM size + ~200 MB for the embeddings model + 1-2 GB for the in-memory index of an average folder. On a 12 GB VRAM GPU, a 7B Q4 + nomic-embed leaves about 5 GB for the index and context.
Context window ≥ 8192
Set the LLM's Context Length to at least 8192, ideally 16384, from the chat's right panel. Below that, LM Studio truncates the retrieved passages and the response loses quality.

#1. Enable Chat with Documents

There is nothing to enable in the strict sense—it is a native chat feature. Simply provide a document in a conversation and LM Studio launches the RAG pipeline in the background.

  1. 01
    Open a new chat
    Open the Chat tab in the sidebar, then click the + button at the top to create a conversation. Make sure an LLM is loaded in the top dropdown menu. If nothing is loaded, select your generation model and wait for the VRAM to fill.
  2. 02
    Drag your file into the input area
    Drag and drop a PDF, DOCX, TXT, or MD directly into the message field at the bottom. A thumbnail appears above the field with the file name and size. You can attach several in succession.
  3. 03
    Ask your question
    Type your question normally, as you would in a standard chat. LM Studio detects the document, chunks it, encodes it, searches for relevant passages, and injects them into the prompt — all in 1–5 seconds depending on the size.
  4. 04
    Read the answer and sources
    The model answers based on the passages it retrieves. Depending on the LLM, it may or may not cite the excerpts. To force citation, add “Cite the exact passages from the document” to your question.
→
Short document vs. long document
If the document fits in the context window (typically fewer than 20 pages of text), LM Studio injects it in full without using RAG—no chunking, no embeddings, just one long prompt. Beyond that, it automatically switches to RAG mode. You don't need to configure anything, but it is useful to understand when interpreting the results.

#2. The embedding model

The embedding model is what turns a piece of text into a vector of numbers. Retrieval quality—and therefore the final quality of RAG—depends directly on this model, much more than on the generation LLM, contrary to what people initially assume.

  1. 01
    Download an embeddings model
    In the Discover tab, filter by “Text Embedding” in the left column. nomic-embed-text-v1.5 is prominently listed—it’s the sensible default. Download the Q4_K_M variant (~80 MB) or Q8_0 (~140 MB) if you insist on maximum quality.
  2. 02
    Verify that it is detected
    Settings → My Models → Embeddings tab. The model should appear there. If LM Studio can’t see it in this tab even though it has been downloaded, it wasn’t recognized as an embeddings model—check the Hugging Face tags or download it again through Discover.
  3. 03
    Select it in the chat
    When a document is attached to a chat, LM Studio displays the name of the embeddings model in use at the bottom of the conversation. Click it to change it. The choice is persisted per chat—handy for comparing nomic and multilingual-e5 on the same document.
nomic-embed-text-v1.5
137M parameters, 768 dimensions, 8192-token context. Excellent English default, adequate French. Recommended by LM Studio.
mxbai-embed-large-v1
335M, 1024 dimensions. Best score on the English MTEB, but 3× slower and 3× heavier. Relevant if retrieval quality is the bottleneck.
multilingual-e5-large
560M, 1024 dimensions. The best choice if your documents are in French or multilingual. Requires prefixing queries with « query: » and passages with « passage: », which LM Studio handles automatically.
bge-large-en-v1.5
335M, 1024 dimensions. Excellent in pure English, avoid for French.
!
The trap of using an English embedding model for French
A nomic-embed or a bge on French content cuts retrieval relevance by 1.5 to 2. You will get passages that are vaguely related to the topic rather than exactly relevant passages. If your documents are in French, choose multilingual-e5-large from the start—the extra latency is negligible compared with the quality gain.

#3. Supported file formats and sizes

LM Studio handles a pragmatic subset of common formats. No OCR for scanned PDFs, no parsing of complex tables, and no direct import from a URL—you need to know these limitations before expecting too much.

PDF (native text)
Supported. PDFs generated from Word, Google Docs, or LaTeX work well. Scanned PDFs without an OCR layer can’t be extracted—run them through ocrmypdf or Tesseract first.
DOCX
Supported. Formatting is ignored, and the text is extracted as plain text. Tables and images are lost.
TXT and MD
Natively supported, it is the ideal format. Markdown preserves its structure (headings, lists), which helps with chunking.
Source code (.py, .js, .ts…)
Treated as text. Works, but to chat with an entire repo, prefer Continue.dev or a tool designed for coding.
CSV, XLSX, JSON, HTML, EPUB
No official support at this time. Convert to TXT or MD with pandoc, csvkit, or a script before importing.
Maximum file size
No strict limit has been announced. In practice, allow 50-100 MB per PDF maximum on the GUI side—beyond that, chunking becomes slow and RAM usage rises quickly.
Prepare unsupported files
# PDF scanné -> PDF avec couche OCR
ocrmypdf scan.pdf scan_ocr.pdf

# HTML -> Markdown propre
pandoc page.html -o page.md

# XLSX -> CSV puis -> TXT lisible
libreoffice --headless --convert-to csv data.xlsx
column -s, -t < data.csv > data.txt

# EPUB -> TXT
pandoc livre.epub -t plain -o livre.txt
→
Several files at once
You can attach up to 5 files to the same chat (this limit has changed across versions; check your build). Beyond that, combine them into one large TXT file with a cat. LM Studio will chunk the whole thing as a single corpus, which works well for Markdown notes or chapters from the same document.

#4. Test with a real document

The best test is a document you know. Take a PDF of about twenty pages—a user manual, report, or contract—and first ask questions whose answers you know to calibrate your confidence in the system.

Typical validation workflow
# Préparer un document test
# - Un PDF de 20-50 pages avec une structure claire
# - Notez 5 faits précis présents dans le document
# - Notez 2 faits absents du document

# Dans LM Studio :
# 1. Chat -> nouvelle conversation
# 2. Charger Qwen 3.5 9B Q4_K_M (ou équivalent)
# 3. Context Length -> 16384 (panneau droit)
# 4. Drag-and-drop du PDF
# 5. Vérifier que nomic-embed (ou e5) est sélectionné
# 6. Poser les 5 questions à réponse connue
# 7. Poser les 2 questions à réponse absente

# Verdict :
# - 5/5 corrects + 2/2 "je ne trouve pas" -> RAG fiable sur ce corpus
# - Hallucinations sur les 2 absents -> baisser la température à 0.1
# - Réponses approximatives -> changer d'embeddings ou monter Top K
Low temperature
For RAG, lower the temperature to 0.1–0.3 in the right-hand panel. The goal is to stick to the passages, not create.
Explicit system prompt
« You answer only from the provided excerpts. If the information is not present, say so clearly. Quote the exact passages in quotation marks. » This simple addition cuts hallucinations by 3.
Structured question
“What is the notice period? Quote the exact clause.” is more useful than “tell me about the notice period.” The quotation forces the model to stay grounded.

#Limitations vs AnythingLLM

LM Studio provides “throwaway” RAG — per conversation, with no persistence between chats and no fine-tuning. That’s intentional: the feature is designed for occasional lookups, not a knowledge base. As soon as you want more, AnythingLLM (free, open source, with a similar GUI approach) takes over.

Persistent workspaces
AnythingLLM stores your documents in reusable workspaces. You index 200 PDFs once, and all your workspace chats can access them. LM Studio reindexes with every new chat.
Clickable citations
AnythingLLM displays the sources with a link to the exact passage. LM Studio merely injects the passage into the prompt—the model may or may not cite it, as it sees fit.
Exposed RAG settings
Top K, chunk size, similarity threshold, embedding model: everything is configurable in AnythingLLM. LM Studio handles it as a black box with reasonable defaults.
External vector databases
AnythingLLM connects to ChromaDB, Qdrant, Pinecone, and Weaviate. LM Studio keeps everything in memory within the process, limiting it to modest corpora.
Multi-utilisateurs
AnythingLLM supports users and workspace sharing. LM Studio is single-user by design.
Source connectors
AnythingLLM can ingest from URLs, Confluence, GitHub, and YouTube transcripts. LM Studio is limited to dragging and dropping files.
i
The right choice for each use case
LM Studio Chat with Documents: test a PDF, summarize a report you received, ask a few questions about a contract, and conduct occasional monitoring. AnythingLLM: a team knowledge base, customer support, and product documentation queried every day. If you’re unsure, start with LM Studio—switching to AnythingLLM takes 30 minutes when the limit becomes a problem.

#When to move to a Python stack

At some point, the GUI hits its ceiling and coding becomes simpler than working around it. Signs that you should move to Python + LlamaIndex (or Haystack):

Structure-aware chunking
Your documents have a strong structure (legal sections, code, scientific articles) that naïve chunking breaks. A Markdown-aware splitter or a docling parser makes a difference.
Hybrid search
You want to combine vector search (semantic) and BM25 (lexical) so you don't miss exact terms (clause numbers, variable names, identifiers). No GUI does this natively today.
Reranking
You add a cross-encoder (BAAI/bge-reranker-v2-m3) to rerank the top 20 results. +15% relevance, but it requires coding.
Metadata and filtering
“Search only within contracts signed in 2024.” Requires metadata per chunk and a filter at query time.
Automated ingestion pipeline
A watched folder, a cron job, a webhook that reindexes whenever a new file arrives. At that point, you’ve definitively left the GUI behind.
Evaluation
Measure retrieval quality on 50 reference questions. Without Python, you're shooting blind.

Good news: LM Studio can stay in front while Python runs the pipeline. The OpenAI-compatible server on localhost:1234 can be consumed by LlamaIndex in two lines—you keep LM Studio for exploratory chat and code the pipeline behind it.

Connect LlamaIndex to LM Studio
from llama_index.core import VectorStoreIndex, SimpleDirectoryReader, Settings
from llama_index.embeddings.huggingface import HuggingFaceEmbedding
from llama_index.llms.openai_like import OpenAILike

# LM Studio sert le LLM via son endpoint OpenAI-compatible
Settings.llm = OpenAILike(
    model="local-model",
    api_base="http://localhost:1234/v1",
    api_key="lm-studio",
    is_chat_model=True,
    context_window=16384,
)

# Embeddings côté Python (LM Studio peut aussi les exposer)
Settings.embed_model = HuggingFaceEmbedding("intfloat/multilingual-e5-large")

docs = SimpleDirectoryReader("./mes_documents").load_data()
index = VectorStoreIndex.from_documents(docs)
index.storage_context.persist("./storage")

qe = index.as_query_engine(similarity_top_k=5)
print(qe.query("Quelles sont les obligations du prestataire ?"))
→
Embeddings via LM Studio too
Since 0.3, LM Studio exposes embeddings through /v1/embeddings (OpenAI-compatible endpoint). You can therefore run both the LLM and the embeddings from LM Studio, using Python only for the pipeline. Convenient for reusing already allocated VRAM.

#Go further

Chat with Documents covers 80% of the needs of personal use or a small team. The logical next steps after this first RAG:

Understanding the mechanisms
The local RAG guide: the introduction covers chunking, embeddings, retrieval, and the pitfalls no GUI reveals. Useful reading before changing the settings.
Choose a better French embedding model
The guide The Best Embedding Models in French compares Solon, multilingual-e5, and BGE on French-language content. The difference is measurable.
Switch to API server
The Transformer LM Studio as an API server guide shows how to expose the OpenAI-compatible endpoint and connect LlamaIndex, Continue.dev, or a custom script without losing your current setup.
Did this guide help you?

Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.