Local GraphRAG: knowledge-graph RAG (guide advanced)
Classic vector RAG handles one-off questions very well, but falls apart when it needs to connect multiple documents to synthesize an answer. GraphRAG tackles this problem by building a knowledge graph—entities, relationships, communities—from your corpus, then querying that graph instead of a simple vector database. This guide shows how to set up a local LLM graph RAG with Ollama, without calling any remote API, and especially when this approach truly outperforms vector search.
#Why GraphRAG?
Imagine the corpus of a law firm: 200 contracts, 500 emails, 80 decisions. You ask: “What are the main legal risks mentioned in our client contracts over the past three years, and with which recurring clients?” A classic vector RAG retrieves 5 or 10 “relevant” chunks, gives them to the LLM, and it responds… with a partial view. It misses cross-cutting patterns.
Rather than returning only raw passages, GraphRAG reasons over a structure: who mentions what, which entities recur, and which relationships link them. For this type of synthetic (“global”) question over a corpus, the graph beats the vector approach — precisely what Microsoft Research’s original 2024 paper demonstrated.
#GraphRAG vs vector RAG: the real difference
Your documents, your AI: a reliable local RAG over your PDFs, notes and mail — nothing leaves your machine.
- Lifetime online access
- PDF + files
- Lifetime updates
The two approaches share one goal: inject external context into the LLM prompt to limit hallucinations. But they do not retrieve the same thing.
- Vector RAG
- Split the corpus into chunks, compute one embedding per chunk, and store them in a vector database (ChromaDB, Qdrant, FAISS). At query time, retrieve the k chunks closest by cosine similarity.
- GraphRAG
- Ask an LLM to extract entities and relationships from each chunk, build a graph, group entities into communities, and summarize each community. At query time, traverse the graph or aggregate community summaries.
- Vector strengths
- Fast to index (a few minutes for 10 MB of text), inexpensive, and excellent for targeted questions (“what is the termination clause in the Acme contract?”).
- GraphRAG strengths
- Excellent for broad questions (“what are the recurring themes?”, “which entities are most connected?”), fine-grained traceability through edges, native multi-hop.
- Vector weakness
- Loses links between documents. A question that requires joining 3 semantically distant documents doesn't retrieve the right chunks.
- GraphRAG weakness
- Heavy indexing: every chunk must pass through an LLM. On a 10 MB corpus, expect several hours and plenty of VRAM, whereas vector search finishes in 5 minutes.
#How it works internally
A complete GraphRAG pipeline chains together 5 steps. All of them use the LLM (except clustering).
- 01ChunkingThe corpus is split into passages of 500 to 1,500 tokens, as in a standard RAG setup. Size has a direct impact on extraction quality: too short, and the LLM misses relationships; too long, and it forgets them.
- 02Entity and relationship extractionEach chunk is passed to an LLM with a structured prompt such as: “Extract all entities (person, organization, place, concept) and the relationships between them. JSON format.” This is the expensive step—one LLM call per chunk.
- 03Building the graphExtracted entities become nodes, and relationships become edges. Identical entities appearing in multiple chunks are merged (entity resolution, often using embeddings or a normalization rule).
- 04Community detectionA clustering algorithm (Leiden at Microsoft, simpler in nano-graphrag) groups strongly connected nodes into communities. These communities are the key to “global” reasoning.
- 05Community summaryThe LLM produces a textual summary of each community from the entities and relationships it contains. These summaries become the retrieval units for global questions.
When queried, GraphRAG distinguishes between two modes: local (searching for a specific entity and its neighborhood) and global (aggregating community summaries). The LLM then combines the retrieved context and the question to produce the final answer.
#Tools available for a local deployment
- Microsoft GraphRAG
- The reference implementation (github.com/microsoft/graphrag). Complete and polished, but heavy: originally designed for Azure OpenAI, the Ollama port requires patience. Very costly in tokens for indexing.
- nano-graphrag
- Minimal implementation (~1000 lines) with native Ollama compatibility (github.com/gusye1234/nano-graphrag). That’s what we’ll use here: 10x less code to understand, with the same idea.
- LightRAG
- More recent variant, optimized for query latency. Compatible with Ollama. Simpler than Microsoft GraphRAG, more structured than nano-graphrag.
- LlamaIndex KnowledgeGraphIndex
- If you’re already using LlamaIndex, integration is immediate, but the approach is more rudimentary (no communities).
#Prerequisites
- Ollama installed and working
- If not, start by following our Ollama installation guide.
- A solid reasoning LLM (14–24B)
- Entity extraction needs horsepower. gpt-oss 20B, Mistral Small 24B, or Qwen 3.5 9B (the lower limit) work well. Below 8B, the generated JSON is often malformed.
- A local embedding model
- nomic-embed-text via Ollama, or bge-m3 / multilingual-e5-large via sentence-transformers.
- 16 GB of VRAM minimum
- 12 GB works with an 8–9B Q4 (Qwen 3.5 9B) but indexing will be slow. 16 GB accommodates gpt-oss 20B or Mistral Small 24B; 24 GB (RTX 4090, M-Max) is comfortable.
- Python 3.10+
- nano-graphrag and most modern RAG frameworks require 3.10 or newer.
#1. Index a corpus with nano-graphrag
First prepare the models on the Ollama side. For this tutorial, we use gpt-oss 20B (MXFP4 quantization by default, ~14 GB) and nomic-embed-text as the embeddings model.
Then install nano-graphrag in a dedicated venv.
The indexing script is about twenty lines long. Connect it to Ollama through the OpenAI-compatible endpoint on port 11434.
Start indexing. Depending on the corpus size and GPU, expect anywhere from a few minutes (1 MB of text) to several hours (50 MB).
At the end, the graphrag_cache/ folder contains the serialized graph, embeddings, and community summaries.
#2. Query the graph
Once indexed, queries are fast (a few seconds per query) because the LLM reads only the context retrieved from the graph, not the entire corpus.
#Local compute cost: what to expect
This is the part that surprises everyone the first time. Indexing 5 MB of text in GraphRAG requires about 50 to 200 times more compute than vector RAG on the same corpus. Here are some concrete benchmarks, in rough orders of magnitude.
- 1 MB corpus (~300 pages)
- gpt-oss 20B (MXFP4) on RTX 4090: ~25 min. for indexing. On RTX 3060 12 GB (partial offload): ~3h. On a 64 GB Mac M3 Max: ~40 min.
- 5 MB corpus (~1500 pages)
- RTX 4090: ~2h. M4 Pro 48 GB: ~3h. Beyond this volume, plan to leave it running overnight.
- VRAM peak
- The gpt-oss 20B model (MXFP4) continuously occupies ~14 GB. Nomic embeddings add ~1 GB. With less than 16 GB of VRAM, expect CPU swapping (partial offload).
- Cost per request
- A few seconds in local mode, 5 to 30 s in global mode (aggregating several communities). Painless compared with indexing.
- Incremental re-indexing
- nano-graphrag currently cannot cleanly update an existing graph. Adding 10% new documents means rerunning a partial or full indexing job. Microsoft GraphRAG handles this better.
#When GraphRAG really beats vector search
GraphRAG is not a universal replacement for vector RAG. It outperforms it for some use cases and significantly underperforms it for others.
- Corpus summarization
- « What are the 5 main themes discussed in our 200 emails about topic X? » → GraphRAG wins hands down. The vector search returns only 5-10 emails; the graph aggregates the communities.
- Multi-hop questions
- “Which vendors work with both Acme and Beta Corp?” → GraphRAG solves it by traversing edges. The vector approach must retrieve the right chunks and rely on the LLM to perform the join.
- Exploring relationships
- “Who are the people most often mentioned in connection with the Atlas project?” → GraphRAG is built for this (centrality, neighborhood). Vector search has no notion of relationships.
- Targeted factual Q&A
- “What is the notice period in the Acme contract dated March 12, 2024?” → Vector RAG wins: faster, more accurate, and less expensive to index.
- Highly scalable corpus
- If you add documents every day, GraphRAG's re-indexing cost becomes prohibitive. Stick with vector search or vector search + BM25.
- Corpus < 500 KB
- No graph needed: the LLM can read everything in context if you have 32k+ tokens. GraphRAG is justified only for corpora too large to fit in context.
#Go further
GraphRAG is an active field: implementations are evolving quickly, and so are the benchmarks. Here are a few ways to continue.
- Local RAG: introduction
- If some vector RAG concepts are still unclear, the introductory guide is a good foundation before taking GraphRAG to production.
- Chunking strategies
- Entity extraction quality depends directly on chunk size and consistency. This guide explores best practices.
- Hybrid BM25 + vector search
- To combine GraphRAG with conventional retrieval, start by looking at hybrid search: the same principle of combining signals.
- Choose your GPU for local AI
- GraphRAG indexing is resource-intensive. If you are running on an 8-12 GB GPU today, moving to 16-24 GB radically changes the throughput.
Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.