Advanced 16 minRAG

Local GraphRAG: knowledge-graph RAG (guide advanced)

Classic vector RAG handles one-off questions very well, but falls apart when it needs to connect multiple documents to synthesize an answer. GraphRAG tackles this problem by building a knowledge graph—entities, relationships, communities—from your corpus, then querying that graph instead of a simple vector database. This guide shows how to set up a local LLM graph RAG with Ollama, without calling any remote API, and especially when this approach truly outperforms vector search.

By Mohamed Meguedmi·Update 2026-08-27·Tested on Windows, macOS, and Linux

#Why GraphRAG?

Imagine the corpus of a law firm: 200 contracts, 500 emails, 80 decisions. You ask: “What are the main legal risks mentioned in our client contracts over the past three years, and with which recurring clients?” A classic vector RAG retrieves 5 or 10 “relevant” chunks, gives them to the LLM, and it responds… with a partial view. It misses cross-cutting patterns.

Rather than returning only raw passages, GraphRAG reasons over a structure: who mentions what, which entities recur, and which relationships link them. For this type of synthetic (“global”) question over a corpus, the graph beats the vector approach — precisely what Microsoft Research’s original 2024 paper demonstrated.

i
Who this guide is for
You already have a working local vector RAG (ChromaDB, LlamaIndex, AnythingLLM…), and you're running into synthesis questions. If not, start with a classic RAG before tackling this one.

#GraphRAG vs vector RAG: the real difference

The Local RAG Kit

Your documents, your AI: a reliable local RAG over your PDFs, notes and mail — nothing leaves your machine.

  • Lifetime online access
  • PDF + files
  • Lifetime updates

The two approaches share one goal: inject external context into the LLM prompt to limit hallucinations. But they do not retrieve the same thing.

Vector RAG
Split the corpus into chunks, compute one embedding per chunk, and store them in a vector database (ChromaDB, Qdrant, FAISS). At query time, retrieve the k chunks closest by cosine similarity.
GraphRAG
Ask an LLM to extract entities and relationships from each chunk, build a graph, group entities into communities, and summarize each community. At query time, traverse the graph or aggregate community summaries.
Vector strengths
Fast to index (a few minutes for 10 MB of text), inexpensive, and excellent for targeted questions (“what is the termination clause in the Acme contract?”).
GraphRAG strengths
Excellent for broad questions (“what are the recurring themes?”, “which entities are most connected?”), fine-grained traceability through edges, native multi-hop.
Vector weakness
Loses links between documents. A question that requires joining 3 semantically distant documents doesn't retrieve the right chunks.
GraphRAG weakness
Heavy indexing: every chunk must pass through an LLM. On a 10 MB corpus, expect several hours and plenty of VRAM, whereas vector search finishes in 5 minutes.
→
The hybrid approach often wins
In practice, the best stacks combine both: a graph for broad questions and navigation, and vector search (or BM25) for targeted queries. See also hybrid BM25 + vector search for another form of combination.

#How it works internally

A complete GraphRAG pipeline chains together 5 steps. All of them use the LLM (except clustering).

  1. 01
    Chunking
    The corpus is split into passages of 500 to 1,500 tokens, as in a standard RAG setup. Size has a direct impact on extraction quality: too short, and the LLM misses relationships; too long, and it forgets them.
  2. 02
    Entity and relationship extraction
    Each chunk is passed to an LLM with a structured prompt such as: “Extract all entities (person, organization, place, concept) and the relationships between them. JSON format.” This is the expensive step—one LLM call per chunk.
  3. 03
    Building the graph
    Extracted entities become nodes, and relationships become edges. Identical entities appearing in multiple chunks are merged (entity resolution, often using embeddings or a normalization rule).
  4. 04
    Community detection
    A clustering algorithm (Leiden at Microsoft, simpler in nano-graphrag) groups strongly connected nodes into communities. These communities are the key to “global” reasoning.
  5. 05
    Community summary
    The LLM produces a textual summary of each community from the entities and relationships it contains. These summaries become the retrieval units for global questions.

When queried, GraphRAG distinguishes between two modes: local (searching for a specific entity and its neighborhood) and global (aggregating community summaries). The LLM then combines the retrieved context and the question to produce the final answer.

#Tools available for a local deployment

Microsoft GraphRAG
The reference implementation (github.com/microsoft/graphrag). Complete and polished, but heavy: originally designed for Azure OpenAI, the Ollama port requires patience. Very costly in tokens for indexing.
nano-graphrag
Minimal implementation (~1000 lines) with native Ollama compatibility (github.com/gusye1234/nano-graphrag). That’s what we’ll use here: 10x less code to understand, with the same idea.
LightRAG
More recent variant, optimized for query latency. Compatible with Ollama. Simpler than Microsoft GraphRAG, more structured than nano-graphrag.
LlamaIndex KnowledgeGraphIndex
If you’re already using LlamaIndex, integration is immediate, but the approach is more rudimentary (no communities).
i
The educational choice
This guide uses nano-graphrag because everything fits into two Python files that you can read, modify, and debug. Once you understand the concept, moving to Microsoft GraphRAG or LightRAG in production becomes trivial.

#Prerequisites

Ollama installed and working
If not, start by following our Ollama installation guide.
A solid reasoning LLM (14–24B)
Entity extraction needs horsepower. gpt-oss 20B, Mistral Small 24B, or Qwen 3.5 9B (the lower limit) work well. Below 8B, the generated JSON is often malformed.
A local embedding model
nomic-embed-text via Ollama, or bge-m3 / multilingual-e5-large via sentence-transformers.
16 GB of VRAM minimum
12 GB works with an 8–9B Q4 (Qwen 3.5 9B) but indexing will be slow. 16 GB accommodates gpt-oss 20B or Mistral Small 24B; 24 GB (RTX 4090, M-Max) is comfortable.
Python 3.10+
nano-graphrag and most modern RAG frameworks require 3.10 or newer.

#1. Index a corpus with nano-graphrag

First prepare the models on the Ollama side. For this tutorial, we use gpt-oss 20B (MXFP4 quantization by default, ~14 GB) and nomic-embed-text as the embeddings model.

Terminal
ollama pull gpt-oss:20b
ollama pull nomic-embed-text
ollama serve  # si pas déjà en service

Then install nano-graphrag in a dedicated venv.

Terminal
python -m venv .venv
source .venv/bin/activate  # Linux/macOS
pip install nano-graphrag

The indexing script is about twenty lines long. Connect it to Ollama through the OpenAI-compatible endpoint on port 11434.

index.py
import asyncio
from nano_graphrag import GraphRAG, QueryParam
from nano_graphrag.llm import ollama_model_if_cache, ollama_embedding

WORKING_DIR = "./graphrag_cache"

async def main():
    rag = GraphRAG(
        working_dir=WORKING_DIR,
        best_model_func=ollama_model_if_cache,
        cheap_model_func=ollama_model_if_cache,
        embedding_func=ollama_embedding,
        best_model_kwargs={"model_name": "gpt-oss:20b"},
        cheap_model_kwargs={"model_name": "gpt-oss:20b"},
    )

    with open("corpus.txt", encoding="utf-8") as f:
        text = f.read()

    await rag.ainsert(text)

if __name__ == "__main__":
    asyncio.run(main())

Start indexing. Depending on the corpus size and GPU, expect anywhere from a few minutes (1 MB of text) to several hours (50 MB).

Terminal
python index.py
!
Be patient: it's inherently slow
On a 5 MB corpus with gpt-oss 20B on RTX 4090, expect about 2 hours. nano-graphrag processes chunks sequentially by default. That's normal—you’re paying for LLM extraction, not embedding computation.

At the end, the graphrag_cache/ folder contains the serialized graph, embeddings, and community summaries.

#2. Query the graph

Once indexed, queries are fast (a few seconds per query) because the LLM reads only the context retrieved from the graph, not the entire corpus.

query.py
import asyncio
from nano_graphrag import GraphRAG, QueryParam
from nano_graphrag.llm import ollama_model_if_cache, ollama_embedding

async def main():
    rag = GraphRAG(
        working_dir="./graphrag_cache",
        best_model_func=ollama_model_if_cache,
        cheap_model_func=ollama_model_if_cache,
        embedding_func=ollama_embedding,
        best_model_kwargs={"model_name": "gpt-oss:20b"},
        cheap_model_kwargs={"model_name": "gpt-oss:20b"},
    )

    # Mode global : synthèse à partir des résumés de communautés
    print(await rag.aquery(
        "Quels sont les thèmes principaux du corpus ?",
        param=QueryParam(mode="global")
    ))

    # Mode local : recherche centrée sur des entités
    print(await rag.aquery(
        "Quelle est la position de l'entreprise X sur le sujet Y ?",
        param=QueryParam(mode="local")
    ))

asyncio.run(main())
→
Choose the right mode
Question starting with “what are the main…,” “what trends…,” or “summarize…” → global mode. Question about a specific entity → local mode. If you are unsure, try both: the answers often differ radically.

#Local compute cost: what to expect

This is the part that surprises everyone the first time. Indexing 5 MB of text in GraphRAG requires about 50 to 200 times more compute than vector RAG on the same corpus. Here are some concrete benchmarks, in rough orders of magnitude.

1 MB corpus (~300 pages)
gpt-oss 20B (MXFP4) on RTX 4090: ~25 min. for indexing. On RTX 3060 12 GB (partial offload): ~3h. On a 64 GB Mac M3 Max: ~40 min.
5 MB corpus (~1500 pages)
RTX 4090: ~2h. M4 Pro 48 GB: ~3h. Beyond this volume, plan to leave it running overnight.
VRAM peak
The gpt-oss 20B model (MXFP4) continuously occupies ~14 GB. Nomic embeddings add ~1 GB. With less than 16 GB of VRAM, expect CPU swapping (partial offload).
Cost per request
A few seconds in local mode, 5 to 30 s in global mode (aggregating several communities). Painless compared with indexing.
Incremental re-indexing
nano-graphrag currently cannot cleanly update an existing graph. Adding 10% new documents means rerunning a partial or full indexing job. Microsoft GraphRAG handles this better.
!
The trap of an LLM that’s too small
Considering Granite 4.2 3B or Qwen 3.5 2B to go faster? JSON extraction will be inconsistent, entities will be misnamed, and the graph will be unusable. This is exactly the mistake where fast indexing produces an unusable graph. Invest in at least an 8B model, ideally gpt-oss 20B or Mistral Small 24B.

#When GraphRAG really beats vector search

GraphRAG is not a universal replacement for vector RAG. It outperforms it for some use cases and significantly underperforms it for others.

Corpus summarization
« What are the 5 main themes discussed in our 200 emails about topic X? » → GraphRAG wins hands down. The vector search returns only 5-10 emails; the graph aggregates the communities.
Multi-hop questions
“Which vendors work with both Acme and Beta Corp?” → GraphRAG solves it by traversing edges. The vector approach must retrieve the right chunks and rely on the LLM to perform the join.
Exploring relationships
“Who are the people most often mentioned in connection with the Atlas project?” → GraphRAG is built for this (centrality, neighborhood). Vector search has no notion of relationships.
Targeted factual Q&A
“What is the notice period in the Acme contract dated March 12, 2024?” → Vector RAG wins: faster, more accurate, and less expensive to index.
Highly scalable corpus
If you add documents every day, GraphRAG's re-indexing cost becomes prohibitive. Stick with vector search or vector search + BM25.
Corpus < 500 KB
No graph needed: the LLM can read everything in context if you have 32k+ tokens. GraphRAG is justified only for corpora too large to fit in context.
→
The rule of thumb
If your users mostly ask questions like “find this specific piece of information,” stick with vector search. If the value lies in summarization, pattern exploration, or cross-sectional analysis of a stable corpus, GraphRAG is worth its indexing cost.

#Go further

GraphRAG is an active field: implementations are evolving quickly, and so are the benchmarks. Here are a few ways to continue.

Local RAG: introduction
If some vector RAG concepts are still unclear, the introductory guide is a good foundation before taking GraphRAG to production.
Chunking strategies
Entity extraction quality depends directly on chunk size and consistency. This guide explores best practices.
Hybrid BM25 + vector search
To combine GraphRAG with conventional retrieval, start by looking at hybrid search: the same principle of combining signals.
Choose your GPU for local AI
GraphRAG indexing is resource-intensive. If you are running on an 8-12 GB GPU today, moving to 16-24 GB radically changes the throughput.
Did this guide help you?

Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.