Home › Catalog › Best local LLM for RAG in 2026

Best local LLM for RAG in 2026

◆ Local RAG — Ask questions to your own documents with a local AI, no cloud · $24 · or all kits $49 →

Ranking updated on 09/10/2026

Ranking of the LLMs best suited to RAG: long context window (≥ 32k tokens to digest multiple documents), synthesis quality on provided sources, and robustness to “distractors” (irrelevant information in the prompt).

Ranking

1

🇨🇳 Qwen 3.6 27B

Alibaba · 27B parameters · Apache 2.0 · 262,144-token context

Dense multimodal 27B released April 22, 2026. 262k ctx (1M YaRN). SWE-bench Verified 77.2%.

Why this ranking 262,144-token context—excellent for large corpora. 27B parameters for high-quality summarization.
ollama run qwen3.6:27b
Q4 VRAM
16 GB
29 GB in Q8
2

🇨🇳 Qwen 3.8 27B

Alibaba · 27B parameters · Apache 2.0 · 262,144-token context

Qwen 3.8 27B: dense multimodal (text + vision), 262k context, ~16 GB Q4 VRAM (18 GB of Ollama weights). Apache 2.0, agentic coding and vision.

Why this ranking 262,144-token context—excellent for large corpora. 27B parameters for high-quality summarization.
ollama run qwen3.8:27b
Q4 VRAM
16 GB
29 GB in Q8
3

🇺🇸 North Mini Code 1.0

Cohere · 30.5B parameters · Apache 2.0 · 488,000 tokens ctx

Cohere 30.5B model focused on agentic coding and reasoning. 488k context, Apache 2.0, ~18 GB VRAM in Q4.

Why this ranking 488,000-token context — excellent for large corpora. 30.5B parameters for high-quality summarization.
ollama run north-mini-code-1.0
Q4 VRAM
18 GB
33 GB in Q8
4

🇺🇸 Gemma 4 26B-A4B MoE

Google · 26B parameters · Apache 2.0 · 128,000 tokens ctx

MoE variant of Gemma 4. 26B/4B active. Full multimodal (text+image+audio).

Why this ranking 128,000-token context—excellent for large corpora. 26B parameters for high-quality summarization.
ollama run gemma4:26b
Q4 VRAM
16 GB
28 GB in Q8
5

🇺🇸 Gemma 4 31B

Google · 31B parameters · Apache 2.0 · 256,000-token context

Dense 31B multimodal (text+image+audio). 140+ languages, 256k context. #3 open model on Chatbot Arena.

Why this ranking 256,000-token context — excellent for large corpora. 31B parameters for high-quality summarization.
ollama run gemma4:31b
Q4 VRAM
18 GB
33 GB in Q8
6

🇨🇳 GLM 4.7 Flash

Zhipu AI · 31B parameters · MIT · 128,000 tokens ctx

GLM-4.7-Flash (MoE 31B, ~3B active): the best code/VRAM ratio in the 30B class. MIT, 128k ctx, very fast on 3090/4090.

Why this ranking 128,000-token context — excellent for large corpora. 31B parameters for high-quality summarization.
ollama run glm-4.7-flash
Q4 VRAM
19 GB
35 GB in Q8
7

🇨🇳 Qwen 3.6 35B-A3B

Alibaba · 35B parameters · Apache 2.0 · 262,000-token context

MoE with 35B/3B active parameters for agentic coding. 73.4% SWE-Bench. Release: April 16, 2026.

Why this ranking 262,000-token context — excellent for large corpora. 35B parameters for high-quality summarization.
ollama run qwen3.6:35b-a3b
Q4 VRAM
21 GB
38 GB in Q8

Comparison table

Rank Model Params Q4 VRAM Context License
#1 Qwen 3.6 27B 27B 16 GB 262 144 Apache 2.0
#2 Qwen 3.8 27B 27B 16 GB 262 144 Apache 2.0
#3 North Mini Code 1.0 30.5B 18 GB 488 000 Apache 2.0
#4 Gemma 4 26B-A4B MoE 26B 16 GB 128 000 Apache 2.0
#5 Gemma 4 31B 31B 18 GB 256 000 Apache 2.0
#6 GLM 4.7 Flash 31B 19 GB 128 000 MIT
#7 Qwen 3.6 35B-A3B 35B 21 GB 262 000 Apache 2.0
The Local RAG Kit

Your documents, your AI: reliable local RAG for your PDFs, notes, and emails—without sending anything to the cloud.

  • Lifetime online access
  • PDF + files
  • Lifetime updates

Ranking methodology

We keep chat/general models with a context of at least 32k tokens (required for useful chunking). The score favors recent models with large context windows and permissive licenses (enterprise deployment).

Criteria considered:

  • Context ≥ 32k tokens
  • Summarization quality
  • Robustness to distractors
  • Free license

The scoring is fully transparent: see our methodology for details on VRAM/tokens/sec calculations.

Frequently asked questions

What context size makes for a good RAG?

At least 32k tokens to handle 5–10 chunks of ~2k tokens plus the prompt. Ideal: 128k tokens (Qwen 3.6 27B has 262 144), enough to handle entire documents without aggressive chunking.

Which model for RAG on 24 GB VRAM (RTX 4090)?

Mistral Small 3.1 24B in Q4_K_M or Qwen 2.5 32B fit in 24 GB in Q4. For Llama 3.3 70B, you need to drop to Q2/Q3 or add a second card.

Which embedding should you pair with these LLMs?

For French: BGE-M3, multilingual-e5-large, or the Mistral embeddings. See the French embeddings guide.

Which local RAG stack is recommended?

Ollama (LLM server) + ChromaDB or Qdrant (vector store) + LlamaIndex or LangChain (orchestration). See the complete guide.

Go further

QuelLLM Kits The reference guide by use case
All kits for life — $49