🇨🇳 Qwen 3.6 27B
Dense multimodal 27B released April 22, 2026. 262k ctx (1M YaRN). SWE-bench Verified 77.2%.
ollama run qwen3.6:27b
Ranking updated on 09/10/2026
Ranking of the LLMs best suited to RAG: long context window (≥ 32k tokens to digest multiple documents), synthesis quality on provided sources, and robustness to “distractors” (irrelevant information in the prompt).
Dense multimodal 27B released April 22, 2026. 262k ctx (1M YaRN). SWE-bench Verified 77.2%.
ollama run qwen3.6:27b
Qwen 3.8 27B: dense multimodal (text + vision), 262k context, ~16 GB Q4 VRAM (18 GB of Ollama weights). Apache 2.0, agentic coding and vision.
ollama run qwen3.8:27b
Cohere 30.5B model focused on agentic coding and reasoning. 488k context, Apache 2.0, ~18 GB VRAM in Q4.
ollama run north-mini-code-1.0
MoE variant of Gemma 4. 26B/4B active. Full multimodal (text+image+audio).
ollama run gemma4:26b
Dense 31B multimodal (text+image+audio). 140+ languages, 256k context. #3 open model on Chatbot Arena.
ollama run gemma4:31b
GLM-4.7-Flash (MoE 31B, ~3B active): the best code/VRAM ratio in the 30B class. MIT, 128k ctx, very fast on 3090/4090.
ollama run glm-4.7-flash
MoE with 35B/3B active parameters for agentic coding. 73.4% SWE-Bench. Release: April 16, 2026.
ollama run qwen3.6:35b-a3b
| Rank | Model | Params | Q4 VRAM | Context | License |
|---|---|---|---|---|---|
| #1 | Qwen 3.6 27B | 27B | 16 GB | 262 144 | Apache 2.0 |
| #2 | Qwen 3.8 27B | 27B | 16 GB | 262 144 | Apache 2.0 |
| #3 | North Mini Code 1.0 | 30.5B | 18 GB | 488 000 | Apache 2.0 |
| #4 | Gemma 4 26B-A4B MoE | 26B | 16 GB | 128 000 | Apache 2.0 |
| #5 | Gemma 4 31B | 31B | 18 GB | 256 000 | Apache 2.0 |
| #6 | GLM 4.7 Flash | 31B | 19 GB | 128 000 | MIT |
| #7 | Qwen 3.6 35B-A3B | 35B | 21 GB | 262 000 | Apache 2.0 |
Your documents, your AI: reliable local RAG for your PDFs, notes, and emails—without sending anything to the cloud.
We keep chat/general models with a context of at least 32k tokens (required for useful chunking). The score favors recent models with large context windows and permissive licenses (enterprise deployment).
Criteria considered:
The scoring is fully transparent: see our methodology for details on VRAM/tokens/sec calculations.
What context size makes for a good RAG?
At least 32k tokens to handle 5–10 chunks of ~2k tokens plus the prompt. Ideal: 128k tokens (Qwen 3.6 27B has 262 144), enough to handle entire documents without aggressive chunking.
Which model for RAG on 24 GB VRAM (RTX 4090)?
Mistral Small 3.1 24B in Q4_K_M or Qwen 2.5 32B fit in 24 GB in Q4. For Llama 3.3 70B, you need to drop to Q2/Q3 or add a second card.
Which embedding should you pair with these LLMs?
For French: BGE-M3, multilingual-e5-large, or the Mistral embeddings. See the French embeddings guide.
Which local RAG stack is recommended?
Ollama (LLM server) + ChromaDB or Qdrant (vector store) + LlamaIndex or LangChain (orchestration). See the complete guide.