🇺🇸 Laguna XS.2
MoE 33B/3B active parameters, Apache 2.0, specializing in agentic coding. 68.2% SWE-Bench Verified, 128k ctx. Runs on a 36 GB Mac. Released April 28, 2026.
ollama run laguna-xs.2
Ranking updated on 09/10/2026
The RTX 4090 (24 GB VRAM, Ada Lovelace architecture) is the mainstream reference for LLM inference in 2026. Here are the models that make the most of it: top-quality Q4/Q5 models that fit in 24 GB, with comfortable throughput (30+ tokens/sec).
RTX 4090: purchasing alternative available for local AI — GMKtec EVO-X2 64 GB / 1 TB (Ryzen AI Max+ 395):
A mini PC is a complete machine: check the required memory and software compatibility. It does not replace macOS/MLX or CUDA.
Why this choice? Our complete overview of the GMKtec EVO-X2 64 GB / 1 TB (Ryzen AI Max+ 395) →
Which PC should you choose for your budget? Our picks from €800 to €3,500 →
Affiliate links — BestLLMfor may earn a commission on purchases, at no extra cost to you, which does not influence the ranking (established independently). As an Amazon Associate, BestLLMfor earns from qualifying purchases.
MoE 33B/3B active parameters, Apache 2.0, specializing in agentic coding. 68.2% SWE-Bench Verified, 128k ctx. Runs on a 36 GB Mac. Released April 28, 2026.
ollama run laguna-xs.2
MoE variant of Gemma 4. 26B/4B active. Full multimodal (text+image+audio).
ollama run gemma4:26b
First open Apache 2.0 dLLM: MoE 16B/1B + 6.2B diffusion decoder. Unified text+vision. Released April 22, 2026.
# HuggingFace : inclusionAI/LLaDA2.0-Uni (Flash Attn 2 + CUDA 12.4 requis)
Dense multimodal 27B released April 22, 2026. 262k ctx (1M YaRN). SWE-bench Verified 77.2%.
ollama run qwen3.6:27b
Qwen 3.8 27B: dense multimodal (text + vision), 262k context, ~16 GB Q4 VRAM (18 GB of Ollama weights). Apache 2.0, agentic coding and vision.
ollama run qwen3.8:27b
GLM-4.7-Flash (MoE 31B, ~3B active): the best code/VRAM ratio in the 30B class. MIT, 128k ctx, very fast on 3090/4090.
ollama run glm-4.7-flash
Dense 31B multimodal (text+image+audio). 140+ languages, 256k context. #3 open model on Chatbot Arena.
ollama run gemma4:31b
| Rank | Model | Params | Q4 VRAM | Context | License | On RTX 4090 |
|---|---|---|---|---|---|---|
| #1 | Laguna XS.2 | 33B | 19 GB | 131 072 | Apache 2.0 | 100 tok/s · Q5_K_M |
| #2 | Gemma 4 26B-A4B MoE | 26B | 16 GB | 128 000 | Apache 2.0 | 60 tok/s · Q5_K_M |
| #3 | LLaDA 2.0 Uni 16B | 16B | 18 GB | 8 192 | Apache 2.0 | 130 tok/s · Q5_K_M |
| #4 | Qwen 3.6 27B | 27B | 16 GB | 262 144 | Apache 2.0 | 32 tok/s · Q5_K_M |
| #5 | Qwen 3.8 27B | 27B | 16 GB | 262 144 | Apache 2.0 | 22 tok/s · Q5_K_M |
| #6 | GLM 4.7 Flash | 31B | 19 GB | 128 000 | MIT | 100 tok/s · Q5_K_M |
| #7 | Gemma 4 31B | 31B | 18 GB | 256 000 | Apache 2.0 | 30 tok/s · Q5_K_M |
Your private, free ChatGPT on your machine in 1 hour — LM Studio, Ollama, Open WebUI, your documents, no cloud.
Free memo
Get the memo VRAM → best coding model → Ollama command (one screen, copy and paste). Then switch to the Copilote Local kit for a setup that actually works.
The Local Copilot kit — the Ollama + Cline + Aider configs are ready to paste, with tuned Modelfiles, troubleshooting, and lifetime online access →No spam. Unsubscribe in 1 click. Your data stays with us (never resold).
Your card → the best coding model to run locally, and the exact Ollama command:
| Your VRAM | Typical GPUs / Macs | Recommended coding model | Command Ollama |
|---|---|---|---|
| 8 GB | RTX 4060 / 3060 · M1-M2 16 GB | Qwen 3.5 9B (Q4, 6.6 GB — 256k context) | ollama run qwen3.5:9b |
| 12 GB | RTX 3060 12 GB / 4070 / 5070 | Qwen 3.5 9B (Q8, 11 GB) or Gemma 4 12B (7.6 GB) | ollama run qwen3.5:9b-q8_0 |
| 16 GB | RTX 5070 Ti / 4080 / 5080 · RX 9070 XT · M4 24 GB | Devstral 24B (Q4, 14 GB) — coding-agent specialist | ollama run devstral:24b |
| 24 GB | RTX 3090 / 4090 · RX 7900 XTX · M4 Pro 48 GB | Qwen 3.8 27B (Q4, 18 GB) — the “close to Copilot” option | ollama run qwen3.8:27b |
| 32 GB | RTX 5090 | Qwen 3.6 35B-A3B (Q4, 23 GB) — fast MoE | ollama run qwen3.6:35b |
| 48 GB+ | Mac M4 Max 64 GB · M2 Ultra 128 GB | Qwen3-Coder 30B-A3B (Q8, 32 GB — 256k context) | ollama run qwen3-coder:30b-a3b-q8_0 |
-base: ollama run qwen2.5-coder:7b-base — it’s still the reference for this specific use case. ⚠️ Qwen 3.8: its reasoning is set very high by default and it “overthinks” simple requests — lower it to low (or turn it off) on first launch. ⚠️ License trap: Codestral 22B = Mistral Non-Production License → prohibited for coding at work. Qwen 3.5/3.8, Gemma 4, and Devstral are Apache 2.0. 💡 Running out of memory? Keep ~1.5 GB of VRAM free for context, or drop down one quantization level.🔌 To connect it to VS Code: Cline (multi-file agent), Aider (CLI) or Tabby/Twinny (FIM autocomplete) — they all connect to Ollama locally. The kit Local Copilot — ready-to-paste configs + tested setup — is available: /copilote-local.
We keep models that fit within 24 GB in Q4_K_M and use at least 40% of the VRAM (otherwise a 7B is sufficient). Bonus points go to models whose VRAM fit is between 60% and 95%—the quality/throughput sweet spot.
Criteria considered:
The scoring is fully transparent: see our methodology for details on VRAM/tokens/sec calculations.
Can you run a 70B on RTX 4090?
In Q4_K_M, a 70B model requires ~40 GB of VRAM—too much for a single 4090. You must either drop to Q2/Q3 (quality loss), offload to CPU RAM (very slow), or add a 2nd card. For a true 70B model, aim for 2× RTX 4090 or a 5090 + DDR5.
Which quantization should you choose on RTX 4090?
Q5_K_M is the sweet spot (less than 1% loss versus FP16 according to benchmarks). Q8 is clearly better than Q5 but uses 50% more VRAM. Use Q4 only if you want a large model that won't fit in Q5.
Mistral Small 3.1 24B or Qwen 2.5 32B on a 4090?
View the comparison. Mistral Small 3.1 is faster (24B < 32B) and better in French. Qwen 2.5 32B is more capable for general tasks and code.
Which inference engine for RTX 4090?
For interactive chat: Ollama (simple) or llama.cpp (maximum control). For server throughput: vLLM or ExLlamaV2. The gain can reach 2–3× on vLLM in batch.
Prices in euros (€) are French market prices including VAT, as checked by BestLLMfor. US prices differ: the Amazon buttons show the current US price.