🇺🇸 Laguna XS.2
MoE 33B/3B active parameters, Apache 2.0, specializing in agentic coding. 68.2% SWE-Bench Verified, 128k ctx. Runs on a 36 GB Mac. Released April 28, 2026.
ollama run laguna-xs.2
Ranking updated on 09/10/2026
The Mac Studio (M2 / M3 / M5 Ultra, up to 512 GB of unified memory, 800 GB/s to 1.2 TB/s) is the most capable consumer workstation for local AI. 70B in Q5, 200B in Q4, frontier 670B in Q3.
Mac Studio : purchasing alternative available for local AI — Mac Studio M5 Max (36 GB / 512 GB) (Other memory tiers: the Apple Store BTO configurator.):
A mini PC is a complete machine: check the required memory and software compatibility. It does not replace macOS/MLX or CUDA.
Why this choice? Our complete guide to Mac Studio M5 Max (36 GB / 512 GB) →
Which PC should you choose for your budget? Our picks from €800 to €3,500 →
Affiliate links — QuelLLM may earn a commission on purchases at no extra cost to you, which does not influence the ranking (established independently). As an Amazon Associate, BestLLMfor earns from qualifying purchases.
MoE 33B/3B active parameters, Apache 2.0, specializing in agentic coding. 68.2% SWE-Bench Verified, 128k ctx. Runs on a 36 GB Mac. Released April 28, 2026.
ollama run laguna-xs.2
GLM-4.7-Flash (MoE 31B, ~3B active): the best code/VRAM ratio in the 30B class. MIT, 128k ctx, very fast on 3090/4090.
ollama run glm-4.7-flash
Mamba-2 + MoE 32B/9B hybrid. ~70% less RAM in long contexts. Apache 2.0.
ollama run granite4:small-h
MoE with 35B/3B active parameters for agentic coding. 73.4% SWE-Bench. Release: April 16, 2026.
ollama run qwen3.6:35b-a3b
MoE 30B/3B active hybrid thinking. MMLU 81.4, AIME24 80.4. 100+ languages.
ollama run qwen3:30b-a3b
MoE with 30B/3B active: thinking mode + instruct. Gold medalist at IMO 2025 and IOI 2025. Fast inference thanks to the 3B active parameters, with 30B-level reasoning capabilities. Released April 2026.
ollama run nemotron-cascade-2
MoE 30B (3.3B active parameters) specialized in agentic coding. Very fast locally, native 256k ctx, the benchmark for 16–24 GB via Ollama.
ollama run qwen3-coder:30b
Vision MoE with 30B/3B active. Vision sweet spot Qwen 3. 256k ctx.
ollama run qwen3-vl:30b
| Rank | Model | Params | Q4 VRAM | Context | License | On Apple M2 Ultra (128 GB) |
|---|---|---|---|---|---|---|
| #1 | Laguna XS.2 | 33B | 19 GB | 131 072 | Apache 2.0 | 100 tok/s · FP16 |
| #2 | GLM 4.7 Flash | 31B | 19 GB | 128 000 | MIT | 100 tok/s · FP16 |
| #3 | Granite 4.0 H-Small 32B-A9B | 32B | 19 GB | 128 000 | Apache 2.0 | 75 tok/s · FP16 |
| #4 | Qwen 3.6 35B-A3B | 35B | 21 GB | 262 000 | Apache 2.0 | 60 tok/s · FP16 |
| #5 | Qwen 3 30B-A3B | 30B | 19 GB | 131 072 | Apache 2.0 | 100 tok/s · FP16 |
| #6 | Nemotron Cascade 2 30B-A3B | 30B | 17 GB | 128 000 | NVIDIA Open Model License | 80 tok/s · FP16 |
| #7 | Qwen3-Coder 30B-A3B | 30B | 19 GB | 262 144 | Apache 2.0 | 100 tok/s · FP16 |
| #8 | Qwen 3 VL 30B-A3B | 30B | 19 GB | 262 144 | Apache 2.0 | 100 tok/s · FP16 |
Local AI on your Mac, fully explored: unified memory, MLX vs. GGUF, the right model for your chip, Ollama and LM Studio tuned for Apple Silicon.
Free memo
Get the memo VRAM → best coding model → Ollama command (one screen, copy and paste). Then switch to the Copilote Local kit for a setup that actually works.
The Local Copilot kit — the Ollama + Cline + Aider configs are ready to paste, with tuned Modelfiles, troubleshooting, and lifetime online access →No spam. Unsubscribe in 1 click. Your data stays with us (never resold).
Your card → the best coding model to run locally, and the exact Ollama command:
| Your VRAM | Typical GPUs / Macs | Recommended coding model | Command Ollama |
|---|---|---|---|
| 8 GB | RTX 4060 / 3060 · M1-M2 16 GB | Qwen 3.5 9B (Q4, 6.6 GB — 256k context) | ollama run qwen3.5:9b |
| 12 GB | RTX 3060 12 GB / 4070 / 5070 | Qwen 3.5 9B (Q8, 11 GB) or Gemma 4 12B (7.6 GB) | ollama run qwen3.5:9b-q8_0 |
| 16 GB | RTX 5070 Ti / 4080 / 5080 · RX 9070 XT · M4 24 GB | Devstral 24B (Q4, 14 GB) — coding-agent specialist | ollama run devstral:24b |
| 24 GB | RTX 3090 / 4090 · RX 7900 XTX · M4 Pro 48 GB | Qwen 3.8 27B (Q4, 18 GB) — the “close to Copilot” option | ollama run qwen3.8:27b |
| 32 GB | RTX 5090 | Qwen 3.6 35B-A3B (Q4, 23 GB) — fast MoE | ollama run qwen3.6:35b |
| 48 GB+ | Mac M4 Max 64 GB · M2 Ultra 128 GB | Qwen3-Coder 30B-A3B (Q8, 32 GB — 256k context) | ollama run qwen3-coder:30b-a3b-q8_0 |
-base : ollama run qwen2.5-coder:7b-base — it’s still the reference for this specific use case. ⚠️ Qwen 3.8: its reasoning is set very high by default and it “overthinks” simple requests — lower it to low (or turn it off) on first launch. ⚠️ License trap: Codestral 22B = Mistral Non-Production License → prohibited for coding at work. Qwen 3.5/3.8, Gemma 4, and Devstral are Apache 2.0. 💡 Running out of memory? Keep ~1.5 GB of VRAM free for context, or drop down one quantization level.🔌 To connect it to VS Code: Cline (multi-file agent), Aider (CLI) or Tabby/Twinny (FIM autocomplete) — they all connect to Ollama locally. The kit Local Copilot — ready-to-paste configs + tested setup — is available: /copilote-local.
Filter: 7–700B (frontier MoE models allowed). Big bonus for 30–200B (peak Studio Ultra) and MoE in general: 800+ GB/s bandwidth really leverages these models.
Criteria considered:
The scoring is fully transparent: see our methodology for details on VRAM/tokens/sec calculations.
Mac Studio M2 Ultra 192 GB: Llama 70B running smoothly?
Yes—Llama 3.3 70B runs at approximately 11 tok/s in Q5_K_M (~48 GB) and 14 tok/s in Q4 (estimates: approximately 70% of the ceiling set by its 800 GB/s). It was the first Apple hardware capable of running a 70B comfortably locally. See the Mac Studio guide.
Mac Studio Ultra 512 GB: DeepSeek 671B?
Yes — DeepSeek R1 671B (37B active parameters, MoE) fits in Q4_K_M (~400 GB) on an M3 Ultra or an M5 Ultra with 512 GB, at a usable speed because only 37B parameters are read per token. No other desktop machine can do this. See DeepSeek R1 671B.
Mac Studio vs. 4× H100 server?
4× H100 (320 GB HBM3) costs ~120 000 € + 2 kW power. A 512 GB Mac Studio costs more than 10 000 €, at approximately 200 W. The H100 is ~5-10× faster in throughput, but the Studio wins on €/GB of memory and silence.
Is MLX mandatory on Studio Ultra?
Recommended. MLX makes better use of unified memory and is often faster than llama.cpp on large models. Ollama (llama.cpp Metal) works, but slightly underutilizes the machine.