🇨🇳 GLM 4.7 Flash
GLM-4.7-Flash (MoE 31B, ~3B active): the best code/VRAM ratio in the 30B class. MIT, 128k ctx, very fast on 3090/4090.
ollama run glm-4.7-flash
Ranking updated on 09/10/2026
64 GB of unified memory is the practical 70B threshold. Llama 3.3 70B Q4_K_M fits in ~40 GB, including 32k context. This is the first threshold where it competes with a local multi-GPU RTX 4090.
Compare prices for MacBook Pro M5 Pro — 24 GB / 1 TB from our partner retailers (verified product pages):
Which PC should you choose for your budget? Our picks from €800 to €3,500 →
Affiliate links — QuelLLM may earn a commission on purchases at no extra cost to you, which does not influence the ranking (established independently). As an Amazon Associate, BestLLMfor earns from qualifying purchases.
GLM-4.7-Flash (MoE 31B, ~3B active): the best code/VRAM ratio in the 30B class. MIT, 128k ctx, very fast on 3090/4090.
ollama run glm-4.7-flash
MoE 33B/3B active parameters, Apache 2.0, specializing in agentic coding. 68.2% SWE-Bench Verified, 128k ctx. Runs on a 36 GB Mac. Released April 28, 2026.
ollama run laguna-xs.2
Mamba-2 + MoE 32B/9B hybrid. ~70% less RAM in long contexts. Apache 2.0.
ollama run granite4:small-h
MoE 30B/3B active hybrid thinking. MMLU 81.4, AIME24 80.4. 100+ languages.
ollama run qwen3:30b-a3b
MoE with 30B/3B active: thinking mode + instruct. Gold medalist at IMO 2025 and IOI 2025. Fast inference thanks to the 3B active parameters, with 30B-level reasoning capabilities. Released April 2026.
ollama run nemotron-cascade-2
MoE 30B (3.3B active parameters) specialized in agentic coding. Very fast locally, native 256k ctx, the benchmark for 16–24 GB via Ollama.
ollama run qwen3-coder:30b
Vision MoE with 30B/3B active. Vision sweet spot Qwen 3. 256k ctx.
ollama run qwen3-vl:30b
Korean agentic MoE with 30B/3B active parameters. Covers KR/EN/JP/ZH/TH/VI. Apache 2.0. MLA attention.
ollama pull hf.co/kakaoai/Kanana-2-30B-GGUF
| Rank | Model | Params | Q4 VRAM | Context | License | On Apple M4 Max (64 GB) |
|---|---|---|---|---|---|---|
| #1 | GLM 4.7 Flash | 31B | 19 GB | 128 000 | MIT | 40 tok/s · Q8 |
| #2 | Laguna XS.2 | 33B | 19 GB | 131 072 | Apache 2.0 | 40 tok/s · Q8 |
| #3 | Granite 4.0 H-Small 32B-A9B | 32B | 19 GB | 128 000 | Apache 2.0 | 30 tok/s · Q8 |
| #4 | Qwen 3 30B-A3B | 30B | 19 GB | 131 072 | Apache 2.0 | 40 tok/s · Q8 |
| #5 | Nemotron Cascade 2 30B-A3B | 30B | 17 GB | 128 000 | NVIDIA Open Model License | 30 tok/s · Q8 |
| #6 | Qwen3-Coder 30B-A3B | 30B | 19 GB | 262 144 | Apache 2.0 | 40 tok/s · Q8 |
| #7 | Qwen 3 VL 30B-A3B | 30B | 19 GB | 262 144 | Apache 2.0 | 40 tok/s · Q8 |
| #8 | Kanana 2 30B-A3B Thinking | 30B | 18 GB | 131 072 | Apache 2.0 | 40 tok/s · Q8 |
Local AI on your Mac, fully explored: unified memory, MLX vs. GGUF, the right model for your chip, Ollama and LM Studio tuned for Apple Silicon.
Free memo
Get the memo VRAM → best coding model → Ollama command (one screen, copy and paste). Then switch to the Copilote Local kit for a setup that actually works.
The Local Copilot kit — the Ollama + Cline + Aider configs are ready to paste, with tuned Modelfiles, troubleshooting, and lifetime online access →No spam. Unsubscribe in 1 click. Your data stays with us (never resold).
Your card → the best coding model to run locally, and the exact Ollama command:
| Your VRAM | Typical GPUs / Macs | Recommended coding model | Command Ollama |
|---|---|---|---|
| 8 GB | RTX 4060 / 3060 · M1-M2 16 GB | Qwen 3.5 9B (Q4, 6.6 GB — 256k context) | ollama run qwen3.5:9b |
| 12 GB | RTX 3060 12 GB / 4070 / 5070 | Qwen 3.5 9B (Q8, 11 GB) or Gemma 4 12B (7.6 GB) | ollama run qwen3.5:9b-q8_0 |
| 16 GB | RTX 5070 Ti / 4080 / 5080 · RX 9070 XT · M4 24 GB | Devstral 24B (Q4, 14 GB) — coding-agent specialist | ollama run devstral:24b |
| 24 GB | RTX 3090 / 4090 · RX 7900 XTX · M4 Pro 48 GB | Qwen 3.8 27B (Q4, 18 GB) — the “close to Copilot” option | ollama run qwen3.8:27b |
| 32 GB | RTX 5090 | Qwen 3.6 35B-A3B (Q4, 23 GB) — fast MoE | ollama run qwen3.6:35b |
| 48 GB+ | Mac M4 Max 64 GB · M2 Ultra 128 GB | Qwen3-Coder 30B-A3B (Q8, 32 GB — 256k context) | ollama run qwen3-coder:30b-a3b-q8_0 |
-base : ollama run qwen2.5-coder:7b-base — it’s still the reference for this specific use case. ⚠️ Qwen 3.8: its reasoning is set very high by default and it “overthinks” simple requests — lower it to low (or turn it off) on first launch. ⚠️ License trap: Codestral 22B = Mistral Non-Production License → prohibited for coding at work. Qwen 3.5/3.8, Gemma 4, and Devstral are Apache 2.0. 💡 Running out of memory? Keep ~1.5 GB of VRAM free for context, or drop down one quantization level.🔌 To connect it to VS Code: Cline (multi-file agent), Aider (CLI) or Tabby/Twinny (FIM autocomplete) — they all connect to Ollama locally. The kit Local Copilot — ready-to-paste configs + tested setup — is available: /copilote-local.
Filter: 7–100B models whose Q4_K_M fits under 48 GB (leaving 16 GB for macOS + context). Bonus: 30–70B (64 GB peak) and MoE.
Criteria considered:
The scoring is fully transparent: see our methodology for details on VRAM/tokens/sec calculations.
64 GB Mac: Llama 70B Q4 runs smoothly?
On M3/M4 Max (400–546 GB/s), yes: 12–18 tokens/sec on Llama 3.3 70B Q4_K_M (~40 GB). On M1/M2 Max (200–400 GB/s), 8–12 tokens/sec—usable but slower. See MBP M4 Max.
64 GB: Llama 70B or Mistral Large 123B?
Llama 70B Q4 (~40 GB) runs smoothly. Mistral Large 123B Q4 (~68 GB) does not fit in 64 GB—you need 96 GB+. Prefer Llama 70B or Mistral Small 3.2 24B Q8 (~26 GB) for high-quality dense models.
Mac 64 GB vs. 2× RTX 3090 (48 GB total VRAM)?
2× 3090 = ~3× faster (936 GB/s per card vs 400 GB/s unified). But a 64 GB Mac = silence + portability + zero cabling. For personal use, the Mac wins on convenience. For real-time professional use, 2× 3090 wins on throughput.
70B MoE on 64 GB?
Yes: Mixtral 8x7B Q4 (~28 GB) or DeepSeek V4 Flash 284B (37B active MoE) Q3_K_S (~140 GB) does NOT fit in 64 GB—you need a Mac Studio with 192+ GB. See Mac Studio.