🇺🇸 Laguna XS.2
MoE 33B/3B active parameters, Apache 2.0, specializing in agentic coding. 68.2% SWE-Bench Verified, 128k ctx. Runs on a 36 GB Mac. Released April 28, 2026.
ollama run laguna-xs.2
Ranking updated on 09/10/2026
For commercial use (SaaS, paid product, internal corporate service), only permissive licenses (Apache 2.0, MIT, BSD) are genuinely hassle-free. We exclude community licenses with user thresholds and non-production licenses.
MoE 33B/3B active parameters, Apache 2.0, specializing in agentic coding. 68.2% SWE-Bench Verified, 128k ctx. Runs on a 36 GB Mac. Released April 28, 2026.
ollama run laguna-xs.2
MoE variant of Gemma 4. 26B/4B active. Full multimodal (text+image+audio).
ollama run gemma4:26b
First open Apache 2.0 dLLM: MoE 16B/1B + 6.2B diffusion decoder. Unified text+vision. Released April 22, 2026.
# HuggingFace : inclusionAI/LLaDA2.0-Uni (Flash Attn 2 + CUDA 12.4 requis)
Dense multimodal 27B released April 22, 2026. 262k ctx (1M YaRN). SWE-bench Verified 77.2%.
ollama run qwen3.6:27b
Qwen 3.8 27B: dense multimodal (text + vision), 262k context, ~16 GB Q4 VRAM (18 GB of Ollama weights). Apache 2.0, agentic coding and vision.
ollama run qwen3.8:27b
GLM-4.7-Flash (MoE 31B, ~3B active): the best code/VRAM ratio in the 30B class. MIT, 128k ctx, very fast on 3090/4090.
ollama run glm-4.7-flash
Dense 31B multimodal (text+image+audio). 140+ languages, 256k context. #3 open model on Chatbot Arena.
ollama run gemma4:31b
Dense 30B Apache 2.0, 12 languages including FR, 131k ctx, GQA 32Q/8KV. OpenAI-compatible tool calling. Released April 29, 2026.
ollama run granite4.1:30b
| Rank | Model | Params | Q4 VRAM | Context | License |
|---|---|---|---|---|---|
| #1 | Laguna XS.2 | 33B | 19 GB | 131 072 | Apache 2.0 |
| #2 | Gemma 4 26B-A4B MoE | 26B | 16 GB | 128 000 | Apache 2.0 |
| #3 | LLaDA 2.0 Uni 16B | 16B | 18 GB | 8 192 | Apache 2.0 |
| #4 | Qwen 3.6 27B | 27B | 16 GB | 262 144 | Apache 2.0 |
| #5 | Qwen 3.8 27B | 27B | 16 GB | 262 144 | Apache 2.0 |
| #6 | GLM 4.7 Flash | 31B | 19 GB | 128 000 | MIT |
| #7 | Gemma 4 31B | 31B | 18 GB | 256 000 | Apache 2.0 |
| #8 | Granite 4.1 30B Instruct | 30B | 17 GB | 131 072 | Apache 2.0 |
Your private, free ChatGPT on your machine in 1 hour — LM Studio, Ollama, Open WebUI, your documents, no cloud.
Strict filter: license contains Apache, MIT, or BSD. No Llama Community (700M MAU threshold), no Mistral NPL, no Gemma (a permissive custom license with acceptable-use clauses).
Criteria considered:
The scoring is fully transparent: see our methodology for details on VRAM/tokens/sec calculations.
Is Llama 3.3 70B commercially free?
No, Llama 3.3 is under the “Llama 3.3 Community License”—free unless your product exceeds 700M MAU. For a future consumer product, that's a risk. Prefer Qwen (Apache 2.0) or Mistral (Apache 2.0 for Open models).
And Google’s Gemma?
Gemma is under the "Gemma License" — permissive but with an acceptable-use clause (not for weapons, mass surveillance, etc.). For most B2B use cases, it’s fine, but read carefully if your product involves defense or security.
Can I fine-tune an Apache 2.0 model and resell it?
Yes — Apache 2.0 allows modification and redistribution, including commercial use. You only need to retain the attribution and license in the redistributed weights. Your own training data remains yours.
Qwen comes from China—is there a legal risk?
The Qwen models are published under Apache 2.0 by Alibaba Cloud. The license is valid internationally. For highly sensitive uses (defense, healthcare), some companies prefer Mistral (🇫🇷) or IBM Granite (🇺🇸) for provenance traceability.