🇨🇳 Qwen 3 VL 30B-A3B
Vision MoE with 30B/3B active. Vision sweet spot Qwen 3. 256k ctx.
ollama run qwen3-vl:30b
Ranking updated on 09/10/2026
Ranking of open-weight LLMs capable of analyzing images as input: OCR, description, VQA, chart analysis, and data extraction from screenshots. All can be self-hosted.
Vision MoE with 30B/3B active. Vision sweet spot Qwen 3. 256k ctx.
ollama run qwen3-vl:30b
MoE variant of Gemma 4. 26B/4B active. Full multimodal (text+image+audio).
ollama run gemma4:26b
First open Apache 2.0 dLLM: MoE 16B/1B + 6.2B diffusion decoder. Unified text+vision. Released April 22, 2026.
# HuggingFace : inclusionAI/LLaDA2.0-Uni (Flash Attn 2 + CUDA 12.4 requis)
Dense multimodal 27B released April 22, 2026. 262k ctx (1M YaRN). SWE-bench Verified 77.2%.
ollama run qwen3.6:27b
Qwen 3.8 27B: dense multimodal (text + vision), 262k context, ~16 GB Q4 VRAM (18 GB of Ollama weights). Apache 2.0, agentic coding and vision.
ollama run qwen3.8:27b
Dense 31B multimodal (text+image+audio). 140+ languages, 256k context. #3 open model on Chatbot Arena.
ollama run gemma4:31b
Dense 8B vision Qwen 3. Best small VLM Qwen generation 3.
ollama run qwen3-vl:8b
| Rank | Model | Params | Q4 VRAM | Context | License |
|---|---|---|---|---|---|
| #1 | Qwen 3 VL 30B-A3B | 30B | 19 GB | 262 144 | Apache 2.0 |
| #2 | Gemma 4 26B-A4B MoE | 26B | 16 GB | 128 000 | Apache 2.0 |
| #3 | LLaDA 2.0 Uni 16B | 16B | 18 GB | 8 192 | Apache 2.0 |
| #4 | Qwen 3.6 27B | 27B | 16 GB | 262 144 | Apache 2.0 |
| #5 | Qwen 3.8 27B | 27B | 16 GB | 262 144 | Apache 2.0 |
| #6 | Gemma 4 31B | 31B | 18 GB | 256 000 | Apache 2.0 |
| #7 | Qwen 3 VL 8B | 8B | 6 GB | 262 144 | Apache 2.0 |
You chose your vision model. The Local AI Kit devotes an entire chapter to running it properly at home: chatting with images locally, step by step (ch. 7).
Filter by the “vision” tag. The score favors newer models (VLMs evolve very quickly) and larger sizes (which render fine details better).
Criteria considered:
The scoring is fully transparent: see our methodology for details on VRAM/tokens/sec calculations.
What is the best open-source LLM for analyzing images?
Qwen 3 VL 30B-A3B is our #1 for 30B parameters. For a lighter setup, Qwen 2.5 VL 7B runs very well with 8–12 GB of VRAM.
Can these models perform OCR?
Yes—the latest VLMs can read text in images (scanned documents, screenshots). For large-scale or industrial OCR, DeepSeek OCR is specialized for the task. For handwritten French text, prefer Qwen 2.5 VL, which handles French well.
Should you use Ollama or llama.cpp for vision?
Both support vision models (Llama 3.2 Vision, LLaVA, Qwen VL). LM Studio does too. Use the API via -images or the parameter images in Ollama.
How much VRAM for a capable VLM?
8 GB for a 7B VL (Qwen 2.5 VL 7B), 12-16 GB for an 11B (Llama 3.2 Vision), and 40+ GB for large 72B models. Vision adds ~1-2 GB of VRAM on top of the text model.
Learn more with our detailed head-to-head matchups of the finalists: