Home › Catalog › Best local LLM with vision in 2026

Best local LLM with vision in 2026

◆ Local AI — Your private ChatGPT, free, on your own machine, in an hour · $24 · or all kits $49 →

Ranking updated on 09/10/2026

Ranking of open-weight LLMs capable of analyzing images as input: OCR, description, VQA, chart analysis, and data extraction from screenshots. All can be self-hosted.

Ranking

1

🇨🇳 Qwen 3 VL 30B-A3B

Alibaba · 30B parameters · Apache 2.0 · 262,144-token context

Vision MoE with 30B/3B active. Vision sweet spot Qwen 3. 256k ctx.

Why this ranking Supports image inputs natively. 30B parameters for detailed analysis.
ollama run qwen3-vl:30b
Q4 VRAM
19 GB
35 GB in Q8
2

🇺🇸 Gemma 4 26B-A4B MoE

Google · 26B parameters · Apache 2.0 · 128,000 tokens ctx

MoE variant of Gemma 4. 26B/4B active. Full multimodal (text+image+audio).

Why this ranking Supports image inputs natively. 26B parameters for accurate analysis.
ollama run gemma4:26b
Q4 VRAM
16 GB
28 GB in Q8
3

🇨🇳 LLaDA 2.0 Uni 16B

Ant Group / inclusionAI · 16B parameters · Apache 2.0 · 8,192-token context

First open Apache 2.0 dLLM: MoE 16B/1B + 6.2B diffusion decoder. Unified text+vision. Released April 22, 2026.

Why this ranking Supports image inputs natively. 16B parameters for solid analysis.
# HuggingFace : inclusionAI/LLaDA2.0-Uni (Flash Attn 2 + CUDA 12.4 requis)
Q4 VRAM
18 GB
30 GB in Q8
4

🇨🇳 Qwen 3.6 27B

Alibaba · 27B parameters · Apache 2.0 · 262,144-token context

Dense multimodal 27B released April 22, 2026. 262k ctx (1M YaRN). SWE-bench Verified 77.2%.

Why this ranking Supports image inputs natively. 27B parameters for adequate analysis.
ollama run qwen3.6:27b
Q4 VRAM
16 GB
29 GB in Q8
5

🇨🇳 Qwen 3.8 27B

Alibaba · 27B parameters · Apache 2.0 · 262,144-token context

Qwen 3.8 27B: dense multimodal (text + vision), 262k context, ~16 GB Q4 VRAM (18 GB of Ollama weights). Apache 2.0, agentic coding and vision.

Why this ranking Supports image inputs natively. 27B parameters for adequate analysis.
ollama run qwen3.8:27b
Q4 VRAM
16 GB
29 GB in Q8
6

🇺🇸 Gemma 4 31B

Google · 31B parameters · Apache 2.0 · 256,000-token context

Dense 31B multimodal (text+image+audio). 140+ languages, 256k context. #3 open model on Chatbot Arena.

Why this ranking Supports image inputs natively. 31B parameters for detailed analysis.
ollama run gemma4:31b
Q4 VRAM
18 GB
33 GB in Q8
7

🇨🇳 Qwen 3 VL 8B

Alibaba · 8B parameters · Apache 2.0 · 262,144-token context

Dense 8B vision Qwen 3. Best small VLM Qwen generation 3.

Why this ranking Natively supports image inputs. 8B parameters for adequate analysis.
ollama run qwen3-vl:8b
Q4 VRAM
6 GB
10 GB in Q8

Comparison table

Rank Model Params Q4 VRAM Context License
#1 Qwen 3 VL 30B-A3B 30B 19 GB 262 144 Apache 2.0
#2 Gemma 4 26B-A4B MoE 26B 16 GB 128 000 Apache 2.0
#3 LLaDA 2.0 Uni 16B 16B 18 GB 8 192 Apache 2.0
#4 Qwen 3.6 27B 27B 16 GB 262 144 Apache 2.0
#5 Qwen 3.8 27B 27B 16 GB 262 144 Apache 2.0
#6 Gemma 4 31B 31B 18 GB 256 000 Apache 2.0
#7 Qwen 3 VL 8B 8B 6 GB 262 144 Apache 2.0
The Local AI Kit

You chose your vision model. The Local AI Kit devotes an entire chapter to running it properly at home: chatting with images locally, step by step (ch. 7).

  • Lifetime online access
  • PDF + files
  • Lifetime updates

Ranking methodology

Filter by the “vision” tag. The score favors newer models (VLMs evolve very quickly) and larger sizes (which render fine details better).

Criteria considered:

  • Image understanding
  • OCR / text extraction
  • VQA (image questions)
  • Ollama / llama.cpp compatible

The scoring is fully transparent: see our methodology for details on VRAM/tokens/sec calculations.

Frequently asked questions

What is the best open-source LLM for analyzing images?

Qwen 3 VL 30B-A3B is our #1 for 30B parameters. For a lighter setup, Qwen 2.5 VL 7B runs very well with 8–12 GB of VRAM.

Can these models perform OCR?

Yes—the latest VLMs can read text in images (scanned documents, screenshots). For large-scale or industrial OCR, DeepSeek OCR is specialized for the task. For handwritten French text, prefer Qwen 2.5 VL, which handles French well.

Should you use Ollama or llama.cpp for vision?

Both support vision models (Llama 3.2 Vision, LLaVA, Qwen VL). LM Studio does too. Use the API via -images or the parameter images in Ollama.

How much VRAM for a capable VLM?

8 GB for a 7B VL (Qwen 2.5 VL 7B), 12-16 GB for an 11B (Llama 3.2 Vision), and 40+ GB for large 72B models. Vision adds ~1-2 GB of VRAM on top of the text model.

Head-to-head comparisons

Learn more with our detailed head-to-head matchups of the finalists:

Go further

QuelLLM Kits The reference guide by use case
All kits for life — $49