Local multimodal vision LLM: analyze images with Ollama
Describe a photo, extract a table from a screenshot, read the total on a scanned invoice: since 2024, open-weight vision LLMs have held their own against GPT-4o on these tasks, and Ollama exposes them with the same simplicity as text models. This guide shows how to install a local vision LLM with Ollama, call it from Python with a base64 image, and make use of Qwen 3.5, Gemma 4, or Qwen 3.8 depending on your VRAM.
#Why run a vision LLM locally?
Sending a customer invoice, a medical screenshot, or an HR document to OpenAI raises a confidentiality issue that legal departments are increasingly refusing to accept. A local vision LLM addresses these concrete use cases: OCR for accounting documents, automatic product-catalog descriptions, image moderation, accessibility (alt text), and data extraction from scanned PDFs.
The difference from conventional OCR (Tesseract, PaddleOCR): a vision LLM understands semantics. It can say “the VAT number is FR12345678901,” whereas OCR gives you a jumble of characters that you then have to parse. On structured documents, reliability exceeds 95% in “JSON extraction” mode.
#Vision models available in Ollama
Your private ChatGPT, free, on your own machine in an hour — LM Studio, Ollama, Open WebUI, your documents, no cloud.
- Lifetime online access
- PDF + files
- Lifetime updates
Ollama has natively supported several vision families since version 0.4 (October 2024), and the 2026 generation has reshuffled the landscape: vision is now integrated into the standard tag for Qwen 3.5 and Gemma 4, so a separate -vl variant is no longer needed. Here are the relevant options in 2026, ranked by French OCR size and quality.
- qwen3.5:9b
- The best quality/VRAM tradeoff in 2026. Excellent French OCR, understands tables and forms, 256k-token context, Apache 2.0 license. ~6.6 GB of weights in Q4.
- qwen3.5:4b
- Lightweight version for 4–6 GB of VRAM or a 16 GB Mac M1/M2. Same Qwen 3.5 multimodal family, decent French OCR, somewhat less rich descriptions. ~3.4 GB.
- qwen3.8:27b
- For RTX 3090/4090/5090 (24 GB): quality close to GPT-4o for complex document analysis. Vision, 262k context, Apache 2.0. ~18 GB of weights in Q4.
- gemma4:12b
- The Google alternative (Gemma 4, April 2026), multimodal and now under Apache 2.0. Very good at general description, slightly less sharp for pure OCR. ~7.6 GB.
- gemma4:26b
- Gemma 4 26B-A4B, multimodal MoE. The high-end tier for intensive professional use: ~19 GB of weights, runs on 24 GB of VRAM or a Mac Studio with unified memory.
- gemma4:e2b-it-qat
- Quantized QAT Gemma 4 E2B (~4.3 GB). A compact option for limited VRAM or edge deployments, under the Apache 2.0 license.
- qwen3.5:2b
- Qwen 3.5 2B multimodal (~1.9 GB, 256k context). For Raspberry Pi 5 or large-scale rapid descriptions. No serious OCR for complex documents.
- llava:7b / llava:13b
- The historical ancestor from 2023. No longer deploy it: its quality is far surpassed by Qwen 3.5 and Gemma 4 in 2026.
#Hardware requirements
The VRAM required for a vision LLM is slightly higher than for its text equivalent because the vision encoder (often a ViT) remains loaded in memory alongside the language model.
- 8 GB VRAM (RTX 3060 12GB, 4060, M1/M2 16GB)
- qwen3.5:4b is comfortable, or qwen3.5:9b (6.6 GB), which just fits at Q4.
- 12 GB VRAM (RTX 3060 12GB, 4070, 5070)
- qwen3.5:9b comfortable, gemma4:12b in Q4_K_M.
- 16 GB VRAM (RTX 4070 Ti Super, 4080, 5070 Ti)
- Everything above plus 32k-token context, or qwen3.5:9b-q8_0 for maximum quality in this range.
- 24 GB VRAM (RTX 3090, 4090, 5090)
- qwen3.8:27b (~18 GB) or gemma4:26b, with near-cloud quality.
- M-series Mac with 24–48 GB unified memory
- qwen3.5:9b or qwen3.8:27b thanks to unified memory. Prefer M3/M4 for bandwidth.
Ollama must be at least version 0.4 for modern vision models. Check with ollama --version. If you are below that, update before going further.
#1. Install a vision model
Installation is identical to that of a text model: ollama pull t downloads the weights and vision encoder in a single command.
The download is approximately 6.6 GB (Q4 LM weights + ViT encoder). Allow 2–3 minutes on fiber. Then verify that the model is listed:
#2. First CLI test
Ollama accepts images directly in the CLI starting with version 0.4: pass the file path after your prompt.
The model loads the image, encodes it through ViT, and responds in streaming mode. On a RTX 4070, allow 2–4 seconds for a short description and 8–12 seconds for a detailed analysis.
#3. Python API with a base64 image
To integrate a vision LLM into an app, Ollama exposes its REST endpoint at http://localhost:11434. The image is passed either as a path (Python binding) or as base64 (raw HTTP).
Recommended method: the official ollama-python library. It handles base64 encoding for you and accepts a path directly.
If you prefer to handle base64 yourself (an image in memory, an S3 blob, or an upload through FastAPI), here's the explicit version:
#4. Practical case: OCR of a French invoice
Structured extraction from a scanned document is the most requested business use case. The pattern that works best: ask for JSON output with an explicit schema in the prompt.
Two tricks that dramatically improve reliability: use format='json' (Ollama then forces valid JSON output) and lower the temperature to 0.1 to limit hallucinations in amounts. On 100 varied French invoices, Qwen 3.5 9B typically achieves 92-96% correct extraction of total amount including tax + date + vendor.
#Qwen 3.5 vs. Gemma 4: which should you choose?
The two families dominate the open-weight vision scene in 2026. Our hands-on comparison on French-language use cases, between qwen3.5:9b and gemma4:12b:
- French OCR (invoices, contracts)
- Qwen 3.5 clearly wins. Gemma 4 tends to "make up" numbers when working with blurry documents. Qwen advantage: +15% accuracy.
- General image description
- A tie. Gemma 4 gives more narrative descriptions, while Qwen 3.5 is more factual. A matter of taste.
- Table comprehension
- Qwen 3.5 excels. It is one of the few models of its size that can read an Excel pivot table without mixing it up.
- Multilingual captioning (FR/EN/ES)
- Qwen 3.5 was trained on more languages. Gemma 4 is a notch better in pure English.
- Latency on RTX 4070
- Comparable: 2–4 seconds for a short response. Gemma 4 12B is 10–15% faster on first inference (lighter vision encoder).
- License
- Both use Apache 2.0 (free commercial use): Gemma 4 switched to Apache 2.0 in April 2026, aligning Google with Qwen. There are no more MAU restrictions to monitor.
#Troubleshooting
- "Error: model does not support images"
- You are calling a text model. Make sure you have pulled a recent multimodal model (qwen3.5, gemma4, or the older llava). ollama list must display the right model.
- The model describes something other than the image
- The base64 is probably corrupted or contains the data:image prefix. Strip data:image/png;base64, before the call. Test in the CLI first to isolate the problem.
- VRAM saturated during loading
- The vision encoder is added to the LM weights. Switch to more aggressive quantization (q3_K_M) or reduce OLLAMA_NUM_CTX to 4096 in the environment.
- OCR misses French accents
- Explicitly specify “Respond in French while preserving accents” in the system prompt. Otherwise, the model may sometimes strip accents (e instead of é).
- Response time > 30 s
- A 4000x3000 image sends the number of visual tokens soaring. Resize it to 1024x1024 max before sending: 90% of the OCR quality for 10% of the compute.
#Go further
You have a working local vision LLM. Three natural directions for going further:
- Integrate into a complete Python app
- The guide Integrate Ollama into a Python application via the REST API covers FastAPI, streaming, and function calling—directly applicable to vision models.
- Build a document extraction pipeline
- The Compta guide, “Invoice extraction,” shows the complete industrialization process (job queue, validation, ERP export) using a local vision LLM.
- Choose the right GPU for your workload
- The Choosing Your GPU for Local AI guide covers the 12/16/24 GB VRAM tiers and their impact on 4B/9B/27B vision models.
Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.