Intermediate 12 minVision

Local multimodal vision LLM: analyze images with Ollama

Describe a photo, extract a table from a screenshot, read the total on a scanned invoice: since 2024, open-weight vision LLMs have held their own against GPT-4o on these tasks, and Ollama exposes them with the same simplicity as text models. This guide shows how to install a local vision LLM with Ollama, call it from Python with a base64 image, and make use of Qwen 3.5, Gemma 4, or Qwen 3.8 depending on your VRAM.

By Mohamed Meguedmi·Update 2026-08-27·Tested on Windows, macOS, and Linux

#Why run a vision LLM locally?

Sending a customer invoice, a medical screenshot, or an HR document to OpenAI raises a confidentiality issue that legal departments are increasingly refusing to accept. A local vision LLM addresses these concrete use cases: OCR for accounting documents, automatic product-catalog descriptions, image moderation, accessibility (alt text), and data extraction from scanned PDFs.

The difference from conventional OCR (Tesseract, PaddleOCR): a vision LLM understands semantics. It can say “the VAT number is FR12345678901,” whereas OCR gives you a jumble of characters that you then have to parse. On structured documents, reliability exceeds 95% in “JSON extraction” mode.

i
What vision means here
A “multimodal vision” LLM accepts images as input in addition to text. It doesn’t generate images (for that, see Stable Diffusion). You give it an image plus a question, and it responds with text.

#Vision models available in Ollama

The Local AI Kit

Your private ChatGPT, free, on your own machine in an hour — LM Studio, Ollama, Open WebUI, your documents, no cloud.

  • Lifetime online access
  • PDF + files
  • Lifetime updates

Ollama has natively supported several vision families since version 0.4 (October 2024), and the 2026 generation has reshuffled the landscape: vision is now integrated into the standard tag for Qwen 3.5 and Gemma 4, so a separate -vl variant is no longer needed. Here are the relevant options in 2026, ranked by French OCR size and quality.

qwen3.5:9b
The best quality/VRAM tradeoff in 2026. Excellent French OCR, understands tables and forms, 256k-token context, Apache 2.0 license. ~6.6 GB of weights in Q4.
qwen3.5:4b
Lightweight version for 4–6 GB of VRAM or a 16 GB Mac M1/M2. Same Qwen 3.5 multimodal family, decent French OCR, somewhat less rich descriptions. ~3.4 GB.
qwen3.8:27b
For RTX 3090/4090/5090 (24 GB): quality close to GPT-4o for complex document analysis. Vision, 262k context, Apache 2.0. ~18 GB of weights in Q4.
gemma4:12b
The Google alternative (Gemma 4, April 2026), multimodal and now under Apache 2.0. Very good at general description, slightly less sharp for pure OCR. ~7.6 GB.
gemma4:26b
Gemma 4 26B-A4B, multimodal MoE. The high-end tier for intensive professional use: ~19 GB of weights, runs on 24 GB of VRAM or a Mac Studio with unified memory.
gemma4:e2b-it-qat
Quantized QAT Gemma 4 E2B (~4.3 GB). A compact option for limited VRAM or edge deployments, under the Apache 2.0 license.
qwen3.5:2b
Qwen 3.5 2B multimodal (~1.9 GB, 256k context). For Raspberry Pi 5 or large-scale rapid descriptions. No serious OCR for complex documents.
llava:7b / llava:13b
The historical ancestor from 2023. No longer deploy it: its quality is far surpassed by Qwen 3.5 and Gemma 4 in 2026.
→
Default recommendation
If you have 8 GB of VRAM or more, start with qwen3.5:9b. It's the best quality-to-size ratio for French in 2026, vision is included in the standard tag, and the Ollama community documents its use cases well.

#Hardware requirements

The VRAM required for a vision LLM is slightly higher than for its text equivalent because the vision encoder (often a ViT) remains loaded in memory alongside the language model.

8 GB VRAM (RTX 3060 12GB, 4060, M1/M2 16GB)
qwen3.5:4b is comfortable, or qwen3.5:9b (6.6 GB), which just fits at Q4.
12 GB VRAM (RTX 3060 12GB, 4070, 5070)
qwen3.5:9b comfortable, gemma4:12b in Q4_K_M.
16 GB VRAM (RTX 4070 Ti Super, 4080, 5070 Ti)
Everything above plus 32k-token context, or qwen3.5:9b-q8_0 for maximum quality in this range.
24 GB VRAM (RTX 3090, 4090, 5090)
qwen3.8:27b (~18 GB) or gemma4:26b, with near-cloud quality.
M-series Mac with 24–48 GB unified memory
qwen3.5:9b or qwen3.8:27b thanks to unified memory. Prefer M3/M4 for bandwidth.

Ollama must be at least version 0.4 for modern vision models. Check with ollama --version. If you are below that, update before going further.

#1. Install a vision model

Installation is identical to that of a text model: ollama pull t downloads the weights and vision encoder in a single command.

Terminal — install Qwen 3.5 9B
ollama pull qwen3.5:9b

The download is approximately 6.6 GB (Q4 LM weights + ViT encoder). Allow 2–3 minutes on fiber. Then verify that the model is listed:

Terminal
ollama list

# NAME            ID            SIZE      MODIFIED
# qwen3.5:9b      abc123...     6.6 GB    2 minutes ago
!
Vision is included in the standard tag
Good news since Qwen 3.5 and Gemma 4: vision is integrated into the default tag (qwen3.5:9b, gemma4:12b). The old-generation trap is gone: you no longer need a separate -vl tag (qwen2.5vl vs qwen2.5). Just make sure you have pulled a recent multimodal model rather than an old text-only model.

#2. First CLI test

Ollama accepts images directly in the CLI starting with version 0.4: pass the file path after your prompt.

Describe an image in French
ollama run qwen3.5:9b "Décris cette image en 3 phrases en français." ./photo.jpg

The model loads the image, encodes it through ViT, and responds in streaming mode. On a RTX 4070, allow 2–4 seconds for a short description and 8–12 seconds for a detailed analysis.

Typical output
L'image montre un chat tigré assis sur un rebord de fenêtre en bois clair.
Le pelage présente des rayures grises et noires caractéristiques d'un tabby.
À l'arrière-plan, on devine un jardin flou avec de la végétation verte.
→
Multiple images at once
You can pass multiple paths: ollama run qwen3.5:9b "Compare these two images" photo1.jpg photo2.jpg. Useful for visual diffing or deduplication.

#3. Python API with a base64 image

To integrate a vision LLM into an app, Ollama exposes its REST endpoint at http://localhost:11434. The image is passed either as a path (Python binding) or as base64 (raw HTTP).

Recommended method: the official ollama-python library. It handles base64 encoding for you and accepts a path directly.

Installation
pip install ollama pillow
vision_local.py — simple call
import ollama

response = ollama.chat(
    model='qwen3.5:9b',
    messages=[{
        'role': 'user',
        'content': 'Quel est le texte visible sur cette image ? Réponds en français.',
        'images': ['./screenshot.png'],
    }],
)

print(response['message']['content'])

If you prefer to handle base64 yourself (an image in memory, an S3 blob, or an upload through FastAPI), here's the explicit version:

With explicit base64
import base64
import requests

with open('facture.png', 'rb') as f:
    img_b64 = base64.b64encode(f.read()).decode('utf-8')

response = requests.post(
    'http://localhost:11434/api/chat',
    json={
        'model': 'qwen3.5:9b',
        'messages': [{
            'role': 'user',
            'content': 'Décris cette image.',
            'images': [img_b64],
        }],
        'stream': False,
    },
    timeout=120,
)

print(response.json()['message']['content'])
i
Expected base64 format
Ollama accepts raw base64, without the data:image/png;base64, prefix. If you retrieve the image from a browser, remember to strip this prefix before the call.

#4. Practical case: OCR of a French invoice

Structured extraction from a scanned document is the most requested business use case. The pattern that works best: ask for JSON output with an explicit schema in the prompt.

ocr_facture.py
import ollama
import json

SYSTEM = '''Tu es un assistant d'extraction de données comptables.
Tu réponds UNIQUEMENT en JSON valide, sans markdown, sans commentaire.'''

PROMPT = '''Extrais les informations de cette facture française au format JSON :
{
  "fournisseur": str,
  "siret": str ou null,
  "numero_facture": str,
  "date": "YYYY-MM-DD",
  "montant_ht": float,
  "tva": float,
  "montant_ttc": float,
  "lignes": [{"description": str, "quantite": float, "prix_unitaire": float}]
}'''

response = ollama.chat(
    model='qwen3.5:9b',
    messages=[
        {'role': 'system', 'content': SYSTEM},
        {'role': 'user', 'content': PROMPT, 'images': ['./facture.png']},
    ],
    format='json',
    options={'temperature': 0.1},
)

data = json.loads(response['message']['content'])
print(f"Fournisseur : {data['fournisseur']}")
print(f"Montant TTC : {data['montant_ttc']} €")

Two tricks that dramatically improve reliability: use format='json' (Ollama then forces valid JSON output) and lower the temperature to 0.1 to limit hallucinations in amounts. On 100 varied French invoices, Qwen 3.5 9B typically achieves 92-96% correct extraction of total amount including tax + date + vendor.

!
Always verify the amounts
Even at 96%, you will have 4 errors per 100 invoices. In production, systematically validate that montant_ht + tva ≈ montant_ttc and flag discrepancies for human review. An LLM is not a certified accountant.

#Qwen 3.5 vs. Gemma 4: which should you choose?

The two families dominate the open-weight vision scene in 2026. Our hands-on comparison on French-language use cases, between qwen3.5:9b and gemma4:12b:

French OCR (invoices, contracts)
Qwen 3.5 clearly wins. Gemma 4 tends to "make up" numbers when working with blurry documents. Qwen advantage: +15% accuracy.
General image description
A tie. Gemma 4 gives more narrative descriptions, while Qwen 3.5 is more factual. A matter of taste.
Table comprehension
Qwen 3.5 excels. It is one of the few models of its size that can read an Excel pivot table without mixing it up.
Multilingual captioning (FR/EN/ES)
Qwen 3.5 was trained on more languages. Gemma 4 is a notch better in pure English.
Latency on RTX 4070
Comparable: 2–4 seconds for a short response. Gemma 4 12B is 10–15% faster on first inference (lighter vision encoder).
License
Both use Apache 2.0 (free commercial use): Gemma 4 switched to Apache 2.0 in April 2026, aligning Google with Qwen. There are no more MAU restrictions to monitor.
→
Practical verdict
For 90% of French-language use cases (OCR, extraction, description), Qwen 3.5 9B is the default choice. Gemma 4 12B is still a good fit if you're already in the Google/Gemma ecosystem (fine-tuning, deployment) or for mainstream English captioning. Need the best document quality on 24 GB? Move up to qwen3.8:27b.

#Troubleshooting

"Error: model does not support images"
You are calling a text model. Make sure you have pulled a recent multimodal model (qwen3.5, gemma4, or the older llava). ollama list must display the right model.
The model describes something other than the image
The base64 is probably corrupted or contains the data:image prefix. Strip data:image/png;base64, before the call. Test in the CLI first to isolate the problem.
VRAM saturated during loading
The vision encoder is added to the LM weights. Switch to more aggressive quantization (q3_K_M) or reduce OLLAMA_NUM_CTX to 4096 in the environment.
OCR misses French accents
Explicitly specify “Respond in French while preserving accents” in the system prompt. Otherwise, the model may sometimes strip accents (e instead of é).
Response time > 30 s
A 4000x3000 image sends the number of visual tokens soaring. Resize it to 1024x1024 max before sending: 90% of the OCR quality for 10% of the compute.
Resize before inference
from PIL import Image

img = Image.open('facture_4k.png')
img.thumbnail((1024, 1024), Image.LANCZOS)
img.save('facture_resized.png', optimize=True)

#Go further

You have a working local vision LLM. Three natural directions for going further:

Integrate into a complete Python app
The guide Integrate Ollama into a Python application via the REST API covers FastAPI, streaming, and function calling—directly applicable to vision models.
Build a document extraction pipeline
The Compta guide, “Invoice extraction,” shows the complete industrialization process (job queue, validation, ERP export) using a local vision LLM.
Choose the right GPU for your workload
The Choosing Your GPU for Local AI guide covers the 12/16/24 GB VRAM tiers and their impact on 4B/9B/27B vision models.
Did this guide help you?

Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.