Intermediate 11 minQwen

Qwen3-VL locally: the leading vision model with Ollama

Qwen3-VL is the strongest open-weight vision-language model to run locally in 2026. With Ollama, you can install it with one command and send it images, screenshots, or scanned documents directly from the terminal or an interface. This guide covers choosing the variant based on your VRAM, step-by-step Qwen3-VL Ollama installation, practical use cases (image description, OCR, and table reading), and its honest limitations compared with cloud models.

By Léa B.·Update 2026-08-06·Tested on Windows, macOS, and Linux

#Why run Qwen3-VL locally

A vision-language model (VLM) accepts both text and images as input. You show it a photo, screenshot, or scan, and ask it a question about it in natural language. Qwen3-VL, developed by Alibaba's Qwen team, has established itself as the open-weight benchmark for this use case: it combines strong visual understanding, robust multilingual OCR (including French), and solid reasoning about what it sees.

The benefit of running it locally is the same as with any self-hosted LLM, but it matters even more here: the images you analyze are often sensitive—work screenshots, administrative documents, invoice photos, and identity documents. With Qwen3-VL on your machine, nothing leaves your disk. No upload to a cloud, no quota, no subscription.

Privacy
The images stay on your machine. Ideal for HR, medical, and legal documents, or any personal scan.
Zero usage cost
No cost per image or token. Process 10 or 10,000 screenshots without a bill that keeps climbing.
Offline
Once the model is downloaded, everything works without an internet connection.
Automatable
Ollama's OpenAI-compatible API lets you script bulk image processing in Python or the shell.

#Requirements and VRAM needed

The Local AI Kit

Your private ChatGPT, free, on your own machine in an hour — LM Studio, Ollama, Open WebUI, your documents, no cloud.

  • Lifetime online access
  • PDF + files
  • Lifetime updates

A VLM is more demanding than a text model of equivalent size: in addition to the language model weights, it loads a vision encoder and must process the “image tokens” generated by each photo. A high-resolution image can represent several thousand tokens, consuming context and memory during inference. Plan for some VRAM headroom beyond the raw weight size.

For Q4_K_M quantization (the best quality/memory compromise for most uses), here are the memory footprint benchmarks by underlying language model size. Add ~1 to 2 GB for the vision encoder and image processing.

~3-4B variant
≈ 2–3 GB of VRAM. Fits comfortably on a RTX 3060 12 GB, or a 16 GB M-series Mac. The entry point for local vision.
~7-8B variant
≈ 5–6 GB of VRAM. Comfortable on RTX 3060/4070 12 GB or RTX 4080 16 GB. The best quality-to-resource ratio for everyday use.
~30–32B variant
≈ 19–20 GB of VRAM. Targets a RTX 4090 24 GB, a professional card, or a Mac with ample unified memory (M4 Pro/Max 32 GB and above).
i
GPU vs. Mac Apple Silicon
On Mac, unified memory serves as both RAM and VRAM: an M4 Pro 24 GB or an M4 Max 48 GB can run large variants that few consumer GPUs can handle. On PC, the GPU's VRAM is what matters — if the model spills over, Ollama offloads part of it to the CPU and throughput collapses.

On the software side, you need Ollama installed (the daemon listens on http://localhost:11434 by default) and a recent version that supports multimodal input. An interface such as Open WebUI or LM Studio makes sending images more convenient than using the command line, but is not required.

#Choose the right variant for your VRAM

Qwen3-VL comes in several sizes. The right approach is not to target the largest variant your machine can technically load, but the one that leaves room for context and image processing. Here's how to decide.

  1. 01
    You have 8–12 GB of VRAM
    Use a 3–4B variant in Q4_K_M. It describes an image accurately, performs OCR on clear text, and responds quickly. Enough for most simple use cases (description, text extraction).
  2. 02
    You have 12–16 GB of VRAM
    Target a 7–8B variant in Q4_K_M or Q5_K_M. That’s the sweet spot: significantly better on complex documents, tables, and visual reasoning while remaining responsive.
  3. 03
    You have 24 GB of VRAM or a well-equipped Mac
    A 30-32B variant in Q4_K_M unlocks the best quality available locally: reliable reading of dense documents, multi-column tables, and fine-grained reasoning over diagrams. This is where the gap with the cloud narrows the most.
  4. 04
    When in doubt
    Start with the 7-8B variant. If the quality is sufficient, you have saved VRAM; if it hits a ceiling on your documents, move up a tier.
→
Quantization matters
For a VLM, stick with Q4_K_M or Q5_K_M. Going too low (Q3 and below) noticeably degrades OCR and the reading of small characters—exactly what you ask a vision model to do. If you're unsure about quantization, a smaller Q4_K_M variant is better than a large Q3 one.

#Install Qwen3-VL with Ollama

Installation takes one command. Ollama downloads the model from its library, including the vision encoder, then makes it available for inference. Replace the tag with the variant selected in the previous step.

Terminal — install and run
# Télécharger la variante (exemple ~8B ; adaptez le tag à votre VRAM)
ollama pull qwen3-vl:8b

# Lancer une session interactive
ollama run qwen3-vl:8b

# Vérifier ce qui est chargé et sur quel matériel (GPU/CPU)
ollama ps

Once you’re in the interactive session, send an image by specifying its path in the prompt. Ollama detects the file path, encodes the image, and adds it to the context. You can then ask as many questions about it as you want.

Session Ollama — analyze an image
>>> Décris cette image en détail : /home/moi/captures/dashboard.png

>>> Quel est le chiffre affiché en haut à droite ?

>>> Y a-t-il une alerte visible ? Résume-la.
!
Recommended absolute path
Depending on the system, a relative path may not be resolved correctly by Ollama. If you get “file not found” even though the image exists, provide the full absolute path. On Windows, escape it or use regular slashes (C:/Users/...).

#Describe images and screenshots

This is the most immediate use case. Qwen3-VL describes the content of a photo, identifies objects, reads text in the image, and reasons about what it observes. It is particularly useful for screenshots—interfaces, dashboards, error messages: it reads the UI, spots elements, and can explain what is displayed.

General description
“What does this image show?” — inventory of the elements, atmosphere, context. Useful for tagging or sorting photos.
Screenshot reading
Interpret a dashboard, transcribe an error message, or explain an unfamiliar interface from a screenshot.
Targeted extraction
“What is the total amount?”, “List the visible names”—the model isolates the requested information instead of describing everything.
Follow-up questions
Once the image is in context, continue asking questions without sending it again. The model retains the visual reference.
→
A precise prompt is better than a vague one
“Describe the image” produces a generic answer. “List only the errors displayed in this screenshot, with their codes” produces a usable result. As with a text LLM, the more tightly scoped the instruction, the better the output—this is even more true in vision, where the model can get distracted by irrelevant details.

#Read scanned documents and tables

OCR and document understanding are where Qwen3-VL really shines and justify moving to a 7-8B variant or larger. It does more than transcribe the text: it understands the structure (headings, columns, boxes), making it possible to extract data from a table or form in a clean format.

Scanned invoice or receipt
Extract the vendor, date, amount excluding/including tax, and line items—and request JSON output that can be used directly.
Image table
Reconstruct a table captured in a photo in Markdown or CSV format, preserving the rows and columns.
Handwritten document
OCR handles printed text easily; handwriting remains more hit-or-miss but works with clear writing.
Administrative form
Identify completed fields and their values, including checked boxes.
Session Ollama — extract a table in Markdown
>>> Voici un tableau scanné : /home/moi/docs/tableau-ventes.png
... Reproduis-le exactement au format Markdown, sans rien ajouter
... ni interpréter. Conserve toutes les colonnes et l'ordre des lignes.
!
Always verify critical figures
No local OCR is 100% reliable. With amounts, references, or numbers, a digit can be misread (an 8 mistaken for a 6, a comma shifted). For accounting or legal use, always review the extracted values—the model accelerates data entry; it is not a source of truth.

#Automate via the API

To process images in bulk—sort a folder of screenshots, extract data from dozens of invoices—the Ollama API lets you script the whole process. Ollama exposes an OpenAI-compatible API on port 11434, with the image transmitted as base64-encoded data. Here is a minimal Python example.

Python — analyze an image by script
import base64
import requests

with open("facture.png", "rb") as f:
    img_b64 = base64.b64encode(f.read()).decode()

resp = requests.post("http://localhost:11434/api/generate", json={
    "model": "qwen3-vl:8b",
    "prompt": "Extrais fournisseur, date et montant TTC au format JSON.",
    "images": [img_b64],
    "stream": False,
})

print(resp.json()["response"])

From there, a simple loop over a folder lets you process an entire batch of images and aggregate the results. Because inference is local, there is no rate limit: the only constraint is your GPU's speed.

→
Request a structured format
For automation, always require an output format (JSON, CSV) in the prompt and specify the expected keys. This avoids having to parse free-form text. Add “reply with JSON only, with no surrounding text” to simplify downstream processing.

#What local vision still can’t do

Let’s be honest: Qwen3-VL running locally is excellent, but it still falls short of the best cloud vision models in some areas, and there are tasks where you need to lower your expectations. Knowing these limitations helps avoid unpleasant surprises.

Small details and low resolution
With a blurry, very dense, or low-resolution image, OCR fails. A clean, properly framed scan makes all the difference. The model does not “guess” what it cannot see clearly.
Precise counting
Precisely counting many identical objects (“how many people are in this photo?”) remains unreliable beyond a few items. This is a known limitation of VLMs.
Fine-grained spatial reasoning
Exact relative positions, measurements, precise geometry: the model gives an idea, not a metrological answer.
Video and real-time
Ollama processes still images. Video analysis requires you to split the video into frames yourself, without native temporal understanding.
Visual hallucinations
Like any LLM, it may claim to see something that isn't there, especially if your question suggests it. Ask neutral, unbiased questions.
i
Local vs. cloud: the real trade-off
Cloud vision models still have the edge on the most difficult cases (very dense documents, complex visual reasoning, counting). But for 90% of common tasks—describing, extracting text, reading a clean table—Qwen3-VL locally gets the job done, for free and without exposing your images. The tradeoff comes down to data sensitivity and the actual difficulty of your images.

#Troubleshooting

“File not found” for an existing image
Relative path resolved incorrectly. Provide the complete absolute path. On Windows, use slashes (C:/...).
Slow or stuttering response
The model is spilling out of VRAM, with part of it running on the CPU. Check with ollama ps; if that's the case, step down one variant or quantization level.
Poor OCR
Image resolution is too low or quantization is too aggressive. Provide a sharper scan and stay at Q4_K_M or higher.
The model ignores the image
You may have accidentally loaded a text variant of Qwen. Make sure the tag contains -vl and that your version of Ollama supports multimodal input.
Out of memory while loading
The variant exceeds your VRAM. Move down to a smaller size — remember to add headroom for the vision encoder and image tokens.
Quick diagnostics
# Le modèle est-il bien chargé et sur GPU ?
ollama ps

# Lister les modèles vision installés
ollama list

# Tester l'API en local
curl http://localhost:11434/api/tags

#Go further

Qwen3-VL uses the same Ollama stack as the rest of your local models. These guides naturally extend your vision setup:

Local multimodal vision LLMs: analyze images with Ollama
An overview of self-hostable VLMs (Qwen-VL, Llama-Vision, Llama 4 Scout) to compare and choose based on your needs.
Install Ollama: Windows, macOS, and Linux
If Ollama is not set up yet, here is the complete installation guide with the GPU prerequisites.
Choose your quantization (Q4, Q5, Q8, FP16)
To determine the right memory footprint for your Qwen3-VL variant based on your VRAM.
Did this guide help you?

Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.