Intermediate 12 minEdge

MiniCPM locally (Ollama): multimodal compact

Direct response

MiniCPM-V is OpenBMB's compact vision-language model series, licensed under Apache 2.0. In Ollama, the minicpm-v tag still loads version 2.6 (8 billion, 5.5 GB), minicpm-v4.5 loads an 8-billion-parameter version weighing 6.1 GB, and minicpm-v4.6 loads a model with about 1.3 billion parameters weighing 1.6 GB. Start with 4.6 on a small machine, then compare the results on your images.

MiniCPM is OpenBMB’s family of compact models, including vision-language models designed to run where larger models cannot. The family has changed significantly: the version loaded by the historical Ollama tag is no longer the latest, and a 1.3-billion-parameter model now exists. This page explains which tag to install, what the stated figures mean, and where a general-purpose model’s OCR capabilities stop.

By Samir K.·Update 2026-09-29·Tested on Windows, macOS, and Linux

#Which MiniCPM version to install in 2026

MiniCPM-V is OpenBMB’s compact vision-language model series, developed by THUNLP (Tsinghua University) and ModelBest, with weights and code released under the Apache 2.0 license. In Ollama, three tags coexist and cause confusion. The minicpm-v tag, without a suffix, still corresponds to MiniCPM-V 2.6, an 8-billion-parameter model weighing 5.5 GB with 32K context. The minicpm-v4.5 tag loads a newer 8-billion-parameter version (6.1 GB, 40K context). The minicpm-v4.6 tag, added to the official Ollama library on June 25, 2026, according to the project repository, weighs 1.6 GB and has approximately 1.3 billion parameters. To explore local vision on a small machine, start with 4.6; keep 4.5 for comparison on dense documents, because nothing guarantees that a 1.3-billion-parameter model will read tightly packed text as well as an 8-billion-parameter model.

The Ollama tags for the MiniCPM family (library cards, September 2026)
Tag OllamaVersionSettingsDownloadAdvertised contextKey takeaway
minicpm-vMiniCPM-V 2.68 billion5.5 GB32K2024 version: multi-image and video; this is what the tag without a suffix loads
minicpm-v4.5MiniCPM-V 4.58 billion6.1 GB40KBuilt on Qwen3-8B and SigLIP2-400M; high-frame-rate video; the publisher reports 77.0 on OpenCompass
minicpm-v4.6MiniCPM-V 4.6about 1.3 billion1.6 GB256KThe lightest and newest; visual-token compression; mobile variants
openbmb/minicpm-o4.5MiniCPM-o 4.59 billion6.1 GB40KOmni version (voice, real-time streams); the Ollama specification sheet lists only text and image as inputs
i
Three names not to confuse
MiniCPM-V refers to vision models, MiniCPM-o to omni models that add voice and real-time streaming, and MiniCPM 5 to the new series of pure-text models for devices (MiniCPM5-1B and MiniCPM5-2B, the latter with about 2.5 billion parameters, available from OpenBMB in Ollama). For image analysis and OCR, MiniCPM-V is the one you need.

#What the series delivers, and what to put into perspective

The Local AI Kit

Your private ChatGPT, free, on your own machine in an hour — LM Studio, Ollama, Open WebUI, your documents, no cloud.

  • Lifetime online access
  • PDF + files
  • Lifetime updates

The series aims for the best capacity-to-size ratio. According to the OpenBMB repository, MiniCPM-V 4.5 scores 77.0 on OpenCompass and, with 8 billion parameters, surpasses GPT-4o-latest, Gemini 2.0 Pro, and Qwen2.5-VL 72B. For version 4.6, the project reports a visual encoder whose compute requirements drop by more than 50%, optional visual-token compression (4x or 16x), and a token throughput approximately 1.5 times that of Qwen3.5-0.8B. The Ollama sheet for the same model reports 2.4 times: both figures come from the project, and neither page details a measurement protocol. These are claims, not independent measurements: benchmark them on your hardware.

Compact multimodal
Text and images, multiple images at once, and video. According to the repository, 4.5 compresses up to 6 video frames into 64 tokens, allowing it to process more frames without additional overhead for the language model.
High-resolution OCR
The 4.5 handles images of any aspect ratio up to 1.8 million pixels (for example, 1344 x 1344). The publisher claims it leads on OCRBench and PDF analysis (OmniDocBench) among general-purpose multimodal models; the 4.6 is said to reach Qwen3.5 2B-level performance on several benchmarks, including OCRBench.
Permissive license
The repository announces weights and code under the Apache 2.0 license, and the Hugging Face cards for versions 4.5 and 4.6 carry the same notice. Check the exact license for the specific version you download.
Many formats
The project offers GGUF, BNB, AWQ, and GPTQ variants, and advertises compatibility with SGLang, vLLM, llama.cpp, and Ollama.

#How much memory, really

Memory reported by OpenBMB and weights shown by Ollama
Model and formatAdvertised memorySource
MiniCPM-V 4.6, Transformers (GPU)4 GBRepository model table
MiniCPM-V 4.6, GGUF (CPU)2 GBRepository model table
MiniCPM-V 4.6, BNB, AWQ, or GPTQ (GPU)3 GBRepository model table
MiniCPM-o 4.5, GGUF10 GBRepository model table
MiniCPM-o 4.5, GPU19 GBRepository model table
minicpm-v4.6 in Ollama1.6 GB downloadOllama fact sheet
minicpm-v4.5 in Ollama6.1 GB downloadOllama fact sheet

These figures describe the weights and basic execution; context and images require additional memory. Images are converted into visual tokens that occupy the context window, and Ollama’s documentation specifies that it applies 4,000 context tokens by default under 24 GiB of VRAM; increasing the value increases the required memory. The advertised 256K context for 4.6 is therefore not allocated by default. Finally, the repository table indicates CPU execution for the GGUF version of 4.6 without publishing a speed: measure it before promising real-time performance.

→
Spilling over to the CPU
If Ollama splits the model between the GPU and CPU, the ollama ps command shows the split in the PROCESSOR column. The culprit is often the image or the context, not the weights: lower the image resolution or the num_ctx value before reducing quantization.

#Install MiniCPM-V with Ollama

Ollama handles the download, visual projector, and API server: there’s nothing else to install. Start by updating Ollama, because 4.6 is recent in the library. The minicpm-v tag page already mentions that it required Ollama 0.3.10 or later.

  1. 01
    Download the model
    ollama pull minicpm-v4.6 récupère les poids du modèle de langage et le projecteur visuel, soit environ 1,6 Go. Pour le 4.5, utilisez minicpm-v4.5 (6,1 Go).
  2. 02
    Start a first chat
    ollama run minicpm-v4.6 ouvre une session interactive. Posez une question texte avant de tester la vision.
  3. 03
    Send an image
    Ollama's documentation shows the file path placed before the question, for example ollama run minicpm-v4.6 ./photo.jpg followed by “Describe this image.”
  4. 04
    Connect an interface
    Once pulled, the model appears in Open WebUI or any client that points to port 11434.
Terminal
# Le plus léger (1,6 Go)
ollama pull minicpm-v4.6

# La version 8 milliards, pour comparer
ollama pull minicpm-v4.5

# Vérifier ce qui est installé, puis discuter
ollama list
ollama run minicpm-v4.6 ./photo.jpg Décris cette image

#Vision and OCR for everyday use

MiniCPM-V is well suited to document tasks: reading an invoice, transcribing a screenshot, extracting table rows, and describing a photo. High-resolution images matter a lot for OCR, because fine text that is unreadable in a thumbnail becomes usable at full resolution. Write precise prompts: “Describe this image” produces a general description, whereas “Faithfully transcribe all visible text, preserving the layout” produces a much cleaner result. To extract data, request an explicit format: JSON, Markdown, or a table.

Ollama API: extract fields from an invoice
import base64, requests

with open("facture.png", "rb") as f:
    img = base64.b64encode(f.read()).decode()

r = requests.post("http://localhost:11434/api/generate", json={
    "model": "minicpm-v4.6",
    "prompt": "Extrais le total, la date et le numéro de facture en JSON.",
    "images": [img],
    "stream": False,
})
print(r.json()["response"])

Ollama's REST API expects the image encoded as base64 in the images field, as its documentation notes; the SDKs also accept a file path. If you use Transformers instead of Ollama, the repository documents a useful setting for 4.6: downsample_mode defaults to 16x, while 4x preserves four times as many visual tokens for finer detail. This is the setting to try when very small text is misread.

!
Always check the OCR
No compact vision model is 100% reliable for critical figures: amounts, reference numbers, dates. A vision model can produce a well-formed value that does not appear on the page. For accounting or legal use, keep human review or a consistency check for sensitive fields.

When should you prefer a dedicated OCR engine? If you process thousands of pages and traceability is paramount, a specialized engine produces visible errors instead of plausible values. PaddleOCR-VL, a model with fewer than one billion parameters, scores 96.33% on OmniDocBench v1.6 according to its publisher, while Tesseract remains the lightweight choice for clean text. MiniCPM-V retains the advantage when you also want to describe the image or answer questions about it.

#Compared with Qwen3-VL and other compact models

Qwen3-VL is MiniCPM-V's direct competitor in the compact vision-model category, available in Ollama, notably in 2-, 4-, and 8-billion-parameter versions. Both cover the same ground: image description, OCR, and questions about a document. OpenBMB's README claims that MiniCPM-V 4.6 outperforms the larger Gemma4-E2B-it model and is more efficient than Qwen3.5-0.8B: these are comparisons published by the vendor using its own evaluations.

Compact vision models available in Ollama
ModelDisplayed sizeWhat we can say about itSite guide
minicpm-v4.61.6 GBThe publisher says so at the Qwen3.5 2B level on several visual benchmarks, including OCRBenchThis page
qwen3-vl:2b1.9 GBSame size class, displayed 256K contextQwen3-VL locally
qwen3-vl:8b6.1 GBEquivalent in size to minicpm-v4.5Qwen3-VL locally
MoondreamSee the guideLightweight vision model, covered in a dedicated guide on the siteMoondream

The choice depends more on your images than on a spec sheet. Both families run under the same Ollama and use the same API: switching between them means changing the model name in the request. You can keep both and route requests based on the task.

Compare on the same image
ollama run minicpm-v4.6 ./ticket.jpg Transcris le texte de ce ticket
ollama run qwen3-vl:2b ./ticket.jpg Transcris le texte de ce ticket

#Edge, mobile, and server: what the project documents

MiniCPM was designed for devices. The repository states that MiniCPM-V 4.6 runs on iOS, Android, and HarmonyOS, with open adaptation code, and shows raw screen recordings on an iPhone 17 Pro Max, a Redmi K70, and a HUAWEI nova 14. An application repository offers the download and deployment guide. To serve multiple users, the project announces compatibility with vLLM and SGLang, in addition to llama.cpp and Ollama.

  1. 01
    Local document digitization
    A mini PC or NAS with an entry-level GPU turns a folder of scans into searchable text without sending a sensitive document to the cloud.
  2. 02
    Offline assistant
    On a Apple Silicon laptop or a modest card, 4.6 reads the screen and responds to screenshots, even without a connection.
  3. 03
    Pipeline preprocessing
    Ahead of a heavier system, it sorts by document type and obvious fields. Only the difficult cases are sent to a large model.
  4. 04
    Bulk alt text
    Generate image descriptions for a website or media library in batches, with no per-call cost once the hardware has paid for itself.
i
MiniCPM-o 4.5: voice requires a different stack
MiniCPM-o 4.5 adds voice conversation and simultaneous video and audio streaming, but the Ollama card announces only text and image input. For voice, the repository points to Transformers or llama.cpp-omni, and its documented limitations are real: speech is sometimes mispronounced in omni mode, and responses sometimes mix English and Chinese.

#Troubleshooting

The model ignores the image
Check the file path and ensure it is accessible from where Ollama is running. In the API, the image must be base64-encoded in the images field.
Sloppy or truncated OCR
The image is probably too compressed or too low-resolution. Provide a sharp, larger version and explicitly request a faithful transcription with the layout preserved. With Transformers, try downsample_mode 4x.
Slow responses
On pure CPU, this is expected. On GPU, check with ollama ps that the model has not been partially offloaded to RAM because of an overly large context or image.
Unstructured output
For reliable JSON, enforce the schema in the prompt and, if your client supports it, use Ollama’s structured format.
Unexpected version
The minicpm-v tag without a suffix still points to 2.6. Explicitly name minicpm-v4.6 or minicpm-v4.5 for reproducible behavior.

#Go further

FAQ
Which Ollama tag should you install for MiniCPM-V?+
For a small machine, minicpm-v4.6: about 1.6 GB and a model with about 1.3 billion parameters. For comparison with an 8-billion-parameter model, minicpm-v4.5 weighs 6.1 GB. Note: the minicpm-v tag without a suffix still loads version 2.6. Always name the version; otherwise, you do not know what you are testing.
Is MiniCPM-V free, including for commercial use?+
The repository announces weights and code under the Apache 2.0 license, and the Hugging Face pages for versions 4.5 and 4.6 include this notice. Apache 2.0 permits commercial use. Nevertheless, review the license for the exact version you download, as an older version may have had different terms.
How much VRAM do you really need?+
The repository table lists 4 GB of GPU memory for 4.6 in Transformers and 2 GB for its GGUF CPU version; 4.5 weighs 6.1 GB in Ollama. Add the context and images: Ollama allocates 4,000 tokens by default with 24 GiB of VRAM. Check with ollama ps.
Does MiniCPM-V read document text well?+
The vendor reports very good OCR scores, with images of up to 1.8 million pixels for 4.5. These are its own evaluations. Like any vision model, it can produce a plausible value that isn't present on the page: keep a human review step for critical figures, or use a dedicated OCR engine for larger volumes.
Can you use it on a phone?+
Yes, the repository states that MiniCPM-V 4.6 runs on iOS, Android, and HarmonyOS, with open adaptation code and an application repository. The published demonstrations are screen recordings. Expect one-off tasks rather than a continuous image stream, and measure latency on your device.
Can MiniCPM-o talk to a model in Ollama?+
Not according to the Ollama spec sheet, which lists only text and image input for minicpm-o4.5. Voice conversation and simultaneous audio and video streaming go through Transformers or llama.cpp-omni, according to the repository. It also requires much more memory: 19 GB on the GPU, 10 GB in GGUF.
Did this guide help you?

Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.