MiniCPM locally (Ollama): multimodal compact
MiniCPM-V is OpenBMB's compact vision-language model series, licensed under Apache 2.0. In Ollama, the minicpm-v tag still loads version 2.6 (8 billion, 5.5 GB), minicpm-v4.5 loads an 8-billion-parameter version weighing 6.1 GB, and minicpm-v4.6 loads a model with about 1.3 billion parameters weighing 1.6 GB. Start with 4.6 on a small machine, then compare the results on your images.
MiniCPM is OpenBMB’s family of compact models, including vision-language models designed to run where larger models cannot. The family has changed significantly: the version loaded by the historical Ollama tag is no longer the latest, and a 1.3-billion-parameter model now exists. This page explains which tag to install, what the stated figures mean, and where a general-purpose model’s OCR capabilities stop.
#Which MiniCPM version to install in 2026
MiniCPM-V is OpenBMB’s compact vision-language model series, developed by THUNLP (Tsinghua University) and ModelBest, with weights and code released under the Apache 2.0 license. In Ollama, three tags coexist and cause confusion. The minicpm-v tag, without a suffix, still corresponds to MiniCPM-V 2.6, an 8-billion-parameter model weighing 5.5 GB with 32K context. The minicpm-v4.5 tag loads a newer 8-billion-parameter version (6.1 GB, 40K context). The minicpm-v4.6 tag, added to the official Ollama library on June 25, 2026, according to the project repository, weighs 1.6 GB and has approximately 1.3 billion parameters. To explore local vision on a small machine, start with 4.6; keep 4.5 for comparison on dense documents, because nothing guarantees that a 1.3-billion-parameter model will read tightly packed text as well as an 8-billion-parameter model.
| Tag Ollama | Version | Settings | Download | Advertised context | Key takeaway |
|---|---|---|---|---|---|
| minicpm-v | MiniCPM-V 2.6 | 8 billion | 5.5 GB | 32K | 2024 version: multi-image and video; this is what the tag without a suffix loads |
| minicpm-v4.5 | MiniCPM-V 4.5 | 8 billion | 6.1 GB | 40K | Built on Qwen3-8B and SigLIP2-400M; high-frame-rate video; the publisher reports 77.0 on OpenCompass |
| minicpm-v4.6 | MiniCPM-V 4.6 | about 1.3 billion | 1.6 GB | 256K | The lightest and newest; visual-token compression; mobile variants |
| openbmb/minicpm-o4.5 | MiniCPM-o 4.5 | 9 billion | 6.1 GB | 40K | Omni version (voice, real-time streams); the Ollama specification sheet lists only text and image as inputs |
#What the series delivers, and what to put into perspective
Your private ChatGPT, free, on your own machine in an hour — LM Studio, Ollama, Open WebUI, your documents, no cloud.
- Lifetime online access
- PDF + files
- Lifetime updates
The series aims for the best capacity-to-size ratio. According to the OpenBMB repository, MiniCPM-V 4.5 scores 77.0 on OpenCompass and, with 8 billion parameters, surpasses GPT-4o-latest, Gemini 2.0 Pro, and Qwen2.5-VL 72B. For version 4.6, the project reports a visual encoder whose compute requirements drop by more than 50%, optional visual-token compression (4x or 16x), and a token throughput approximately 1.5 times that of Qwen3.5-0.8B. The Ollama sheet for the same model reports 2.4 times: both figures come from the project, and neither page details a measurement protocol. These are claims, not independent measurements: benchmark them on your hardware.
- Compact multimodal
- Text and images, multiple images at once, and video. According to the repository, 4.5 compresses up to 6 video frames into 64 tokens, allowing it to process more frames without additional overhead for the language model.
- High-resolution OCR
- The 4.5 handles images of any aspect ratio up to 1.8 million pixels (for example, 1344 x 1344). The publisher claims it leads on OCRBench and PDF analysis (OmniDocBench) among general-purpose multimodal models; the 4.6 is said to reach Qwen3.5 2B-level performance on several benchmarks, including OCRBench.
- Permissive license
- The repository announces weights and code under the Apache 2.0 license, and the Hugging Face cards for versions 4.5 and 4.6 carry the same notice. Check the exact license for the specific version you download.
- Many formats
- The project offers GGUF, BNB, AWQ, and GPTQ variants, and advertises compatibility with SGLang, vLLM, llama.cpp, and Ollama.
#How much memory, really
| Model and format | Advertised memory | Source |
|---|---|---|
| MiniCPM-V 4.6, Transformers (GPU) | 4 GB | Repository model table |
| MiniCPM-V 4.6, GGUF (CPU) | 2 GB | Repository model table |
| MiniCPM-V 4.6, BNB, AWQ, or GPTQ (GPU) | 3 GB | Repository model table |
| MiniCPM-o 4.5, GGUF | 10 GB | Repository model table |
| MiniCPM-o 4.5, GPU | 19 GB | Repository model table |
| minicpm-v4.6 in Ollama | 1.6 GB download | Ollama fact sheet |
| minicpm-v4.5 in Ollama | 6.1 GB download | Ollama fact sheet |
These figures describe the weights and basic execution; context and images require additional memory. Images are converted into visual tokens that occupy the context window, and Ollama’s documentation specifies that it applies 4,000 context tokens by default under 24 GiB of VRAM; increasing the value increases the required memory. The advertised 256K context for 4.6 is therefore not allocated by default. Finally, the repository table indicates CPU execution for the GGUF version of 4.6 without publishing a speed: measure it before promising real-time performance.
#Install MiniCPM-V with Ollama
Ollama handles the download, visual projector, and API server: there’s nothing else to install. Start by updating Ollama, because 4.6 is recent in the library. The minicpm-v tag page already mentions that it required Ollama 0.3.10 or later.
- 01Download the modelollama pull minicpm-v4.6 récupère les poids du modèle de langage et le projecteur visuel, soit environ 1,6 Go. Pour le 4.5, utilisez minicpm-v4.5 (6,1 Go).
- 02Start a first chatollama run minicpm-v4.6 ouvre une session interactive. Posez une question texte avant de tester la vision.
- 03Send an imageOllama's documentation shows the file path placed before the question, for example ollama run minicpm-v4.6 ./photo.jpg followed by “Describe this image.”
- 04Connect an interfaceOnce pulled, the model appears in Open WebUI or any client that points to port 11434.
#Vision and OCR for everyday use
MiniCPM-V is well suited to document tasks: reading an invoice, transcribing a screenshot, extracting table rows, and describing a photo. High-resolution images matter a lot for OCR, because fine text that is unreadable in a thumbnail becomes usable at full resolution. Write precise prompts: “Describe this image” produces a general description, whereas “Faithfully transcribe all visible text, preserving the layout” produces a much cleaner result. To extract data, request an explicit format: JSON, Markdown, or a table.
Ollama's REST API expects the image encoded as base64 in the images field, as its documentation notes; the SDKs also accept a file path. If you use Transformers instead of Ollama, the repository documents a useful setting for 4.6: downsample_mode defaults to 16x, while 4x preserves four times as many visual tokens for finer detail. This is the setting to try when very small text is misread.
When should you prefer a dedicated OCR engine? If you process thousands of pages and traceability is paramount, a specialized engine produces visible errors instead of plausible values. PaddleOCR-VL, a model with fewer than one billion parameters, scores 96.33% on OmniDocBench v1.6 according to its publisher, while Tesseract remains the lightweight choice for clean text. MiniCPM-V retains the advantage when you also want to describe the image or answer questions about it.
- PaddleOCR: OCR that understands the page
- Tesseract OCR: read a scan locally
- Extract invoice data end to end
#Compared with Qwen3-VL and other compact models
Qwen3-VL is MiniCPM-V's direct competitor in the compact vision-model category, available in Ollama, notably in 2-, 4-, and 8-billion-parameter versions. Both cover the same ground: image description, OCR, and questions about a document. OpenBMB's README claims that MiniCPM-V 4.6 outperforms the larger Gemma4-E2B-it model and is more efficient than Qwen3.5-0.8B: these are comparisons published by the vendor using its own evaluations.
| Model | Displayed size | What we can say about it | Site guide |
|---|---|---|---|
| minicpm-v4.6 | 1.6 GB | The publisher says so at the Qwen3.5 2B level on several visual benchmarks, including OCRBench | This page |
| qwen3-vl:2b | 1.9 GB | Same size class, displayed 256K context | Qwen3-VL locally |
| qwen3-vl:8b | 6.1 GB | Equivalent in size to minicpm-v4.5 | Qwen3-VL locally |
| Moondream | See the guide | Lightweight vision model, covered in a dedicated guide on the site | Moondream |
The choice depends more on your images than on a spec sheet. Both families run under the same Ollama and use the same API: switching between them means changing the model name in the request. You can keep both and route requests based on the task.
#Edge, mobile, and server: what the project documents
MiniCPM was designed for devices. The repository states that MiniCPM-V 4.6 runs on iOS, Android, and HarmonyOS, with open adaptation code, and shows raw screen recordings on an iPhone 17 Pro Max, a Redmi K70, and a HUAWEI nova 14. An application repository offers the download and deployment guide. To serve multiple users, the project announces compatibility with vLLM and SGLang, in addition to llama.cpp and Ollama.
- 01Local document digitizationA mini PC or NAS with an entry-level GPU turns a folder of scans into searchable text without sending a sensitive document to the cloud.
- 02Offline assistantOn a Apple Silicon laptop or a modest card, 4.6 reads the screen and responds to screenshots, even without a connection.
- 03Pipeline preprocessingAhead of a heavier system, it sorts by document type and obvious fields. Only the difficult cases are sent to a large model.
- 04Bulk alt textGenerate image descriptions for a website or media library in batches, with no per-call cost once the hardware has paid for itself.
- SGLang: serving a local LLM to multiple users
- vLLM: what is it, who is it for, and when should you use it?
#Troubleshooting
- The model ignores the image
- Check the file path and ensure it is accessible from where Ollama is running. In the API, the image must be base64-encoded in the images field.
- Sloppy or truncated OCR
- The image is probably too compressed or too low-resolution. Provide a sharp, larger version and explicitly request a faithful transcription with the layout preserved. With Transformers, try downsample_mode 4x.
- Slow responses
- On pure CPU, this is expected. On GPU, check with ollama ps that the model has not been partially offloaded to RAM because of an overly large context or image.
- Unstructured output
- For reliable JSON, enforce the schema in the prompt and, if your client supports it, use Ollama’s structured format.
- Unexpected version
- The minicpm-v tag without a suffix still points to 2.6. Explicitly name minicpm-v4.6 or minicpm-v4.5 for reproducible behavior.
#Go further
- Install Ollama in 5 minutes
- A local multimodal vision LLM with Ollama
- Choose your quantization
- Open WebUI with Ollama
- Source: official MiniCPM-V repository from OpenBMB
- Source: Ollama page from minicpm-v4.6
- Source: MiniCPM-V 4.6 Hugging Face model card
- Source: Ollama documentation on vision
Which Ollama tag should you install for MiniCPM-V?+
Is MiniCPM-V free, including for commercial use?+
How much VRAM do you really need?+
Does MiniCPM-V read document text well?+
Can you use it on a phone?+
Can MiniCPM-o talk to a model in Ollama?+
Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.