Intermediate 12 minVision

Moondream: the vision model that fits partout

Direct response

Moondream is a family of open vision-language models that describe images, answer questions, detect, point, and segment. Moondream 2 (1.9 billion parameters, Apache 2.0) requires about 2 GB of VRAM in int4; the 0.5B version requires 816 MiB. Generations 3 and 3.1 (9 billion total, 2 active) require about 10 GB of VRAM and a license that only prohibits reselling Moondream itself as a hosted service.

Impressive multimodal models weigh tens of gigabytes and occupy an entire graphics card. Moondream tackles the problem from the other end: a family of compact vision models, the smallest version of which fits in less than one gigabyte, for describing an image, answering a question, or locating an object on an ordinary computer. This guide, updated September 28, 2026, distinguishes the four variants, quantifies their resource use based on their publisher's figures, explains the licenses, and says what not to ask of it.

By Mohamed Meguedmi·Update 2026-09-29·Tested on Windows, macOS, and Linux

#A family, not a model

Moondream is published by M87 Labs. Its original creator, Vikhyat Korrapati, describes it in the GitHub repository as “a tiny vision language model that kicks ass and runs anywhere.” The name Moondream now refers to four variants with very different characteristics, and older guides often confuse them.

Moondream variants, according to M87 Labs and Hugging Face
VariantArchitectureWeights and memoryContextLicense
Moondream 2 (0,5B)500 million parameters, designed as a distillation target for embedded useint4: 375 MiB to download, 816 MiB in memoryNot specifiedApache 2.0 for the code repository (model license unspecified)
Moondream 2Dense model, 1.9 billion parameters, updated since March 20243.85 GB file; int8: 2,624 MiB of VRAM; int4: 2,002 MiB2,000 tokensApache 2.0
Moondream 3 PreviewMixture of experts, 9 billion total, 2 billion active per token; 64 experts, 8 activatedAbout 18.5 GB in BF16, or 10.51 GB in FP832,000 tokensBusiness Source License 1.1 with a “No Third-Party Service” clause
Moondream 3.1 9B A2BMixture of experts, 9.0 billion total, 2.0 active; 8 experts, 2 activated, 1 sharedFP8 experts, at least 10 GB of VRAM according to the specifications32,768 tokensMoondream Model License 1.0, effective since July 7, 2026

The difference in scale changes the nature of the use cases. A 2-billion-parameter model can run on a laptop without a dedicated GPU or on an embedded device. A 9-billion-parameter model overall, even if it activates only 2 parameters per token, must keep all its weights in memory: that is why Generation 3 requires around ten gigabytes, with a 10.51 GB FP8 file for the Preview. Moondream’s documentation specifies that a base Mac mini M4 with 16 GB can run Moondream 2, but Moondream 3’s weights exceed its unified memory.

i
A warning about version 3.1
The version 3.1 model card labels its benchmark figures “PRELIMINARY”; they may change before public release. This guide therefore reproduces none of those scores. The memory figures cited later come from technical documentation and published file sizes.

#What it can do

The Local AI Kit

Your private ChatGPT, free, on your own machine in an hour — LM Studio, Ollama, Open WebUI, your documents, no cloud.

  • Lifetime online access
  • PDF + files
  • Lifetime updates

Moondream models expose native capabilities, with structured outputs instead of free-form text to interpret. According to the version 3.1 model card, each one has a fixed output contract.

Query (answer a question)
Natural-language question about the image, free-form answer. This is the general-purpose input. The model card specifies that there is no calibrated confidence score: ask verifiable questions.
Caption (describe)
Caption in three lengths: short, normal, or long. The documentation warns that a long caption trades recall for a higher risk of hallucination and should not be analyzed to derive counts or coordinates.
Detect
Bounding boxes for any expression, with coordinates normalized from 0 to 1 starting at the upper-left corner. Up to 200 boxes per call, with a default threshold of 0.30.
Point (pointer)
The center of each instance, up to 64 points: more economical than a box and suited to counting.
Segment (to segment)
One mask per corresponding object, up to 50 instances; confidence scores are not provided for the masks.

Moondream 2 gained these capabilities in stages. Its change log, published by the author, mentions long captions, open-vocabulary labeling, and improved counting in March 2025 (CountBenchQA from 80 to 86.4). In April 2025, document and chart reading improved (ChartQA from 74.8 to 77.5). In June 2025, a grounded reasoning mode arrived, enabled by reasoning=True in the query skill, trading speed for accuracy, along with 20 to 40% faster text generation according to the author. These figures come from the author’s own evaluations: use them to track the evolution, not to compare it with another model.

One usage note: Moondream doesn’t process the entire image; it breaks it into pieces. According to its FAQ, it resizes the image so that a grid of tiles roughly 378 × 378 pixels covers it, with at most 12 local tiles plus one global view—about 1.7 million pixels for the local tiles. The version 3.1 specification confirms adaptive tiling into at most 12 tiles for PNG, JPEG, or WebP images no more than 4,096 pixels on the long side. A 12-megapixel photo is therefore heavily reduced before analysis: to read a detail, crop the area first.

#What it consumes

Memory and speed figures reported by M87 Labs for the smaller variants
MetricMoondream 2 (0,5B)Moondream 2 (1.9B)
Runtime memory, int8996 MiB2,624 MiB
Runtime memory, int4816 MiB2 002 MiB
Download, int4375 MiBNot specified
Speed, RTX 3090, int4 with compile()Not specified184 tokens per second

These figures come from Moondream’s model page: context matters—they refer to a RTX 3090 card, in int4 with compile(), and to text-generation speed, not time per image. Time per image depends on the output size and number of tiles. The same page states that int4 quantization reduces memory by 42% for a 0.6% loss in accuracy, without specifying the evaluation; treat this as the publisher’s claim.

For 3.1 generation, the hardware supported by the Photon engine is an NVIDIA Ampere or newer card on Linux or Windows, or a Mac with Apple Silicon running macOS 13 or later, with Python 3.10 through 3.14. The model page also lists NVIDIA SM75 and later, with at least 10 GB of VRAM, a GGUF version for x86 or ARM processors, and MLX for Apple Silicon. The documentation also lists Jetson cards, including the 8 GB Jetson Orin Nano, among the supported platforms.

#Three ways to run it

  1. 01
    With the moondream library and Photon
    Install the package, then load the model: the weights download from Hugging Face on the first call, and no API key is required for base models, according to the documentation. Photon only works on NVIDIA Ampere or newer GPUs, or on Mac Apple Silicon.
  2. 02
    With Transformers
    The vikhyatk/moondream2 repository loads with AutoModelForCausalLM. Specify the revision: the author warns that the model is updated frequently and recommends pinning the version in production. The trust_remote_code=True option executes code provided by the repository: enable it only if you trust the source.
  3. 03
    With Ollama
    The Ollama moondream page lists 1,7 GB, a context of 2 000 tokens, and a last update two years ago, so it predates the 2025 versions of Moondream 2 and does not cover generation 3. Make sure it offers the capabilities you need before adopting it.
Python, according to Moondream's documentation
# pip install --upgrade moondream
import moondream as md
from PIL import Image

model = md.photon("moondream2")   # ou "moondream3.1-9B-A2B" avec une carte de 10 Go ou plus
image = Image.open("photo.jpg")

print(model.caption(image, length="short")["caption"])
print(model.query(image, "Combien de personnes sont visibles ?")["answer"])
for obj in model.detect(image, "casque")["objects"]:
    print(obj)

Remember that the detect capability returns normalized coordinates: multiply them by the width and height of the original image to recover pixel coordinates.

#Where it wins

  1. 01
    Index a photo library
    Generate a caption for each image to make a photo library searchable by text. The cost per image determines feasibility, and this is where a small model wins. Use short captions: long ones may invent details.
  2. 02
    Filter before a large model
    Roughly filter images so that only those that warrant it are sent to the expensive model. The point skill counts, while the detect skill automatically crops the area to analyze next.
  3. 03
    Describe images for accessibility
    Generate alternative text in batches for an existing site, locally and without sending visuals to a third-party service. Review a sample: a caption remains a hypothesis.
  4. 04
    Monitor a stream
    Ask a closed-ended question about successive images on a modest machine running continuously. The 0.5B model is designed to be fine-tuned for a narrow task, which greatly improves its accuracy, according to M87 Labs.

A best practice emerges from the version 3.1 documentation: ask closed-ended, precise questions. A small model handles “is there a headset on the person’s head?” well but loses focus on “analyze this scene.”

#What not to ask it to do

The version 3.1 documentation devotes an entire section to its limitations, which is rare and valuable. It describes the most common evaluation failures, and the thresholds it provides serve as useful reference points for models in the same family.

Short texts
Text smaller than about 10 pixels at the tile’s resolution is often misread or ignored: crop and enlarge the area first.
Small objects
Objects smaller than 12 pixels on the long side are missed by detect: use point on a cropped area.
Loaded scenes
Beyond roughly 50 instances, recall drops and adjacent identical objects merge. Counts above 50 are unreliable.
Blurry text rendering
With blurry, stylized, or partially obscured text, the model may produce a plausible but incorrect string without any low-confidence signal. For an entire document, an OCR engine is safer.
High-stakes use cases
Medical interpretation and legal or financial decisions based on documents without human review are considered out of scope. The output of a small model must be checked.

#Licenses: what you can do

Moondream model licenses, based on their model cards and licenses
ModelLicenseWhat it enables, in summary
Moondream 2Apache 2.0Free to use, including commercially, under the usual Apache 2.0 terms
Moondream 3 PreviewBusiness Source License 1.1 with a “No Third-Party Service” clausePersonal, research, and most commercial use; no competing paid offering from M87 Labs, including hosted access or a paid SDK that bundles the weights
Moondream 3.1Moondream Model License 1.0, source-available, effective July 7, 2026Commercial use, fine-tuning, quantization, and redistribution are permitted; a separate license is required to offer Moondream itself as a general hosted service

To index your own photo library or describe images on your site, none of these licenses poses a problem: internal use or use as a product component is allowed. The restriction applies only to someone who wants to sell access to Moondream itself, such as a general hosted vision API. If you are unsure about commercial use, the license is authoritative, not this summary.

#Moondream or a large vision model

The question to ask isn’t “which one is best?” but “how many images, and for what decision?” A handful of images per day and open-ended questions: choose the largest model your card can handle. Thousands of images and a repeated closed-ended question: a small model makes the project feasible.

A cascade architecture combines both: the small model handles everything and flags interesting cases, while the large model intervenes only for those cases. For a substantial collection, this provides the best balance between quality and compute time. Moondream’s detection capability serves as the filter: it crops the object, and the large model analyzes the cropped area.

#FAQ

FAQ
Is Moondream free?+
Yes for local execution. Moondream 2 is Apache 2.0-licensed, and generations 3 and 3.1 use source-available licenses that allow commercial use. There is no usage cost locally, and no API key is required for the base models with Photon. Only offering a general hosted service around Moondream requires a separate commercial license.
Do you need a graphics card?+
Not necessarily for the smaller variants: the 0.5B version lists 816 MiB of memory in int4, and Moondream 2 runs on GPU, CPU, mobile, and Raspberry Pi according to M87 Labs. The Photon engine, however, requires an NVIDIA Ampere card or a Apple Silicon Mac. For generation 3.1, plan on more than 10 GB of VRAM.
Can it read a scanned document?+
It reads short text, panels, and labels, and is improving on documents. But for a dense page, an OCR engine is safer: the version 3.1 documentation warns that the model may produce a plausible but incorrect string from blurry text without any warning signal. It does not hallucinate fewer values than a large model, and text smaller than 10 pixels is often misread.
Moondream or a larger vision model?+
For a few images and open-ended questions, choose the largest model your card can handle. For thousands of images and the same closed-ended question repeated, Moondream makes the project feasible. A cascade combines both: the small model filters and crops, while the large model handles only the interesting cases.
Can you use it on many images in succession?+
This is its ideal use case. The Photon engine advertises automatic batch processing and prefix-cache management. Keep in mind that each image is reduced to about 1.7 million local-tile pixels: for fine detail, crop first. Test on a sample before processing thousands of images.
Can you fine-tune it for a specific task?+
Yes. License 3.1 permits fine-tuning, quantization, and redistribution, and M87 Labs presents the 0.5B version as a base to fine-tune for a narrow task, with an accuracy gain according to the publisher. For everyday use, start with the base model: fine-tuning requires annotated data and a testing protocol before it is worth the cost.
Did this guide help you?

Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.