Moondream: the vision model that fits partout
Moondream is a family of open vision-language models that describe images, answer questions, detect, point, and segment. Moondream 2 (1.9 billion parameters, Apache 2.0) requires about 2 GB of VRAM in int4; the 0.5B version requires 816 MiB. Generations 3 and 3.1 (9 billion total, 2 active) require about 10 GB of VRAM and a license that only prohibits reselling Moondream itself as a hosted service.
Impressive multimodal models weigh tens of gigabytes and occupy an entire graphics card. Moondream tackles the problem from the other end: a family of compact vision models, the smallest version of which fits in less than one gigabyte, for describing an image, answering a question, or locating an object on an ordinary computer. This guide, updated September 28, 2026, distinguishes the four variants, quantifies their resource use based on their publisher's figures, explains the licenses, and says what not to ask of it.
#A family, not a model
Moondream is published by M87 Labs. Its original creator, Vikhyat Korrapati, describes it in the GitHub repository as “a tiny vision language model that kicks ass and runs anywhere.” The name Moondream now refers to four variants with very different characteristics, and older guides often confuse them.
| Variant | Architecture | Weights and memory | Context | License |
|---|---|---|---|---|
| Moondream 2 (0,5B) | 500 million parameters, designed as a distillation target for embedded use | int4: 375 MiB to download, 816 MiB in memory | Not specified | Apache 2.0 for the code repository (model license unspecified) |
| Moondream 2 | Dense model, 1.9 billion parameters, updated since March 2024 | 3.85 GB file; int8: 2,624 MiB of VRAM; int4: 2,002 MiB | 2,000 tokens | Apache 2.0 |
| Moondream 3 Preview | Mixture of experts, 9 billion total, 2 billion active per token; 64 experts, 8 activated | About 18.5 GB in BF16, or 10.51 GB in FP8 | 32,000 tokens | Business Source License 1.1 with a “No Third-Party Service” clause |
| Moondream 3.1 9B A2B | Mixture of experts, 9.0 billion total, 2.0 active; 8 experts, 2 activated, 1 shared | FP8 experts, at least 10 GB of VRAM according to the specifications | 32,768 tokens | Moondream Model License 1.0, effective since July 7, 2026 |
The difference in scale changes the nature of the use cases. A 2-billion-parameter model can run on a laptop without a dedicated GPU or on an embedded device. A 9-billion-parameter model overall, even if it activates only 2 parameters per token, must keep all its weights in memory: that is why Generation 3 requires around ten gigabytes, with a 10.51 GB FP8 file for the Preview. Moondream’s documentation specifies that a base Mac mini M4 with 16 GB can run Moondream 2, but Moondream 3’s weights exceed its unified memory.
#What it can do
Your private ChatGPT, free, on your own machine in an hour — LM Studio, Ollama, Open WebUI, your documents, no cloud.
- Lifetime online access
- PDF + files
- Lifetime updates
Moondream models expose native capabilities, with structured outputs instead of free-form text to interpret. According to the version 3.1 model card, each one has a fixed output contract.
- Query (answer a question)
- Natural-language question about the image, free-form answer. This is the general-purpose input. The model card specifies that there is no calibrated confidence score: ask verifiable questions.
- Caption (describe)
- Caption in three lengths: short, normal, or long. The documentation warns that a long caption trades recall for a higher risk of hallucination and should not be analyzed to derive counts or coordinates.
- Detect
- Bounding boxes for any expression, with coordinates normalized from 0 to 1 starting at the upper-left corner. Up to 200 boxes per call, with a default threshold of 0.30.
- Point (pointer)
- The center of each instance, up to 64 points: more economical than a box and suited to counting.
- Segment (to segment)
- One mask per corresponding object, up to 50 instances; confidence scores are not provided for the masks.
Moondream 2 gained these capabilities in stages. Its change log, published by the author, mentions long captions, open-vocabulary labeling, and improved counting in March 2025 (CountBenchQA from 80 to 86.4). In April 2025, document and chart reading improved (ChartQA from 74.8 to 77.5). In June 2025, a grounded reasoning mode arrived, enabled by reasoning=True in the query skill, trading speed for accuracy, along with 20 to 40% faster text generation according to the author. These figures come from the author’s own evaluations: use them to track the evolution, not to compare it with another model.
One usage note: Moondream doesn’t process the entire image; it breaks it into pieces. According to its FAQ, it resizes the image so that a grid of tiles roughly 378 × 378 pixels covers it, with at most 12 local tiles plus one global view—about 1.7 million pixels for the local tiles. The version 3.1 specification confirms adaptive tiling into at most 12 tiles for PNG, JPEG, or WebP images no more than 4,096 pixels on the long side. A 12-megapixel photo is therefore heavily reduced before analysis: to read a detail, crop the area first.
#What it consumes
| Metric | Moondream 2 (0,5B) | Moondream 2 (1.9B) |
|---|---|---|
| Runtime memory, int8 | 996 MiB | 2,624 MiB |
| Runtime memory, int4 | 816 MiB | 2 002 MiB |
| Download, int4 | 375 MiB | Not specified |
| Speed, RTX 3090, int4 with compile() | Not specified | 184 tokens per second |
These figures come from Moondream’s model page: context matters—they refer to a RTX 3090 card, in int4 with compile(), and to text-generation speed, not time per image. Time per image depends on the output size and number of tiles. The same page states that int4 quantization reduces memory by 42% for a 0.6% loss in accuracy, without specifying the evaluation; treat this as the publisher’s claim.
For 3.1 generation, the hardware supported by the Photon engine is an NVIDIA Ampere or newer card on Linux or Windows, or a Mac with Apple Silicon running macOS 13 or later, with Python 3.10 through 3.14. The model page also lists NVIDIA SM75 and later, with at least 10 GB of VRAM, a GGUF version for x86 or ARM processors, and MLX for Apple Silicon. The documentation also lists Jetson cards, including the 8 GB Jetson Orin Nano, among the supported platforms.
#Three ways to run it
- 01With the moondream library and PhotonInstall the package, then load the model: the weights download from Hugging Face on the first call, and no API key is required for base models, according to the documentation. Photon only works on NVIDIA Ampere or newer GPUs, or on Mac Apple Silicon.
- 02With TransformersThe vikhyatk/moondream2 repository loads with AutoModelForCausalLM. Specify the revision: the author warns that the model is updated frequently and recommends pinning the version in production. The trust_remote_code=True option executes code provided by the repository: enable it only if you trust the source.
- 03With OllamaThe Ollama moondream page lists 1,7 GB, a context of 2 000 tokens, and a last update two years ago, so it predates the 2025 versions of Moondream 2 and does not cover generation 3. Make sure it offers the capabilities you need before adopting it.
Remember that the detect capability returns normalized coordinates: multiply them by the width and height of the original image to recover pixel coordinates.
- Overview of local multimodal models
- Source: official documentation for local execution
- Source: Moondream 2's Hugging Face model card
#Where it wins
- 01Index a photo libraryGenerate a caption for each image to make a photo library searchable by text. The cost per image determines feasibility, and this is where a small model wins. Use short captions: long ones may invent details.
- 02Filter before a large modelRoughly filter images so that only those that warrant it are sent to the expensive model. The point skill counts, while the detect skill automatically crops the area to analyze next.
- 03Describe images for accessibilityGenerate alternative text in batches for an existing site, locally and without sending visuals to a third-party service. Review a sample: a caption remains a hypothesis.
- 04Monitor a streamAsk a closed-ended question about successive images on a modest machine running continuously. The 0.5B model is designed to be fine-tuned for a narrow task, which greatly improves its accuracy, according to M87 Labs.
A best practice emerges from the version 3.1 documentation: ask closed-ended, precise questions. A small model handles “is there a headset on the person’s head?” well but loses focus on “analyze this scene.”
#What not to ask it to do
The version 3.1 documentation devotes an entire section to its limitations, which is rare and valuable. It describes the most common evaluation failures, and the thresholds it provides serve as useful reference points for models in the same family.
- Short texts
- Text smaller than about 10 pixels at the tile’s resolution is often misread or ignored: crop and enlarge the area first.
- Small objects
- Objects smaller than 12 pixels on the long side are missed by detect: use point on a cropped area.
- Loaded scenes
- Beyond roughly 50 instances, recall drops and adjacent identical objects merge. Counts above 50 are unreliable.
- Blurry text rendering
- With blurry, stylized, or partially obscured text, the model may produce a plausible but incorrect string without any low-confidence signal. For an entire document, an OCR engine is safer.
- High-stakes use cases
- Medical interpretation and legal or financial decisions based on documents without human review are considered out of scope. The output of a small model must be checked.
#Licenses: what you can do
| Model | License | What it enables, in summary |
|---|---|---|
| Moondream 2 | Apache 2.0 | Free to use, including commercially, under the usual Apache 2.0 terms |
| Moondream 3 Preview | Business Source License 1.1 with a “No Third-Party Service” clause | Personal, research, and most commercial use; no competing paid offering from M87 Labs, including hosted access or a paid SDK that bundles the weights |
| Moondream 3.1 | Moondream Model License 1.0, source-available, effective July 7, 2026 | Commercial use, fine-tuning, quantization, and redistribution are permitted; a separate license is required to offer Moondream itself as a general hosted service |
To index your own photo library or describe images on your site, none of these licenses poses a problem: internal use or use as a product component is allowed. The restriction applies only to someone who wants to sell access to Moondream itself, such as a general hosted vision API. If you are unsure about commercial use, the license is authoritative, not this summary.
#Moondream or a large vision model
The question to ask isn’t “which one is best?” but “how many images, and for what decision?” A handful of images per day and open-ended questions: choose the largest model your card can handle. Thousands of images and a repeated closed-ended question: a small model makes the project feasible.
A cascade architecture combines both: the small model handles everything and flags interesting cases, while the large model intervenes only for those cases. For a substantial collection, this provides the best balance between quality and compute time. Moondream’s detection capability serves as the filter: it crops the object, and the large model analyzes the cropped area.
- Local multimodal models with Ollama
- Qwen3-VL locally
- PaddleOCR: OCR that understands the page
- Source: Moondream’s GitHub repository
- Source: Moondream's model page
#FAQ
Is Moondream free?+
Do you need a graphics card?+
Can it read a scanned document?+
Moondream or a larger vision model?+
Can you use it on many images in succession?+
Can you fine-tune it for a specific task?+
Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.