Muse Glimmer 30B: Meta's return to open-weight models, running locally with Ollama
Meta surprised everyone on August 10, 2026, by releasing Muse Glimmer 30B under the Apache 2.0 license, after two years without a major open-weight release. The model is multimodal, designed for agent use cases, and fits on a single 24–32 GB GPU thanks to the officially distributed GGUF Q4_K_M. This guide shows how to install it with Ollama on release day, enable DFlash speculative decoding, and read the announced benchmarks with the necessary perspective.
#Why Muse Glimmer 30B
Muse Glimmer 30B marks Meta's return to the open-weight arena. Three things make it an interesting model to host yourself: the Apache 2.0 license (commercial use without a user-threshold clause, unlike the older Llama licenses), native text + image multimodality, and an architecture designed for agent loops—reliable tool calls, structured outputs, and an adjustable reasoning budget.
The real argument is still its size. With 30 billion parameters and Meta's published GGUF Q4_K_M, the model runs on a single consumer GPU with 24 to 32 GB. No need for multi-GPU or disk offloading: it is in the same accessibility category as a Qwen 32B, but with vision included.
- License
- Apache 2.0 — free commercial use; redistribution and fine-tuning permitted without volume restrictions.
- Terms
- Text and image input, text output. Designed for reading screenshots, diagrams, and documents.
- Target
- Agents and tooling: function calling, strict JSON, and long context for execution traces.
- Official format
- Day-0 GGUF Q4_K_M release on the Ollama library, plus an MLX port for Apple Silicon.
#Prerequisites and VRAM (24-32 GB)
Your private ChatGPT, free, on your own machine in an hour — LM Studio, Ollama, Open WebUI, your documents, no cloud.
- Lifetime online access
- PDF + files
- Lifetime updates
A 30B in Q4_K_M weighs around 18–19 GB in weights. But with Muse Glimmer, the image projector, and the context cache, the actual footprint is higher—allow 24 to 32 GB of VRAM for comfortable multimodal use with a reasonable working context. For text-only use and a small context, a 24 GB GPU is more than enough.
- Ideal target
- RTX 4090 24 GB or RTX 5090 32 GB on the NVIDIA side—all in VRAM, with no fallback to the CPU.
- Apple Silicon
- Mac M4 Pro/Max with 32 GB of unified memory or more. The shared memory handles the model and image cache without trouble.
- Entry-level
- A RTX 4080 16 GB runs the model but spills over in VRAM: some layers move to RAM, and inference slows significantly.
- Software
- Ollama ≥ 0.6 (day-0 architecture support), up-to-date GPU drivers (CUDA on the NVIDIA side), 30 GB of free disk space.
#Step-by-step Ollama installation
If Ollama is not already installed, the procedure takes two minutes. The daemon listens on http://localhost:11434 by default, and it is the one that downloads and serves the model.
- 01Install OllamaOn Linux, a single command. On macOS and Windows, download the application from ollama.com. Then verify that the daemon is running with `ollama --version`.
- 02Run Muse Glimmer 30BThe official tag points to the Q4_K_M GGUF. The download is approximately 18-19 GB: plan for the bandwidth and disk space.
- 03First exchangeRun the model interactively. On the first startup, Ollama loads the weights into VRAM (a few seconds depending on the GPU), then you get a prompt.
- 04Expose the APIOnce the model has been pulled, the OpenAI-compatible endpoint is already available on port 11434. No additional configuration is needed to connect an app or Open WebUI to it.
For programmatic use, the REST API returns JSON. Here's a minimal text-only request:
#Use multimodal capabilities
Muse Glimmer’s strength is reasoning over images: reading an error screenshot, extracting a table from a scan, or describing an architecture diagram. In the CLI, simply paste the image path into the prompt.
Via the API, pass the base64-encoded image in the `images` field of the message. This is the standard format for Ollama vision models, so any existing client works without modification.
For agent use cases, Muse Glimmer accepts tool definitions in OpenAI format. You describe your functions in the `tools` field, and the model returns a structured call that your code executes before returning the result. This is where the model has received the most work: few hallucinated arguments and reliably valid JSON.
#DFlash speculative decoding
Meta ships Muse Glimmer with DFlash, its speculative decoding method: a small « draft » model proposes several tokens ahead, which the main model validates in a single pass. When the proposals are good—which is common with code and structured text—we generate several tokens per step instead of one, resulting in higher throughput with no loss of quality, while the output remains identical to standard decoding.
- The principle
- A lightweight draft model guesses what comes next; the 30B model verifies it in batches and accepts the correct tokens.
- The gain
- Higher throughput on predictable content (code, JSON, agent traces); zero to negative on highly creative text where guessing fails.
- Cost
- The draft model uses a little extra VRAM—a reason to aim for 32 GB if you want DFlash active in multimodal mode.
- Activation
- Depending on the Ollama version, DFlash is configured through a Modelfile parameter or a service option. Check the tag's official release notes.
#On Mac: the MLX port
Meta also released an MLX version, the Apple framework optimized for the unified memory of M-series chips. On a recent Mac, MLX makes better use of the integrated GPU than llama.cpp's Metal backend for this model, with a more predictable memory footprint.
In practice: if you are on Apple Silicon and want maximum speed, test the MLX port alongside Ollama. Ollama remains the simplest option for integration and an OpenAI-compatible API; MLX targets raw throughput on Mac. The choice depends on whether you prioritize convenience or pure performance.
#The reported scores, in perspective
Meta reports a score of 51.2 on SWE-Bench Pro, the benchmark for resolving real software tickets. On paper, that is excellent for a 30B—enough to compete with much larger models. But a marketing figure is not a usage verdict.
- Context matters
- A SWE-Bench score depends heavily on the agent harness, prompt, and number of allowed attempts. Two setups can separate the same model by 15 points.
- Q4 is not FP16
- Official scores are measured at full precision. The GGUF Q4_K_M you run loses some fidelity—the gap is real on edge-case tasks.
- Your tasks ≠ the benchmark
- SWE-Bench is open-source Python. Your own tickets—another language, a proprietary codebase, French—may not behave the same way.
The right approach: treat 51.2 as an encouraging signal, not a promise. Set up five to ten representative tasks from your actual work, measure the success rate in Q4_K_M on your machine, and compare it with the model you already use. This is the only benchmark that matters for your decision.
#Troubleshooting
- VRAM overflow in multimodal use
- Reduce the size of the images sent and the context (`num_ctx`), or disable DFlash to free up draft-model memory.
- Tag not found
- Support is day-0 but requires a recent version of Ollama. Run `ollama --version` and update if the pull fails with an architecture error.
- Very slow generation
- Check with `ollama ps` that the model is running 100% on the GPU. If it shows a CPU percentage, VRAM is overflowing—this is typical with 16 GB.
- Invalid tool JSON
- Lower the temperature and provide explicit tool schemas. Muse Glimmer is reliable for structured output, but a high temperature breaks consistency.
#Go further
Muse Glimmer installs like any other Ollama model: the guides below cover the fundamentals if you are getting started or want to optimize.
- Install Ollama properly
- The “Installing Ollama on Linux” guide explains systemd and NVIDIA/AMD GPU configuration if you’re starting from scratch.
- Choose the right quantization
- “Choosing your quantization (Q4, Q5, Q8, FP16)” explains the quality/memory tradeoff—useful for deciding whether Q4_K_M is enough or whether to target Q5.
- GPU sizing
- Before buying for a 30B multimodal model, “Choosing Your GPU for Local AI” keeps you from underestimating the required VRAM.
Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.