Intermediate 12 minMeta

Muse Glimmer 30B: Meta's return to open-weight models, running locally with Ollama

Meta surprised everyone on August 10, 2026, by releasing Muse Glimmer 30B under the Apache 2.0 license, after two years without a major open-weight release. The model is multimodal, designed for agent use cases, and fits on a single 24–32 GB GPU thanks to the officially distributed GGUF Q4_K_M. This guide shows how to install it with Ollama on release day, enable DFlash speculative decoding, and read the announced benchmarks with the necessary perspective.

By Mohamed Meguedmi·Update 2026-08-12·Tested on Windows, macOS, and Linux

#Why Muse Glimmer 30B

Muse Glimmer 30B marks Meta's return to the open-weight arena. Three things make it an interesting model to host yourself: the Apache 2.0 license (commercial use without a user-threshold clause, unlike the older Llama licenses), native text + image multimodality, and an architecture designed for agent loops—reliable tool calls, structured outputs, and an adjustable reasoning budget.

The real argument is still its size. With 30 billion parameters and Meta's published GGUF Q4_K_M, the model runs on a single consumer GPU with 24 to 32 GB. No need for multi-GPU or disk offloading: it is in the same accessibility category as a Qwen 32B, but with vision included.

License
Apache 2.0 — free commercial use; redistribution and fine-tuning permitted without volume restrictions.
Terms
Text and image input, text output. Designed for reading screenshots, diagrams, and documents.
Target
Agents and tooling: function calling, strict JSON, and long context for execution traces.
Official format
Day-0 GGUF Q4_K_M release on the Ollama library, plus an MLX port for Apple Silicon.
i
Local multimodal use: two files
Like all vision models under Ollama, Muse Glimmer consists of the language model plus an image projector (mmproj). Ollama handles both automatically when you pull the official tag: you do not need to assemble anything manually.

#Prerequisites and VRAM (24-32 GB)

The Local AI Kit

Your private ChatGPT, free, on your own machine in an hour — LM Studio, Ollama, Open WebUI, your documents, no cloud.

  • Lifetime online access
  • PDF + files
  • Lifetime updates

A 30B in Q4_K_M weighs around 18–19 GB in weights. But with Muse Glimmer, the image projector, and the context cache, the actual footprint is higher—allow 24 to 32 GB of VRAM for comfortable multimodal use with a reasonable working context. For text-only use and a small context, a 24 GB GPU is more than enough.

Ideal target
RTX 4090 24 GB or RTX 5090 32 GB on the NVIDIA side—all in VRAM, with no fallback to the CPU.
Apple Silicon
Mac M4 Pro/Max with 32 GB of unified memory or more. The shared memory handles the model and image cache without trouble.
Entry-level
A RTX 4080 16 GB runs the model but spills over in VRAM: some layers move to RAM, and inference slows significantly.
Software
Ollama ≥ 0.6 (day-0 architecture support), up-to-date GPU drivers (CUDA on the NVIDIA side), 30 GB of free disk space.
!
16 GB is just enough
On a 16 GB GPU, multimodal use frequently overflows VRAM as soon as you send a somewhat large image. If you are limited to 16 GB, stick to text-only use and reduce the context, or target a 14B model instead.

#Step-by-step Ollama installation

If Ollama is not already installed, the procedure takes two minutes. The daemon listens on http://localhost:11434 by default, and it is the one that downloads and serves the model.

  1. 01
    Install Ollama
    On Linux, a single command. On macOS and Windows, download the application from ollama.com. Then verify that the daemon is running with `ollama --version`.
  2. 02
    Run Muse Glimmer 30B
    The official tag points to the Q4_K_M GGUF. The download is approximately 18-19 GB: plan for the bandwidth and disk space.
  3. 03
    First exchange
    Run the model interactively. On the first startup, Ollama loads the weights into VRAM (a few seconds depending on the GPU), then you get a prompt.
  4. 04
    Expose the API
    Once the model has been pulled, the OpenAI-compatible endpoint is already available on port 11434. No additional configuration is needed to connect an app or Open WebUI to it.
Terminal — install Ollama (Linux)
curl -fsSL https://ollama.com/install.sh | sh
ollama --version
Terminal — pull and run the model
# Télécharge le GGUF Q4_K_M officiel (~18-19 Go)
ollama pull muse-glimmer:30b

# Chat interactif
ollama run muse-glimmer:30b

For programmatic use, the REST API returns JSON. Here's a minimal text-only request:

Terminal — API call
curl http://localhost:11434/api/chat -d '{
  "model": "muse-glimmer:30b",
  "messages": [
    { "role": "user", "content": "Résume en trois points ce que change la licence Apache 2.0 pour un projet commercial." }
  ],
  "stream": false
}'
→
A graphical interface as a bonus
If you prefer a ChatGPT-style experience over the terminal, connect Open WebUI to the same daemon Ollama. You can drag your images directly into the chat, which is more convenient for testing multimodal capabilities.

#Use multimodal capabilities

Muse Glimmer’s strength is reasoning over images: reading an error screenshot, extracting a table from a scan, or describing an architecture diagram. In the CLI, simply paste the image path into the prompt.

Terminal — question about an image
ollama run muse-glimmer:30b
>>> Décris cette capture et liste les erreurs visibles: ./capture-console.png

Via the API, pass the base64-encoded image in the `images` field of the message. This is the standard format for Ollama vision models, so any existing client works without modification.

Python — vision via the API
import base64, requests

with open("schema.png", "rb") as f:
    img = base64.b64encode(f.read()).decode()

resp = requests.post("http://localhost:11434/api/chat", json={
    "model": "muse-glimmer:30b",
    "messages": [{
        "role": "user",
        "content": "Quels composants communiquent avec la base de données ?",
        "images": [img],
    }],
    "stream": False,
})
print(resp.json()["message"]["content"])

For agent use cases, Muse Glimmer accepts tool definitions in OpenAI format. You describe your functions in the `tools` field, and the model returns a structured call that your code executes before returning the result. This is where the model has received the most work: few hallucinated arguments and reliably valid JSON.

#DFlash speculative decoding

Meta ships Muse Glimmer with DFlash, its speculative decoding method: a small « draft » model proposes several tokens ahead, which the main model validates in a single pass. When the proposals are good—which is common with code and structured text—we generate several tokens per step instead of one, resulting in higher throughput with no loss of quality, while the output remains identical to standard decoding.

The principle
A lightweight draft model guesses what comes next; the 30B model verifies it in batches and accepts the correct tokens.
The gain
Higher throughput on predictable content (code, JSON, agent traces); zero to negative on highly creative text where guessing fails.
Cost
The draft model uses a little extra VRAM—a reason to aim for 32 GB if you want DFlash active in multimodal mode.
Activation
Depending on the Ollama version, DFlash is configured through a Modelfile parameter or a service option. Check the tag's official release notes.
i
Don't expect a universal gain
Speculative decoding mainly speeds up predictable outputs. When VRAM is already full, loading the draft model can instead trigger an overflow and slow everything down. Measure tokens/s before and after on your actual workload before drawing conclusions.

#On Mac: the MLX port

Meta also released an MLX version, the Apple framework optimized for the unified memory of M-series chips. On a recent Mac, MLX makes better use of the integrated GPU than llama.cpp's Metal backend for this model, with a more predictable memory footprint.

In practice: if you are on Apple Silicon and want maximum speed, test the MLX port alongside Ollama. Ollama remains the simplest option for integration and an OpenAI-compatible API; MLX targets raw throughput on Mac. The choice depends on whether you prioritize convenience or pure performance.

#The reported scores, in perspective

Meta reports a score of 51.2 on SWE-Bench Pro, the benchmark for resolving real software tickets. On paper, that is excellent for a 30B—enough to compete with much larger models. But a marketing figure is not a usage verdict.

Context matters
A SWE-Bench score depends heavily on the agent harness, prompt, and number of allowed attempts. Two setups can separate the same model by 15 points.
Q4 is not FP16
Official scores are measured at full precision. The GGUF Q4_K_M you run loses some fidelity—the gap is real on edge-case tasks.
Your tasks ≠ the benchmark
SWE-Bench is open-source Python. Your own tickets—another language, a proprietary codebase, French—may not behave the same way.

The right approach: treat 51.2 as an encouraging signal, not a promise. Set up five to ten representative tasks from your actual work, measure the success rate in Q4_K_M on your machine, and compare it with the model you already use. This is the only benchmark that matters for your decision.

#Troubleshooting

VRAM overflow in multimodal use
Reduce the size of the images sent and the context (`num_ctx`), or disable DFlash to free up draft-model memory.
Tag not found
Support is day-0 but requires a recent version of Ollama. Run `ollama --version` and update if the pull fails with an architecture error.
Very slow generation
Check with `ollama ps` that the model is running 100% on the GPU. If it shows a CPU percentage, VRAM is overflowing—this is typical with 16 GB.
Invalid tool JSON
Lower the temperature and provide explicit tool schemas. Muse Glimmer is reliable for structured output, but a high temperature breaks consistency.

#Go further

Muse Glimmer installs like any other Ollama model: the guides below cover the fundamentals if you are getting started or want to optimize.

Install Ollama properly
The “Installing Ollama on Linux” guide explains systemd and NVIDIA/AMD GPU configuration if you’re starting from scratch.
Choose the right quantization
“Choosing your quantization (Q4, Q5, Q8, FP16)” explains the quality/memory tradeoff—useful for deciding whether Q4_K_M is enough or whether to target Q5.
GPU sizing
Before buying for a 30B multimodal model, “Choosing Your GPU for Local AI” keeps you from underestimating the required VRAM.
Did this guide help you?

Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.