Beginner 8 minConcepts

MoE explained: why a 30B-A3B runs like a small model model

Since 2025, most major open-weight releases have one thing in common: they are Mixture of Experts, or MoE, models. Names such as Qwen3 30B-A3B, gpt-oss 120B, or Llama 4 Scout 17B-A2B keep appearing, with this two-number notation that raises questions. Once stated, the idea is simple: an MoE model contains many parameters but activates only a small fraction of them for each token. As a result, a “30-billion-parameter” model can run at the speed of a 3-billion-parameter model. This guide explains the mechanism, decodes the notation, and details the real impact on your VRAM and throughput.

By Mohamed Meguedmi·Update 2026-08-02·Tested on Windows, macOS, and Linux

#The problem MoE solves

A conventional language model is called “dense”: to generate each word, it passes your text through all of its parameters. A dense 32B model uses its 32 billion parameters for every token. That makes it powerful, but also slow and resource-hungry: the more parameters there are, the more compute and memory it requires, with no workaround.

The local dilemma is right there. You want the knowledge and nuance of a large model, but the speed and efficiency of a small one. On a consumer card, a 70B dense model is either out of reach or painfully slow. Mixture of Experts breaks this trade-off by separating two things we thought were linked: model size and the cost of each token.

i
The intuition in one sentence
A dense model puts everyone to work on every question. An MoE wakes only the few relevant specialists and lets the others sleep.

#The principle: experts activated on demand

The Local AI Kit

Your private ChatGPT, free, on your own machine in an hour — LM Studio, Ollama, Open WebUI, your documents, no cloud.

  • Lifetime online access
  • PDF + files
  • Lifetime updates

In a Mixture-of-Experts model, the large compute layers are no longer a single block but a collection of subnetworks called “experts.” A model can contain 64, 128, or sometimes more. For each token processed, a small dispatcher—the “router”—looks at the context and selects only 2, 4, or 8 experts to activate. The remaining experts perform no computation for that token.

Contrary to what the word suggests, these experts are not thematic: there is no “French expert” or “code expert.” The router is trained at the same time as the model and distributes the workload according to statistical patterns that no one controls manually. The selection changes from one token to the next. Over an entire sentence, nearly all the experts end up being used, but never all at the same time.

Experts
Specialized subnetworks, all present in memory, but only a minority of them working on each token.
Router (gating)
The lightweight component that decides, token by token, which experts to activate.
Active experts
The number of experts activated per token, often 2 to 8. This number determines the speed.
Active parameters
The total number of parameters actually activated per token—the fraction that matters for computation.
→
A mental image
Imagine a practice with 128 doctors. A triage coordinator receives the patient and directs them to the 4 relevant specialists. The practice is enormous (a lot of knowledge available), but each consultation involves only 4 people (low cost per patient).

#Decoding the 30B-A3B notation

The notation that confuses beginners is actually easy to read once you know the key. Take Qwen3 30B-A3B: the first number is the model's total number of parameters, and the second (after the « A », for Active) is the number of active parameters per token.

30B
The model contains 30 billion parameters in total. This is what determines the memory required to load it.
A3B
Only ~3 billion parameters are active for each token. That is what determines generation speed.

In other words, 30B-A3B means “30 billion in reserve, 3 billion at work.” The model has the knowledge breadth of a 30B, but generates at roughly the speed of a dense 3B. The same logic applies to all other recent releases.

Qwen3 30B-A3B
30 billion total, 3 billion active—the quintessential consumer MoE.
Llama 4 Scout 17B-A2B
17 billion in total, approximately 2 billion active per token, plus multimodal capabilities.
gpt-oss 120B
120 billion in total, but only ~5 billion active—hence its surprising speed for its size.
DeepSeek V3.2 671B-A37B
671 billion in reserve, 37 billion active: a giant that “costs” only as much as an average model per token.
!
The reading trap
The second number doesn’t reduce memory: a 30B-A3B takes up the space of a 30B, not a 3B. “A3B” speeds up computation; it doesn’t compress the model. This is confusion n°1 among beginners.

#VRAM: what really changes

This is the most misunderstood point, so let's be clear. All experts in an MoE must be loaded into memory because the router can call any of them for the next token. VRAM requirements therefore depend on the total number of parameters, not the number of active parameters. A 30B-A3B is sized like a dense 30B model.

Qwen3 30B-A3B in Q4_K_M
≈ 18–19 GB of VRAM, like a dense 32B model. It therefore targets a 24 GB RTX 4090, or spills into RAM on a smaller one.
Llama 4 Scout 17B-A2B in Q4
≈ 10–11 GB, within the range of a RTX 4070 12 GB or a 3060 12 GB.
gpt-oss 120B
Several dozen GB: reserved for high-end configurations or RAM/SSD offloading.

The good news about local MoE is therefore not memory, but the quality-to-speed ratio at equal VRAM. Where a 30B dense model struggles on your card, a 30B-A3B MoE occupies the same space while generating much faster. And because inactive experts are not used for every token, MoE models often tolerate offloading some of the weights to RAM well: what remains on the GPU gets priority.

i
VRAM benchmark for Q4
3B ≈ 2 GB · 7B ≈ 5 GB · 14B ≈ 9 GB · 32B ≈ 19 GB · 70B ≈ 40 GB. For an MoE, apply these reference points to the first number (the total), never the active number.

#Speed: why it's fast

An LLM’s generation speed depends mainly on the number of parameters traversed for each token. Because an MoE activates only a small fraction of them, the computation per token drops dramatically. In practice, a 30B-A3B generates tokens at a rate close to a 3B dense model, while responding with the depth of a 30B model.

This is exactly what explains the MoE wave on modest configurations: you get the quality of a large model without paying its speed penalty. The price is paid elsewhere—in memory, as we have seen, and in a heavier disk footprint during download, since all experts must be stored.

What speeds things up
Few active parameters per token → less computation → more tokens/second.
What doesn't change
VRAM and file size, which track the total parameter count.
The net gain
With the same memory, an MoE provides more knowledge at a speed comparable to that of a much smaller dense model.

#Limitations to know

MoE is not magic, and it does not make dense models obsolete. With the same number of active parameters, a dense model generally remains one level ahead in reasoning finesse: a 30B-A3B is excellent, but a true 30B dense model that activated all 30 billion parameters for every token would be stronger—at the cost of prohibitively slow local performance.

Unreduced memory
You need the VRAM for the total. An MoE does not fit a large model into a small card.
Quality per active token
With comparable active parameters, a well-trained dense model often retains the edge on the most demanding reasoning tasks.
Larger download
All experts are stored: the GGUF file weighs the total, not the active one.
Sensitivity to quantization
The router and some experts do not tolerate overly aggressive quantization well; stick to Q4_K_M or higher.
→
How to choose in practice
Tight on VRAM and need speed with strong knowledge? An MoE shines. Prioritizing the finest reasoning with comfortable VRAM? A dense model with an equivalent number of active parameters may offer better value.

#The leading MoE models to try locally

Here are the Mixture of Experts models most worth testing at home in 2026, from the most accessible to the most demanding.

Qwen3 30B-A3B
The best entry point. Solid general-purpose quality and coding, remarkable speed, ~18 GB in Q4. The benchmark for discovering MoE.
Llama 4 Scout 17B-A2B
Multimodal MoE (text + image) with a long context, using less VRAM. Ideal if you have a 12 GB card.
gpt-oss 120B
OpenAI's open-weight model, with only ~5B active parameters: astonishingly fast for its size, but it requires a powerful setup.
DeepSeek V3.2 671B-A37B
The monster. Reserved for setups with lots of RAM and aggressive offloading, for those curious about the top tier.

If you're just getting started, begin with Qwen3 30B-A3B: it's the MoE that best demonstrates the value of the architecture on a single consumer GPU.

#Try an MoE in 3 minutes

If Ollama is already installed and listening on its default port, two commands are enough to feel the difference between an MoE and a dense model.

  1. 01
    Download and run a MoE
    Get Qwen3 30B-A3B and start chatting with it immediately. The first launch downloads the model (expect several GB).
  2. 02
    Monitor the speed
    Ask a slightly long question and note the throughput. Despite its 30 billion parameters, it responds instantly, at the speed of a small model.
  3. 03
    Compare it with a dense model
    Run a dense model of a similar size and ask the same question. The MoE should generate noticeably faster with comparable VRAM.
Terminal
# Lancer un MoE et discuter avec
ollama run qwen3:30b-a3b

# Pour comparer avec un dense de taille voisine
ollama run qwen3:32b
i
Check throughput
Add the --verbose flag (ollama run qwen3:30b-a3b --verbose) to see tokens per second at the end of the response and objectively compare MoE and dense models on your machine.

#Go further

MoE is easier to understand when connected to the other fundamentals of local AI. These guides directly build on this one:

Choose your quantization (Q4, Q5, Q8, FP16)
MoE is sensitive to overly aggressive quantization: this guide helps you find the right quality/memory tradeoff.
Install Ollama: Windows, macOS, and Linux
To set up the local daemon on port 11434 before pulling your first MoE.
Llama 4 Scout locally: installation and initial tests with Ollama
A detailed step-by-step multimodal MoE, showing the theory in this guide in action.
Did this guide help you?

Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.