Beginner 12 minModels

Qwen 3 locally: full test and benchmarks real

Qwen 3 has become the open-weight standard for many people running an LLM locally. But between the dense versions (8B, 14B, 32B) and the 30B-A3B MoE version, it's easy to get lost. This Qwen3 Ollama benchmark runs three variants on three very different machines—RTX 4090, RTX 3060 12 GB, and a MacBook Pro M4 Pro—to provide real numbers and a straightforward verdict.

By Mohamed Meguedmi·Update 2026-08-27·Tested on Windows, macOS, and Linux
i
In brief
Qwen 3 comes in dense versions (8B, 14B, 32B) and an MoE version (30B-A3B, 3B active parameters out of 30). · On RTX 4090, expect about 95 tok/s for 8B, 78 tok/s for 14B, and 110 tok/s for 30B-A3B—the MoE is faster than the dense 8B. · On RTX 3060 12 GB, the 14B (about 32 tok/s) remains the best compromise; the 30B-A3B does not fit in this card's VRAM. · On a Mac M4 Pro, expect 38, 24, and 45 tok/s, respectively.

#The Qwen 3 lineup: dense vs. MoE

Alibaba released Qwen 3 in two distinct families. Dense models (Qwen3-1.7B, 4B, 8B, 14B, 32B) follow the standard architecture: all parameters are used for every token. The MoE family (Qwen3-30B-A3B, Qwen3-235B-A22B) activates only a fraction of the experts for each token—hence the A3B suffix = 3 billion active parameters out of 30 total.

In practical terms, for a local user, two things change: the required VRAM corresponds to the total parameters (the entire model must fit in memory), but speed corresponds to the active parameters. A 30B-A3B MoE fits in 19 GB in Q4 like a dense 32B, but generates tokens at the speed of a 3B. That’s what makes it so interesting.

i
The three test variants
We focus on Qwen3-8B (dense), Qwen3-14B (dense), and Qwen3-30B-A3B (MoE). These are the three sizes that most 12–24 GB configurations can run comfortably.

#Tested setup and hardware

The Local AI Kit

Your private ChatGPT, free, on your own machine in an hour — LM Studio, Ollama, Open WebUI, your documents, no cloud.

  • Lifetime online access
  • PDF + files
  • Lifetime updates

Three machines cover roughly the full spectrum of local setups in 2026:

RTX 4090 24 GB
Windows 11 tower, Ryzen 9 7950X, 64 GB DDR5. High-end consumer hardware.
RTX 3060 12 GB
Linux tower running Ubuntu 24.04, Ryzen 5 5600X, 32 GB DDR4. The most popular budget configuration.
MacBook Pro M4 Pro 48 GB
macOS 15, 16-core GPU, unified memory. Representative of high-end laptops Apple.

Identical backend everywhere: Ollama (daemon on localhost:11434). All models in Q4_K_M by default, 8192-token context, standardized prompt of about 500 tokens. Speed measured over 10 generations averaging 500 tokens.

Retrieve the three models
ollama pull qwen3:8b
ollama pull qwen3:14b
ollama pull qwen3:30b-a3b
→
Measure on your system
To reproduce this test, add --verbose to ollama run. Ollama displays the eval rate (generation tokens/sec) and prompt eval rate (prefill speed) at the end of each response.

#Qwen3-8B: the entry-level dense model

Dense model with 8 billion parameters, ~5 GB in Q4_K_M. Designed to run everywhere: 8 GB of VRAM is more than enough. It is the natural candidate when you discover Qwen 3 on a RTX 3060 or laptop.

RTX 4090
~95 tokens/sec. The card is largely underutilized — all available VRAM is underused at this level.
RTX 3060 12 GB
~52 tokens/sec. Very comfortable, well above the fluidity threshold (20 t/s).
M4 Pro 48 GB
~38 tokens/sec. Slower than RTX cards but highly usable, with zero fan noise.

In terms of quality, Qwen3-8B handles everyday tasks very well: summarization, rewriting, simple code, and Q&A. It is clearly outclassed once you get into multi-step reasoning or nontrivial code—that's where you can tell it's an 8B model. Its French is decent but not excellent, with some anglicisms and awkward phrasing.

#Qwen3-14B: the dense sweet spot

Dense 14B, about 9 GB in Q4_K_M. Easily fits on 12 GB of VRAM with a reasonable context. This is the model we recommend to 90% of people with a RTX 3060 12 GB, 4070, or equivalent.

RTX 4090
~78 tokens/sec. Excellent speed—we’re still well below the GPU ceiling.
RTX 3060 12 GB
~32 tokens/sec. Comfortable for chat use; 8k context works without offloading.
M4 Pro 48 GB
~24 tokens/sec. At the edge of comfortable but usable every day.

The quality jump over the 8B is clear. Correct Python code for 50-80-line functions, stronger reasoning, and noticeably better French (natural phrasing, few gender errors). It is the first model in the lineup that genuinely resembles a low-end cloud assistant in terms of consistency.

!
Default context Ollama
Ollama defaults to a 2048-token context limit, which is ridiculous for Qwen 3 (which natively supports 32k). Force num_ctx through a Modelfile or with /set parameter num_ctx 8192 in the session; otherwise, you will cut off long conversations.

#Qwen3-30B-A3B: the MoE that changes everything

This is where it gets interesting. The 30B-A3B weighs about 19 GB in Q4_K_M (= you need 24 GB of VRAM to run it comfortably with context), but generates using only 3B active parameters. Result: 30B quality, 3B speed.

RTX 4090
~110 tokens/sec. Yes, faster than the dense 8B model. That's the magic of MoE: less computation per token.
RTX 3060 12 GB
Doesn’t fit in VRAM. Partial CPU offload ≈ 8 t/s — usable for testing but frustrating.
M4 Pro 48 GB
~45 tokens/sec. Unified memory shines here: everything fits in RAM, and the MoE offsets the limited FLOPS.

The quality is very close to a dense Qwen 32B. The 30B-A3B is particularly strong at coding (in the same league as a Qwen3-Coder 30B-A3B on intermediate Python tasks) and structured reasoning. For written French, it is genuinely good—roughly at the level of a Mistral Small 24B, for reference.

→
The surprise verdict
If you have 24 GB of VRAM or a Mac with 32+ GB of unified memory, Qwen3-30B-A3B remains an excellent general-purpose MoE: better quality AND better speed than the dense 14B, with no tradeoff. Since this test, the next generation has taken over in this range—Qwen 3.8 27B (the “Copilot-like” option of 2026, with 262k context) and Qwen 3.6 35B-A3B, the safe choice for 24–32 GB configurations. The 30B-A3B measured here retains all of its MoE DNA.

#Thinking mode: useful or not?

Qwen 3 introduces a toggleable thinking mode. When enabled, the model first generates a <think>...</think> block with its reasoning before producing the final answer. It is inspired by DeepSeek R1 but is more modular: you can enable or disable it on the fly.

Enable or disable thinking
# Dans la session Ollama
/set parameter think true
/set parameter think false

# Ou via prompt direct (Qwen 3 reconnaît les balises)
>>> /think Combien fait 17 × 23 ?
>>> /no_think Salut, ça va ?

What we observe from 50 tested prompts:

Simple tasks (chat, paraphrasing)
Thinking mode = wasted time. It adds 2 to 5 seconds of latency with zero quality gain. Disable it.
Math, logic, multi-step reasoning
Net accuracy gain: +20 to +30% correct answers on GSM8K-style problems. Keep it enabled.
Complex code (refactoring, debugging)
Moderate gain (+10%) in quality, but doubled latency. Test it based on your patience.
Creative generation (long-form text, dialogue)
Often counterproductive: the model overthinks and loses spontaneity.
i
Token cost
A typical <think> block adds 500 to 2000 additional tokens before the response. On a 30B-A3B at 100 t/s, this is invisible. On a 14B at 30 t/s, it adds 10 to 60 seconds of latency— weigh your priorities.

#French quality vs. Llama 4

We compared the French outputs of Qwen3-14B and Qwen3-30B-A3B with Llama 4 Scout (the equivalent Meta model, released around the same time) across three exercises: writing a professional email, explaining a technical concept in plain language, and analyzing a short literary text.

Qwen3-14B
Correct French, natural phrasing, but a few persistent gender errors (“la problème”). Vocabulary is somewhat flat.
Qwen3-30B-A3B
Very good written French. No structural errors, varied vocabulary, and register appropriate to the context. Slightly behind Mistral Small 24B on highly idiomatic nuances.
Llama 4 Scout
Flawless French grammatically, but with a more “translated from English” tone than Qwen. Personal preferences.

Pragmatic verdict: for professional French, Qwen3-30B-A3B and Llama 4 Scout are neck and neck. The choice comes down to other criteria (speed, size, license). For literary or highly idiomatic French, Mistral Small remains a notch ahead.


#Which Qwen 3 should you choose for your setup?

6-8 GB of VRAM (RTX 3050, 3060 Ti, 4060)
Qwen3-8B in Q4_K_M. The 14B runs in Q3, but quality deteriorates — stick with the clean 8B.
12 GB of VRAM (RTX 3060 12, 4070, 5070)
Qwen3-14B in Q4_K_M is the undisputed sweet spot. Comfortable 16k context.
16 GB (RTX 4080, 5070 Ti, 5080)
Qwen3-14B in Q5 or Q6 for comfort, or a 30B-A3B in Q3_K_M if you want to test the MoE.
24 GB (RTX 3090, 4090, RX 7900 XTX)
Qwen3-30B-A3B in Q4_K_M. No debate—it’s the best general-purpose local model available today.
Mac M-series with 32-48 GB unified memory
Qwen3-30B-A3B in Q4 or Q5. The MoE is particularly well suited to unified memory Apple.
Mac Studio Ultra 128 GB+
Jump straight to Qwen3-235B-A22B (~120 GB in Q4) if you want the top of the range. Otherwise, the 30B-A3B in Q8 remains excellent.
→
And what about thinking mode?
By default, leave it disabled. Enable it on demand when tackling a math, logic, or complex reasoning problem. For everyday chat, it only adds latency. The same applies to the next generation: Qwen 3.8 27B tends to overthink with its default reasoning setting — switch it to low for everyday conversations.

#Go further

If Qwen 3 is your starting point, know that the next generation has since taken over in our rankings — Qwen 3.5 9B has become the default choice for 8 GB configurations (256k context, vision), and Qwen 3.8 27B for 24 GB configurations; the settings below remain valid unchanged for these models. A few ways to go further:

Choose your quantization
To understand exactly what you give up between Q4_K_M and Q5_K_M on Qwen 3, and when moving up in quantization makes sense.
Customize with a Modelfile
To set the thinking mode, context, and a default FR system prompt on your Qwen 3 instance—instead of configuring them again each session.
Open WebUI with Ollama
To move from the terminal to a real chat interface with history, Markdown, and attachments without giving up local operation.
Did this guide help you?

Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.