Qwen 3 locally: full test and benchmarks real
Qwen 3 has become the open-weight standard for many people running an LLM locally. But between the dense versions (8B, 14B, 32B) and the 30B-A3B MoE version, it's easy to get lost. This Qwen3 Ollama benchmark runs three variants on three very different machines—RTX 4090, RTX 3060 12 GB, and a MacBook Pro M4 Pro—to provide real numbers and a straightforward verdict.
#The Qwen 3 lineup: dense vs. MoE
Alibaba released Qwen 3 in two distinct families. Dense models (Qwen3-1.7B, 4B, 8B, 14B, 32B) follow the standard architecture: all parameters are used for every token. The MoE family (Qwen3-30B-A3B, Qwen3-235B-A22B) activates only a fraction of the experts for each token—hence the A3B suffix = 3 billion active parameters out of 30 total.
In practical terms, for a local user, two things change: the required VRAM corresponds to the total parameters (the entire model must fit in memory), but speed corresponds to the active parameters. A 30B-A3B MoE fits in 19 GB in Q4 like a dense 32B, but generates tokens at the speed of a 3B. That’s what makes it so interesting.
#Tested setup and hardware
Your private ChatGPT, free, on your own machine in an hour — LM Studio, Ollama, Open WebUI, your documents, no cloud.
- Lifetime online access
- PDF + files
- Lifetime updates
Three machines cover roughly the full spectrum of local setups in 2026:
- RTX 4090 24 GB
- Windows 11 tower, Ryzen 9 7950X, 64 GB DDR5. High-end consumer hardware.
- RTX 3060 12 GB
- Linux tower running Ubuntu 24.04, Ryzen 5 5600X, 32 GB DDR4. The most popular budget configuration.
- MacBook Pro M4 Pro 48 GB
- macOS 15, 16-core GPU, unified memory. Representative of high-end laptops Apple.
Identical backend everywhere: Ollama (daemon on localhost:11434). All models in Q4_K_M by default, 8192-token context, standardized prompt of about 500 tokens. Speed measured over 10 generations averaging 500 tokens.
#Qwen3-8B: the entry-level dense model
Dense model with 8 billion parameters, ~5 GB in Q4_K_M. Designed to run everywhere: 8 GB of VRAM is more than enough. It is the natural candidate when you discover Qwen 3 on a RTX 3060 or laptop.
- RTX 4090
- ~95 tokens/sec. The card is largely underutilized — all available VRAM is underused at this level.
- RTX 3060 12 GB
- ~52 tokens/sec. Very comfortable, well above the fluidity threshold (20 t/s).
- M4 Pro 48 GB
- ~38 tokens/sec. Slower than RTX cards but highly usable, with zero fan noise.
In terms of quality, Qwen3-8B handles everyday tasks very well: summarization, rewriting, simple code, and Q&A. It is clearly outclassed once you get into multi-step reasoning or nontrivial code—that's where you can tell it's an 8B model. Its French is decent but not excellent, with some anglicisms and awkward phrasing.
#Qwen3-14B: the dense sweet spot
Dense 14B, about 9 GB in Q4_K_M. Easily fits on 12 GB of VRAM with a reasonable context. This is the model we recommend to 90% of people with a RTX 3060 12 GB, 4070, or equivalent.
- RTX 4090
- ~78 tokens/sec. Excellent speed—we’re still well below the GPU ceiling.
- RTX 3060 12 GB
- ~32 tokens/sec. Comfortable for chat use; 8k context works without offloading.
- M4 Pro 48 GB
- ~24 tokens/sec. At the edge of comfortable but usable every day.
The quality jump over the 8B is clear. Correct Python code for 50-80-line functions, stronger reasoning, and noticeably better French (natural phrasing, few gender errors). It is the first model in the lineup that genuinely resembles a low-end cloud assistant in terms of consistency.
#Qwen3-30B-A3B: the MoE that changes everything
This is where it gets interesting. The 30B-A3B weighs about 19 GB in Q4_K_M (= you need 24 GB of VRAM to run it comfortably with context), but generates using only 3B active parameters. Result: 30B quality, 3B speed.
- RTX 4090
- ~110 tokens/sec. Yes, faster than the dense 8B model. That's the magic of MoE: less computation per token.
- RTX 3060 12 GB
- Doesn’t fit in VRAM. Partial CPU offload ≈ 8 t/s — usable for testing but frustrating.
- M4 Pro 48 GB
- ~45 tokens/sec. Unified memory shines here: everything fits in RAM, and the MoE offsets the limited FLOPS.
The quality is very close to a dense Qwen 32B. The 30B-A3B is particularly strong at coding (in the same league as a Qwen3-Coder 30B-A3B on intermediate Python tasks) and structured reasoning. For written French, it is genuinely good—roughly at the level of a Mistral Small 24B, for reference.
#Thinking mode: useful or not?
Qwen 3 introduces a toggleable thinking mode. When enabled, the model first generates a <think>...</think> block with its reasoning before producing the final answer. It is inspired by DeepSeek R1 but is more modular: you can enable or disable it on the fly.
What we observe from 50 tested prompts:
- Simple tasks (chat, paraphrasing)
- Thinking mode = wasted time. It adds 2 to 5 seconds of latency with zero quality gain. Disable it.
- Math, logic, multi-step reasoning
- Net accuracy gain: +20 to +30% correct answers on GSM8K-style problems. Keep it enabled.
- Complex code (refactoring, debugging)
- Moderate gain (+10%) in quality, but doubled latency. Test it based on your patience.
- Creative generation (long-form text, dialogue)
- Often counterproductive: the model overthinks and loses spontaneity.
#French quality vs. Llama 4
We compared the French outputs of Qwen3-14B and Qwen3-30B-A3B with Llama 4 Scout (the equivalent Meta model, released around the same time) across three exercises: writing a professional email, explaining a technical concept in plain language, and analyzing a short literary text.
- Qwen3-14B
- Correct French, natural phrasing, but a few persistent gender errors (“la problème”). Vocabulary is somewhat flat.
- Qwen3-30B-A3B
- Very good written French. No structural errors, varied vocabulary, and register appropriate to the context. Slightly behind Mistral Small 24B on highly idiomatic nuances.
- Llama 4 Scout
- Flawless French grammatically, but with a more “translated from English” tone than Qwen. Personal preferences.
Pragmatic verdict: for professional French, Qwen3-30B-A3B and Llama 4 Scout are neck and neck. The choice comes down to other criteria (speed, size, license). For literary or highly idiomatic French, Mistral Small remains a notch ahead.
#Which Qwen 3 should you choose for your setup?
- 6-8 GB of VRAM (RTX 3050, 3060 Ti, 4060)
- Qwen3-8B in Q4_K_M. The 14B runs in Q3, but quality deteriorates — stick with the clean 8B.
- 12 GB of VRAM (RTX 3060 12, 4070, 5070)
- Qwen3-14B in Q4_K_M is the undisputed sweet spot. Comfortable 16k context.
- 16 GB (RTX 4080, 5070 Ti, 5080)
- Qwen3-14B in Q5 or Q6 for comfort, or a 30B-A3B in Q3_K_M if you want to test the MoE.
- 24 GB (RTX 3090, 4090, RX 7900 XTX)
- Qwen3-30B-A3B in Q4_K_M. No debate—it’s the best general-purpose local model available today.
- Mac M-series with 32-48 GB unified memory
- Qwen3-30B-A3B in Q4 or Q5. The MoE is particularly well suited to unified memory Apple.
- Mac Studio Ultra 128 GB+
- Jump straight to Qwen3-235B-A22B (~120 GB in Q4) if you want the top of the range. Otherwise, the 30B-A3B in Q8 remains excellent.
#Go further
If Qwen 3 is your starting point, know that the next generation has since taken over in our rankings — Qwen 3.5 9B has become the default choice for 8 GB configurations (256k context, vision), and Qwen 3.8 27B for 24 GB configurations; the settings below remain valid unchanged for these models. A few ways to go further:
- Choose your quantization
- To understand exactly what you give up between Q4_K_M and Q5_K_M on Qwen 3, and when moving up in quantization makes sense.
- Customize with a Modelfile
- To set the thinking mode, context, and a default FR system prompt on your Qwen 3 instance—instead of configuring them again each session.
- Open WebUI with Ollama
- To move from the terminal to a real chat interface with history, Markdown, and attachments without giving up local operation.
Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.