Magistral Small 24B vs Qwen 3 32B: mini-Mistral 2026
Le Magistral Small 24B is Mistral AI's compact reasoning model, designed to run locally on a single consumer GPU. Compared with it, the Qwen 3 32B Alibaba's offering has established itself as one of the leading open-weight models in the medium-sized dense-model category since its launch in 2025. This article compares the two in detail: VRAM requirements by quantization, inference speed on real hardware, licenses, results on public benchmarks, and concrete usage scenarios. Whether you're building a reasoning-focused workstation or looking for a versatile model for long-text processing, this page gives you the information you need to choose.
Overview of the two models
Magistral Small 24B
Magistral Small 24B belongs to the Magistral family from Mistral AI, alongside a larger variant (Magistral Medium, ~123 billion parameters). Its main characteristic is its focus structured reasoning : the model generates an intermediate chain of thought before producing its final answer, distinguishing it from Mistral models of previous generations. With 24 billion parameters, it targets the category of LLMs accessible on a 24 GB GPU in Q4 quantization.
The official technical datasheet is available at HuggingFace — mistralai/Magistral-Small-2506 and in theannounces Mistral AI.
Qwen 3 32B
The Qwen 3 32B is a model dense (non-MoE) produced by Alibaba Research, from the third generation of the Qwen family. Its dense architecture makes VRAM usage more predictable than the family's MoE variants, such as the Qwen 3 235B-A22B which activates only 22 billion parameters at inference despite its 235 billion total parameters. The Qwen 3 32B covers a broad spectrum: coding, mathematical reasoning, multilingual processing, and long context.
Official documentation on HuggingFace — Qwen/Qwen3-32B and on the blog Qwen.
Technical specifications and VRAM
For VRAM estimates, the rule of thumb is: parameters × 0.5 to 0.55 GB in Q4, × 1 GB in Q8, × 2 GB in FP16. The values below are estimates based on this method and community feedback—check them against your backend (llama.cpp, llamafile, LM Studio).
Magistral Small 24B
- Settings : 24 billion
- Architecture : dense, transformer decoder (architecture details to be confirmed)
- Q4_K_M : ~13 GB VRAM
- Q5_K_M : ~16 GB of VRAM (estimated)
- Q8_0 : ~24 GB VRAM
- FP16 : ~48 GB of VRAM
- Context : to be confirmed depending on the published variant; recent Mistral models target 32 768 to 128 000 tokens
In Q4_K_M, the model fits comfortably on a RTX 4090 (24 GB), with room for the KV cache at short context lengths. In Q8, the card reaches its limit; an Ada RTX 6000 (48 GB) or an A5000 is preferable for long sessions.
Qwen 3 32B
- Settings : 32 billion
- Architecture : dense, transformer decoder
- Q4_K_M : ~18 GB VRAM
- Q5_K_M : ~22 GB VRAM (estimated)
- Q8_0 : ~32 GB VRAM
- FP16 : ~64 GB VRAM
- Context : 128,000 tokens
The Qwen 3 32B runs in Q4 on a RTX 4090, but with a long context (64 K tokens+), the KV cache pushes the total beyond the 24 GB available. For comfortable use with extended context, a 48 GB card (RTX 6000 Ada, A40) is recommended.
To go further with VRAM calculations by quantization level, see the quelllm.fr VRAM guide.
Inference speed
The following figures are estimates based on community feedback about llama.cpp and LM Studio. They vary depending on the hardware, prompt length, backend, and exact quantization used.
On RTX 4090 (24 GB, Q4_K_M)
Magistral Small 24B - Generation : ~25–40 tokens/sec (estimated, short prompt) - Prefill : ~1,500–2,500 tokens/sec (estimated)
Qwen 3 32B - Generation : ~18–28 tokens/sec (estimated, short prompt) - Prefill : ~1,000–1,800 tokens/sec (estimated)
The lighter model (24B) generates noticeably faster on the same GPU—around 30 to 50% more tokens per second in Q4 (estimated). The gap narrows on a 48 GB GPU or in a multi-request batch configuration.
On Apple M3 Max (128 GB unified RAM)
- Magistral Small 24B Q4 : ~20–30 tokens/sec (estimated)
- Qwen 3 32B Q4 : ~14–22 tokens/sec (estimated)
Apple Silicon Macs can run both models comfortably, including the 24B model in Q8 with 64 GB of RAM.
Licenses
Magistral Small 24B Mistral AI has released its open models under Apache 2.0 or a Modified MIT license, depending on the version. The exact license for Magistral Small 24B must be checked on the official Hugging Face model card before any commercial use or redistribution. Do not assume the license based solely on the model family.
Qwen 3 32B Published under Apache 2.0, allowing commercial use, modification, and redistribution without royalties, provided that legal notices are retained. This is one of the practical advantages of Qwen 3 32B over models with more restrictive community licenses.
To filter catalog models by license type, go to quelllm.fr/catalogue.
Benchmarks and performance comparison
Benchmark figures vary depending on the model version, evaluation harness, and generation settings. The values below come from public reports or are estimates — check official leaderboards such as theHuggingFace Open LLM Leaderboard before basing deployment decisions on it.
MMLU (general knowledge)
- Magistral Small 24B : ~70–76% (to be confirmed)
- Qwen 3 32B : ~80–85% (to be confirmed, instruct version)
The Qwen 3 32B benefits from Alibaba's extensive training data, particularly dense in academic and multilingual content, which is reflected in MMLU.
HumanEval (coding)
- Magistral Small 24B : ~60–72% (estimated, with reasoning mode enabled)
- Qwen 3 32B : ~75–85% (estimated)
On pure coding benchmarks, Qwen 3 32B retains the advantage. However, Magistral Small's reasoning mode can make up the difference on complex algorithmic problems requiring multiple decomposition steps.
AIME (mathematical reasoning)
The Magistral Small 24B aims to stand out on reasoning benchmarks such as AIME. Its chain-of-thought-oriented design enables it to tackle competitive mathematics problems with a structured methodology rather than pattern matching. The precise scores still need to be confirmed on public leaderboards.
Positioning in the quelllm.fr catalog
To place these two models in the broader ecosystem:
- Le Mistral Medium 3.5 128B (128B, ~74 GB Q4) represents the top tier in the Mistral lineup, beyond the reach of a single consumer GPU.
- Le Mistral Small 4 (119B, ~72 GB Q4) offers significantly greater Mistral power but also requires a cluster or an 80 GB professional card.
- Le Qwen 2.5 72B Instruct remains a proven reference in the Qwen family for those with ~42 GB of VRAM in Q4.
Concrete use cases
When to choose Magistral Small 24B
- Step-by-step reasoning : multi-step code analysis, debugging, and logical or mathematical problem-solving.
- Lightweight autonomous agent : its structured reasoning capability makes it a good engine for agent pipelines that need to justify their decisions.
- Tight VRAM budget : on a machine with 16–24 GB, it's the most accessible Mistral option that includes advanced reasoning capabilities.
- Speed priority : generate quickly (~30 t/s estimated) with reasoning quality superior to what its size would suggest.
When to choose Qwen 3 32B
- Coding and code review : its HumanEval scores make it a better choice for everyday development assistance.
- Multilingual processing : the Qwen family is natively strong in Chinese, Japanese, Korean, and European languages.
- Long context : 128K native tokens for ingesting entire documents or long conversation histories.
- Unambiguous Apache 2.0 license : direct commercial deployment without prior verification.
- General versatility : writing, summarization, Q&A, data analysis—the range is broad.
FAQ
Q: Does Magistral Small 24B run on a RTX 4090?
Yes. In Q4_K_M, it requires about 13 GB of VRAM—a RTX 4090 (24 GB) runs it with enough headroom for the KV cache on short contexts. In Q8, all 24 GB available are used; very long contexts (32 K filled tokens) can saturate memory. In practice, Q5_K_M is a good quality/VRAM compromise on this card.
Q: What is the license for Qwen 3 32B?
The Qwen 3 32B is released under Apache 2.0, a permissive license that allows commercial use, modification, and redistribution. No royalty is required to integrate it into a product or service, provided that the original legal notices are retained. It is one of the most favorable licenses in the open-weight ecosystem.
Q: Magistral Small 24B or Qwen 3 32B for RAG (Retrieval-Augmented Generation)?
For RAG, Qwen 3 32B has an advantage thanks to its native 128 K-token context and strong document-understanding performance. Magistral Small 24B can produce better-structured, better-argued answers to complex queries thanks to its reasoning mode. The choice mainly depends on the length of the ingested documents and the complexity of the questions asked.
Q: Is there a significant speed difference between the two models?
Yes. With 8 billion fewer parameters, the Magistral Small 24B is significantly faster at generation on the same GPU—roughly 30 to 50% more tokens per second in Q4 on RTX 4090 (estimated). This difference fades on 80 GB GPUs or in batch deployments, where memory bandwidth is no longer the primary bottleneck.
Q: Does Magistral Small 24B support thinking mode?
Yes, that is its main characteristic compared with Mistral LLMs from the previous generation. It generates an intermediate reasoning process before the final answer. Activation occurs through the prompt format or a dedicated option, depending on the backend used (llama.cpp, LM Studio, etc.). Implementation details must be checked in the official HuggingFace documentation.
Q: How do these models compare with other LLMs in the quelllm.fr catalog?
Le quelllm.fr catalog lists 249 models with filterable VRAM, license, and tokens/sec specs. For a direct comparison with another model of similar size, the built-in comparator lets you compare the entries side by side according to your criteria (available VRAM, license, use case).
Conclusion
Le Magistral Small 24B and the Qwen 3 32B embody two philosophies for a similar form factor: structured reasoning on one side, dense versatility on the other. If your priority is inference speed and step-by-step reasoning quality on a 24 GB GPU, the Magistral Small 24B deserves serious consideration. If you prioritize coding, long context, and an unambiguous Apache 2.0 license, the Qwen 3 32B is the solid choice. Explore all models filtered by VRAM and license on quelllm.fr/catalogue, or leave the configurator sort through them based on your exact configuration.