Magistral Small 24B vs Qwen 3 32B: mini-Mistral 2026

Le Magistral Small 24B is Mistral AI's compact reasoning model, designed to run locally on a single consumer GPU. Compared with it, the Qwen 3 32B Alibaba's offering has established itself as one of the leading open-weight models in the medium-sized dense-model category since its launch in 2025. This article compares the two in detail: VRAM requirements by quantization, inference speed on real hardware, licenses, results on public benchmarks, and concrete usage scenarios. Whether you're building a reasoning-focused workstation or looking for a versatile model for long-text processing, this page gives you the information you need to choose.


Overview of the two models

Magistral Small 24B

Magistral Small 24B belongs to the Magistral family from Mistral AI, alongside a larger variant (Magistral Medium, ~123 billion parameters). Its main characteristic is its focus structured reasoning : the model generates an intermediate chain of thought before producing its final answer, distinguishing it from Mistral models of previous generations. With 24 billion parameters, it targets the category of LLMs accessible on a 24 GB GPU in Q4 quantization.

The official technical datasheet is available at HuggingFace — mistralai/Magistral-Small-2506 and in theannounces Mistral AI.

Qwen 3 32B

The Qwen 3 32B is a model dense (non-MoE) produced by Alibaba Research, from the third generation of the Qwen family. Its dense architecture makes VRAM usage more predictable than the family's MoE variants, such as the Qwen 3 235B-A22B which activates only 22 billion parameters at inference despite its 235 billion total parameters. The Qwen 3 32B covers a broad spectrum: coding, mathematical reasoning, multilingual processing, and long context.

Official documentation on HuggingFace — Qwen/Qwen3-32B and on the blog Qwen.


Technical specifications and VRAM

For VRAM estimates, the rule of thumb is: parameters × 0.5 to 0.55 GB in Q4, × 1 GB in Q8, × 2 GB in FP16. The values below are estimates based on this method and community feedback—check them against your backend (llama.cpp, llamafile, LM Studio).

Magistral Small 24B

In Q4_K_M, the model fits comfortably on a RTX 4090 (24 GB), with room for the KV cache at short context lengths. In Q8, the card reaches its limit; an Ada RTX 6000 (48 GB) or an A5000 is preferable for long sessions.

Qwen 3 32B

The Qwen 3 32B runs in Q4 on a RTX 4090, but with a long context (64 K tokens+), the KV cache pushes the total beyond the 24 GB available. For comfortable use with extended context, a 48 GB card (RTX 6000 Ada, A40) is recommended.

To go further with VRAM calculations by quantization level, see the quelllm.fr VRAM guide.


Inference speed

The following figures are estimates based on community feedback about llama.cpp and LM Studio. They vary depending on the hardware, prompt length, backend, and exact quantization used.

On RTX 4090 (24 GB, Q4_K_M)

Magistral Small 24B - Generation : ~25–40 tokens/sec (estimated, short prompt) - Prefill : ~1,500–2,500 tokens/sec (estimated)

Qwen 3 32B - Generation : ~18–28 tokens/sec (estimated, short prompt) - Prefill : ~1,000–1,800 tokens/sec (estimated)

The lighter model (24B) generates noticeably faster on the same GPU—around 30 to 50% more tokens per second in Q4 (estimated). The gap narrows on a 48 GB GPU or in a multi-request batch configuration.

On Apple M3 Max (128 GB unified RAM)

Apple Silicon Macs can run both models comfortably, including the 24B model in Q8 with 64 GB of RAM.


Licenses

Magistral Small 24B Mistral AI has released its open models under Apache 2.0 or a Modified MIT license, depending on the version. The exact license for Magistral Small 24B must be checked on the official Hugging Face model card before any commercial use or redistribution. Do not assume the license based solely on the model family.

Qwen 3 32B Published under Apache 2.0, allowing commercial use, modification, and redistribution without royalties, provided that legal notices are retained. This is one of the practical advantages of Qwen 3 32B over models with more restrictive community licenses.

To filter catalog models by license type, go to quelllm.fr/catalogue.


Benchmarks and performance comparison

Benchmark figures vary depending on the model version, evaluation harness, and generation settings. The values below come from public reports or are estimates — check official leaderboards such as theHuggingFace Open LLM Leaderboard before basing deployment decisions on it.

MMLU (general knowledge)

The Qwen 3 32B benefits from Alibaba's extensive training data, particularly dense in academic and multilingual content, which is reflected in MMLU.

HumanEval (coding)

On pure coding benchmarks, Qwen 3 32B retains the advantage. However, Magistral Small's reasoning mode can make up the difference on complex algorithmic problems requiring multiple decomposition steps.

AIME (mathematical reasoning)

The Magistral Small 24B aims to stand out on reasoning benchmarks such as AIME. Its chain-of-thought-oriented design enables it to tackle competitive mathematics problems with a structured methodology rather than pattern matching. The precise scores still need to be confirmed on public leaderboards.

Positioning in the quelllm.fr catalog

To place these two models in the broader ecosystem:


Concrete use cases

When to choose Magistral Small 24B

When to choose Qwen 3 32B


FAQ

Q: Does Magistral Small 24B run on a RTX 4090?

Yes. In Q4_K_M, it requires about 13 GB of VRAM—a RTX 4090 (24 GB) runs it with enough headroom for the KV cache on short contexts. In Q8, all 24 GB available are used; very long contexts (32 K filled tokens) can saturate memory. In practice, Q5_K_M is a good quality/VRAM compromise on this card.

Q: What is the license for Qwen 3 32B?

The Qwen 3 32B is released under Apache 2.0, a permissive license that allows commercial use, modification, and redistribution. No royalty is required to integrate it into a product or service, provided that the original legal notices are retained. It is one of the most favorable licenses in the open-weight ecosystem.

Q: Magistral Small 24B or Qwen 3 32B for RAG (Retrieval-Augmented Generation)?

For RAG, Qwen 3 32B has an advantage thanks to its native 128 K-token context and strong document-understanding performance. Magistral Small 24B can produce better-structured, better-argued answers to complex queries thanks to its reasoning mode. The choice mainly depends on the length of the ingested documents and the complexity of the questions asked.

Q: Is there a significant speed difference between the two models?

Yes. With 8 billion fewer parameters, the Magistral Small 24B is significantly faster at generation on the same GPU—roughly 30 to 50% more tokens per second in Q4 on RTX 4090 (estimated). This difference fades on 80 GB GPUs or in batch deployments, where memory bandwidth is no longer the primary bottleneck.

Q: Does Magistral Small 24B support thinking mode?

Yes, that is its main characteristic compared with Mistral LLMs from the previous generation. It generates an intermediate reasoning process before the final answer. Activation occurs through the prompt format or a dedicated option, depending on the backend used (llama.cpp, LM Studio, etc.). Implementation details must be checked in the official HuggingFace documentation.

Q: How do these models compare with other LLMs in the quelllm.fr catalog?

Le quelllm.fr catalog lists 249 models with filterable VRAM, license, and tokens/sec specs. For a direct comparison with another model of similar size, the built-in comparator lets you compare the entries side by side according to your criteria (available VRAM, license, use case).


Conclusion

Le Magistral Small 24B and the Qwen 3 32B embody two philosophies for a similar form factor: structured reasoning on one side, dense versatility on the other. If your priority is inference speed and step-by-step reasoning quality on a 24 GB GPU, the Magistral Small 24B deserves serious consideration. If you prioritize coding, long context, and an unambiguous Apache 2.0 license, the Qwen 3 32B is the solid choice. Explore all models filtered by VRAM and license on quelllm.fr/catalogue, or leave the configurator sort through them based on your exact configuration.

Article published and updated on by Mohamed Meguedmi · Data source: /api/models.json · Content license: CC BY 4.0.

An error or update to report? Contribute.