🇨🇳 DeepSeek R1 Distill 32B
The best accessible open-weight reasoner.
ollama run deepseek-r1:32b
Ranking updated on 09/10/2026
Ranking of specialized reasoning LLMs: those that produce an explicit chain of thought before answering. Excellent at math, formal logic, debugging, and scientific questions requiring multiple steps.
The best accessible open-weight reasoner.
ollama run deepseek-r1:32b
Apache 2.0 RL reasoner. AIME24 79.5, MATH-500 90.6. Direct competitor to DeepSeek R1.
ollama run qwq:32b
MIT 14B reasoner. Beats R1-Distill-Llama-70B on AIME/GPQA with 50× fewer parameters.
ollama run phi4-reasoning:14b
Distilled R1 Qwen 14B. AIME24 69.7, MATH-500 93.9. Outperforms o1-mini on many benchmarks.
ollama run deepseek-r1:14b
Dense multimodal 27B released April 22, 2026. 262k ctx (1M YaRN). SWE-bench Verified 77.2%.
ollama run qwen3.6:27b
Qwen 3.8 27B: dense multimodal (text + vision), 262k context, ~16 GB Q4 VRAM (18 GB of Ollama weights). Apache 2.0, agentic coding and vision.
ollama run qwen3.8:27b
| Rank | Model | Params | Q4 VRAM | Context | License |
|---|---|---|---|---|---|
| #1 | DeepSeek R1 Distill 32B | 32B | 19 GB | 32 768 | MIT |
| #2 | QwQ 32B | 32B | 19 GB | 131 072 | Apache 2.0 |
| #3 | Phi-4 Reasoning 14B | 14B | 9 GB | 32 768 | MIT |
| #4 | DeepSeek R1 Distill Qwen 14B | 14B | 9 GB | 131 072 | MIT |
| #5 | Qwen 3.6 27B | 27B | 16 GB | 262 144 | Apache 2.0 |
| #6 | Qwen 3.8 27B | 27B | 16 GB | 262 144 | Apache 2.0 |
Your private, free ChatGPT on your machine in 1 hour — LM Studio, Ollama, Open WebUI, your documents, no cloud.
We filter by the “reasoning” tag—models explicitly trained to deploy a chain of thought (tokens
Criteria considered:
The scoring is fully transparent: see our methodology for details on VRAM/tokens/sec calculations.
What is a reasoning LLM?
A model that first generates a chain of thought (a sequence of internal steps) before producing the final answer. It is slower and more verbose, but much more accurate on multi-step problems (math, logic, complex debugging).
DeepSeek R1 or QwQ-32B: which is better?
View the comparison. The two are very close, DeepSeek R1 is slightly ahead on MATH, QwQ on GPQA. Choose based on license (both are MIT/Apache) and VRAM (same ~32B size).
Can these models be used in real time (chat)?
Possible but not recommended: they generate 2-5x more tokens than a normal model (the chain of thought is visible). Better: reserve them for questions that warrant it, and use a fast model (Mistral 7B, Llama 3.1 8B) for basic chat.
How much VRAM for DeepSeek R1 32B?
19 GB in Q4_K_M, 23 GB in Q5, 35 GB in Q8. An RTX 4090 (24 GB) fits comfortably in Q5. An RTX 3060 12 GB is out of the running at this size.
Learn more with our detailed head-to-head matchups of the finalists: