Install DeepSeek R1 with Ollama
DeepSeek R1 is a chain-of-thought reasoning model released under the MIT license by DeepSeek. The full model has 671 billion parameters and remains out of reach for a personal machine, but the distilled versions (1.5B to 70B) run comfortably on consumer hardware. This tutorial shows how to install DeepSeek R1 with Ollama, choose the right size, and take advantage of reasoning mode.
#Why DeepSeek R1?
R1 isn't a "classic" LLM: before answering, it generates a visible chain of thought enclosed in <think>...</think> tags. This phase enables it to solve math, logic, or coding problems that direct-generation models often miss. On AIME 2024 and MATH-500, the 14B and 32B distillates decisively outperform Llama 3.1 70B Instruct, which is nevertheless 2 to 5 times larger.
The other advantage is the MIT license, with no commercial-use or geographic restrictions. You can integrate R1 into a product, agent, or internal tool without asking anyone for permission. That's rare for a reasoning model at this level.
#Available distilled versions
Your private ChatGPT, free, on your own machine in an hour — LM Studio, Ollama, Open WebUI, your documents, no cloud.
- Lifetime online access
- PDF + files
- Lifetime updates
Ollama distributes seven sizes under the same deepseek-r1 name. The default tag points to the distilled 7B (Qwen 7B base).
- deepseek-r1:1.5b
- Qwen 2.5 Math 1.5B base model. Lightweight, perfect for testing on a laptop without a GPU. Limited quality on complex problems.
- deepseek-r1:7b
- Base Qwen 2.5 Math 7B. Default tag. A good fit for 8 GB of VRAM. An excellent starting point.
- deepseek-r1:8b
- Base Llama 3.1 8B. Slightly more verbose than the 7B, better in French.
- deepseek-r1:14b
- Qwen 2.5 14B base. Noticeable jump in math quality. Sweet spot for 12 GB of VRAM.
- deepseek-r1:32b
- Base Qwen 2.5 32B. o1-mini-level performance on most reasoning benchmarks. Requires 24 GB of VRAM in Q4.
- deepseek-r1:70b
- Base Llama 3.3 70B. The most capable distillate. RTX 5090 32 GB or Mac Studio 64 GB+.
- deepseek-r1:671b
- The full model (MoE 671B, 37B active). Beyond a personal machine, but possible on a Mac Studio Ultra 512 GB.
#Requirements and VRAM by size
Before installing DeepSeek R1 with Ollama, verify that you have Ollama installed (the daemon listens by default on http://localhost:11434) and that your GPU has enough memory for the target size. Here are the VRAM requirements for Q4_K_M quantization, which is what Ollama downloads by default.
- 1.5B Q4
- ~1.1 GB VRAM. Runs on any GPU or even on pure CPU (15–25 tok/s).
- 7B Q4
- ~4.7 GB VRAM. RTX 3050 8 GB, 4060 8 GB, Mac M1/M2 16 GB.
- 8B Q4
- ~5.2 GB VRAM. Same target as the 7B, with tighter headroom on 6 GB.
- 14B Q4
- ~9 GB VRAM. RTX 3060 12 GB, 4070 12 GB, M2 Pro 16 GB.
- 32B Q4
- ~19 GB VRAM. RTX 3090/4090 24 GB, M3 Max 36 GB.
- 70B Q4
- ~40 GB VRAM. RTX 5090 32 GB with offload, Mac Studio Ultra 64 GB+.
#1. Install the model
A single command is enough. Ollama downloads the weights from its official registry and loads the model into memory on first use.
For a specific size, add the tag:
The first launch downloads the model (from 1 GB for the 1.5B to 40 GB for the 70B). Subsequent launches are instant as long as the model remains in the Ollama cache.
#2. Choose your size
The size choice depends on three variables: VRAM, task type, and your tolerance for waiting. R1 produces an average of 500 to 2000 reasoning tokens before the final answer—so that's 3 to 10 times more tokens than a conventional chat model with direct generation. Speed really matters.
- Test R1 without a serious GPU
- 1.5B or 7B. Usable chat latency even on a CPU.
- High-school math / simple Python
- 7B or 8B. Enough 80% of the time, disappointing for advanced algebra.
- Serious math, logic, and long-form reasoning
- 14B minimum. A clear jump between 8B and 14B on AIME and MATH-500.
- o1-mini level, complex debugging, demonstrations
- 32B. This is the ideal target if you have 24 GB of VRAM.
- The maximum possible locally
- 70B Q4 on RTX 5090 with offload or a Mac Studio with 64 GB+.
#3. First reasoning prompt
R1 shines on problems that require multiple steps. Let’s test it on an AIME classic, approachable with the 14B:
Pay close attention to the structure: everything between <think> and </think> is the model's chain of thought. You can display it to the end user, hide it, or log it for debugging. It's an excellent explainability lever in an app.
#4. Q4 vs Q8: choosing the quantization
By default, Ollama downloads Q4_K_M (the recommended quality/memory compromise). For reasoning models, some prefer moving up to Q8_0 — quality degradation in the chain of thought is more noticeable than in standard chat.
- Q4_K_M
- Default compromise. -2 to -4% quality vs. FP16 on math benchmarks. This is the one to choose first.
- Q5_K_M
- +1% quality vs. Q4, +25% memory. Interesting if you have spare VRAM but can’t move up to the next size.
- Q8_0
- Near-FP16 quality (-0.5%). Twice the memory of Q4. Recommended if you want the best for the 14B or 32B and have the VRAM.
- FP16
- Full-precision reference. No practical benefit over Q8 for distillates, and 4× heavier than Q4.
#DeepSeek R1 vs Qwen 3.8 27B
Qwen 3.8 27B, released on August 14, 2026, by the Qwen team, is the benchmark open-weight general-purpose model of 2026: 262k context tokens, vision, an Apache 2.0 license, and above all an integrated reasoning mode. It is the natural alternative to R1 when you want a single model capable of reasoning AND doing everything else. Which one should you choose?
- DeepSeek R1 32B
- A specialist in pure reasoning. Best on math benchmarks (AIME, MATH-500). Reasoning is always visible and easy to parse via <think>...</think>. That is all it does, but it does it very well.
- Qwen 3.8 27B
- General-purpose 2026 model with built-in reasoning. 262k context, vision, Apache 2.0, ~18 GB in Q4. By default, it overthinks: even for a simple question, it produces an unnecessarily long chain of thought. Set its reasoning level to low to get direct answers again.
- Practical verdict
- R1 32B remains ahead on pure math/logic and for a stable, loggable reasoning format. Qwen 3.8 27B wins as soon as you want a versatile model (code, chat, vision, long context) that reasons on demand. For French, both are solid.
- Resources
- Comparable. R1 32B fits in Q4 on 24 GB (RTX 3090/4090). Qwen 3.8 27B is a little lighter (~18 GB in Q4) and leaves more room for context on the same card.
#Troubleshooting
- Model that does not display <think>
- You may be using a client that hides XML tags. Test with ollama run en CLI to verify that the model generates them correctly. If the JSON API strips them, the issue is on the client side.
- Truncated responses
- The default context for Ollama is 4096 tokens. R1 consumes a lot of them for reasoning. Increase it with /set parameter num_ctx 16384 in the session, or via the options.num_ctx API parameter.
- Very slow despite a good GPU
- Check with ollama ps that the model is using 100% GPU. If the PROCESSOR column shows CPU, you are spilling into RAM. Move down one size or switch to more aggressive quantization.
- The model loops on its chain of thought
- This is typical of small distillates (1.5B, 7B) on problems that are too difficult. Either use a larger model, ask a simpler question, or add “answer in fewer than 500 reasoning tokens” to your prompt.
- “insufficient memory” error
- Model too large for your VRAM. ollama rm deepseek-r1:32b then ollama pull deepseek-r1:14b. You can also force partial offload with OLLAMA_GPU_LAYERS.
#Go further
A few natural next steps after this installation:
- Understanding quantization
- The site's Q4/Q5/Q8 guide details the precise trade-offs, which is useful before moving up in size.
- Choosing a GPU for R1
- If you're seriously targeting 32B, the GPU buying guide gives the 24/32 GB thresholds to aim for in 2026.
- Connect R1 to an interface
- Open WebUI or LM Studio natively handle <think> tags and let you collapse/expand the reasoning.
Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.