Configure DeepSeek with Ollama: Deployment Guide complet
DeepSeek covers two very different needs: generating code (Coder) and reasoning step by step (R1). This guide shows how to configure DeepSeek with Ollama properly—choose the right variant, set the temperature and context window, fit it into its VRAM, and serve multiple requests in parallel. The goal: stable, fast inference without trial and error with the parameters.
#Why configure DeepSeek with Ollama
Ollama handles downloading the weights, distributing the workload between GPU and CPU, and the local API for you. You get a server listening on http://localhost:11434 and a single command to launch any variant of DeepSeek. Everything else—temperature, context size, memory—is controlled through a few parameters that are better understood than endured.
A good configuration changes everything: an DeepSeek R1 launched at the wrong temperature can loop or truncate its reasoning; an DeepSeek Coder with too short a context forgets half your file. The goal of this guide is to set these options once and for all.
#DeepSeek Coder vs Chat vs R1: which one to install
Your private ChatGPT, free, on your own machine in an hour — LM Studio, Ollama, Open WebUI, your documents, no cloud.
- Lifetime online access
- PDF + files
- Lifetime updates
Three families coexist under the name DeepSeek in the Ollama library. They are not configured the same way because they are not intended for the same use.
- deepseek-coder
- Code-specialized: completion, function generation, refactoring. Available in small sizes (1.3B, 6.7B, 33B). Ideal when connected to an editor for local autocompletion.
- deepseek-v3 / deepseek-llm (Chat)
- General-purpose conversational model. Direct answers, without exposed chain of thought. For a versatile assistant in French and English.
- deepseek-r1
- Reasoning model: it lays out a chain of thought (<think> tags) before answering. Excellent at math, logic, and multi-step problems. Distilled versions 1.5B / 7B / 8B / 14B / 32B / 70B to run locally.
Simple rule: code as you type → Coder; an assistant that chats → Chat/V3; a problem that requires reasoning → R1. There's nothing stopping you from installing all three; Ollama stores them side by side.
#Requirements and VRAM by size
Ollama must be installed and its service must be active. The default quantization for DeepSeek models on Ollama is Q4_K_M, the best quality-to-memory tradeoff for most cards. Here is the VRAM to plan for in Q4, weights only—allow 1 to 2 GB extra for the context.
- 1.5B–3B (distilled R1)
- ≈ 2 GB · runs even without a dedicated GPU, or effortlessly on a RTX 3060 12 GB.
- 7B – 8B
- ≈ 5 GB · RTX 3060 12 GB, RTX 4070 12 GB. The sweet spot for daily use.
- 14B
- ≈ 9 GB · RTX 4070 12 GB is tight, RTX 4080 16 GB is comfortable.
- 32B / 33B (Coder)
- ≈ 19 GB · RTX 4090 24 GB, or a Mac with 24–48 GB of unified memory.
- 70B
- ≈ 40 GB · two 24 GB GPUs, or an M4 Pro Mac with 48 GB or more.
#Install and run DeepSeek
- 01Verify that Ollama is runningThe service must be running. `ollama --version` confirms the installation; the server listens on port 11434.
- 02Download the modelChoose the variant and size with `ollama pull`. The tag after the colon sets the size (e.g., :7b, :14b, :32b). Without a tag, Ollama uses the default size.
- 03Start a session`ollama run` downloads the model if needed, then opens an interactive prompt. Type /bye to exit.
- 04Test the APIThe same model is immediately available through the local REST API, ready to connect to Open WebUI, an editor, or your scripts.
#Adjust temperature and context window
The two parameters with the greatest impact on quality are temperature (creativity vs. determinism) and num_ctx (context window size). The default values of Ollama are not always ideal, especially for R1.
- temperature (R1)
- 0.5 to 0.7. DeepSeek recommends ~0.6 for R1: too low, and reasoning freezes; too high, and it rambles. Avoid 0 on a reasoning model.
- temperature (Coder)
- 0.1 to 0.3. For code, you want deterministic and reproducible output, not creativity.
- num_ctx
- Context size in tokens. Often 4096 by default. Increase it to 8192 or 16384 for long files or long reasoning chains—at the cost of more VRAM.
- top_p
- About 0.95. Leave it as is in most cases; adjust the temperature first.
- repeat_penalty
- 1.1 by default. Useful if R1 repeats itself in a loop in its chain of thought.
#Locking in your configuration with a Modelfile
Adjusting parameters for every session is tedious. A Modelfile creates a named variant that bundles your settings and a system prompt. You get a ready-to-use model that you invoke like any other.
The same approach applies to R1: a Modelfile with temperature 0.6 and num_ctx 16384 gives you a calibrated reasoning model without having to think about it at every launch.
#Optimizing VRAM and memory
If a model spills out of VRAM, Ollama places the excess layers on the CPU — and tokens per second collapse. Three levers for staying full-GPU.
- Lower quantization
- Q4_K_M by default. Stick with it: it’s already the right compromise. Move up to Q5_K_M or Q8_0 only if your VRAM allows it and you’re looking for a step up in quality.
- Choose the right size
- A full-GPU 14B beats a 32B spilling onto the CPU, both in speed and usability. Aim for the largest size that fits entirely in VRAM.
- Mastering num_ctx
- The KV cache grows with the context. A reasonable context (8192) frees up room for the model weights.
- Unload after use
- OLLAMA_KEEP_ALIVE controls how long a model stays in memory after the last request. Lower it to free the GPU faster between models.
#Parallel execution and concurrency
Ollama can serve multiple requests simultaneously and keep multiple models loaded at once. Useful for a shared assistant or for running Coder and R1 at the same time. Two environment variables control this behavior.
- OLLAMA_NUM_PARALLEL
- Number of requests processed in parallel by the same model. Each request consumes its own share of the context, and therefore VRAM. Increase cautiously (2, 4) depending on the GPU.
- OLLAMA_MAX_LOADED_MODELS
- Number of distinct models kept in memory simultaneously. Set it to 2 to keep Coder and R1 hot at the same time, if VRAM allows.
#Common troubleshooting
- R1 loops endlessly inside <think>
- Temperature is too low or repeat_penalty is missing. Raise temperature to 0.6 and add repeat_penalty 1.1.
- Very slow generation
- The model spills over to the CPU. Check `ollama ps`: if PROCESSOR isn’t 100% GPU, reduce the size, quantization, or num_ctx.
- Truncated response
- num_predict (maximum number of generated tokens) is too low, or the context is full. Increase num_ctx and leave num_predict unrestricted (-1).
- The code comes out with the reasoning tags
- You are using R1 for code. Switch to deepseek-coder, which does not produce a chain of thought.
#Go further
With DeepSeek configured and calibrated, here are the natural next steps: dive deeper into R1, refine your variants, and choose the quantization based on your GPU.
- Install DeepSeek R1 with Ollama
- The dedicated guide to the reasoning model, distilled versions, and chain of thought, going beyond configuration.
- Customize a model with the Ollama Modelfile
- A system of templates, system prompts, and parameters—to industrialize the variants covered here.
- Choose your quantization (Q4, Q5, Q8, FP16)
- Understand the quality/VRAM tradeoffs to fit the largest possible DeepSeek on your GPU.
Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.