Intermediate 15 minTools

Configure DeepSeek with Ollama: Deployment Guide complet

DeepSeek covers two very different needs: generating code (Coder) and reasoning step by step (R1). This guide shows how to configure DeepSeek with Ollama properly—choose the right variant, set the temperature and context window, fit it into its VRAM, and serve multiple requests in parallel. The goal: stable, fast inference without trial and error with the parameters.

By Mohamed Meguedmi·Update 2026-08-28·Tested on Windows, macOS, and Linux

#Why configure DeepSeek with Ollama

Ollama handles downloading the weights, distributing the workload between GPU and CPU, and the local API for you. You get a server listening on http://localhost:11434 and a single command to launch any variant of DeepSeek. Everything else—temperature, context size, memory—is controlled through a few parameters that are better understood than endured.

A good configuration changes everything: an DeepSeek R1 launched at the wrong temperature can loop or truncate its reasoning; an DeepSeek Coder with too short a context forgets half your file. The goal of this guide is to set these options once and for all.

#DeepSeek Coder vs Chat vs R1: which one to install

The Local AI Kit

Your private ChatGPT, free, on your own machine in an hour — LM Studio, Ollama, Open WebUI, your documents, no cloud.

  • Lifetime online access
  • PDF + files
  • Lifetime updates

Three families coexist under the name DeepSeek in the Ollama library. They are not configured the same way because they are not intended for the same use.

deepseek-coder
Code-specialized: completion, function generation, refactoring. Available in small sizes (1.3B, 6.7B, 33B). Ideal when connected to an editor for local autocompletion.
deepseek-v3 / deepseek-llm (Chat)
General-purpose conversational model. Direct answers, without exposed chain of thought. For a versatile assistant in French and English.
deepseek-r1
Reasoning model: it lays out a chain of thought (<think> tags) before answering. Excellent at math, logic, and multi-step problems. Distilled versions 1.5B / 7B / 8B / 14B / 32B / 70B to run locally.
i
R1: they are distillations
The 7B to 70B sizes of deepseek-r1 on Ollama are distillations based on Qwen and Llama, not the complete 671B model. That’s what makes them runnable on a consumer-grade card—with some of their larger sibling’s reasoning ability.

Simple rule: code as you type → Coder; an assistant that chats → Chat/V3; a problem that requires reasoning → R1. There's nothing stopping you from installing all three; Ollama stores them side by side.

#Requirements and VRAM by size

Ollama must be installed and its service must be active. The default quantization for DeepSeek models on Ollama is Q4_K_M, the best quality-to-memory tradeoff for most cards. Here is the VRAM to plan for in Q4, weights only—allow 1 to 2 GB extra for the context.

1.5B–3B (distilled R1)
≈ 2 GB · runs even without a dedicated GPU, or effortlessly on a RTX 3060 12 GB.
7B – 8B
≈ 5 GB · RTX 3060 12 GB, RTX 4070 12 GB. The sweet spot for daily use.
14B
≈ 9 GB · RTX 4070 12 GB is tight, RTX 4080 16 GB is comfortable.
32B / 33B (Coder)
≈ 19 GB · RTX 4090 24 GB, or a Mac with 24–48 GB of unified memory.
70B
≈ 40 GB · two 24 GB GPUs, or an M4 Pro Mac with 48 GB or more.
→
Verify before downloading
Run `ollama ps` after an initial test: the PROCESSOR column shows 100% GPU if the model fits entirely in VRAM, or a GPU/CPU split if it spills over. A split means a noticeable slowdown; move down one model size or quantization level.

#Install and run DeepSeek

  1. 01
    Verify that Ollama is running
    The service must be running. `ollama --version` confirms the installation; the server listens on port 11434.
  2. 02
    Download the model
    Choose the variant and size with `ollama pull`. The tag after the colon sets the size (e.g., :7b, :14b, :32b). Without a tag, Ollama uses the default size.
  3. 03
    Start a session
    `ollama run` downloads the model if needed, then opens an interactive prompt. Type /bye to exit.
  4. 04
    Test the API
    The same model is immediately available through the local REST API, ready to connect to Open WebUI, an editor, or your scripts.
Terminal — download and run
# Modèle de raisonnement, distillation 7B
ollama pull deepseek-r1:7b

# Modèle de code, 6.7B
ollama pull deepseek-coder:6.7b

# Lancer une session interactive
ollama run deepseek-r1:7b

# Voir ce qui est chargé et où (GPU/CPU)
ollama ps
Terminal — REST API call
curl http://localhost:11434/api/generate -d '{
  "model": "deepseek-r1:7b",
  "prompt": "Explique la récursivité en une phrase.",
  "stream": false
}'

#Adjust temperature and context window

The two parameters with the greatest impact on quality are temperature (creativity vs. determinism) and num_ctx (context window size). The default values of Ollama are not always ideal, especially for R1.

temperature (R1)
0.5 to 0.7. DeepSeek recommends ~0.6 for R1: too low, and reasoning freezes; too high, and it rambles. Avoid 0 on a reasoning model.
temperature (Coder)
0.1 to 0.3. For code, you want deterministic and reproducible output, not creativity.
num_ctx
Context size in tokens. Often 4096 by default. Increase it to 8192 or 16384 for long files or long reasoning chains—at the cost of more VRAM.
top_p
About 0.95. Leave it as is in most cases; adjust the temperature first.
repeat_penalty
1.1 by default. Useful if R1 repeats itself in a loop in its chain of thought.
Interactive session — tune on the fly
ollama run deepseek-r1:7b
>>> /set parameter temperature 0.6
>>> /set parameter num_ctx 8192
>>> Résous : un train part à 60 km/h...
>>> /bye
API—per-request parameters
{
  "model": "deepseek-coder:6.7b",
  "prompt": "Écris une fonction Python de tri fusion.",
  "options": {
    "temperature": 0.2,
    "num_ctx": 8192,
    "top_p": 0.95
  },
  "stream": false
}
!
num_ctx costs memory
Doubling the context window significantly increases the VRAM used by the KV cache. On a barely adequate card, going from 4096 to 16384 can push part of the model onto the CPU. Increase the context only if you have a real use for it.

#Locking in your configuration with a Modelfile

Adjusting parameters for every session is tedious. A Modelfile creates a named variant that bundles your settings and a system prompt. You get a ready-to-use model that you invoke like any other.

Modelfile — deepseek-coder-fr
FROM deepseek-coder:6.7b

PARAMETER temperature 0.2
PARAMETER num_ctx 8192
PARAMETER top_p 0.95

SYSTEM """Tu es un assistant de programmation.
Réponds en français, commente le code, et privilégie la clarté."""
Terminal — create and use the variant
# Créer le modèle à partir du fichier
ollama create deepseek-coder-fr -f ./Modelfile

# L'utiliser comme n'importe quel modèle
ollama run deepseek-coder-fr

The same approach applies to R1: a Modelfile with temperature 0.6 and num_ctx 16384 gives you a calibrated reasoning model without having to think about it at every launch.

#Optimizing VRAM and memory

If a model spills out of VRAM, Ollama places the excess layers on the CPU — and tokens per second collapse. Three levers for staying full-GPU.

Lower quantization
Q4_K_M by default. Stick with it: it’s already the right compromise. Move up to Q5_K_M or Q8_0 only if your VRAM allows it and you’re looking for a step up in quality.
Choose the right size
A full-GPU 14B beats a 32B spilling onto the CPU, both in speed and usability. Aim for the largest size that fits entirely in VRAM.
Mastering num_ctx
The KV cache grows with the context. A reasonable context (8192) frees up room for the model weights.
Unload after use
OLLAMA_KEEP_ALIVE controls how long a model stays in memory after the last request. Lower it to free the GPU faster between models.
Terminal — free the GPU quickly
# Garder les modèles chargés 2 minutes seulement (défaut : 5 min)
export OLLAMA_KEEP_ALIVE=2m

# Décharger immédiatement un modèle précis
ollama stop deepseek-r1:7b

#Parallel execution and concurrency

Ollama can serve multiple requests simultaneously and keep multiple models loaded at once. Useful for a shared assistant or for running Coder and R1 at the same time. Two environment variables control this behavior.

OLLAMA_NUM_PARALLEL
Number of requests processed in parallel by the same model. Each request consumes its own share of the context, and therefore VRAM. Increase cautiously (2, 4) depending on the GPU.
OLLAMA_MAX_LOADED_MODELS
Number of distinct models kept in memory simultaneously. Set it to 2 to keep Coder and R1 hot at the same time, if VRAM allows.
Terminal — start with concurrency
# Servir 4 requêtes en parallèle, 2 modèles chargés
export OLLAMA_NUM_PARALLEL=4
export OLLAMA_MAX_LOADED_MODELS=2

# Redémarrer le service pour prendre en compte les variables
ollama serve
!
Parallelism comes at a VRAM cost
Each parallel request and each loaded model uses its own memory. On a barely adequate card, increasing these values quickly pushes work onto the CPU. Measure with `ollama ps` after raising the settings, and lower them again if PROCESSOR no longer shows 100% GPU.

#Common troubleshooting

R1 loops endlessly inside <think>
Temperature is too low or repeat_penalty is missing. Raise temperature to 0.6 and add repeat_penalty 1.1.
Very slow generation
The model spills over to the CPU. Check `ollama ps`: if PROCESSOR isn’t 100% GPU, reduce the size, quantization, or num_ctx.
Truncated response
num_predict (maximum number of generated tokens) is too low, or the context is full. Increase num_ctx and leave num_predict unrestricted (-1).
The code comes out with the reasoning tags
You are using R1 for code. Switch to deepseek-coder, which does not produce a chain of thought.

#Go further

With DeepSeek configured and calibrated, here are the natural next steps: dive deeper into R1, refine your variants, and choose the quantization based on your GPU.

Install DeepSeek R1 with Ollama
The dedicated guide to the reasoning model, distilled versions, and chain of thought, going beyond configuration.
Customize a model with the Ollama Modelfile
A system of templates, system prompts, and parameters—to industrialize the variants covered here.
Choose your quantization (Q4, Q5, Q8, FP16)
Understand the quality/VRAM tradeoffs to fit the largest possible DeepSeek on your GPU.
Did this guide help you?

Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.