Advanced 16 minDeepSeek

DeepSeek V4 Pro locally: VRAM, hardware, installation

DeepSeek V4 Pro is the world's most powerful open-weight model at its release on April 24, 2026. It has 1.6 trillion parameters in a MoE architecture (49 billion active per token), an MIT license, a 1M-token context, hybrid CSA+HCA attention, and three reasoning modes. On published benchmarks, it rivals GPT-5 and Claude 4.7 — with no cloud server, no API, and no token limit. This guide explains how to run it locally and what hardware you need.

By Mohamed Meguedmi·Update 2026-04-25·Tested on Windows, macOS, and Linux

#DeepSeek V4 Pro at a glance

Output
April 24, 2026 — DeepSeek-AI on HuggingFace, MIT-licensed weights, paper published simultaneously.
Architecture
Dense Mixture-of-Experts: 1.6T total parameters, 49B active per token, 256 experts with top-8 routing.
Warning
Hybrid CSA (Compressed Sparse Attention) + HCA (Hybrid Constrained Attention)—the first implementation at this scale.
Context
1,000,000 native tokens (with extended RoPE and flash attention v3).
Pretraining
32T+ tokens, Muon optimizer (replaces AdamW), mixed FP4+FP8 precision (the most efficient training known).
Thinking modes
3 levels: Non / High / Max. Max mode approaches o4 performance on AIME and GPQA.
License
MIT — commercial use, redistribution, and modification allowed without restriction. More permissive than the Llama license.
→
Why this is an event
DeepSeek V4 Pro is the first open-weights model to surpass GPT-4o on every public benchmark, at a reported training cost of $12M (vs. ~$80M for GPT-4). Released under the MIT license, it can be self-hosted by a government, bank, or hospital without an agreement with a cloud provider. This is a paradigm shift for AI sovereignty.

#1. CSA+HCA and 1.6T MoE architecture

DeepSeek V4 Pro combines three major innovations, each published separately over the last 12 months and combined here for the first time at trillion scale.

The Local AI Kit

Your private ChatGPT, free, on your own machine in an hour — LM Studio, Ollama, Open WebUI, your documents, no cloud.

  • Lifetime online access
  • PDF + files
  • Lifetime updates
1.6T MoE / 49B active
256 experts per layer, top-8 routing. For each token, only 49 billion parameters compute—hence speed comparable to a dense 50B model despite the total size.
CSA — Compressed Sparse Attention
Compresses the old keys/values in the KV cache using low-rank decomposition. Cuts attention memory by a factor of 4 for contexts >128k.
HCA — Hybrid Constrained Attention
Combines local attention (4k sliding window) and global attention (sparse), depending on the layer. 2× throughput on long contexts vs. dense attention.
mHC — multi-Head Compressor
Cross-head compression mechanism — reduces activation VRAM with no measurable loss in quality.
Muon optimizer
AdamW's successor, converging 1.8× faster during the pretraining of very large models. Adopted by DeepSeek since V3.5.
FP4+FP8 training
FP4 mix on attention layers, FP8 on MLPs. -45% training memory vs. pure FP8, with quality preserved through a proprietary scaling scheme.
i
What this changes in practice
At comparable total parameter counts (1.6T), DeepSeek V4 Pro is ~3× faster at inference than an equivalent dense model and uses 4× less VRAM on long contexts. This makes local Q4 execution feasible on an 8×H100 80 GB cluster (instead of having to rent a TPU pod).

#2. The 3 thinking modes (Non / High / Max)

DeepSeek V4 Pro's thinking system is explicit and toggleable, unlike the opaque modes of competing proprietary models.

No mode (default)
Direct responses with no visible chain-of-thought. Minimal latency (~2–5s at 8 tok/s). For chat, translation, and simple text generation.
High mode
Inserts 200-2000 reasoning tokens before the response. +30-40% quality on math, code, and analysis. Triggered via the /think flag in Ollama, or the thinking_mode header in the API.
Max mode
Exhaustive reasoning (up to 16k chain-of-thought tokens). Approaches o4 performance on AIME 2025 (87% vs. 91% for o4). Expensive: 10–30 minutes per request locally.
Enable thinking mode via Ollama (V4 Flash, identical in Pro)
# Mode High
ollama run deepseek-v4 "Résous : x² + 5x - 14 = 0 /think"

# Mode Max
ollama run deepseek-v4 "Démontre la conjecture de Collatz pour n<100 /think_max"

#3. Hardware required by quantization

Total VRAM required (model + 32k-token KV cache)
QuantizationModel VRAM+ KV 32kPossible configurations
FP16 (reference)3,200 GB~3,350 GBCloud: 16× H100 80GB or 8× H200 141GB. Not accessible locally.
Q8_01,700 GB~1,800 GB8× H200 141GB cluster. ~250 k$ in hardware.
Q5_K_M1 150 GB~1,230 GBMac Studio 512 GB cluster ×3, or 4× H200 141GB.
Q4_K_M960 GB~1 040 GB8× A100 80GB cluster, or 2× Mac Studio 512 GB in a network swarm.
IQ3_M (compressed)720 GB~800 GBSingle Mac Studio 512 + partial SSD offload. Slow but possible.
IQ2_XS (extreme)480 GB~560 GBMac Studio 512 GB only; quality is severely degraded.
!
Local realism
DeepSeek V4 Pro Q4 remains at 960 GB—this is a data-center model, not for individual workstations. If you want a local single-machine frontier in 2026, look at DeepSeek V4 Flash (170 GB in Q4) or wait for the Pro→70B distillations arriving in the next few weeks.

#4. Local installation

Three options depending on your infrastructure: llama.cpp for a multi-machine swarm, vLLM for a homogeneous GPU cluster, or MLX-distributed for a Mac Studio swarm.

#Option A — distributed llama.cpp (recommended for multiple machines)

2× Mac Studio 512 GB swarm setup
# Sur chaque machine
git clone https://github.com/ggerganov/llama.cpp && cd llama.cpp
make GGML_METAL=1 GGML_RPC=1 GGML_CUDA=0

# Téléchargement modèle (split GGUF, ~960 Go)
huggingface-cli download deepseek-ai/DeepSeek-V4-Pro-GGUF \
  --include "*Q4_K_M*.gguf" --local-dir ./models

# Machine 1 (rpc-server, écoute sur :50052)
./build/bin/rpc-server -H 0.0.0.0 -p 50052

# Machine 2 (rpc-server également)
./build/bin/rpc-server -H 0.0.0.0 -p 50052

# Coordinator (peut être l'une des deux)
./build/bin/llama-cli \
  -m ./models/DeepSeek-V4-Pro-Q4_K_M-00001-of-00020.gguf \
  --rpc 192.168.1.10:50052,192.168.1.11:50052 \
  -ngl 99 -fa -c 32768 \
  -ctk q8_0 -ctv q8_0

#Option B — vLLM on an H100/H200 cluster

vLLM 8× H200 141GB
pip install vllm>=0.7

vllm serve deepseek-ai/DeepSeek-V4-Pro \
  --tensor-parallel-size 8 \
  --pipeline-parallel-size 1 \
  --max-model-len 1000000 \
  --quantization fp8 \
  --enable-prefix-caching \
  --port 8000

# Endpoint OpenAI-compatible sur :8000
curl http://localhost:8000/v1/chat/completions -d '{...}'

#Option C — Ollama (when the port becomes available)

As of April 25, 2026, Ollama has not yet integrated DeepSeek V4 Pro (community Q4_K_M quants are being released this weekend). Track progress at ollama.com/library.

→
Real electricity cost
8× H200 inference cluster DeepSeek V4 Pro Q8: ~5.2 kW. An 8-hour session costs ~€6.30 in electricity (France 2026). Compared with the DeepSeek API hosted at ~ $0.28/M output tokens, local becomes cost-effective beyond ~50M tokens/month.

#5. Benchmarks vs GPT-5, Claude 4.7, Gemini 3

Scores published by DeepSeek (paper dated April 24, 2026), standard public benchmarks
BenchmarkDeepSeek V4 Pro MaxGPT-5Claude 4.7 OpusGemini 3 Ultra
MMLU-Pro84.285.184.783.9
GPQA Diamond78.576.879.277.1
AIME 202587.082.485.680.2
LiveCodeBench v579.878.275.476.1
SWE-bench Verified59.462.163.855.7
LongBench v2 (1M)72.6—68.470.2
Humanity's Last Exam26.823.425.122.9
i
Independent benchmark coming soon
These figures are announced by DeepSeek. Independent benchmarks (LMSYS Arena, ARC-AGI, HLE) will be published within 2-4 weeks. Watch for confirmation: the model should appear in the Chatbot Arena top 5 within 10 days.

#6. Use the 1M-token context

The 1M context is native (not extended via YaRN), thanks to HCA. It's one of the three or four open-weight models in the world to offer 1M, and the only one to do so with this level of recall quality.

Needle-in-Haystack 1M
98% recovery across the entire context (vs. 76% for Gemini 1.5 Pro at equivalent settings).
Use case 1 — Entire codebases
Linux kernel net/* (~800k tokens) ingested in one go, with cross-file refactoring in a single request.
Use case 2 — Legal research
Civil Code + recent decisions (~600k tokens) in context, with case-law extraction in one pass.
Use case 3 — Accounting analysis
10 years of financial statements + appendices (~400k tokens), with multi-year anomaly detection.
1M KV cache VRAM
With KV Q4 and CSA enabled: ~280 GB. At 1M without compression: ~2 TB (impractical).

#7. Ideal use cases

Scientific research
Math, theoretical physics, computational chemistry. Max mode beats o4 on certain AIME/Putnam benchmarks.
Security / sovereignty
Banks, defense, healthcare, government: 100% self-hosted frontier model, MIT-licensed and auditable. No equivalent other than LLaMA 4 (more restrictive license).
Complex long-horizon agents
The 1M context, Max reasoning, and 49B active parameters make it a better agent than smaller dense models, without the inference cost of GPT-5.
Enterprise-scale RAG
Knowledge base with several million tokens (codebases, internal docs, ISO standards). Direct search without retrieval thanks to the context.

#Limitations and alternatives

Prohibitive hardware for individuals
Minimum 960 GB VRAM for Q4 = an approximately €80k used cluster. Out of reach for an individual.
Tooling immature at J+1
Community GGUF quants are in progress, vLLM requires the dev branch, and MLX-distributed is still experimental. Wait ~2 weeks for stability.
Carbon footprint
8× H200 cluster 24/7 = ~45 MWh/year. Weigh that against API usage (which amortizes costs better).
More accessible alternatives
DeepSeek V4 Flash 284B (170 GB Q4, fits Mac Studio) · DeepSeek V3.2 671B (240 GB Q2 on Mac Studio 512) · Llama 4 Behemoth (close in quality).

#FAQ

Is DeepSeek V4 Pro really better than GPT-5?+
According to the benchmarks published by DeepSeek: broadly comparable, sometimes better (AIME, GPQA, code). Worse on SWE-bench (full coding agent). Independent Chatbot Arena benchmarks (within 10 days) will settle the question. In any case, it is the first open-weights model for which the question can be taken seriously.
What is the exact license for DeepSeek V4 Pro?+
Pure MIT license — the most permissive. Commercial use, redistribution, forking, modification, and integration into a closed product are all allowed without restrictions or a strong attribution clause. More permissive than Llama (which requires >700M MAU for free use) or Apache 2.0 (which requires attribution).
Why 1.6T parameters if only 49B are active?+
That's the MoE trick: store many specialized experts (256 per layer), but activate only 8 per token (including 1 shared expert). You gain quality (each expert specializes) without paying the inference cost of a dense 1.6T model. Memory cost remains that of 1.6T, but compute cost is that of a 49B model.
Can you run DeepSeek V4 Pro on a Mac Studio Ultra 512 GB alone?+
No, not in Q4 (960 GB required). Possible in IQ2_XS (~480 GB) with degraded quality and a speed of 1–2 tok/s. For a practical single-machine setup, you must either wait for 70B/100B distillations or choose the V4 Flash 284B variant (170 GB Q4).
What is the difference between the 3 thinking modes?+
No = no explicit reasoning, minimal latency, quality ~Mistral Large. High = 200-2000 tokens of chain-of-thought before the answer, +30-40% quality on math/code/analysis. Max = up to 16k reasoning tokens, quality ~o4, but slow (10-30 min/request on a local cluster).
How do you enable thinking mode in Ollama / vLLM / API?+
Ollama: /think or /think_max suffix in the prompt. vLLM: X-Thinking-Mode: high|max header. Official API DeepSeek: thinking_mode field in the JSON body. The <think>...</think> tag in the response delimits the chain of thought.
Can DeepSeek V4 Pro handle vision (multimodal input)?+
No, V4 Pro is text-only. A VL (Vision-Language) variant is announced for late May 2026, based on the same architecture but with an integrated vision encoder. For open-weights vision today: Qwen3.5-VL or Llama 4 Vision.
Are distilled models (70B, 32B, 14B) coming?+
DeepSeek confirmed in the paper that distillations to Llama-3.3-70B and Qwen-2.5-32B base would be released within 4-6 weeks. Previously: DeepSeek-R1 had been distilled into 6 versions (1.5B to 70B), some of which outperform the original models. Expect some real bangers.
Did this guide help you?

Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.