DeepSeek V4 Pro locally: VRAM, hardware, installation
DeepSeek V4 Pro is the world's most powerful open-weight model at its release on April 24, 2026. It has 1.6 trillion parameters in a MoE architecture (49 billion active per token), an MIT license, a 1M-token context, hybrid CSA+HCA attention, and three reasoning modes. On published benchmarks, it rivals GPT-5 and Claude 4.7 — with no cloud server, no API, and no token limit. This guide explains how to run it locally and what hardware you need.
#DeepSeek V4 Pro at a glance
- Output
- April 24, 2026 — DeepSeek-AI on HuggingFace, MIT-licensed weights, paper published simultaneously.
- Architecture
- Dense Mixture-of-Experts: 1.6T total parameters, 49B active per token, 256 experts with top-8 routing.
- Warning
- Hybrid CSA (Compressed Sparse Attention) + HCA (Hybrid Constrained Attention)—the first implementation at this scale.
- Context
- 1,000,000 native tokens (with extended RoPE and flash attention v3).
- Pretraining
- 32T+ tokens, Muon optimizer (replaces AdamW), mixed FP4+FP8 precision (the most efficient training known).
- Thinking modes
- 3 levels: Non / High / Max. Max mode approaches o4 performance on AIME and GPQA.
- License
- MIT — commercial use, redistribution, and modification allowed without restriction. More permissive than the Llama license.
#1. CSA+HCA and 1.6T MoE architecture
DeepSeek V4 Pro combines three major innovations, each published separately over the last 12 months and combined here for the first time at trillion scale.
Your private ChatGPT, free, on your own machine in an hour — LM Studio, Ollama, Open WebUI, your documents, no cloud.
- Lifetime online access
- PDF + files
- Lifetime updates
- 1.6T MoE / 49B active
- 256 experts per layer, top-8 routing. For each token, only 49 billion parameters compute—hence speed comparable to a dense 50B model despite the total size.
- CSA — Compressed Sparse Attention
- Compresses the old keys/values in the KV cache using low-rank decomposition. Cuts attention memory by a factor of 4 for contexts >128k.
- HCA — Hybrid Constrained Attention
- Combines local attention (4k sliding window) and global attention (sparse), depending on the layer. 2× throughput on long contexts vs. dense attention.
- mHC — multi-Head Compressor
- Cross-head compression mechanism — reduces activation VRAM with no measurable loss in quality.
- Muon optimizer
- AdamW's successor, converging 1.8× faster during the pretraining of very large models. Adopted by DeepSeek since V3.5.
- FP4+FP8 training
- FP4 mix on attention layers, FP8 on MLPs. -45% training memory vs. pure FP8, with quality preserved through a proprietary scaling scheme.
#2. The 3 thinking modes (Non / High / Max)
DeepSeek V4 Pro's thinking system is explicit and toggleable, unlike the opaque modes of competing proprietary models.
- No mode (default)
- Direct responses with no visible chain-of-thought. Minimal latency (~2–5s at 8 tok/s). For chat, translation, and simple text generation.
- High mode
- Inserts 200-2000 reasoning tokens before the response. +30-40% quality on math, code, and analysis. Triggered via the /think flag in Ollama, or the thinking_mode header in the API.
- Max mode
- Exhaustive reasoning (up to 16k chain-of-thought tokens). Approaches o4 performance on AIME 2025 (87% vs. 91% for o4). Expensive: 10–30 minutes per request locally.
#3. Hardware required by quantization
| Quantization | Model VRAM | + KV 32k | Possible configurations |
|---|---|---|---|
| FP16 (reference) | 3,200 GB | ~3,350 GB | Cloud: 16× H100 80GB or 8× H200 141GB. Not accessible locally. |
| Q8_0 | 1,700 GB | ~1,800 GB | 8× H200 141GB cluster. ~250 k$ in hardware. |
| Q5_K_M | 1 150 GB | ~1,230 GB | Mac Studio 512 GB cluster ×3, or 4× H200 141GB. |
| Q4_K_M | 960 GB | ~1 040 GB | 8× A100 80GB cluster, or 2× Mac Studio 512 GB in a network swarm. |
| IQ3_M (compressed) | 720 GB | ~800 GB | Single Mac Studio 512 + partial SSD offload. Slow but possible. |
| IQ2_XS (extreme) | 480 GB | ~560 GB | Mac Studio 512 GB only; quality is severely degraded. |
#4. Local installation
Three options depending on your infrastructure: llama.cpp for a multi-machine swarm, vLLM for a homogeneous GPU cluster, or MLX-distributed for a Mac Studio swarm.
#Option A — distributed llama.cpp (recommended for multiple machines)
#Option B — vLLM on an H100/H200 cluster
#Option C — Ollama (when the port becomes available)
As of April 25, 2026, Ollama has not yet integrated DeepSeek V4 Pro (community Q4_K_M quants are being released this weekend). Track progress at ollama.com/library.
#5. Benchmarks vs GPT-5, Claude 4.7, Gemini 3
| Benchmark | DeepSeek V4 Pro Max | GPT-5 | Claude 4.7 Opus | Gemini 3 Ultra |
|---|---|---|---|---|
| MMLU-Pro | 84.2 | 85.1 | 84.7 | 83.9 |
| GPQA Diamond | 78.5 | 76.8 | 79.2 | 77.1 |
| AIME 2025 | 87.0 | 82.4 | 85.6 | 80.2 |
| LiveCodeBench v5 | 79.8 | 78.2 | 75.4 | 76.1 |
| SWE-bench Verified | 59.4 | 62.1 | 63.8 | 55.7 |
| LongBench v2 (1M) | 72.6 | — | 68.4 | 70.2 |
| Humanity's Last Exam | 26.8 | 23.4 | 25.1 | 22.9 |
#6. Use the 1M-token context
The 1M context is native (not extended via YaRN), thanks to HCA. It's one of the three or four open-weight models in the world to offer 1M, and the only one to do so with this level of recall quality.
- Needle-in-Haystack 1M
- 98% recovery across the entire context (vs. 76% for Gemini 1.5 Pro at equivalent settings).
- Use case 1 — Entire codebases
- Linux kernel net/* (~800k tokens) ingested in one go, with cross-file refactoring in a single request.
- Use case 2 — Legal research
- Civil Code + recent decisions (~600k tokens) in context, with case-law extraction in one pass.
- Use case 3 — Accounting analysis
- 10 years of financial statements + appendices (~400k tokens), with multi-year anomaly detection.
- 1M KV cache VRAM
- With KV Q4 and CSA enabled: ~280 GB. At 1M without compression: ~2 TB (impractical).
#7. Ideal use cases
- Scientific research
- Math, theoretical physics, computational chemistry. Max mode beats o4 on certain AIME/Putnam benchmarks.
- Security / sovereignty
- Banks, defense, healthcare, government: 100% self-hosted frontier model, MIT-licensed and auditable. No equivalent other than LLaMA 4 (more restrictive license).
- Complex long-horizon agents
- The 1M context, Max reasoning, and 49B active parameters make it a better agent than smaller dense models, without the inference cost of GPT-5.
- Enterprise-scale RAG
- Knowledge base with several million tokens (codebases, internal docs, ISO standards). Direct search without retrieval thanks to the context.
#Limitations and alternatives
- Prohibitive hardware for individuals
- Minimum 960 GB VRAM for Q4 = an approximately €80k used cluster. Out of reach for an individual.
- Tooling immature at J+1
- Community GGUF quants are in progress, vLLM requires the dev branch, and MLX-distributed is still experimental. Wait ~2 weeks for stability.
- Carbon footprint
- 8× H200 cluster 24/7 = ~45 MWh/year. Weigh that against API usage (which amortizes costs better).
- More accessible alternatives
- DeepSeek V4 Flash 284B (170 GB Q4, fits Mac Studio) · DeepSeek V3.2 671B (240 GB Q2 on Mac Studio 512) · Llama 4 Behemoth (close in quality).
#FAQ
Is DeepSeek V4 Pro really better than GPT-5?+
What is the exact license for DeepSeek V4 Pro?+
Why 1.6T parameters if only 49B are active?+
Can you run DeepSeek V4 Pro on a Mac Studio Ultra 512 GB alone?+
What is the difference between the 3 thinking modes?+
How do you enable thinking mode in Ollama / vLLM / API?+
Can DeepSeek V4 Pro handle vision (multimodal input)?+
Are distilled models (70B, 32B, 14B) coming?+
Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.