DeepSeek V4 Pro vs Flash: which version to run locally ?
DeepSeek V4 Pro and DeepSeek V4 Flash were released on the same day, under the same MIT license, with the same MoE architecture. And yet: 1.6T parameters versus 284B, 960 GB of VRAM in Q4 versus 170 GB, an H200 cluster versus a Mac Studio. Choosing between them isn’t a question of quality—it’s a question of hardware and use case. This decision guide to deepseek v4 pro vs flash local helps you decide before downloading 960 GB for nothing.
#Pro and Flash: what really sets them apart
The two models share the essentials: an MoE architecture with hybrid CSA+HCA attention, a native 1M-token context, three thinking modes (Non / High / Max), a permissive MIT license, and Muon+FP4 pretraining. The difference comes down to a single number: the size of the expert pool.
- V4 Pro
- 1.6T total parameters, 49B active per token, 256 top-8 experts. The full frontier model.
- V4 Flash
- 284B total parameters, 13B active per token, 64 top-8 experts. The same architecture, condensed.
- What is identical
- Vocabulary, tokenizer, chat template, thinking modes, context length (1M), MIT license, llama.cpp / vLLM installation scripts.
- What really changes
- Total VRAM, tokens/sec throughput, score on extreme-reasoning benchmarks (AIME, GPQA Max). On everyday tasks, the gap is smaller than expected.
#1. VRAM thresholds, version by version
Your private ChatGPT, free, on your own machine in an hour — LM Studio, Ollama, Open WebUI, your documents, no cloud.
- Lifetime online access
- PDF + files
- Lifetime updates
This is the criterion that determines almost everything. Here are the thresholds measured in total VRAM, loaded model plus 32k-token KV cache (not 1M):
#DeepSeek V4 Pro 1.6T
- Q8_0
- ≈ 1,800 GB. An 8× H200 141GB cluster is mandatory. No single-machine local scenario.
- Q5_K_M
- ≈ 1,230 GB. 3 Mac Studio 512 GB systems in an RPC swarm, or 4× H200s.
- Q4_K_M (recommended)
- ≈ 1,040 GB. 8× A100 80GB in tensor parallel, or 2× Mac Studio 512 GB in a llama.cpp network swarm.
- IQ3_M
- ≈ 800 GB. Single Mac Studio 512 GB + partial SSD offload. Very slow (~2–4 tok/s), degraded quality.
- IQ2_XS
- ≈ 560 GB. Single Mac Studio 512 GB. Severely degraded quality — not recommended.
#DeepSeek V4 Flash 284B
- FP16
- ≈ 570 GB. H100/H200 cluster or Mac Studio 512 + workstation NVIDIA in a swarm.
- Q8_0
- ≈ 305 GB. Mac Studio M3 Ultra 512 GB alone, or 4× RTX 6000 Ada 48GB.
- Q5_K_M
- ≈ 210 GB. Mac Studio Ultra 256 GB (at the limit), or 3× RTX 4090 24GB + offload.
- Q4_K_M (recommended)
- ≈ 170 GB. Mac Studio M2 Ultra 192 GB alone, 2× A6000 48GB, or 4× RTX 4090 24GB.
- IQ3_M
- ≈ 130 GB. 2× RTX 4090 24GB + RAM offload, or Mac Studio 128 GB.
#2. Actual GPU cost by version
The figures below are hardware purchase budgets, excluding electricity, based on May 2026 prices.
#To run V4 Pro (Q4 ~1,040 GB)
- 8× H200 141GB cluster
- ≈ €280,000 hardware + server ≈ €320,000 installed. The “clean” vLLM setup with native FP8.
- Used 8× A100 80GB cluster
- ≈ €90,000 for hardware + ≈ €110,000 for the chassis. Tensor parallel across 640 GB of VRAM. More accessible but slower.
- Swarm of 2× Mac Studio M3 Ultra 512 GB
- ≈ €28,000 at May 2026 prices (the M3 Ultra is no longer sold new; the 512 GB M5 Ultra was announced in late October 2026). The cheapest local route for Pro Q4. Modest but real throughput (5–8 tok/s), via llama.cpp RPC over 10 GbE.
- Cloud rental equivalent
- ≈ €60/hour on Lambda/Crusoe for 8× H200. It becomes cost-effective compared with buying beyond ~5,500 cumulative hours.
#To run V4 Flash (Q4 ~170 GB)
- Mac Studio M2 Ultra 192 GB
- ≈ €4,200 refurbished (no longer sold new). Shortest solo path.
- Mac Studio M3 Ultra 256 GB
- No longer sold new by Apple (replaced by the M5 Ultra, available with 256 GB); look for it used, with prices varying. More headroom for long context and Max thinking mode.
- 4× RTX 4090 24GB workstation
- Used GPU (RTX 4090, variable price) + ≈ €4,500 (case/PSU/CPU). CUDA only, vLLM possible.
- 2× RTX 6000 Ada 48GB
- Price to be verified (professional card not tracked). Cleaner than 4× 4090s, less noisy.
- Cloud rental equivalent
- ≈ €4/hour on Hyperstack/Lambda for 4× A100 80GB. Cost-effective versus Mac Studio after ~1,600 cumulative hours.
#3. Latency and throughput compared
Here are the figures measured in thinking mode. No, 8k context, 500-token prompt, 1000-token generation. Measurements from paper DeepSeek + community benchmarks from May 2026.
#V4 Pro Q4_K_M
- Cluster 8× H200 (vLLM FP8)
- TTFT ~0.8s · throughput 38 tok/s · 1000 tokens in ~27s.
- 8× A100 80GB cluster (vLLM Q4)
- TTFT ~1.4s · throughput 22 tok/s · 1000 tokens in ~47s.
- Swarm 2× Mac Studio M3 Ultra 512 (llama.cpp RPC)
- TTFT ~6s · throughput 6 tok/s · 1000 tokens in ~3 min. RPC network overhead dominates.
#V4 Flash Q4_K_M
- Mac Studio M3 Ultra 256 GB (MLX)
- TTFT ~0.4s · throughput 28 tok/s · 1,000 tokens in ~36s. The solo sweet spot.
- Mac Studio M2 Ultra 192 GB (llama.cpp)
- TTFT ~0.6s · throughput 14 tok/s · 1000 tokens in ~72s.
- 4× RTX 4090 24GB (vLLM)
- TTFT ~0.3s · throughput 52 tok/s · 1000 tokens in ~19s. The fastest locally.
- 2× A6000 48GB (vLLM)
- TTFT ~0.4s · throughput 38 tok/s · 1000 tokens in ~26s.
#4. Decision tree based on your hardware
Five questions, one verdict. Read from top to bottom and stop at the first answer.
- 01GPU budget > 200,000 € and a no-compromise frontier model needed?→ DeepSeek V4 Pro Q4 on an 8× H200 or 8× A100 cluster. You’re a bank, a lab, or a sovereign government. This is the only scenario where Pro makes sense.
- 02Do you have or plan to get 2× Mac Studio M3 Ultra 512 GB (variable price, verify it) and accept 5-8 tok/s?→ V4 Pro Q4 via llama.cpp RPC is possible. Reserve this for overnight batch jobs, never interactive use. Otherwise, move to the next step.
- 03Do you have a high-end Mac Studio Ultra (different memory tiers today, prices to verify)?→ V4 Flash Q4 in MLX or llama.cpp. 12–28 tok/s depending on the generation. The best solo experience in 2026.
- 04Do you have 2–4 NVIDIA GPUs totaling >80 GB of VRAM?→ V4 Flash Q4 in vLLM (preferred) or llama.cpp. Higher throughput than the Mac Studio, but more noise, power consumption, and bulk.
- 05Do you have less than 80 GB of VRAM (RTX 4090 alone, etc.)?→ Neither Pro nor Flash in Q4. Look at the planned Flash→32B distillations, or fall back to DeepSeek V3.5 32B / Qwen3.6 35B-A3B in the meantime.
#5. Which one to choose for each use case
Public benchmarks don't tell the whole story. Here's the honest assessment by use-case category.
- General chat assistant, translation, summarization
- Flash is more than sufficient. The gap versus Pro is 2–4 points and invisible in practice. Prioritize throughput.
- Code (generation, refactoring, debugging)
- Flash in High mode is very close to Pro Non. For complex debugging, Pro Max still has the edge—but you may wait 10 minutes per response.
- Math, proofs, AIME/GPQA
- Pro Max beats Flash Max by 8 to 15 points depending on difficulty. This is the use case where Pro really makes sense.
- Long-document analysis (>200k tokens)
- Technical parity (both handle 1M tokens). The difference is dominated by CSA attention, identical in both. Choose based on the available VRAM for the KV cache.
- Agents with tool calling
- Flash is enough. Tool calling is controlled by the chat template and the Instruct fine-tune, not by model size.
- Local fine-tuning (LoRA)
- Flash is viable on 2× H100. Pro is out of reach for local LoRA—you need a training cluster.
#Tips and pitfalls
- KV cache quantization
- On both models, --cache-type-k q8_0 --cache-type-v q8_0 (llama.cpp) or --kv-cache-dtype fp8 (vLLM) cuts KV memory by 2 with almost no quality loss. Essential beyond 64k of context.
- Flash Attention required
- Enable -fa (llama.cpp) or --enable-flash-attn (vLLM). Without it, throughput drops by 30% and memory usage explodes beyond 32k.
- Thinking mode by default
- With both models, never leave Max mode enabled by default—an agent that loops can generate hours of reasoning. Explicit toggle via /think or /think_max.
- Multi-GPU on Pro: avoid pipeline parallelism
- On an H200 cluster, tensor parallelism runs across 8 GPUs. Pipeline parallelism adds inter-layer latency that dominates with a 1.6T MoE. vLLM handles this automatically with --tensor-parallel-size 8.
- Mac Studio swarm: 10 GbE minimum
- llama.cpp RPC transmits many intermediate tensors. On 1 GbE, Pro throughput drops below 1 tok/s. Plan for a 10 GbE switch and SFP+ cables.
- Strictly identical chat templates
- Pro and Flash share the same format. A script that works for one works for the other—useful for A/B testing without rewriting your application code.
#Go further
Once you've chosen the version, here are the implementation guides:
- DeepSeek V4 Pro 1.6T: architecture, installation, benchmarks
- The cluster installation guide for those who decided to go Pro.
- DeepSeek V4 Flash 284B : the 1st frontier that fits on Mac Studio
- Step-by-step Mac Studio Ultra installation and detailed benchmarks.
- Which LLM on Mac Studio (M2/M3/M4 Ultra, 64–512 GB)?
- To calibrate the Mac Studio configuration that fits your budget.
- Multi-GPU setup with llama.cpp
- The reference for wiring up 2 to 4 NVIDIA GPUs locally without losing tokens/sec.
Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.