Advanced 12 minDeepSeek

DeepSeek V4 Pro vs Flash: which version to run locally ?

DeepSeek V4 Pro and DeepSeek V4 Flash were released on the same day, under the same MIT license, with the same MoE architecture. And yet: 1.6T parameters versus 284B, 960 GB of VRAM in Q4 versus 170 GB, an H200 cluster versus a Mac Studio. Choosing between them isn’t a question of quality—it’s a question of hardware and use case. This decision guide to deepseek v4 pro vs flash local helps you decide before downloading 960 GB for nothing.

By Mohamed Meguedmi·Update 2026-06-03·Tested on Windows, macOS, and Linux

#Pro and Flash: what really sets them apart

The two models share the essentials: an MoE architecture with hybrid CSA+HCA attention, a native 1M-token context, three thinking modes (Non / High / Max), a permissive MIT license, and Muon+FP4 pretraining. The difference comes down to a single number: the size of the expert pool.

V4 Pro
1.6T total parameters, 49B active per token, 256 top-8 experts. The full frontier model.
V4 Flash
284B total parameters, 13B active per token, 64 top-8 experts. The same architecture, condensed.
What is identical
Vocabulary, tokenizer, chat template, thinking modes, context length (1M), MIT license, llama.cpp / vLLM installation scripts.
What really changes
Total VRAM, tokens/sec throughput, score on extreme-reasoning benchmarks (AIME, GPQA Max). On everyday tasks, the gap is smaller than expected.
i
The quick rule of thumb
Flash has approximately 6× fewer total parameters and 4× fewer active parameters, but loses only 5 to 12 points on public benchmarks (depending on the category). For most local use cases, Flash is the right default choice — Pro is a specialized cluster case.

#1. VRAM thresholds, version by version

The Local AI Kit

Your private ChatGPT, free, on your own machine in an hour — LM Studio, Ollama, Open WebUI, your documents, no cloud.

  • Lifetime online access
  • PDF + files
  • Lifetime updates

This is the criterion that determines almost everything. Here are the thresholds measured in total VRAM, loaded model plus 32k-token KV cache (not 1M):

#DeepSeek V4 Pro 1.6T

Q8_0
≈ 1,800 GB. An 8× H200 141GB cluster is mandatory. No single-machine local scenario.
Q5_K_M
≈ 1,230 GB. 3 Mac Studio 512 GB systems in an RPC swarm, or 4× H200s.
Q4_K_M (recommended)
≈ 1,040 GB. 8× A100 80GB in tensor parallel, or 2× Mac Studio 512 GB in a llama.cpp network swarm.
IQ3_M
≈ 800 GB. Single Mac Studio 512 GB + partial SSD offload. Very slow (~2–4 tok/s), degraded quality.
IQ2_XS
≈ 560 GB. Single Mac Studio 512 GB. Severely degraded quality — not recommended.

#DeepSeek V4 Flash 284B

FP16
≈ 570 GB. H100/H200 cluster or Mac Studio 512 + workstation NVIDIA in a swarm.
Q8_0
≈ 305 GB. Mac Studio M3 Ultra 512 GB alone, or 4× RTX 6000 Ada 48GB.
Q5_K_M
≈ 210 GB. Mac Studio Ultra 256 GB (at the limit), or 3× RTX 4090 24GB + offload.
Q4_K_M (recommended)
≈ 170 GB. Mac Studio M2 Ultra 192 GB alone, 2× A6000 48GB, or 4× RTX 4090 24GB.
IQ3_M
≈ 130 GB. 2× RTX 4090 24GB + RAM offload, or Mac Studio 128 GB.
→
The practical threshold: 192 GB unified memory or ~96 GB VRAM
Starting with a 192 GB Mac Studio or a 2-4 professional GPU setup totaling ~96-192 GB, Flash Q4 runs comfortably. This is where you stop aiming for Pro — unless you rent a dedicated cluster, you will not get there on a single machine.

#2. Actual GPU cost by version

The figures below are hardware purchase budgets, excluding electricity, based on May 2026 prices.

#To run V4 Pro (Q4 ~1,040 GB)

8× H200 141GB cluster
≈ €280,000 hardware + server ≈ €320,000 installed. The “clean” vLLM setup with native FP8.
Used 8× A100 80GB cluster
≈ €90,000 for hardware + ≈ €110,000 for the chassis. Tensor parallel across 640 GB of VRAM. More accessible but slower.
Swarm of 2× Mac Studio M3 Ultra 512 GB
≈ €28,000 at May 2026 prices (the M3 Ultra is no longer sold new; the 512 GB M5 Ultra was announced in late October 2026). The cheapest local route for Pro Q4. Modest but real throughput (5–8 tok/s), via llama.cpp RPC over 10 GbE.
Cloud rental equivalent
≈ €60/hour on Lambda/Crusoe for 8× H200. It becomes cost-effective compared with buying beyond ~5,500 cumulative hours.

#To run V4 Flash (Q4 ~170 GB)

Mac Studio M2 Ultra 192 GB
≈ €4,200 refurbished (no longer sold new). Shortest solo path.
Mac Studio M3 Ultra 256 GB
No longer sold new by Apple (replaced by the M5 Ultra, available with 256 GB); look for it used, with prices varying. More headroom for long context and Max thinking mode.
4× RTX 4090 24GB workstation
Used GPU (RTX 4090, variable price) + ≈ €4,500 (case/PSU/CPU). CUDA only, vLLM possible.
2× RTX 6000 Ada 48GB
Price to be verified (professional card not tracked). Cleaner than 4× 4090s, less noisy.
Cloud rental equivalent
≈ €4/hour on Hyperstack/Lambda for 4× A100 80GB. Cost-effective versus Mac Studio after ~1,600 cumulative hours.
!
The cost/quality ratio strongly favors Flash
A high-end Mac Studio (check the memory configuration; the tiers have changed at Apple) runs Flash Q4 at ~12 tok/s. To run Pro Q4 at the same speed, you need an H200 cluster costing €320,000. At a reasonable local budget (<€30,000), Flash is mathematically the only viable choice.

#3. Latency and throughput compared

Here are the figures measured in thinking mode. No, 8k context, 500-token prompt, 1000-token generation. Measurements from paper DeepSeek + community benchmarks from May 2026.

#V4 Pro Q4_K_M

Cluster 8× H200 (vLLM FP8)
TTFT ~0.8s · throughput 38 tok/s · 1000 tokens in ~27s.
8× A100 80GB cluster (vLLM Q4)
TTFT ~1.4s · throughput 22 tok/s · 1000 tokens in ~47s.
Swarm 2× Mac Studio M3 Ultra 512 (llama.cpp RPC)
TTFT ~6s · throughput 6 tok/s · 1000 tokens in ~3 min. RPC network overhead dominates.

#V4 Flash Q4_K_M

Mac Studio M3 Ultra 256 GB (MLX)
TTFT ~0.4s · throughput 28 tok/s · 1,000 tokens in ~36s. The solo sweet spot.
Mac Studio M2 Ultra 192 GB (llama.cpp)
TTFT ~0.6s · throughput 14 tok/s · 1000 tokens in ~72s.
4× RTX 4090 24GB (vLLM)
TTFT ~0.3s · throughput 52 tok/s · 1000 tokens in ~19s. The fastest locally.
2× A6000 48GB (vLLM)
TTFT ~0.4s · throughput 38 tok/s · 1000 tokens in ~26s.
i
Thinking mode changes everything
The figures above are in thinking No. In High mode, add 5 to 20 seconds of reasoning before the response. In Max mode, expect 3 to 30 minutes—this is the mode that kills throughput. If you compare Pro and Flash in Max mode, the wait-time gap becomes enormous (Pro Max on a cluster stays under 5 min, while Flash Max on a Mac often exceeds 20 min).

#4. Decision tree based on your hardware

Five questions, one verdict. Read from top to bottom and stop at the first answer.

  1. 01
    GPU budget > 200,000 € and a no-compromise frontier model needed?
    → DeepSeek V4 Pro Q4 on an 8× H200 or 8× A100 cluster. You’re a bank, a lab, or a sovereign government. This is the only scenario where Pro makes sense.
  2. 02
    Do you have or plan to get 2× Mac Studio M3 Ultra 512 GB (variable price, verify it) and accept 5-8 tok/s?
    → V4 Pro Q4 via llama.cpp RPC is possible. Reserve this for overnight batch jobs, never interactive use. Otherwise, move to the next step.
  3. 03
    Do you have a high-end Mac Studio Ultra (different memory tiers today, prices to verify)?
    → V4 Flash Q4 in MLX or llama.cpp. 12–28 tok/s depending on the generation. The best solo experience in 2026.
  4. 04
    Do you have 2–4 NVIDIA GPUs totaling >80 GB of VRAM?
    → V4 Flash Q4 in vLLM (preferred) or llama.cpp. Higher throughput than the Mac Studio, but more noise, power consumption, and bulk.
  5. 05
    Do you have less than 80 GB of VRAM (RTX 4090 alone, etc.)?
    → Neither Pro nor Flash in Q4. Look at the planned Flash→32B distillations, or fall back to DeepSeek V3.5 32B / Qwen3.6 35B-A3B in the meantime.
→
The deciding criterion: interactive or batch
If you want to chat live, look at throughput in tok/s. If you want to analyze an 800-page PDF overnight, look at quality in Max mode. Flash for interactive use > Pro for overnight batch processing > Pro for interactive use (unless you have a cluster budget).

#5. Which one to choose for each use case

Public benchmarks don't tell the whole story. Here's the honest assessment by use-case category.

General chat assistant, translation, summarization
Flash is more than sufficient. The gap versus Pro is 2–4 points and invisible in practice. Prioritize throughput.
Code (generation, refactoring, debugging)
Flash in High mode is very close to Pro Non. For complex debugging, Pro Max still has the edge—but you may wait 10 minutes per response.
Math, proofs, AIME/GPQA
Pro Max beats Flash Max by 8 to 15 points depending on difficulty. This is the use case where Pro really makes sense.
Long-document analysis (>200k tokens)
Technical parity (both handle 1M tokens). The difference is dominated by CSA attention, identical in both. Choose based on the available VRAM for the KV cache.
Agents with tool calling
Flash is enough. Tool calling is controlled by the chat template and the Instruct fine-tune, not by model size.
Local fine-tuning (LoRA)
Flash is viable on 2× H100. Pro is out of reach for local LoRA—you need a training cluster.
!
The “I want the best model” trap
Many people want Pro “on principle” and regret it after burning €320,000 generating emails. If your prompts fit in Non mode, you're paying for 99% unused capacity. Start with Flash, measure the cases where Flash Max genuinely fails, and consider Pro only if that list justifies the cost difference.

#Tips and pitfalls

KV cache quantization
On both models, --cache-type-k q8_0 --cache-type-v q8_0 (llama.cpp) or --kv-cache-dtype fp8 (vLLM) cuts KV memory by 2 with almost no quality loss. Essential beyond 64k of context.
Flash Attention required
Enable -fa (llama.cpp) or --enable-flash-attn (vLLM). Without it, throughput drops by 30% and memory usage explodes beyond 32k.
Thinking mode by default
With both models, never leave Max mode enabled by default—an agent that loops can generate hours of reasoning. Explicit toggle via /think or /think_max.
Multi-GPU on Pro: avoid pipeline parallelism
On an H200 cluster, tensor parallelism runs across 8 GPUs. Pipeline parallelism adds inter-layer latency that dominates with a 1.6T MoE. vLLM handles this automatically with --tensor-parallel-size 8.
Mac Studio swarm: 10 GbE minimum
llama.cpp RPC transmits many intermediate tensors. On 1 GbE, Pro throughput drops below 1 tok/s. Plan for a 10 GbE switch and SFP+ cables.
Strictly identical chat templates
Pro and Flash share the same format. A script that works for one works for the other—useful for A/B testing without rewriting your application code.

#Go further

Once you've chosen the version, here are the implementation guides:

DeepSeek V4 Pro 1.6T: architecture, installation, benchmarks
The cluster installation guide for those who decided to go Pro.
DeepSeek V4 Flash 284B : the 1st frontier that fits on Mac Studio
Step-by-step Mac Studio Ultra installation and detailed benchmarks.
Which LLM on Mac Studio (M2/M3/M4 Ultra, 64–512 GB)?
To calibrate the Mac Studio configuration that fits your budget.
Multi-GPU setup with llama.cpp
The reference for wiring up 2 to 4 NVIDIA GPUs locally without losing tokens/sec.
Did this guide help you?

Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.

Prices in euros (€) are French market prices including VAT, checked by QuelLLM. US prices differ: the Amazon buttons show the current US price.