DeepSeek V3.2 locally: 671B MoE accessible to mortels
Doing a deepseek v3.2 local installation means facing reality: 671 billion total parameters, 37 billion active per token, and a new sparse attention mechanism that changes the game for long contexts. This guide covers the real hardware requirements (not the marketing claims), explains ik_llama's Q2_K_XL quantization, demonstrates the selective offload trick using --override-tensor, and ends with tokens/sec measured on a home workstation in 2026.
#Why install DeepSeek V3.2 locally
DeepSeek V3.2 is the intermediate update released by DeepSeek AI in late 2025 under the MIT license. It introduces two significant changes compared with V3: DeepSeek Sparse Attention (DSA), a sparse-attention mechanism that makes the memory cost of long context subquadratic, and a refinement of multi-token prediction training. On paper, it preserves V3 quality while improving efficiency at 128k+ tokens.
Installing it locally addresses three concrete needs: total privacy (no document is sent to DeepSeek), reproducibility (the weights do not disappear overnight), and unrestricted experimentation (no rate limits and no mandatory moderation filter).
#MoE 671B / 37B active + DeepSeek Sparse Attention
Your private ChatGPT, free, on your own machine in an hour — LM Studio, Ollama, Open WebUI, your documents, no cloud.
- Lifetime online access
- PDF + files
- Lifetime updates
DeepSeek V3.2 retains V3’s MoE architecture: a router selects 8 experts out of 256 per layer, so only ~37 billion parameters are involved for each generated token. The remaining 634 billion are idle, but must remain addressable; otherwise, the router loses access to experts it has not loaded.
- Total parameters
- ≈ 671 billion (the bit rate does not matter; this is the amount that has to fit somewhere: RAM, VRAM, or an SSD via mmap).
- Active parameters per token
- ≈ 37 billion. This figure dictates theoretical speed: a 671B MoE incurs the compute cost of a dense 37B model.
- Experts per layer
- 256 MoE experts + 1 shared expert, top-8 routing. The more RAM you have, the more expert layers remain resident and the more stable generation becomes.
- Warning
- Multi-Head Latent Attention (MLA) inherited from V3, now combined with DeepSeek Sparse Attention (DSA) for contexts > 32k tokens.
- Context
- 128k tokens advertised, usable in practice thanks to DSA—the KV cache no longer grows linearly as it does on a conventional dense model such as Qwen 3.5 or Gemma 4.
#Honest hardware requirements
No consumer-grade setup can handle DeepSeek V3.2 without contortions. Here are the three realistic profiles for 2026, ranked by the speed/budget tradeoff.
- Profile A — high-end DDR5 (192–384 GB)
- Workstation with Threadripper 7960X / Xeon W or EPYC 9354P, 256 GB DDR5 ECC (8 channels), no GPU required. Q2_K_XL fits entirely in RAM. Speed: 5–9 tok/s on CPU alone.
- Profile B — RTX 5090 + 192 GB DDR5 (recommended)
- The 2026 sweet spot: RTX 5090 (32 GB GDDR7, 1792 GB/s bandwidth) + 192 GB of DDR5 on a recent consumer platform. Keep the attention layers and shared expert on the GPU, with the rest in RAM. Speed: 10–15 tok/s during generation.
- Profile C — modest RAM + Gen 4/5 NVMe SSD
- 96–128 GB DDR5 + NVMe SSD at 7 GB/s minimum, ≥ 1 TB free. Experts are memory-mapped from the SSD. Speed: 1.5–3 tok/s — usable for overnight batch jobs, not interactively.
#1. Choose the quantization for V3.2
The GGUF ecosystem for DeepSeek V3.2 is driven by two main players: Unsloth (which publishes UD variants = “Unsloth Dynamic,” including the well-known Q2_K_XL) and bartowski. UDs use an importance matrix to preserve sensitive layers (attention, shared expert) at higher precision than the rest—a real practical difference in final quality.
- IQ1_S (~150 GB)
- Dynamic 1-bit quantization. Fits in 192 GB of RAM with room to spare. Degraded but usable quality for general-purpose chat. Avoid for code.
- Q2_K_XL (~220 GB)
- The ik_llama / Unsloth Dynamic reference. Keeps critical layers in Q4–Q5 and compresses the experts to Q2. ~95% of Q4 quality on reasoning benchmarks. Fits on 256 GB of RAM or 192 GB plus a little SSD offload.
- Q3_K_S (~290 GB)
- For 384 GB workstations. Near-Q4 quality on long contexts. Recommended if you plan to seriously use 128k tokens.
- Q4_K_M (~400 GB)
- Full throttle. Reserved for dual-socket EPYC servers with 512 GB+ of DDR5. Beyond that, the quality gains become marginal for local use.
#2. Compile ik_llama.cpp (the fork that supports V3.2)
At the time of writing, the main ggerganov/llama.cpp branch supports DeepSeek V3.2 but has no optimizations specific to DeepSeek's MoE routing. The ik_llama.cpp fork (Iwan Kawrakow) includes accelerated CUDA kernels for expert tensors and dynamic DSA support—a 30 to 50% throughput gain on V3.2, depending on the configuration.
On Mac Apple Silicon, replace GGML_CUDA with GGML_METAL. On AMD, use GGML_HIP with ROCm 6.3+. Allow 10 to 15 minutes for the build on a recent machine—it’s C++ with CUDA generation, so don’t run it on a laptop without AC power.
#3. Download the GGUF V3.2 files
DeepSeek V3.2’s GGUF weights are published on Hugging Face. For Q2_K_XL Unsloth (the recommended option for most configurations), a single download of ~220 GB is split across several files.
The download takes several hours on a typical consumer connection—plan for a stable fiber connection and a target drive with ≥ 250 GB free. GGUFs of this size are split into 5 to 7 files (split-00001-of-N.gguf), and ik_llama.cpp automatically detects the splits when you point it to the first one.
#4. Launching with --override-tensor (the key to hybrid performance)
The secret to a good local deepseek v3.2 installation comes down to one option: --override-tensor (alias -ot). It lets you use regex to target which layers stay on CPU/RAM and which go to the GPU. With an MoE, you absolutely do not want to load everything onto the GPU (VRAM will never be enough), but you absolutely want attention and shared layers to be accelerated.
The regex \.ffn_(up|down|gate)_exps\. targets the three FFN tensors per expert—that is, the overwhelming majority of the 671B parameters—and forces them to stay on the CPU. What runs on the GPU: attention (MLA), the shared expert, embeddings, and the head. On RTX 5090 32 GB, it uses ~22 GB of VRAM; the rest is used for the KV cache.
#5. Tokens/sec measured on a home workstation
Benchmarks measured with the May 2026 ik_llama.cpp build, Q2_K_XL, 32k context, mixed French/code prompt. The figures are stable within ±15% depending on the prompt and memory-page warm-up.
- RTX 5090 + Ryzen 9 7950X3D + 192 GB DDR5-6000
- Prefill ~110 tok/s, generation 12–15 tok/s. The recommended home setup for 2026.
- RTX 4090 + Threadripper 7960X + 256 GB DDR5 ECC
- Prefill ~90 tok/s, generation 9–12 tok/s. The 4090 matches the 5090 on this hybrid profile because the bottleneck is RAM, not the GPU.
- EPYC 9354P + 384 GB DDR5 ECC (12 channels), no GPU
- Prefill ~70 tok/s, generation 7–9 tok/s. Total memory bandwidth (~460 GB/s) makes up for the lack of a GPU.
- Mac Studio M3 Ultra 192 GB
- Prefill ~55 tok/s, generation 6–8 tok/s in Q2_K_XL. Unified memory (~820 GB/s) saves the day.
- Ryzen 9 7950X + 128 GB DDR5 + Samsung 990 Pro SSD
- Prefill ~25 tok/s, generation 1.5–2.5 tok/s. Honestly usable only in batch.
#DeepSeek V3.2 vs. DeepSeek R1: which one should you install?
A recurring question, because the two models share the same 671B / 37B-active MoE architecture and the same MIT license. The difference is in post-training, not in the underlying skeleton.
- DeepSeek V3.2
- General-purpose model (chat), direct answers. Excellent in French, very strong at coding. Adds DSA for long context. Preferred for everyday assistance, summarization, and writing.
- DeepSeek R1
- Reasoning model (explicit chain of thought à la o1). Emits many internal tokens (<think>...</think>) before the final response. Best for math, logic, and complex algorithmic debugging.
- Inference cost
- R1 generates 3 to 10 times more tokens (because of thinking) to produce the same final answer. At the same throughput, R1 takes 5 minutes where V3.2 takes 30 seconds. Critical locally, where every tok/s counts.
- Long context
- V3.2, thanks to DSA, handles 128k tokens with a reasonable KV cache. R1 lacks DSA and runs out of memory beyond 64k.
- VRAM/RAM required
- Identical. Both models share the 671B/37B architecture and fit in the same ~220 GB Q2_K_XL quantization.
#Troubleshooting
- “unknown model architecture: deepseek2”
- Your llama.cpp is too old or was compiled without deepseek2 support. Update ik_llama.cpp (main branch after April 2026) and recompile. The standard pre-built binary is not always sufficient.
- CUDA OOM during loading
- Your -ot regex does not offload enough experts to the CPU. Check the actual tensor distribution with ollama-style and ik_llama-bench. The correct regex for V3.2 properly targets ffn_(up|down|gate)_exps, not just ffn_exps.
- Generation at 1 tok/s on 192 GB RAM
- The kernel will page in from the GGUF until it has loaded all hot pages. An initial warm-up prompt of 200 tokens preloads the experts. Otherwise, increase --threads up to the number of physical cores (not logical cores).
- Responses that switch to Chinese for no reason
- Incorrect chat template. Make sure --chat-template is set to auto and that the GGUF actually includes the DeepSeek Jinja template. Otherwise, explicitly pass --chat-template deepseek3.
- DSA appears inactive (linear KV cache)
- DSA is enabled only from a configurable minimum context length. For contexts < 16k, the model uses standard dense attention—as expected, and this does not affect quality.
- Crash on long prompt after warmup
- You probably exhausted the swap. DeepSeek V3.2 does not like swap: if physical RAM is insufficient, disable swap (sudo swapoff -a) and let llama.cpp handle paging through mmap; it is faster and more stable.
#Go further
DeepSeek V3.2 locally is a starting point for exploring the class of “self-hostable frontier models.” Three complementary avenues:
- Push the llama.cpp build
- The guide “Compile llama.cpp with CUDA” details the advanced flags (Flash Attention, MMQ, offload tensors) that genuinely change MoE performance.
- Understanding modern quantization
- The “Choosing your quantization (Q4, Q5, Q8, FP16)” guide explains why Q2_K_XL works where standard Q2 fails, and when to move up to Q3/Q4.
- Go one step further in size
- The "Kimi K2 locally" guide applies the same method to a 1T-parameter MoE—useful for anticipating where the ecosystem is headed in 2026-2027.
Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.