Advanced 20 minllama.cpp

DeepSeek V3.2 locally: 671B MoE accessible to mortels

Doing a deepseek v3.2 local installation means facing reality: 671 billion total parameters, 37 billion active per token, and a new sparse attention mechanism that changes the game for long contexts. This guide covers the real hardware requirements (not the marketing claims), explains ik_llama's Q2_K_XL quantization, demonstrates the selective offload trick using --override-tensor, and ends with tokens/sec measured on a home workstation in 2026.

By Mohamed Meguedmi·Update 2026-08-27·Tested on Windows, macOS, and Linux

#Why install DeepSeek V3.2 locally

DeepSeek V3.2 is the intermediate update released by DeepSeek AI in late 2025 under the MIT license. It introduces two significant changes compared with V3: DeepSeek Sparse Attention (DSA), a sparse-attention mechanism that makes the memory cost of long context subquadratic, and a refinement of multi-token prediction training. On paper, it preserves V3 quality while improving efficiency at 128k+ tokens.

Installing it locally addresses three concrete needs: total privacy (no document is sent to DeepSeek), reproducibility (the weights do not disappear overnight), and unrestricted experimentation (no rate limits and no mandatory moderation filter).

i
Let's be honest about the target
A local DeepSeek V3.2 deployment is not a real-time replacement for ChatGPT. The goal is instead useful throughput (5–15 tok/s) on a powerful home workstation, for batch reasoning, long-document summarization, or complex code where quality matters more than latency.

#MoE 671B / 37B active + DeepSeek Sparse Attention

The Local AI Kit

Your private ChatGPT, free, on your own machine in an hour — LM Studio, Ollama, Open WebUI, your documents, no cloud.

  • Lifetime online access
  • PDF + files
  • Lifetime updates

DeepSeek V3.2 retains V3’s MoE architecture: a router selects 8 experts out of 256 per layer, so only ~37 billion parameters are involved for each generated token. The remaining 634 billion are idle, but must remain addressable; otherwise, the router loses access to experts it has not loaded.

Total parameters
≈ 671 billion (the bit rate does not matter; this is the amount that has to fit somewhere: RAM, VRAM, or an SSD via mmap).
Active parameters per token
≈ 37 billion. This figure dictates theoretical speed: a 671B MoE incurs the compute cost of a dense 37B model.
Experts per layer
256 MoE experts + 1 shared expert, top-8 routing. The more RAM you have, the more expert layers remain resident and the more stable generation becomes.
Warning
Multi-Head Latent Attention (MLA) inherited from V3, now combined with DeepSeek Sparse Attention (DSA) for contexts > 32k tokens.
Context
128k tokens advertised, usable in practice thanks to DSA—the KV cache no longer grows linearly as it does on a conventional dense model such as Qwen 3.5 or Gemma 4.
→
What DSA really changes
DeepSeek Sparse Attention selects a subset of positions to attend to for each token, instead of scanning the entire context. At 128k, it cuts KV-cache memory usage by a factor of 3 to 5 depending on the setting, and speeds up prefill by the same factor. This is the only technical argument for targeting V3.2 rather than V3 if you want to take advantage of long contexts.

#Honest hardware requirements

No consumer-grade setup can handle DeepSeek V3.2 without contortions. Here are the three realistic profiles for 2026, ranked by the speed/budget tradeoff.

Profile A — high-end DDR5 (192–384 GB)
Workstation with Threadripper 7960X / Xeon W or EPYC 9354P, 256 GB DDR5 ECC (8 channels), no GPU required. Q2_K_XL fits entirely in RAM. Speed: 5–9 tok/s on CPU alone.
Profile B — RTX 5090 + 192 GB DDR5 (recommended)
The 2026 sweet spot: RTX 5090 (32 GB GDDR7, 1792 GB/s bandwidth) + 192 GB of DDR5 on a recent consumer platform. Keep the attention layers and shared expert on the GPU, with the rest in RAM. Speed: 10–15 tok/s during generation.
Profile C — modest RAM + Gen 4/5 NVMe SSD
96–128 GB DDR5 + NVMe SSD at 7 GB/s minimum, ≥ 1 TB free. Experts are memory-mapped from the SSD. Speed: 1.5–3 tok/s — usable for overnight batch jobs, not interactively.
!
192 GB DDR5, the bare minimum
Below 192 GB of RAM, Q2_K_XL quantization (~220 GB for V3.2) does not fit in physical memory, and you constantly page from the SSD. Even with a fast Gen 5 NVMe drive, you will drop below 2 tok/s. If your RAM budget is constrained, move down to IQ1_S rather than saturating the SSD.

#1. Choose the quantization for V3.2

The GGUF ecosystem for DeepSeek V3.2 is driven by two main players: Unsloth (which publishes UD variants = “Unsloth Dynamic,” including the well-known Q2_K_XL) and bartowski. UDs use an importance matrix to preserve sensitive layers (attention, shared expert) at higher precision than the rest—a real practical difference in final quality.

IQ1_S (~150 GB)
Dynamic 1-bit quantization. Fits in 192 GB of RAM with room to spare. Degraded but usable quality for general-purpose chat. Avoid for code.
Q2_K_XL (~220 GB)
The ik_llama / Unsloth Dynamic reference. Keeps critical layers in Q4–Q5 and compresses the experts to Q2. ~95% of Q4 quality on reasoning benchmarks. Fits on 256 GB of RAM or 192 GB plus a little SSD offload.
Q3_K_S (~290 GB)
For 384 GB workstations. Near-Q4 quality on long contexts. Recommended if you plan to seriously use 128k tokens.
Q4_K_M (~400 GB)
Full throttle. Reserved for dual-socket EPYC servers with 512 GB+ of DDR5. Beyond that, the quality gains become marginal for local use.
→
Why Q2_K_XL instead of Q2_K_S
On a 671B MoE, sensitivity to quantization is highly heterogeneous: attention layers and routers handle Q2 poorly, while FFN experts handle it well. Q2_K_XL respects this distribution (mixed precision), whereas Q2_K_S applies the same bit rate everywhere. For ~5% more memory, you gain ~10% quality on coding tasks.

#2. Compile ik_llama.cpp (the fork that supports V3.2)

At the time of writing, the main ggerganov/llama.cpp branch supports DeepSeek V3.2 but has no optimizations specific to DeepSeek's MoE routing. The ik_llama.cpp fork (Iwan Kawrakow) includes accelerated CUDA kernels for expert tensors and dynamic DSA support—a 30 to 50% throughput gain on V3.2, depending on the configuration.

Clone and compile ik_llama.cpp (CUDA + Flash Attention)
git clone https://github.com/ikawrakow/ik_llama.cpp
cd ik_llama.cpp
cmake -B build \
  -DGGML_CUDA=ON \
  -DGGML_CUDA_FA_ALL_QUANTS=ON \
  -DGGML_CUDA_F16=ON
cmake --build build --config Release -j $(nproc)

On Mac Apple Silicon, replace GGML_CUDA with GGML_METAL. On AMD, use GGML_HIP with ROCm 6.3+. Allow 10 to 15 minutes for the build on a recent machine—it’s C++ with CUDA generation, so don’t run it on a laptop without AC power.

i
Why ik_llama instead of a prebuilt binary
Linux/Windows releases of ik_llama.cpp exist, but the CUDA kernels generated at compile time are sensitive to your GPU’s compute capability (sm_120 for RTX 5090, sm_89 for RTX 4090, etc.). A generic binary falls back to generic code and loses 20–30% throughput. At this scale, the fifteen minutes of compilation are worth it.

#3. Download the GGUF V3.2 files

DeepSeek V3.2’s GGUF weights are published on Hugging Face. For Q2_K_XL Unsloth (the recommended option for most configurations), a single download of ~220 GB is split across several files.

Q2_K_XL download from Hugging Face
pip install -U "huggingface_hub[cli]"
huggingface-cli download \
  unsloth/DeepSeek-V3.2-GGUF \
  --include "*UD-Q2_K_XL*" \
  --local-dir ./models/deepseek-v32-q2kxl

The download takes several hours on a typical consumer connection—plan for a stable fiber connection and a target drive with ≥ 250 GB free. GGUFs of this size are split into 5 to 7 files (split-00001-of-N.gguf), and ik_llama.cpp automatically detects the splits when you point it to the first one.

!
Verify checksums
A GGUF truncated by 1 byte produces coherent responses for 20 tokens, then spirals into an infinite loop. It is the most painful bug to diagnose. Always verify the SHAs provided in the Hugging Face repo before the first inference—it takes 30 seconds per file with sha256sum.

#4. Launching with --override-tensor (the key to hybrid performance)

The secret to a good local deepseek v3.2 installation comes down to one option: --override-tensor (alias -ot). It lets you use regex to target which layers stay on CPU/RAM and which go to the GPU. With an MoE, you absolutely do not want to load everything onto the GPU (VRAM will never be enough), but you absolutely want attention and shared layers to be accelerated.

Launch RTX 5090 + 192 GB DDR5 (profile B)
./build/bin/llama-server \
  --model ./models/deepseek-v32-q2kxl/DeepSeek-V3.2-UD-Q2_K_XL-00001-of-00005.gguf \
  --ctx-size 65536 \
  --cache-type-k q8_0 \
  --cache-type-v q8_0 \
  --n-gpu-layers 999 \
  -ot "\.ffn_(up|down|gate)_exps\.=CPU" \
  --threads 16 \
  --flash-attn \
  --host 0.0.0.0 --port 8080

The regex \.ffn_(up|down|gate)_exps\. targets the three FFN tensors per expert—that is, the overwhelming majority of the 671B parameters—and forces them to stay on the CPU. What runs on the GPU: attention (MLA), the shared expert, embeddings, and the head. On RTX 5090 32 GB, it uses ~22 GB of VRAM; the rest is used for the KV cache.

Pure CPU launch (profile A, 256 GB DDR5)
./build/bin/llama-server \
  --model ./models/deepseek-v32-q2kxl/DeepSeek-V3.2-UD-Q2_K_XL-00001-of-00005.gguf \
  --ctx-size 32768 \
  --cache-type-k q8_0 \
  --cache-type-v q8_0 \
  --threads 32 \
  --flash-attn \
  --host 0.0.0.0 --port 8080
→
Quantize the KV cache even at 32k
Sur V3.2, --cache-type-k q8_0 --cache-type-v q8_0 divisent par 2 la consommation du cache KV avec une perte de qualité imperceptible. C'est ce qui vous permet de tenir 64k de contexte sur un GPU 32 Go au lieu de cracher en OOM à 24k.

#5. Tokens/sec measured on a home workstation

Benchmarks measured with the May 2026 ik_llama.cpp build, Q2_K_XL, 32k context, mixed French/code prompt. The figures are stable within ±15% depending on the prompt and memory-page warm-up.

RTX 5090 + Ryzen 9 7950X3D + 192 GB DDR5-6000
Prefill ~110 tok/s, generation 12–15 tok/s. The recommended home setup for 2026.
RTX 4090 + Threadripper 7960X + 256 GB DDR5 ECC
Prefill ~90 tok/s, generation 9–12 tok/s. The 4090 matches the 5090 on this hybrid profile because the bottleneck is RAM, not the GPU.
EPYC 9354P + 384 GB DDR5 ECC (12 channels), no GPU
Prefill ~70 tok/s, generation 7–9 tok/s. Total memory bandwidth (~460 GB/s) makes up for the lack of a GPU.
Mac Studio M3 Ultra 192 GB
Prefill ~55 tok/s, generation 6–8 tok/s in Q2_K_XL. Unified memory (~820 GB/s) saves the day.
Ryzen 9 7950X + 128 GB DDR5 + Samsung 990 Pro SSD
Prefill ~25 tok/s, generation 1.5–2.5 tok/s. Honestly usable only in batch.
i
Why prefill is fast and generation is slow
Prefill processes the prompt in a batch and benefits from matrix parallelism—that's where memory bandwidth and the GPU shine. Generation produces one token at a time, and each token must reread the ~37B routed parameters. This is intrinsic to MoE; no llama.cpp fork will work miracles.

#DeepSeek V3.2 vs. DeepSeek R1: which one should you install?

A recurring question, because the two models share the same 671B / 37B-active MoE architecture and the same MIT license. The difference is in post-training, not in the underlying skeleton.

DeepSeek V3.2
General-purpose model (chat), direct answers. Excellent in French, very strong at coding. Adds DSA for long context. Preferred for everyday assistance, summarization, and writing.
DeepSeek R1
Reasoning model (explicit chain of thought à la o1). Emits many internal tokens (<think>...</think>) before the final response. Best for math, logic, and complex algorithmic debugging.
Inference cost
R1 generates 3 to 10 times more tokens (because of thinking) to produce the same final answer. At the same throughput, R1 takes 5 minutes where V3.2 takes 30 seconds. Critical locally, where every tok/s counts.
Long context
V3.2, thanks to DSA, handles 128k tokens with a reasonable KV cache. R1 lacks DSA and runs out of memory beyond 64k.
VRAM/RAM required
Identical. Both models share the 671B/37B architecture and fit in the same ~220 GB Q2_K_XL quantization.
→
Practical verdict
Install V3.2 by default, and switch to R1 occasionally for tasks that justify an explicit chain of thought (math proofs, complex debugging). Many local users keep both GGUFs on their NVMe and use a simple routing wrapper.

#Troubleshooting

“unknown model architecture: deepseek2”
Your llama.cpp is too old or was compiled without deepseek2 support. Update ik_llama.cpp (main branch after April 2026) and recompile. The standard pre-built binary is not always sufficient.
CUDA OOM during loading
Your -ot regex does not offload enough experts to the CPU. Check the actual tensor distribution with ollama-style and ik_llama-bench. The correct regex for V3.2 properly targets ffn_(up|down|gate)_exps, not just ffn_exps.
Generation at 1 tok/s on 192 GB RAM
The kernel will page in from the GGUF until it has loaded all hot pages. An initial warm-up prompt of 200 tokens preloads the experts. Otherwise, increase --threads up to the number of physical cores (not logical cores).
Responses that switch to Chinese for no reason
Incorrect chat template. Make sure --chat-template is set to auto and that the GGUF actually includes the DeepSeek Jinja template. Otherwise, explicitly pass --chat-template deepseek3.
DSA appears inactive (linear KV cache)
DSA is enabled only from a configurable minimum context length. For contexts < 16k, the model uses standard dense attention—as expected, and this does not affect quality.
Crash on long prompt after warmup
You probably exhausted the swap. DeepSeek V3.2 does not like swap: if physical RAM is insufficient, disable swap (sudo swapoff -a) and let llama.cpp handle paging through mmap; it is faster and more stable.

#Go further

DeepSeek V3.2 locally is a starting point for exploring the class of “self-hostable frontier models.” Three complementary avenues:

Push the llama.cpp build
The guide “Compile llama.cpp with CUDA” details the advanced flags (Flash Attention, MMQ, offload tensors) that genuinely change MoE performance.
Understanding modern quantization
The “Choosing your quantization (Q4, Q5, Q8, FP16)” guide explains why Q2_K_XL works where standard Q2 fails, and when to move up to Q3/Q4.
Go one step further in size
The "Kimi K2 locally" guide applies the same method to a 1T-parameter MoE—useful for anticipating where the ecosystem is headed in 2026-2027.
Did this guide help you?

Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.