Install KTransformers: 70B+ LLMs on consumer GPUs
The framework KTransformers installation allows MoE (Mixture of Experts) models with more than 600 billion parameters to run locally on a single RTX 4090 or 3090 card by offloading inactive experts to CPU RAM. This KTransformers installation relies on selective offloading of MoE experts and uses the CPU's AVX-512 / AMX instructions to offload the GPU. This guide covers the hardware requirements, step-by-step installation procedure, compatible models (including DeepSeek V3 consumer GPU), measured performance, and common deployment pitfalls for a Local 671B LLM.
Why KTransformers changes the game for MoE LLMs
KTransformers is a Python framework developed by the KVCache.AI Lab at Tsinghua University, released under the Apache 2.0 license on GitHub. Its distinguishing feature: it loads only the active MoE experts into VRAM during inference, keeping the others in system RAM. For a model like DeepSeek V3 671B which activates only ~37B parameters per token out of its 671B total, making the VRAM savings massive.
In practice, where standard inference (vLLM, transformers) would require ~400 GB of Q4 VRAM for DeepSeek V3, KTransformers lets you run the same model with:
- GPU : 24 GB VRAM (RTX 4090 or 3090)
- CPU RAM : 380 to 512 GB DDR5
- CPU : Intel Xeon with AMX or AMD EPYC Genoa (estimated optimal)
This approach directly competes with llama.cpp and its mechanism --n-gpu-layers, but KTransformers goes further with MoE architectures by routing each expert to the right device instead of offloading entire layers.
To put the scale in perspective, here are a few MoE models that become accessible through this approach:
- DeepSeek V3 671B : ~400 GB Q4 VRAM → runnable on a 24 GB GPU + 384 GB RAM
- DeepSeek R1 671B : same profile, plus reasoning
- Kimi K2.5 : 1000B parameters, ~32B active
- Qwen 3 235B-A22B : 235B with 22B active
Hardware and software requirements
Before launching the KTransformers installation, validate your configuration. The framework specifically targets MoE architectures and requires a particular RAM/VRAM balance.
Minimum hardware for DeepSeek V3 Q4_K_M :
- GPU NVIDIA : RTX 3090 / 4090 / 5090 (24 GB) or A6000 (48 GB). Compute capability ≥ 8.0
- DDR5 RAM : 384 GB minimum, 512 GB recommended for long contexts
- CPU : Intel Xeon Sapphire Rapids (AMX), Xeon Emerald Rapids, or AMD EPYC 9004 (estimated). Desktop CPUs such as the Core i9-14900K work, but without AMX
- Storage : 450 GB free on Gen4 NVMe for GGUF weights
- OS : Ubuntu 22.04 LTS or Debian 12 (Windows WSL2 supported but 20-30% slower, to be confirmed)
Software stack :
- Python 3.11
- CUDA 12.4 or 12.6
- PyTorch 2.4+ compiled with CUDA
- gcc-11 minimum (gcc-13 for AMX)
- CMake 3.25+
To compare this approach with a traditional all-GPU setup, see our vLLM guide which details tensor parallelism on H100 clusters.
Step-by-step installation procedure
La KTransformers installation follows a strict sequence. Any deviation in CUDA/PyTorch versions breaks CPU kernel compilation.
Step 1 — Isolated Python environment :
conda create -n ktrans python=3.11
conda activate ktrans
conda install -c conda-forge libstdcxx-ng
Step 2 — PyTorch with CUDA :
pip install torch==2.4.0 --index-url https://download.pytorch.org/whl/cu124
pip install packaging ninja cpufeature numpy
Step 3 — Clone and compile :
git clone https://github.com/kvcache-ai/ktransformers.git
cd ktransformers
git submodule update --init --recursive
bash install.sh
The script install.sh compiles the CPU kernels (llamafile, AMX if detected) and installs the local wheel. Compilation takes 10 to 20 minutes depending on the CPU.
Step 4 — Downloading the weights :
KTransformers consumes GGUF files (quantization of llama.cpp). Retrieve, for example DeepSeek-V3-Q4_K_M.gguf depuis Hugging Face :
huggingface-cli download unsloth/DeepSeek-V3-GGUF \
--include "DeepSeek-V3-Q4_K_M*" \
--local-dir ./models/deepseek-v3-q4
Step 5 — Launching inference :
python -m ktransformers.local_chat \
--model_path deepseek-ai/DeepSeek-V3 \
--gguf_path ./models/deepseek-v3-q4 \
--cpu_infer 48 \
--max_new_tokens 1024
The parameter --cpu_infer corresponds to the number of CPU threads dedicated to offloaded experts. Set it to (cores physiques - 2).
Performance measured on consumer GPUs
The figures below come from the official KTransformers v0.2 benchmark and community feedback. For a Local 671B LLM quantified Q4_K_M:
- RTX 4090 + Xeon Gold 6454S (AMX) : 13.6 tok/s prefill, 14 tok/s decode (DeepSeek V3)
- RTX 4090 + i9-14900K (without AMX) : estimated 8 tok/s decode
- RTX 3090 + EPYC 9354 : estimated 10 tok/s decode (to be confirmed based on RAM bandwidth)
For comparison, DeepSeek R1 Distill Llama 70B running fully on the GPU across 2× RTX 4090 reaches 25–30 tok/s — faster in absolute terms, but the distilled model underperforms the full 671B model on reasoning benchmarks.
Quality benchmarks (official reports DeepSeek) :
- DeepSeek V3 671B : MMLU 88.5, HumanEval 82.6, MATH-500 90.2
- DeepSeek R1 671B : AIME 2024 79.8, MATH-500 97.3, Codeforces 96.3 percentile
- Qwen 3 235B-A22B : MMLU-Pro 73 estimated, HumanEval 88 to be confirmed
For a deeper look at the quality/speed tradeoffs, see DeepSeek R1 vs. Llama 3.3 70B.
Compatible models and limitations
Not all MoE LLMs are natively supported. KTransformers maintains a list of YAML optimizations by architecture. In the current catalog:
Native support confirmed :
- DeepSeek V3 671B (MoE, 37B active)
- DeepSeek R1 671B (MoE, 37B active)
- DeepSeek V3.2 (685B, MoE)
- Mixtral 8x22B Instruct (141B, 8 experts × 22B)
- Qwen 3 235B-A22B
Support in progress or experimental :
- Kimi K2.5 et Kimi K2.6 (DeepSeek-like architecture, should work)
- GLM-5.1 (744B MoE, to be confirmed)
- ERNIE 4.5 300B-A47B
Unsupported (dense architectures, no benefit from offloading):
- Llama 3.1 405B Instruct — use vLLM multi-GPU instead
- Qwen 2.5 72B Instruct — full-GPU on 2× RTX 4090
The practical maximum context remains limited by the RAM available for the KV cache: allow ~30 GB extra for 32K tokens on DeepSeek V3. Beyond 64K tokens, throughput drops sharply.
FAQ
Q: Does KTransformers work on Mac Apple Silicon?
No. The framework relies on CUDA for GPU kernels and AVX-512/AMX for x86 CPU kernels. On Mac M-series, use llama.cpp with Metal, or MLX. Our LLM guide for Mac details the alternatives. An M3 Ultra 512 GB can run DeepSeek V3 Q4 natively without offloading, at an estimated 18–20 tok/s.
Q: How does it differ from llama.cpp in offload mode?
llama.cpp offloads entire layers (--n-gpu-layers N), which works for dense models such as Llama 3.3 70B Instruct. KTransformers offloads them individual experts of an MoE, which is much more efficient for DeepSeek V3, where only 8 of 256 experts are activated per token. Typical gain: 3 to 5× faster than llama.cpp on DeepSeek V3 with equivalent hardware.
Q: Can you fine-tune with KTransformers?
No, the framework targets inference exclusively. To fine-tune an MoE, look at DeepSpeed-MoE or Unsloth (limited to dense models < 70B). KTransformers supports neither dynamic LoRA nor QLoRA. If you're looking to adapt a 70B model, use Llama 3.1 Nemotron 70B with Unsloth in 4-bit QLoRA on RTX 4090.
Q: How much does a DeepSeek V3 consumer-GPU setup cost?
Estimated budget (2026): used RTX 4090 €1,500, W790 server motherboard or EPYC SP5 €800, Xeon w7-2495X CPU €2,200, 384 GB DDR5 ECC €3,000, 2 TB NVMe €200, PSU + case €400. Total: ~€8,000 to run a Local 671B LLM. Compare that with the ~€30,000 cost of a single H100 80 GB, which would not even be sufficient at full precision.
Q: Is OpenAI-compatible server mode available?
Yes, since v0.2. Run ktransformers --server_port 10002 --model_path ... --gguf_path ... and point your OpenAI client to http://localhost:10002/v1. Compatible with Open WebUI, Continue.dev, and LiteLLM. SSE streaming works, but function calling remains partial (to be confirmed depending on the version).
Q: KTransformers vs. SGLang for DeepSeek V3?
SGLang targets high-performance multi-GPU deployments (H100/H200) and reaches 60+ tok/s on DeepSeek V3 with 8× H100s. KTransformers targets single-GPU consumer systems with RAM offloading, at 10–15 tok/s. Simple choice: SGLang if you have a cluster, KTransformers if you have a workstation.
Conclusion
La KTransformers installation opens access to frontier MoE models (DeepSeek V3, R1, Kimi K2.5) on an €8,000 workstation instead of a €200,000 H100 cluster. The trade-off: 10–15 tok/s instead of 60+, but with response quality equivalent to the full-precision model. To identify the MoE model suited to your VRAM and RAM, use the quelllm.fr configurator, or explore the catalog of 249 indexed LLMs to compare licenses and sizes.