Install KTransformers: 70B+ LLMs on consumer GPUs

The framework KTransformers installation allows MoE (Mixture of Experts) models with more than 600 billion parameters to run locally on a single RTX 4090 or 3090 card by offloading inactive experts to CPU RAM. This KTransformers installation relies on selective offloading of MoE experts and uses the CPU's AVX-512 / AMX instructions to offload the GPU. This guide covers the hardware requirements, step-by-step installation procedure, compatible models (including DeepSeek V3 consumer GPU), measured performance, and common deployment pitfalls for a Local 671B LLM.

Why KTransformers changes the game for MoE LLMs

KTransformers is a Python framework developed by the KVCache.AI Lab at Tsinghua University, released under the Apache 2.0 license on GitHub. Its distinguishing feature: it loads only the active MoE experts into VRAM during inference, keeping the others in system RAM. For a model like DeepSeek V3 671B which activates only ~37B parameters per token out of its 671B total, making the VRAM savings massive.

In practice, where standard inference (vLLM, transformers) would require ~400 GB of Q4 VRAM for DeepSeek V3, KTransformers lets you run the same model with:

This approach directly competes with llama.cpp and its mechanism --n-gpu-layers, but KTransformers goes further with MoE architectures by routing each expert to the right device instead of offloading entire layers.

To put the scale in perspective, here are a few MoE models that become accessible through this approach:

Hardware and software requirements

Before launching the KTransformers installation, validate your configuration. The framework specifically targets MoE architectures and requires a particular RAM/VRAM balance.

Minimum hardware for DeepSeek V3 Q4_K_M :

Software stack :

To compare this approach with a traditional all-GPU setup, see our vLLM guide which details tensor parallelism on H100 clusters.

Step-by-step installation procedure

La KTransformers installation follows a strict sequence. Any deviation in CUDA/PyTorch versions breaks CPU kernel compilation.

Step 1 — Isolated Python environment :

conda create -n ktrans python=3.11
conda activate ktrans
conda install -c conda-forge libstdcxx-ng

Step 2 — PyTorch with CUDA :

pip install torch==2.4.0 --index-url https://download.pytorch.org/whl/cu124
pip install packaging ninja cpufeature numpy

Step 3 — Clone and compile :

git clone https://github.com/kvcache-ai/ktransformers.git
cd ktransformers
git submodule update --init --recursive
bash install.sh

The script install.sh compiles the CPU kernels (llamafile, AMX if detected) and installs the local wheel. Compilation takes 10 to 20 minutes depending on the CPU.

Step 4 — Downloading the weights :

KTransformers consumes GGUF files (quantization of llama.cpp). Retrieve, for example DeepSeek-V3-Q4_K_M.gguf depuis Hugging Face :

huggingface-cli download unsloth/DeepSeek-V3-GGUF \
  --include "DeepSeek-V3-Q4_K_M*" \
  --local-dir ./models/deepseek-v3-q4

Step 5 — Launching inference :

python -m ktransformers.local_chat \
  --model_path deepseek-ai/DeepSeek-V3 \
  --gguf_path ./models/deepseek-v3-q4 \
  --cpu_infer 48 \
  --max_new_tokens 1024

The parameter --cpu_infer corresponds to the number of CPU threads dedicated to offloaded experts. Set it to (cores physiques - 2).

Performance measured on consumer GPUs

The figures below come from the official KTransformers v0.2 benchmark and community feedback. For a Local 671B LLM quantified Q4_K_M:

For comparison, DeepSeek R1 Distill Llama 70B running fully on the GPU across 2× RTX 4090 reaches 25–30 tok/s — faster in absolute terms, but the distilled model underperforms the full 671B model on reasoning benchmarks.

Quality benchmarks (official reports DeepSeek) :

For a deeper look at the quality/speed tradeoffs, see DeepSeek R1 vs. Llama 3.3 70B.

Compatible models and limitations

Not all MoE LLMs are natively supported. KTransformers maintains a list of YAML optimizations by architecture. In the current catalog:

Native support confirmed :

Support in progress or experimental :

Unsupported (dense architectures, no benefit from offloading):

The practical maximum context remains limited by the RAM available for the KV cache: allow ~30 GB extra for 32K tokens on DeepSeek V3. Beyond 64K tokens, throughput drops sharply.

FAQ

Q: Does KTransformers work on Mac Apple Silicon?

No. The framework relies on CUDA for GPU kernels and AVX-512/AMX for x86 CPU kernels. On Mac M-series, use llama.cpp with Metal, or MLX. Our LLM guide for Mac details the alternatives. An M3 Ultra 512 GB can run DeepSeek V3 Q4 natively without offloading, at an estimated 18–20 tok/s.

Q: How does it differ from llama.cpp in offload mode?

llama.cpp offloads entire layers (--n-gpu-layers N), which works for dense models such as Llama 3.3 70B Instruct. KTransformers offloads them individual experts of an MoE, which is much more efficient for DeepSeek V3, where only 8 of 256 experts are activated per token. Typical gain: 3 to 5× faster than llama.cpp on DeepSeek V3 with equivalent hardware.

Q: Can you fine-tune with KTransformers?

No, the framework targets inference exclusively. To fine-tune an MoE, look at DeepSpeed-MoE or Unsloth (limited to dense models < 70B). KTransformers supports neither dynamic LoRA nor QLoRA. If you're looking to adapt a 70B model, use Llama 3.1 Nemotron 70B with Unsloth in 4-bit QLoRA on RTX 4090.

Q: How much does a DeepSeek V3 consumer-GPU setup cost?

Estimated budget (2026): used RTX 4090 €1,500, W790 server motherboard or EPYC SP5 €800, Xeon w7-2495X CPU €2,200, 384 GB DDR5 ECC €3,000, 2 TB NVMe €200, PSU + case €400. Total: ~€8,000 to run a Local 671B LLM. Compare that with the ~€30,000 cost of a single H100 80 GB, which would not even be sufficient at full precision.

Q: Is OpenAI-compatible server mode available?

Yes, since v0.2. Run ktransformers --server_port 10002 --model_path ... --gguf_path ... and point your OpenAI client to http://localhost:10002/v1. Compatible with Open WebUI, Continue.dev, and LiteLLM. SSE streaming works, but function calling remains partial (to be confirmed depending on the version).

Q: KTransformers vs. SGLang for DeepSeek V3?

SGLang targets high-performance multi-GPU deployments (H100/H200) and reaches 60+ tok/s on DeepSeek V3 with 8× H100s. KTransformers targets single-GPU consumer systems with RAM offloading, at 10–15 tok/s. Simple choice: SGLang if you have a cluster, KTransformers if you have a workstation.

Conclusion

La KTransformers installation opens access to frontier MoE models (DeepSeek V3, R1, Kimi K2.5) on an €8,000 workstation instead of a €200,000 H100 cluster. The trade-off: 10–15 tok/s instead of 60+, but with response quality equivalent to the full-precision model. To identify the MoE model suited to your VRAM and RAM, use the quelllm.fr configurator, or explore the catalog of 249 indexed LLMs to compare licenses and sizes.

Article published on by Mohamed Meguedmi · Data source: /api/models.json · Content license: CC BY 4.0.

An error or update to report? Contribute.