Kimi K2 locally: 1 trillion MoE parameters from soi
Running Kimi K2 locally with llama.cpp means running a one-trillion-parameter MoE model on a personal workstation. The feat comes down to two things: only 32B parameters are active per token (the rest sleep), and llama.cpp can leave inactive weights on an SSD via mmap. This guide covers the hardware configuration, Q2_K_S quantization, disk offload, and the real-world figures you can expect.
#Why run Kimi K2 locally
Kimi K2 is Moonshot AI’s flagship model, released with open weights (Modified MIT) and a MoE (Mixture of Experts) architecture with 1 trillion total parameters and about 32B active per token. On reasoning and coding benchmarks, it competes with closed frontier models, and its long context (128k tokens advertised) makes it a serious candidate for large-scale document analysis.
Running a model this size locally was a fantasy 18 months ago. Three developments changed that: the widespread availability of high-capacity DDR5 (192–384 GB for a few hundred euros), Gen 4 NVMe SSDs capable of 7 GB/s sequential reads, and the llama.cpp team’s work on highly aggressive quantizations such as Q2_K_S and IQ1_M.
#Understanding the 1T / 32B-active MoE
Your private ChatGPT, free, on your own machine in an hour — LM Studio, Ollama, Open WebUI, your documents, no cloud.
- Lifetime online access
- PDF + files
- Lifetime updates
The MoE architecture replaces dense FFN blocks with a layer of experts (often 256+) from which a router selects 8 to 12 per token. In terms of computation, only the parameters of the selected experts are processed—hence the ~32B “active” parameters. In terms of memory, however, all experts must be addressable; otherwise, the router loses access to 95% of the model.
- Total parameters
- ≈ 1,000 billion (1T). This is what sits in RAM or on disk.
- Active parameters per token
- ≈ 32 billion. This is what determines compute and speed.
- Experts per layer
- Several hundred, with a fixed number routed per token (top-k routing).
- Shared layers (attention)
- Always active. They dominate memory cost at small context sizes.
- KV cache
- Grows linearly with context. At 128k, the KV can exceed 30 GB even when quantized.
#Memory budget to plan for
Memory requirements for a 1T MoE break down into three distinct components. Understanding this breakdown is crucial before investing in RAM or an SSD.
- Model size (quantized)
- Q2_K_S ≈ 245 GB, Q3_K_S ≈ 320 GB, Q4_K_M ≈ 480 GB, Q8_0 ≈ 1 TB. This is the dominant and most compressible component.
- KV cache (context)
- Allow ~0.25 MB per token in FP16, ~0.12 MB in Q8. That’s 32 GB for 128k tokens in Q8. Configurable via --cache-type-k/-v.
- Activation buffers
- A few GB per GPU for intermediate computations. Marginal, but don't forget it on 24 GB GPUs.
#Hardware requirements
Three machine profiles can run Kimi K2 locally, with very different speed trade-offs.
- Profile A — high-end DDR5 (192–384 GB RAM)
- Threadripper/Xeon W workstation or EPYC platform. 256 GB DDR5 ECC, no GPU required. The model fits entirely in RAM in Q2_K_S. Expected speed: 4–8 tok/s on pure CPU.
- Profile B — DDR5 + 1 24 GB GPU
- 96–128 GB RAM + RTX 4090/3090. Offload the shared layers to the GPU; the experts remain in RAM. Expected speed: 6–12 tok/s, depending on how many experts fit in VRAM.
- Profile C — modest RAM + NVMe SSD
- 64–96 GB RAM + NVMe Gen 4 (7 GB/s read, ≥ 1 TB free). The weights are mmaped from the SSD. Expected speed: 1–3 tok/s, heavily dependent on the actual SSD throughput.
#1. Compile llama.cpp with the correct backend
Kimi K2 requires a recent version of llama.cpp (a recent build after December 2025) that supports its MoE architecture and IQ quantizations. We compile it from the official repo.
On Apple Silicon Macs, replace GGML_CUDA with GGML_METAL. On AMD with ROCm, use GGML_HIP. The GGML_CUDA_FA_ALL_QUANTS flag enables Flash Attention for all quantizations—essential for fitting 128k context without exhausting VRAM.
#2. Choose the GGUF quantization
For a model this size, forget Q4_K_M and above unless you have 512 GB of RAM. Extreme quantization—Q2_K_S, IQ2_XXS, or even IQ1_M—is what makes local deployment viable, at the cost of measurable but often acceptable quality loss on an MoE.
- Q2_K_S (~245 GB)
- The sweet spot. ~5% degradation on code benchmarks, ~3% on reasoning. Recommended if you have 256 GB of RAM or a fast SSD.
- IQ2_XXS (~210 GB)
- Even more aggressive via importance matrix. Fits in 192 GB of RAM with some offloading. One notch lower in quality.
- IQ1_M (~155 GB)
- Hybrid 1-bit quantization. For highly constrained configurations. Degradation is visible, but the model remains coherent.
- Q3_K_S (~320 GB)
- Near-Q4 quality but requires 384 GB of RAM. For well-equipped EPYC workstations.
GGUFs of this size are split across multiple files (split-00001-of-00007.gguf etc). llama.cpp automatically detects the splits if you point it to the first file.
#3. Run Kimi K2 with mmap and offload
The basic command uses llama-server, the HTTP server provided by llama.cpp. It exposes an OpenAI-compatible endpoint on port 8080 by default.
With a 24 GB GPU, offload the shared layers (attention) to the GPU and leave the MoE experts in RAM or on the SSD. The -ot flag lets you target precisely what goes to the GPU versus the CPU.
The regex passed to -ot means: “all FFN layers of the experts (the bulk of the MoE model) stay on CPU/RAM; the rest goes to the GPU.” This trick makes the hybrid approach viable: you benefit from the GPU for attention without having to load everything onto it.
#4. Measured real-world tokens/sec
Here are some real-world throughput benchmarks. The numbers vary depending on the prompt (prefill is slow, generation is more stable) and the state of the OS page cache.
- EPYC 9354P, 384 GB DDR5, Q2_K_S, entirely in RAM
- Prefill ~80 tok/s, generation ~6-8 tok/s. Reference configuration for reasoning batches.
- Threadripper 7960X, 256 GB DDR5, Q2_K_S
- Prefill ~50 tok/s, generation ~4-6 tok/s. DDR5 memory latency dominates.
- Ryzen 9 7950X, 128 GB DDR5 + Gen 4 NVMe SSD
- Prefill ~15 tok/s, generation ~1.5–3 tok/s. The SSD becomes the limiting factor as soon as you exceed resident RAM.
- Mac Studio M2 Ultra 192 GB
- Generation at ~4–7 tok/s in IQ2_XXS. Unified memory bandwidth (800 GB/s) helps enormously with MoEs.
- 256 GB workstation + RTX 4090 (hybrid)
- Prefill ~60 tok/s, generation ~7-10 tok/s. Running attention on the GPU cuts prefill time by 2.
#5. 128k-context use case
Kimi K2's greatest strength compared with local 7B-70B models is its 128k-token context window. In practice, you can feed it an entire annual report, a medium-sized codebase, or about a hundred pages of PDF, then ask for an overall summary that sees the whole picture—not chunk-based RAG.
- Long-document summarization
- A 300-page report fits within 120k tokens. Kimi K2 produces a structured summary in one pass, whereas RAG would mosaic one together.
- Codebase refactoring
- Load 50 Python files (~80k tokens) and request a coherent architecture review. Very useful for cross-cutting technical debt.
- Correlated log analysis
- Paste 50 MB of logs (after filtering) and ask for an incident timeline. The model sees all the correlations, not just the sliding windows.
- Translating large documents
- Maintains terminological consistency across an entire book, where chunk-based translation loses the thread.
#Troubleshooting
- “failed to load model” or magic number error
- Your llama.cpp is too old for this GGUF. Update to a build after December 2025 and check the architecture version announced by the GGUF repackager.
- Generation at 0.2 tok/s even though the RAM would be sufficient
- The kernel will page from the GGUF until it has touched every page. An initial “warm-up run” on a short prompt preloads the hot pages. Otherwise, increase --threads to saturate the bandwidth sooner.
- Brutal OOM on a long prompt
- The KV cache exploded. Set --ctx-size to 32768 or enable Q8 cache quantization. KV size grows linearly with context, not with model size.
- Incoherent output or broken language
- Often an incorrect chat template. Check --chat-template or make sure the GGUF includes the official Moonshot template (jinja). A wrong end-of-sequence token causes looping drift.
- SSD hits 80°C, throughput collapses
- High-end NVMe drives throttle without a heatsink. During a long session, add a heatsink or reduce the load by switching to a smaller quantization that fits in RAM.
#Go further
Kimi K2 locally is a playground for very large models. Three ways to push further:
- Mastering the llama.cpp build
- The “Compiling llama.cpp with CUDA” guide details the less obvious flags (offload tensors, MMQ, Flash Attention) that change performance on this type of model.
- Understanding quantization
- The “Choosing Your Quantization” guide lays out the useful conceptual foundations before diving into IQ2_XXS or Q3_K_S, and clarifies what actually degrades.
- Push compression further
- The TurboQuant guide details a method newer than K-quants for fitting frontier models on consumer hardware, with different trade-offs.
Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.