Advanced 18 minllama.cpp

Kimi K2 locally: 1 trillion MoE parameters from soi

Running Kimi K2 locally with llama.cpp means running a one-trillion-parameter MoE model on a personal workstation. The feat comes down to two things: only 32B parameters are active per token (the rest sleep), and llama.cpp can leave inactive weights on an SSD via mmap. This guide covers the hardware configuration, Q2_K_S quantization, disk offload, and the real-world figures you can expect.

By Mohamed Meguedmi·Update 2026-06-05·Tested on Windows, macOS, and Linux

#Why run Kimi K2 locally

Kimi K2 is Moonshot AI’s flagship model, released with open weights (Modified MIT) and a MoE (Mixture of Experts) architecture with 1 trillion total parameters and about 32B active per token. On reasoning and coding benchmarks, it competes with closed frontier models, and its long context (128k tokens advertised) makes it a serious candidate for large-scale document analysis.

Running a model this size locally was a fantasy 18 months ago. Three developments changed that: the widespread availability of high-capacity DDR5 (192–384 GB for a few hundred euros), Gen 4 NVMe SSDs capable of 7 GB/s sequential reads, and the llama.cpp team’s work on highly aggressive quantizations such as Q2_K_S and IQ1_M.

i
Not real-time, but usable
Let's be clear: at 2–6 tokens/sec, Kimi K2 locally is not an interactive chatbot. It's a model you query in batches, leave to process a long context, and use when response quality matters more than latency.

#Understanding the 1T / 32B-active MoE

The Local AI Kit

Your private ChatGPT, free, on your own machine in an hour — LM Studio, Ollama, Open WebUI, your documents, no cloud.

  • Lifetime online access
  • PDF + files
  • Lifetime updates

The MoE architecture replaces dense FFN blocks with a layer of experts (often 256+) from which a router selects 8 to 12 per token. In terms of computation, only the parameters of the selected experts are processed—hence the ~32B “active” parameters. In terms of memory, however, all experts must be addressable; otherwise, the router loses access to 95% of the model.

Total parameters
≈ 1,000 billion (1T). This is what sits in RAM or on disk.
Active parameters per token
≈ 32 billion. This is what determines compute and speed.
Experts per layer
Several hundred, with a fixed number routed per token (top-k routing).
Shared layers (attention)
Always active. They dominate memory cost at small context sizes.
KV cache
Grows linearly with context. At 128k, the KV can exceed 30 GB even when quantized.
→
Why MoE fits where a 1T dense model never would
A dense 1T-parameter model would require ~2 TB in FP16 and ~250 GB even in Q2. At equivalent quality, a 1T / 32B-active MoE can run on 200–250 GB with aggressive quantization, and most of that memory is read, not written—which is why mmap on SSD is viable.

#Memory budget to plan for

Memory requirements for a 1T MoE break down into three distinct components. Understanding this breakdown is crucial before investing in RAM or an SSD.

Model size (quantized)
Q2_K_S ≈ 245 GB, Q3_K_S ≈ 320 GB, Q4_K_M ≈ 480 GB, Q8_0 ≈ 1 TB. This is the dominant and most compressible component.
KV cache (context)
Allow ~0.25 MB per token in FP16, ~0.12 MB in Q8. That’s 32 GB for 128k tokens in Q8. Configurable via --cache-type-k/-v.
Activation buffers
A few GB per GPU for intermediate computations. Marginal, but don't forget it on 24 GB GPUs.
!
RAM ≠ model storage
With mmap, llama.cpp can address weights that do not fit in RAM: the kernel pages them in from the SSD on demand. But each nonresident page costs a disk read during inference—which is why a fast NVMe drive matters. RAM = speed, SSD = capacity.

#Hardware requirements

Three machine profiles can run Kimi K2 locally, with very different speed trade-offs.

Profile A — high-end DDR5 (192–384 GB RAM)
Threadripper/Xeon W workstation or EPYC platform. 256 GB DDR5 ECC, no GPU required. The model fits entirely in RAM in Q2_K_S. Expected speed: 4–8 tok/s on pure CPU.
Profile B — DDR5 + 1 24 GB GPU
96–128 GB RAM + RTX 4090/3090. Offload the shared layers to the GPU; the experts remain in RAM. Expected speed: 6–12 tok/s, depending on how many experts fit in VRAM.
Profile C — modest RAM + NVMe SSD
64–96 GB RAM + NVMe Gen 4 (7 GB/s read, ≥ 1 TB free). The weights are mmaped from the SSD. Expected speed: 1–3 tok/s, heavily dependent on the actual SSD throughput.
→
The SSD is the new bottleneck
If you plan to rely on disk offloading, check the SSD’s random latency, not just its marketing sequential throughput. A Samsung 990 Pro or WD SN850X does the job properly. An entry-level QLC SSD collapses to 200 MB/s for random reads—unusable for this use case.

#1. Compile llama.cpp with the correct backend

Kimi K2 requires a recent version of llama.cpp (a recent build after December 2025) that supports its MoE architecture and IQ quantizations. We compile it from the official repo.

Clone and compile (CUDA + mmap)
git clone https://github.com/ggerganov/llama.cpp
cd llama.cpp
cmake -B build -DGGML_CUDA=ON -DGGML_CUDA_FA_ALL_QUANTS=ON
cmake --build build --config Release -j

On Apple Silicon Macs, replace GGML_CUDA with GGML_METAL. On AMD with ROCm, use GGML_HIP. The GGML_CUDA_FA_ALL_QUANTS flag enables Flash Attention for all quantizations—essential for fitting 128k context without exhausting VRAM.

i
Precompiled binaries
If you want to avoid compilation, the GitHub releases provide binaries for Linux/Windows/macOS. Make sure the version supports the deepseek2 or kimi architecture, depending on the GGUF repackaging used.

#2. Choose the GGUF quantization

For a model this size, forget Q4_K_M and above unless you have 512 GB of RAM. Extreme quantization—Q2_K_S, IQ2_XXS, or even IQ1_M—is what makes local deployment viable, at the cost of measurable but often acceptable quality loss on an MoE.

Q2_K_S (~245 GB)
The sweet spot. ~5% degradation on code benchmarks, ~3% on reasoning. Recommended if you have 256 GB of RAM or a fast SSD.
IQ2_XXS (~210 GB)
Even more aggressive via importance matrix. Fits in 192 GB of RAM with some offloading. One notch lower in quality.
IQ1_M (~155 GB)
Hybrid 1-bit quantization. For highly constrained configurations. Degradation is visible, but the model remains coherent.
Q3_K_S (~320 GB)
Near-Q4 quality but requires 384 GB of RAM. For well-equipped EPYC workstations.
Download a GGUF Q2_K_S from Hugging Face
pip install -U "huggingface_hub[cli]"
huggingface-cli download \
  unsloth/Kimi-K2-Instruct-GGUF \
  --include "*Q2_K_S*" \
  --local-dir ./models/kimi-k2-q2ks

GGUFs of this size are split across multiple files (split-00001-of-00007.gguf etc). llama.cpp automatically detects the splits if you point it to the first file.

!
Hash and provenance
Download your GGUF files from a trusted repository (unsloth, bartowski, and mradermacher are the most widely followed). Verify the provided SHA checksums. A poorly quantized or corrupted GGUF may produce coherent text for the first prompt and then go off the rails—a painful diagnosis.

#3. Run Kimi K2 with mmap and offload

The basic command uses llama-server, the HTTP server provided by llama.cpp. It exposes an OpenAI-compatible endpoint on port 8080 by default.

CPU launch + SSD offload via mmap
./build/bin/llama-server \
  --model ./models/kimi-k2-q2ks/kimi-k2-Q2_K_S-00001-of-00007.gguf \
  --ctx-size 32768 \
  --cache-type-k q8_0 \
  --cache-type-v q8_0 \
  --threads 16 \
  --no-mmap=false \
  --host 0.0.0.0 --port 8080

With a 24 GB GPU, offload the shared layers (attention) to the GPU and leave the MoE experts in RAM or on the SSD. The -ot flag lets you target precisely what goes to the GPU versus the CPU.

Hybrid CPU + GPU 24 GB
./build/bin/llama-server \
  --model ./models/kimi-k2-q2ks/kimi-k2-Q2_K_S-00001-of-00007.gguf \
  --ctx-size 65536 \
  --cache-type-k q8_0 \
  --cache-type-v q8_0 \
  --n-gpu-layers 999 \
  -ot "\.ffn_.*_exps\.=CPU" \
  --threads 12 \
  --host 0.0.0.0 --port 8080

The regex passed to -ot means: “all FFN layers of the experts (the bulk of the MoE model) stay on CPU/RAM; the rest goes to the GPU.” This trick makes the hybrid approach viable: you benefit from the GPU for attention without having to load everything onto it.

→
mmap is enabled by default
llama.cpp uses mmap automatically. If physical RAM is insufficient, the kernel will page from the GGUF file on disk — so place the GGUF on your fastest NVMe drive, never on an HDD. Avoid --no-mmap unless you have a specific reason.

#4. Measured real-world tokens/sec

Here are some real-world throughput benchmarks. The numbers vary depending on the prompt (prefill is slow, generation is more stable) and the state of the OS page cache.

EPYC 9354P, 384 GB DDR5, Q2_K_S, entirely in RAM
Prefill ~80 tok/s, generation ~6-8 tok/s. Reference configuration for reasoning batches.
Threadripper 7960X, 256 GB DDR5, Q2_K_S
Prefill ~50 tok/s, generation ~4-6 tok/s. DDR5 memory latency dominates.
Ryzen 9 7950X, 128 GB DDR5 + Gen 4 NVMe SSD
Prefill ~15 tok/s, generation ~1.5–3 tok/s. The SSD becomes the limiting factor as soon as you exceed resident RAM.
Mac Studio M2 Ultra 192 GB
Generation at ~4–7 tok/s in IQ2_XXS. Unified memory bandwidth (800 GB/s) helps enormously with MoEs.
256 GB workstation + RTX 4090 (hybrid)
Prefill ~60 tok/s, generation ~7-10 tok/s. Running attention on the GPU cuts prefill time by 2.
i
Prefill vs generation
For an MoE of this size, prefill (reading the prompt) benefits from batch parallelism and remains acceptable. Token-by-token generation is slower because each new token rereads the routed experts. This is intrinsic to the architecture, not a llama.cpp defect.

#5. 128k-context use case

Kimi K2's greatest strength compared with local 7B-70B models is its 128k-token context window. In practice, you can feed it an entire annual report, a medium-sized codebase, or about a hundred pages of PDF, then ask for an overall summary that sees the whole picture—not chunk-based RAG.

Long-document summarization
A 300-page report fits within 120k tokens. Kimi K2 produces a structured summary in one pass, whereas RAG would mosaic one together.
Codebase refactoring
Load 50 Python files (~80k tokens) and request a coherent architecture review. Very useful for cross-cutting technical debt.
Correlated log analysis
Paste 50 MB of logs (after filtering) and ask for an incident timeline. The model sees all the correlations, not just the sliding windows.
Translating large documents
Maintains terminological consistency across an entire book, where chunk-based translation loses the thread.
→
Quantize the KV cache to fit 128k
With --cache-type-k q8_0 --cache-type-v q8_0, the KV cache at 128k drops from about 60 GB to ~30 GB. The quality loss is imperceptible in practice, and you gain the headroom needed to avoid saturating RAM.

#Troubleshooting

“failed to load model” or magic number error
Your llama.cpp is too old for this GGUF. Update to a build after December 2025 and check the architecture version announced by the GGUF repackager.
Generation at 0.2 tok/s even though the RAM would be sufficient
The kernel will page from the GGUF until it has touched every page. An initial “warm-up run” on a short prompt preloads the hot pages. Otherwise, increase --threads to saturate the bandwidth sooner.
Brutal OOM on a long prompt
The KV cache exploded. Set --ctx-size to 32768 or enable Q8 cache quantization. KV size grows linearly with context, not with model size.
Incoherent output or broken language
Often an incorrect chat template. Check --chat-template or make sure the GGUF includes the official Moonshot template (jinja). A wrong end-of-sequence token causes looping drift.
SSD hits 80°C, throughput collapses
High-end NVMe drives throttle without a heatsink. During a long session, add a heatsink or reduce the load by switching to a smaller quantization that fits in RAM.

#Go further

Kimi K2 locally is a playground for very large models. Three ways to push further:

Mastering the llama.cpp build
The “Compiling llama.cpp with CUDA” guide details the less obvious flags (offload tensors, MMQ, Flash Attention) that change performance on this type of model.
Understanding quantization
The “Choosing Your Quantization” guide lays out the useful conceptual foundations before diving into IQ2_XXS or Q3_K_S, and clarifies what actually degrades.
Push compression further
The TurboQuant guide details a method newer than K-quants for fitting frontier models on consumer hardware, with different trade-offs.
Did this guide help you?

Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.