Kimi K3 at home: the truth about the hardware required (2.8T parameters)
Kimi K3 does not run on a desktop workstation: its native weights weigh 1.56 TB, and the smallest published GGUF quantization (Unsloth, dynamic 1-bit) still uses 594 GB, with 610 GB of memory recommended. No 512 GB Mac Studio and no single RTX 5090 are sufficient. To try it, use Moonshot’s API or kimi-k3:cloud on Ollama; locally, target a smaller model.
The weights for Kimi K3 have been available since late July 2026, and searches for “ollama kimi k3,” “kimi k3 lmstudio,” and “kimi k3 on 5090” show that many readers are wondering whether they can install it at home. This guide gives the figures published by Moonshot and Unsloth, calculates what your machine can load, and explains what to do instead.
#Kimi K3 locally: what Moonshot actually released
Moonshot's Hugging Face model card describes Kimi K3 as an open-weight multimodal model with 2.8 trillion parameters and a one-million-token context window. It is a Mixture-of-Experts model: 896 experts, 16 of which are selected for each token, for 104 billion active parameters. The weights are distributed in MXFP4 (weights) with MXFP8 activations, using quantization-aware training from the supervised fine-tuning phase onward. The code and weights are covered by the “Kimi K3 License”: read this text before any commercial use, because it is not a standard MIT or Apache license. The model appeared on Hugging Face around July 27, 2026, according to the specialist press.
- Total / active parameters
- 2.8T total, 104B active per token (Moonshot specification sheet).
- Experts
- 896 experts, 16 selected per token, plus 2 shared experts.
- Published format
- MXFP4 for weights, MXFP8 for activations. This is not a GGUF.
- Context
- 1,048,576 tokens, with a KV cache added to the weights' memory footprint.
- Recommended engines
- vLLM, SGLang, and TokenSpeed. llama.cpp uses community GGUFs.
#What the weights really weigh
Your private ChatGPT, free, on your own machine in an hour — LM Studio, Ollama, Open WebUI, your documents, no cloud.
- Lifetime online access
- PDF + files
- Lifetime updates
Two figures are circulating, and they need to be distinguished. Some articles estimate “around 1.4 TB” by multiplying 2.8T parameters by half a byte. The file size is higher: Unsloth and Runpod give 1.56 TB for the native version because the layers outside the experts (attention, routers, shared experts) remain at higher precision. The unquantized BF16 version would reach approximately 5.6 TB, according to Runpod. The version of this guide that stated 1.4 TB was therefore about 10% too low.
| Format | Size | Recommended total memory | Measured fidelity |
|---|---|---|---|
| Native MXFP4 / UD-Q8_K_XL | 1.56 TB | 1.6 TB | Lossless |
| Unsloth UD-Q2_K_XL (dynamic 2-bit) | 861.3 GB | 880 GB | About 90% top-1 agreement |
| Unsloth UD-IQ2_XXS | 711.1 GB | 726 GB | 84.1% top-1 agreement |
| Unsloth UD-IQ1_M | 648.9 GB | 665 GB | 81.2% top-1 agreement |
| Unsloth UD-IQ1_S (dynamic 1-bit) | 594 GB | 610 GB | 78.9% top-1 agreement |
| BF16 (theoretical) | approximately 5.6 TB | Out of reach | Reference |
The table disproves a misconception from the older version: a Q2 is not “always on the order of a terabyte.” Unsloth does offer GGUFs under 900 GB. But 594 GB is still almost fifteen times the memory of a 32 GB RTX 5090, and the quality gap is real: the dynamic 1-bit version reproduces the original model’s choice in only 78.9% of cases on Unsloth’s benchmark. To understand how bits affect quality, see our quantization guide.
#Hardware: what's recommended and what's possible
Moonshot does not publish a minimum configuration on the model page: it points to the vLLM, SGLang, and TokenSpeed recipes and to its API. The documented reference for native self-hosting comes from Runpod: a node with eight B300 GPUs of 288 GB each, or about 2.3 TB, or sixteen B200s spread across two nodes. The figure “64 accelerators or more” that this page previously repeated does not appear in any primary source consulted; it has been removed.
With an Unsloth GGUF, the rule given in their documentation is simple: total RAM plus VRAM should roughly equal the quantization size; otherwise, the model works but spills to disk, which is much slower. Unsloth also cites about 20 tokens per second during generation when the model fits on B200s. That is a vendor throughput figure on datacenter hardware, not a measurement applicable to a home machine.
#What about a Mac? Compute over hearsay
The figure of “16 seconds per token on an M1 Max MacBook Pro” reported on this page could not be traced to any verifiable source, so we do not present it as fact. What the sources establish is clearer. Kingy AI notes that the starting point (a 1.56 TB checkpoint) already exceeded a 512 GB Mac Studio, and that even the 553.2 GiB file for the 1-bit version exceeds this machine before overhead. A 128 GB Mac has about one-fifth of the recommended 610 GB.
You can estimate the order of magnitude yourself, though. If memory is insufficient, each token rereads the experts it activates from the SSD: 104 billion parameters at about 4 bits, or roughly 50 GB read per token. With an SSD delivering 3 GB/s, you get about 17 seconds per token; at 7 GB/s, about 7 seconds. This is a theoretical ceiling calculation, not a measurement, but it explains why reports of running “from disk” are measured in seconds per token, not tokens per second.
#Ollama, LM Studio, llama.cpp: what each tool does
- Ollama
- The official library lists only one variant, kimi-k3:cloud: the model runs on Ollama's servers, not on your machine. The ollama run kimi-k3:cloud command lets you try the model with the usual interface, subject to the privacy and billing limits of a remote service.
- LM Studio
- It loads local GGUF files. For K3, this means downloading several hundred GB of Unsloth files and having the corresponding memory available. On a consumer machine, the answer is no.
- llama.cpp
- This is the engine targeted by Unsloth GGUFs, which rely on a llama.cpp-derived branch with vision support. It handles offloading experts to the CPU and multipart files: the realistic path for a workstation with several hundred GB of RAM.
- vLLM and SGLang
- The two engines recommended by Moonshot for a multi-GPU service using the native format.
| Machine | Available memory | Verdict |
|---|---|---|
| RTX 5090 (32 GB) + 64 GB of RAM | approximately 96 GB | Impossible: 6 times too little, even at 1 bit |
| Mac mini or MacBook, 16 to 64 GB | 16 to 64 GB | Impossible locally; API or cloud only |
| Mac Studio 128 to 256 GB | 128 to 256 GB | Insufficient for the recommended 610 GB |
| Mac Studio 512 GB | 512 GB | Below the 1-bit file size, before context |
| 768 GB to 1 TB RAM workstation + GPU | 600 to 900 GB | Feasible in 1- to 2-bit GGUF, modest throughput |
| 8 B300 GPUs (about 2.3 TB) | 2.3 TB | Documented configuration for the native format |
#Realistic options for trying Kimi K3
- 011. Moonshot's official APIThe model is called kimi-k3 on platform.kimi.ai, with an API compatible with OpenAI and Anthropic. It is the fastest way to judge quality, with a reasoning_effort parameter adjustable to low, high, or max.
- 022. Ollama cloudollama run kimi-k3:cloud uses your usual Ollama workflow with remote inference. The library page lists prices per million tokens: check them before automating. Our Ollama Cloud guide details the limits.
- 033. Renting a GPURent a multi-GPU node by the hour for a test, using vLLM. First calculate the break-even point (hourly cost divided by sustained tokens per second), as Runpod explains in its FAQ.
- 044. A high-RAM workstationOnly if you already have 700 GB of memory: GGUF UD-IQ1_S and llama.cpp, accepting reduced quality and low throughput.
#Open weights does not mean you can run it at home
Open weights guarantee the right to download, audit, fine-tune, and, depending on the license, redistribute them. They do not guarantee that anyone can run them. A 30-billion-parameter model on a 24 GB card genuinely belongs to you; a 2.8T model that only a cluster can load remains, in practice, a service. Kimi K3 is good news for auditability and for organizations that already have a cluster, without changing what an individual can do at home.
#Alternatives your hardware can actually load
The QuelLLM catalog estimates memory requirements in Q4, excluding context: DeepSeek V4 Flash 284B about 170 GB, GLM 5.2 753B-A40B about 437 GB, Kimi K3 about 1,624 GB. For everyday use, a 30- to 70-billion-parameter model in Q4 remains the best balance between quality and feasibility.
- DeepSeek V4 Flash 284B
- Approximately 170 GB in Q4 according to the catalog: possible on a high-memory Mac Studio or a workstation. See the dedicated guide.
- GLM 5.2 753B-A40B
- About 437 GB in Q4: it’s the closest form factor to K3 accessible to a well-equipped workstation.
- Kimi K2.5 and K2.7
- Around 600 GB in Q4: smaller than K3, but still workstation-class infrastructure.
- 30- to 70-billion-parameter models
- On a 24–32 GB card or a 64 GB Mac: the sensible choice for real-world use. The VRAM calculator gives you the exact figure.
#Verdict: Kimi K3 locally, almost never
Native weights of 1.56 TB, quantizations from 594 to 861 GB, and 610 GB of memory for the smallest one: Kimi K3 is a server model. To evaluate it, use the API or kimi-k3:cloud. For real use at home, choose a model that fits in your memory with room for the context.
Can Kimi K3 be installed with Ollama?+
Does Kimi K3 run on a RTX 5090?+
Can a Mac mini or Mac Studio run Kimi K3?+
Can LM Studio load Kimi K3?+
Which quantization should you choose for Kimi K3?+
Why is Kimi K3 so large with only 104 billion active parameters?+
#Go further
- DeepSeek V4 Flash 284B: the first frontier model on Mac Studio
- GLM-5.2 locally: Ollama and LM Studio
- Choose your quantization (Q4, Q5, Q8, FP16)
- MoE explained: why a 30B-A3B runs like a small model
- Ollama Cloud: pricing, reviews, and limitations
- VRAM Calculator
- Source: Kimi K3 Hugging Face model card
- Source: Unsloth, Kimi K3 locally
- Source: Runpod technical FAQ
- Source: Ollama library, kimi-k3
Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.