Intermediate 11 minNews

Kimi K3 at home: the truth about the hardware required (2.8T parameters)

Direct response

Kimi K3 does not run on a desktop workstation: its native weights weigh 1.56 TB, and the smallest published GGUF quantization (Unsloth, dynamic 1-bit) still uses 594 GB, with 610 GB of memory recommended. No 512 GB Mac Studio and no single RTX 5090 are sufficient. To try it, use Moonshot’s API or kimi-k3:cloud on Ollama; locally, target a smaller model.

The weights for Kimi K3 have been available since late July 2026, and searches for “ollama kimi k3,” “kimi k3 lmstudio,” and “kimi k3 on 5090” show that many readers are wondering whether they can install it at home. This guide gives the figures published by Moonshot and Unsloth, calculates what your machine can load, and explains what to do instead.

By Mohamed Meguedmi·Update 2026-09-30·Tested on Windows, macOS, and Linux

#Kimi K3 locally: what Moonshot actually released

Moonshot's Hugging Face model card describes Kimi K3 as an open-weight multimodal model with 2.8 trillion parameters and a one-million-token context window. It is a Mixture-of-Experts model: 896 experts, 16 of which are selected for each token, for 104 billion active parameters. The weights are distributed in MXFP4 (weights) with MXFP8 activations, using quantization-aware training from the supervised fine-tuning phase onward. The code and weights are covered by the “Kimi K3 License”: read this text before any commercial use, because it is not a standard MIT or Apache license. The model appeared on Hugging Face around July 27, 2026, according to the specialist press.

Total / active parameters
2.8T total, 104B active per token (Moonshot specification sheet).
Experts
896 experts, 16 selected per token, plus 2 shared experts.
Published format
MXFP4 for weights, MXFP8 for activations. This is not a GGUF.
Context
1,048,576 tokens, with a KV cache added to the weights' memory footprint.
Recommended engines
vLLM, SGLang, and TokenSpeed. llama.cpp uses community GGUFs.
i
Active does not mean resident
Only 104 billion parameters work on each token, but the router can choose any expert, so all the weights must remain accessible. Memory, not computing power, dictates the hardware. The principle is explained in detail in our guide to MoE.

#What the weights really weigh

The Local AI Kit

Your private ChatGPT, free, on your own machine in an hour — LM Studio, Ollama, Open WebUI, your documents, no cloud.

  • Lifetime online access
  • PDF + files
  • Lifetime updates

Two figures are circulating, and they need to be distinguished. Some articles estimate “around 1.4 TB” by multiplying 2.8T parameters by half a byte. The file size is higher: Unsloth and Runpod give 1.56 TB for the native version because the layers outside the experts (attention, routers, shared experts) remain at higher precision. The unquantized BF16 version would reach approximately 5.6 TB, according to Runpod. The version of this guide that stated 1.4 TB was therefore about 10% too low.

Published sizes for Kimi K3 (sources: Unsloth, Runpod)
FormatSizeRecommended total memoryMeasured fidelity
Native MXFP4 / UD-Q8_K_XL1.56 TB1.6 TBLossless
Unsloth UD-Q2_K_XL (dynamic 2-bit)861.3 GB880 GBAbout 90% top-1 agreement
Unsloth UD-IQ2_XXS711.1 GB726 GB84.1% top-1 agreement
Unsloth UD-IQ1_M648.9 GB665 GB81.2% top-1 agreement
Unsloth UD-IQ1_S (dynamic 1-bit)594 GB610 GB78.9% top-1 agreement
BF16 (theoretical)approximately 5.6 TBOut of reachReference

The table disproves a misconception from the older version: a Q2 is not “always on the order of a terabyte.” Unsloth does offer GGUFs under 900 GB. But 594 GB is still almost fifteen times the memory of a 32 GB RTX 5090, and the quality gap is real: the dynamic 1-bit version reproduces the original model’s choice in only 78.9% of cases on Unsloth’s benchmark. To understand how bits affect quality, see our quantization guide.

#Hardware: what's recommended and what's possible

Moonshot does not publish a minimum configuration on the model page: it points to the vLLM, SGLang, and TokenSpeed recipes and to its API. The documented reference for native self-hosting comes from Runpod: a node with eight B300 GPUs of 288 GB each, or about 2.3 TB, or sixteen B200s spread across two nodes. The figure “64 accelerators or more” that this page previously repeated does not appear in any primary source consulted; it has been removed.

With an Unsloth GGUF, the rule given in their documentation is simple: total RAM plus VRAM should roughly equal the quantization size; otherwise, the model works but spills to disk, which is much slower. Unsloth also cites about 20 tokens per second during generation when the model fits on B200s. That is a vendor throughput figure on datacenter hardware, not a measurement applicable to a home machine.

#What about a Mac? Compute over hearsay

The figure of “16 seconds per token on an M1 Max MacBook Pro” reported on this page could not be traced to any verifiable source, so we do not present it as fact. What the sources establish is clearer. Kingy AI notes that the starting point (a 1.56 TB checkpoint) already exceeded a 512 GB Mac Studio, and that even the 553.2 GiB file for the 1-bit version exceeds this machine before overhead. A 128 GB Mac has about one-fifth of the recommended 610 GB.

You can estimate the order of magnitude yourself, though. If memory is insufficient, each token rereads the experts it activates from the SSD: 104 billion parameters at about 4 bits, or roughly 50 GB read per token. With an SSD delivering 3 GB/s, you get about 17 seconds per token; at 7 GB/s, about 7 seconds. This is a theoretical ceiling calculation, not a measurement, but it explains why reports of running “from disk” are measured in seconds per token, not tokens per second.

!
What “it runs” means
At 10 seconds per token, a 300-token response takes nearly fifty minutes, and K3’s reasoning mode, which is always active, adds even more reasoning tokens before the response. Technically demanding, practically unusable.

#Ollama, LM Studio, llama.cpp: what each tool does

Ollama
The official library lists only one variant, kimi-k3:cloud: the model runs on Ollama's servers, not on your machine. The ollama run kimi-k3:cloud command lets you try the model with the usual interface, subject to the privacy and billing limits of a remote service.
LM Studio
It loads local GGUF files. For K3, this means downloading several hundred GB of Unsloth files and having the corresponding memory available. On a consumer machine, the answer is no.
llama.cpp
This is the engine targeted by Unsloth GGUFs, which rely on a llama.cpp-derived branch with vision support. It handles offloading experts to the CPU and multipart files: the realistic path for a workstation with several hundred GB of RAM.
vLLM and SGLang
The two engines recommended by Moonshot for a multi-GPU service using the native format.
Your machine versus Kimi K3
MachineAvailable memoryVerdict
RTX 5090 (32 GB) + 64 GB of RAMapproximately 96 GBImpossible: 6 times too little, even at 1 bit
Mac mini or MacBook, 16 to 64 GB16 to 64 GBImpossible locally; API or cloud only
Mac Studio 128 to 256 GB128 to 256 GBInsufficient for the recommended 610 GB
Mac Studio 512 GB512 GBBelow the 1-bit file size, before context
768 GB to 1 TB RAM workstation + GPU600 to 900 GBFeasible in 1- to 2-bit GGUF, modest throughput
8 B300 GPUs (about 2.3 TB)2.3 TBDocumented configuration for the native format

#Realistic options for trying Kimi K3

  1. 01
    1. Moonshot's official API
    The model is called kimi-k3 on platform.kimi.ai, with an API compatible with OpenAI and Anthropic. It is the fastest way to judge quality, with a reasoning_effort parameter adjustable to low, high, or max.
  2. 02
    2. Ollama cloud
    ollama run kimi-k3:cloud uses your usual Ollama workflow with remote inference. The library page lists prices per million tokens: check them before automating. Our Ollama Cloud guide details the limits.
  3. 03
    3. Renting a GPU
    Rent a multi-GPU node by the hour for a test, using vLLM. First calculate the break-even point (hourly cost divided by sustained tokens per second), as Runpod explains in its FAQ.
  4. 04
    4. A high-RAM workstation
    Only if you already have 700 GB of memory: GGUF UD-IQ1_S and llama.cpp, accepting reduced quality and low throughput.

#Open weights does not mean you can run it at home

Open weights guarantee the right to download, audit, fine-tune, and, depending on the license, redistribute them. They do not guarantee that anyone can run them. A 30-billion-parameter model on a 24 GB card genuinely belongs to you; a 2.8T model that only a cluster can load remains, in practice, a service. Kimi K3 is good news for auditability and for organizations that already have a cluster, without changing what an individual can do at home.

#Alternatives your hardware can actually load

The QuelLLM catalog estimates memory requirements in Q4, excluding context: DeepSeek V4 Flash 284B about 170 GB, GLM 5.2 753B-A40B about 437 GB, Kimi K3 about 1,624 GB. For everyday use, a 30- to 70-billion-parameter model in Q4 remains the best balance between quality and feasibility.

DeepSeek V4 Flash 284B
Approximately 170 GB in Q4 according to the catalog: possible on a high-memory Mac Studio or a workstation. See the dedicated guide.
GLM 5.2 753B-A40B
About 437 GB in Q4: it’s the closest form factor to K3 accessible to a well-equipped workstation.
Kimi K2.5 and K2.7
Around 600 GB in Q4: smaller than K3, but still workstation-class infrastructure.
30- to 70-billion-parameter models
On a 24–32 GB card or a 64 GB Mac: the sensible choice for real-world use. The VRAM calculator gives you the exact figure.

#Verdict: Kimi K3 locally, almost never

Native weights of 1.56 TB, quantizations from 594 to 861 GB, and 610 GB of memory for the smallest one: Kimi K3 is a server model. To evaluate it, use the API or kimi-k3:cloud. For real use at home, choose a model that fits in your memory with room for the context.

FAQ
Can Kimi K3 be installed with Ollama?+
Not locally. The Ollama library only offers the kimi-k3:cloud variant, which runs on remote servers rather than on your machine. To load a model locally for real, you would need a GGUF several hundred GB in size through llama.cpp, which rules out a typical workstation, even a very well-equipped one.
Does Kimi K3 run on a RTX 5090?+
No. A RTX 5090 offers 32 GB of VRAM, while the smallest Unsloth GGUF weighs 594 GB and requires 610 GB of total memory. Even with 128 GB of system RAM, the machine would still fall far short, and offloading to disk would make generation unusable.
Can a Mac mini or Mac Studio run Kimi K3?+
Not in practice. A 512 GB Mac Studio is smaller than the 1-bit file size (553.2 GiB), even before accounting for context, and a Mac mini falls even further short. On these machines, choose a smaller model, or use the model's API or Ollama cloud.
Can LM Studio load Kimi K3?+
In theory, yes, because LM Studio reads GGUF files and Unsloth publishes them. In practice, you need the corresponding memory: at least 610 GB for the 1-bit version. On a consumer machine, the application cannot load the model; it is better to target a more modest model.
Which quantization should you choose for Kimi K3?+
Unsloth recommends UD-IQ1_S (594 GB) as a balance between size and quality: 78.9% top-1 agreement with the original according to its measurements. The 2-bit UD-Q2_K_XL, at 861 GB, reaches about 90%. In both cases, this already requires a workstation with more than 600 GB of memory.
Why is Kimi K3 so large with only 104 billion active parameters?+
Because the router can activate any of the 896 experts: only 16 compute at each token, but all must remain available in memory. The compute cost per token is moderate; memory capacity, not processing power, is what requires server-grade hardware.

#Go further

Did this guide help you?

Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.