Choose your quantization (Q4, Q5, Q8, FP16)
Choose Q4_K_M by default: at 4.89 bits per weight, it reduces the model size by more than threefold compared with FP16, with little quality loss. Move up to Q5_K_M or Q6_K if memory allows, use Q8_0 only if you have room left, and avoid anything below 4 bits without a good reason. What matters is your available memory: the file, context, and some headroom must fit.
On Hugging Face and in Ollama, the same model exists in dozens of variants: Q4_K_M, Q5_K_S, Q6_K, Q8_0, IQ3_XXS, FP16. This guide gives you measured sizes, a calculation for estimating the memory of any model, and a decision rule based on your VRAM.
#Quantization: what you trade for memory
A model is a set of billions of numbers, called weights. In FP16, each occupies 16 bits: an 8-billion-parameter model weighs about 16 GB. Quantization stores each weight in fewer bits. According to the llama-quantize tool documentation, this process reduces model size and can speed up inference, at the cost of a loss of precision measured by perplexity or Kullback-Leibler divergence. The GGUF format used by llama.cpp, Ollama and LM Studio groups these variants. The right choice depends on three quantities: your card’s or Mac’s memory, the model size, and the context you plan to use. Q4_K_M offers the best trade-off in most cases because the quality loss remains small while the size drops to about 30% of FP16.
#Decoding the names: Q4_K_M, Q5_K_S, Q8_0, IQ3_XXS
Your private ChatGPT, free, on your own machine in an hour — LM Studio, Ollama, Open WebUI, your documents, no cloud.
- Lifetime online access
- PDF + files
- Lifetime updates
- The number (Q2 to Q8)
- The nominal number of bits per weight. The actual value is higher because sensitive tensors remain at higher precision: Q4_K_M measures 4.89 bits per weight.
- _0 and _1
- Older formats with a uniform precision level. Q8_0 is still used for its near-zero loss, but Q4_0 and Q5_0 have been superseded by the K variants.
- _K
- The newer K-quants distribute bits according to tensor importance.
- S, M, L
- Small, Medium, Large: for the same number of bits, M keeps more bits in the important tensors than S, so it weighs a little more and loses a little less.
- IQ
- I-quants, using an importance matrix: at the same size, they preserve quality better at very low bitrates (below 4 bits), and require slightly more decoding compute. They rely on an imatrix file.
#Decision table: measured sizes and losses
The llama-quantize README lists, for Llama 3.1 8B, the number of bits per weight and the size of each format. The Ollama library shows the size of the distributed files for the same model. The two sources broadly agree and provide an estimate applicable to other models in the same family.
| Format | Bits per weight | llama.cpp size | Size Ollama | Perplexity loss at 7B | Typical use |
|---|---|---|---|---|---|
| Q2_K | 3,16 | 2.95 GiB | not listed | +0.87 (extreme loss) | Avoid |
| Q3_K_M | 4,00 | 3.74 GiB | not listed | +0,24 | Last resort |
| Q4_K_M | 4,89 | 4.58 GiB | 4.9 GB | +0,05 | Default choice |
| Q5_K_M | 5,70 | 5,33 GiB | 5.7 GB | +0,014 | Available VRAM headroom |
| Q6_K | 6,56 | 6.14 GiB | 6.6 GB | +0,004 | Maximum reasonable precision |
| Q8_0 | 8,50 | 7.95 GiB | 8.5 GB | +0,0004 | Near-lossless reference |
| FP16 | 16,00 | 14.96 GiB | 16 GB | 0 | Training, conversion |
One nuance about the last column: the perplexity losses come from a discussion in the llama.cpp repository that reproduces the tool's help table, created in 2023 for a 7-billion-parameter model. They show the relative ordering of the formats, not the exact behavior of current models, some of which are more sensitive.
#Calculate a model's memory in seconds
The size of a GGUF file follows this formula: parameters in billions multiplied by bits per weight, divided by eight, gives gigabytes. For a 70B in Q4_K_M: 70 × 4.89 ÷ 8 ≈ 42.8 GB, as confirmed by the llama.cpp table for Llama 3.1 70B (43.1 GB). Then add the KV cache, which grows with the context, plus a margin for the system.
| Model | Q4_K_M | Q5_K_M | Q8_0 | FP16 |
|---|---|---|---|---|
| 8B | 4,9 | 5,7 | 8,5 | 16 |
| 14B | 8,6 | 10,0 | 14,9 | 28 |
| 32B | 19,6 | 22,8 | 34,0 | 64 |
| 70B | 42,8 | 49,9 | 74,4 | 140 |
#What quantization really does to quality
Perplexity metrics rank formats consistently: more bits, less loss. But perplexity is a poor predictor of specific capabilities. A study published on arXiv in January 2026 compared llama.cpp formats on Llama-3.1-8B-Instruct, with reasoning, knowledge, and instruction-following tests: perplexity on WikiText-2 ranged from 7.32 for FP16 to 8.96 for Q3_K_S, and the authors noted that perplexity is not a complete predictor of downstream behavior. Q5_0’s average score even slightly exceeded FP16’s (69.92% versus 69.47%), which reflects measurement noise, not an improved model.
The useful result concerns mathematical reasoning: on GSM8K, the score drops from 77.63 in FP16 to 68.31 in Q3_K_S. The authors recommend avoiding aggressive 3-bit quantization and preferring Q4_K_S or Q5_0 when the workload involves multistep reasoning. The tasks penalized first by quantization are therefore computation, coding, and long chains of reasoning—not everyday conversation.
- From Q8 to Q6
- No measurable difference in practice in the published tables.
- From Q6 to Q4_K_M
- Small loss, most noticeable on long-reasoning tasks, code, and underrepresented languages.
- Under 4 bits
- Sharp drop in reasoning: avoid without a memory-related reason.
#Why a smaller file also generates faster
The llama.cpp README measurements, taken on Llama 3.1 8B, illustrate a point the guides omit: text generation reads all active weights at every token, so it depends mainly on the number of bytes to read. In this table, generation throughput rises from about 51 tokens per second in Q8_0 to about 72 in Q4_K_M, and 90 in Q2_K_S. Prompt processing, however, remains stable at around 800 tokens per second because it is compute-bound rather than memory-bound. These figures come from specific hardware in the repository and are not representative of your machine; remember the trend: fewer bits, faster generation.
Two practical consequences. First, switching from Q4_K_M to Q8_0 for an imperceptible quality gain costs about 30% of generation speed in these measurements. Second, weight quantization and KV-cache quantization are two separate settings: the latter reduces long-context memory usage and is covered in its own guide.
#Choose using three rules
- 01Start with the memory you have availableVRAM on an NVIDIA or AMD card, or a Mac's unified memory, minus about 25% for the system. If the model exceeds that, computation spills over to the CPU and speed drops, often by a factor of several times.
- 02Choose the largest model that fits in Q4_K_MWith the same amount of memory, a larger model in 4-bit is generally better than a smaller one in 8-bit. This is a widely shared rule of thumb, but verify it on your own tasks, especially for code.
- 03Then increase precision with the remaining headroomIf VRAM remains after reserving space for the context, move to Q5_K_M and then Q6_K. The gain is smaller than moving up to the next size.
The labels in the Ollama library indicate the format and size of each variant; when you omit the suffix, Ollama applies the model's default label. On Hugging Face, GGUF files include the quantization in their name.
#I-quants and sub-4-bit formats
The llama.cpp table shows the actual savings. On Llama 3.1 8B, IQ4_XS weighs 4.17 GiB versus 4.58 GiB for Q4_K_M, or about 9% less; IQ3_XXS weighs 3.04 GiB; IQ2_XXS weighs 2.23 GiB. In the repository's measurements, generation throughput for these formats remains close to that of the K-quants, within a few tokens per second: decoding overhead is therefore modest. I-quants, however, depend on the quality of the imatrix and the model: they only make sense when Q4_K_M does not fit.
- IQ4_XS
- Similar in size to Q4_K_M, minus ~9%. Useful when VRAM is tight.
- IQ3_XXS and IQ3_S
- A 13B at about 5 GB in size, with a noticeable loss in reasoning.
- IQ2 and IQ1
- Only for gigantic models, when nothing else fits. With this kind of quantization, fidelity can be measured: see the Kimi K3 example, where dynamic 1-bit achieves only 78.9% agreement with the original.
#Models published directly in 4-bit
Some recent models are no longer trained in 16-bit and then compressed: their weights are produced in 4-bit with quantization-aware training. Kimi K3 is an example: its specifications list MXFP4 weights and MXFP8 activations, with quantization-aware training. For these models, the native file is already the reference, and further quantization to a lower format degrades it more than quantizing a 16-bit model. Always read the model specifications before converting it.
#What to choose based on your memory
| Available memory | Reasonable choice |
|---|---|
| 6 to 8 GB | 7-8B in Q4_K_M; IQ4_XS if it doesn’t fit |
| 12 GB | 8B in Q6_K or Q8_0, or 14B in Q4_K_M (8.6 GB) |
| 16 GB | 14B in Q5_K_M (10 GB) with a comfortable context |
| 24 GB | 32B in Q4_K_M (19.6 GB) with a moderate context |
| 32 GB | 32B in Q5_K_M (22.8 GB) or Q6_K |
| 64 GB and more | 70B in Q4_K_M (42.8 GB) |
Which quantization should you choose: Q4, Q5, or Q8?+
What does the M in Q4_K_M mean?+
How much VRAM does a Q4_K_M model require?+
Does quantization degrade quality in French?+
Is a 14B model in Q4 better than an 8B model in Q8?+
Is Q8_0 really lossless?+
#Go further
- Q4_K_M, Q5_K_M, and Q6_K in practice
- Understanding the context window
- Quantize the KV cache: save VRAM
- GGUF, safetensors: understanding the formats
- VRAM Calculator
- Source: llama-quantize README (llama.cpp)
- Source: arXiv study on llama.cpp quantization (2026)
- Source: Ollama labels from Llama 3.1
- Source: llama.cpp discussion on quantization methods
Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.