GGUF quantization in 2026: Q4_K_M vs Q5_K_M vs Q6_K in pratique
The GGUF Q5_K_M vs. Q4_K_M quantization comparison has been circulating on forums for three years without clear figures. In 2026, with Qwen3, Llama 4, and DeepSeek V3.2, the question comes up again: is Q4_K_M still the reasonable default, or should you move up to Q5_K_M or even Q6_K? We measured perplexity, French quality, speed, and file size across three model families. Here's what we take away in practice.
#Why this comparison
Quantization is the art of compressing a model's weights from 16 or 32 bits per parameter down to 8, 5, 4, or even 2 bits. Fewer bits mean less VRAM and more speed, but degraded quality. llama.cpp's GGUF format offers about a dozen variants, and 90% of users default to Q4_K_M without knowing whether it's the right choice for their model or use case.
The problem: the recommendations circulating date from 2024, don’t account for importance matrices (imatrix), which have since become widespread, and confuse raw perplexity with actual quality in French. We reran the tests properly on three 2026 families: Qwen3-14B, Llama 4 Scout (17B-A2B MoE), and DeepSeek V3.2-Lite (16B). The goal: an honest, usable comparison of GGUF Q5_K_M vs. Q4_K_M quantization.
#Quick reminder: GGUF and K-quants
Your private ChatGPT, free, on your own machine in an hour — LM Studio, Ollama, Open WebUI, your documents, no cloud.
- Lifetime online access
- PDF + files
- Lifetime updates
GGUF is llama.cpp’s file format. Inside, each weight tensor can be quantized independently according to a grid of named schemes: Q2_K, Q3_K_S, Q3_K_M, Q3_K_L, Q4_K_S, Q4_K_M, Q5_K_S, Q5_K_M, Q6_K, Q8_0—and the IQ variants (i-quants) for very low precisions.
- The number
- Average number of bits per weight. Q4 ≈ 4 bits, Q5 ≈ 5 bits, Q8 ≈ 8 bits.
- _K
- K-quant scheme: weights grouped into blocks with fine-grained scaling. Much better than the older Q4_0 / Q4_1.
- _S / _M / _L
- Small / Medium / Large. The “important” tensors (attention, embed) receive higher precision in M and L. _M is the reasonable default.
- Q6_K
- No S/M/L: a single format. Very close to FP16 in quality, at 60% of the size.
- Q8_0
- Nearly lossless. Keep it for very small models (< 3B) or when VRAM is tight.
#Test protocol
We quantized each model in 7 variants (Q2_K, Q3_K_M, Q4_K_M, Q5_K_M, Q6_K, Q8_0, reference FP16) using an importance matrix (imatrix) calibrated on 4 MB of mixed FR/EN/code text—French Wikipedia, Python code, and technical fiction. The conversions were performed with llama.cpp release b5500+ on RTX 4090.
For each variant, we measured four metrics: (1) perplexity on English wikitext-103 and a French corpus of 200 articles; (2) score on 50 French prompts covering reasoning, translation, code, and summarization, blindly rated by three reviewers; (3) speed in tokens per second at 4k and 32k context; (4) file size in GB.
#Perplexity table by quantization
Perplexity (PPL) measures how “surprised” the model is by a reference text. Lower is better. We compare it as a percentage degradation versus FP16—that’s what’s readable, not the raw value.
Results on Qwen3-14B (English wikitext / French corpus perplexity, difference vs FP16):
- FP16 (reference)
- PPL EN 5.42 / FR 7.18 — baseline, 28 GB size
- Q8_0
- +0.04% EN / +0.05% FR—14.9 GB in size. Indistinguishable from FP16 on normal text.
- Q6_K
- +0.15% EN / +0.21% FR — 11.5 GB in size. Excellent, already very difficult to distinguish.
- Q5_K_M
- +0.48% EN / +0.61% FR — 9.9 GB size. The quality sweet spot.
- Q4_K_M
- +1.12% EN / +1.48% FR — size 8.4 GB. The classic default: visible but moderate degradation.
- Q3_K_M
- +3.85% EN / +5.12% FR — 6.6 GB size. First real step down in quality.
- Q2_K
- +11.4% EN / +16.7% FR — 5.3 GB size. Avoid except in a VRAM emergency.
On Llama 4 Scout (17B-A2B MoE), the losses are more pronounced at low quantization: Q4_K_M degrades by +1.9% in FR, Q3_K_M by +6.4%. MoE models handle aggressive quants less well because each expert sees only a fraction of the tokens and has less redundancy. Q5_K_M is clearly preferable on Scout.
On DeepSeek V3.2-Lite, typical behavior: Q4_K_M holds at +1.3% FR, Q5_K_M at +0.5%. Q6_K is nearly free in terms of quality.
#Quality degradation in French
This is the angle most often overlooked in English-language benchmarks. LLMs see fewer French tokens during training, so their French representations are more fragile under compression. The empirical rule we've confirmed: French degradation is ~30 to 40% more pronounced than English degradation at the same quantization level.
Qualitative French scores on 50 prompts (average score out of 10, three blind judges), Qwen3-14B:
- FP16
- 8.4 / 10—the benchmark
- Q8_0
- 8.4 / 10 — strictly identical in practice
- Q6_K
- 8.3 / 10 — imperceptible difference
- Q5_K_M
- 8.1 / 10 — a few turns of phrase are slightly less elegant, substance unchanged
- Q4_K_M
- 7.7 / 10 — correct answers but noticeable simplifications on complex questions
- Q3_K_M
- 6.9 / 10 — hallucinations that appear, impoverished French vocabulary
- Q2_K
- 5.2 / 10 — agreement errors, mistranslations, and occasional unintended shifts into English
#Speed: speed / quality tradeoff
On RTX 4090 with full GPU offload, Qwen3-14B context 4k:
- Q4_K_M
- 78 tok/s — the fastest viable K-quant
- Q5_K_M
- 65 tok/s (-17%) — a real but not dramatic penalty
- Q6_K
- 54 tok/s (-31%) — you start to notice it
- Q8_0
- 42 tok/s (-46%) — VRAM usage barely changes
- FP16
- 26 tok/s (-67%)—not competitive on this card
On a Mac M4 Pro with 48 GB of unified memory, the profile changes: Q4_K_M and Q5_K_M are nearly tied (32 vs 30 tok/s) because the bottleneck is memory bandwidth, not compute. On Mac, the “penalty” for Q5_K_M is negligible—might as well use it.
On RTX 3060 12 GB, VRAM becomes the dominant factor: Q5_K_M for Qwen3-14B just fits with 8k context, while Q4_K_M leaves room for 16k. Here, the choice is dictated by the context you want, not by quality.
#imatrix vs static: the difference is real
“Static” quantization (without imatrix) assigns precision uniformly. Quantization with imatrix uses a calibration corpus to identify important weights and preserve more precision for them. This has become the standard at Bartowski, mradermacher, and most serious releases on Hugging Face in 2026.
Measured gain on Qwen3-14B Q4_K_M, FR:
- Static (without imatrix)
- French PPL +2.1% vs FP16, qualitative score 7.4/10
- imatrix EN-only
- French PPL +1.6%, score 7.6/10—better but suboptimal for French
- Mixed FR/EN/code imatrix
- PPL FR +1.48%, score 7.7/10—the best practice
The gap between imatrix and static is larger at lower quants (Q2, Q3) than at higher ones (Q6, Q8). With Q4_K_M, imatrix gives you roughly 0.5–1% better quality; with Q3_K_M, it’s 2–3%. With Q2_K, it’s the only way to get a usable result.
#Recommendation by use case
There is no universal answer. Here are the choices we stand behind, ranked by profile:
- Comfortable VRAM headroom (model uses < 60% of the card)
- Choose Q6_K. The difference from FP16 is imperceptible, and the file is 40% smaller. There is no reason to go lower.
- Just enough VRAM for the model
- Q4_K_M remains the reasonable default. Save VRAM for the KV cache and long context.
- Demanding French-language use (writing, legal, medical)
- Move up to Q5_K_M at a minimum. The Q4 penalty for French is noticeable. If possible, use Q6_K.
- MoE model (Llama 4, DeepSeek, Qwen3-A3B)
- Q5_K_M rather than Q4_K_M. MoE models suffer more from aggressive quantization.
- Small model (1B–3B) for edge / CPU
- Always use Q8_0. Small models have little redundancy, so losing bits hurts them badly.
- Code model (Qwen3-Coder, Devstral)
- Q5_K_M or Q6_K. The code is intolerant of subtle hallucinations; Q4’s VRAM savings are not worth it.
- Very limited VRAM (8 GB), large model desired
- IQ3_M or IQ3_XS with imatrix. Better than Q3_K_M at the same size.
- No GPU, CPU only
- Q4_K_M. IQ quants slow down on pure CPU, Q5 uses too much RAM, and Q4_K_M is the best throughput/quality compromise.
#Common pitfalls
- Compare perplexity across model families
- Not useful. PPL is comparable only with the model held constant, on the same corpus. Compare Qwen3 Q4 with Qwen3 Q5, not Qwen3 Q4 with Llama 4 Q4.
- Q4_0 or Q4_1
- These old non-K schemes have no reason to exist anymore. If you come across a recent GGUF Q4_0, it is probably a lazy upload—move on.
- KV cache quantization
- That is a separate topic (the --cache-type-k q8_0 parameter in llama.cpp). It is often confused with model quantization. Doing both at once cuts VRAM usage but compounds the quality loss.
- Critical tensors at low precision
- Some tools let you keep the embed and output layers at Q8 even in a Q4 GGUF. That's what Q4_K_M does by default. If you see Q4_K_S labeled “pure Q4” by an uploader, be careful: the actual quality is lower.
- Believing that Q5_K_M uses 25% more VRAM than Q4_K_M
- False. The actual file-size difference is ~18%. And once in VRAM, the KV cache (context) takes up the same amount of space in both cases.
#Go further
This comparison assumes you already know how to work with GGUFs and run llama.cpp. If anything remains unclear, here are the related guides:
- Choose your quantization (Q4, Q5, Q8, FP16)
- Our introductory guide if you want the fundamentals before this advanced comparison.
- TurboQuant: quantizing large models
- The next step toward fitting a frontier model (>100B) on consumer hardware.
- Compile llama.cpp with CUDA
- Essential if you want to quantize freshly published weights yourself with imatrix.
Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.