Intermediate 14 minQuantization

GGUF quantization in 2026: Q4_K_M vs Q5_K_M vs Q6_K in pratique

The GGUF Q5_K_M vs. Q4_K_M quantization comparison has been circulating on forums for three years without clear figures. In 2026, with Qwen3, Llama 4, and DeepSeek V3.2, the question comes up again: is Q4_K_M still the reasonable default, or should you move up to Q5_K_M or even Q6_K? We measured perplexity, French quality, speed, and file size across three model families. Here's what we take away in practice.

By Marie L.·Update 2026-06-11·Tested on Windows, macOS, and Linux

#Why this comparison

Quantization is the art of compressing a model's weights from 16 or 32 bits per parameter down to 8, 5, 4, or even 2 bits. Fewer bits mean less VRAM and more speed, but degraded quality. llama.cpp's GGUF format offers about a dozen variants, and 90% of users default to Q4_K_M without knowing whether it's the right choice for their model or use case.

The problem: the recommendations circulating date from 2024, don’t account for importance matrices (imatrix), which have since become widespread, and confuse raw perplexity with actual quality in French. We reran the tests properly on three 2026 families: Qwen3-14B, Llama 4 Scout (17B-A2B MoE), and DeepSeek V3.2-Lite (16B). The goal: an honest, usable comparison of GGUF Q5_K_M vs. Q4_K_M quantization.

i
Who this guide is for
You already know what quantization is (if not, see our introductory Q4/Q5/Q8 guide). You use llama.cpp directly, or Ollama / LM Studio in the background. You want to optimize a local setup rather than use the default suggested quantization.

#Quick reminder: GGUF and K-quants

The Local AI Kit

Your private ChatGPT, free, on your own machine in an hour — LM Studio, Ollama, Open WebUI, your documents, no cloud.

  • Lifetime online access
  • PDF + files
  • Lifetime updates

GGUF is llama.cpp’s file format. Inside, each weight tensor can be quantized independently according to a grid of named schemes: Q2_K, Q3_K_S, Q3_K_M, Q3_K_L, Q4_K_S, Q4_K_M, Q5_K_S, Q5_K_M, Q6_K, Q8_0—and the IQ variants (i-quants) for very low precisions.

The number
Average number of bits per weight. Q4 ≈ 4 bits, Q5 ≈ 5 bits, Q8 ≈ 8 bits.
_K
K-quant scheme: weights grouped into blocks with fine-grained scaling. Much better than the older Q4_0 / Q4_1.
_S / _M / _L
Small / Medium / Large. The “important” tensors (attention, embed) receive higher precision in M and L. _M is the reasonable default.
Q6_K
No S/M/L: a single format. Very close to FP16 in quality, at 60% of the size.
Q8_0
Nearly lossless. Keep it for very small models (< 3B) or when VRAM is tight.
→
What about IQ quants?
IQ2_XS, IQ3_S, IQ4_XS… use codebooks (vector quantization). Better quality than K-quants at the same bit depth, but slower for pure CPU inference and slower for RAM/SSD offload. For full GPU offload, they’re often worth it. We’ll look at them below.

#Test protocol

We quantized each model in 7 variants (Q2_K, Q3_K_M, Q4_K_M, Q5_K_M, Q6_K, Q8_0, reference FP16) using an importance matrix (imatrix) calibrated on 4 MB of mixed FR/EN/code text—French Wikipedia, Python code, and technical fiction. The conversions were performed with llama.cpp release b5500+ on RTX 4090.

Example: generating a GGUF Q5_K_M with imatrix
# 1. Calibration imatrix sur un corpus mixte
./llama-imatrix \
  -m qwen3-14b-f16.gguf \
  -f calibration_mixte_fr_en_code.txt \
  -o qwen3-14b.imatrix

# 2. Conversion avec imatrix
./llama-quantize \
  --imatrix qwen3-14b.imatrix \
  qwen3-14b-f16.gguf \
  qwen3-14b-Q5_K_M.gguf \
  Q5_K_M

For each variant, we measured four metrics: (1) perplexity on English wikitext-103 and a French corpus of 200 articles; (2) score on 50 French prompts covering reasoning, translation, code, and summarization, blindly rated by three reviewers; (3) speed in tokens per second at 4k and 32k context; (4) file size in GB.

#Perplexity table by quantization

Perplexity (PPL) measures how “surprised” the model is by a reference text. Lower is better. We compare it as a percentage degradation versus FP16—that’s what’s readable, not the raw value.

Results on Qwen3-14B (English wikitext / French corpus perplexity, difference vs FP16):

FP16 (reference)
PPL EN 5.42 / FR 7.18 — baseline, 28 GB size
Q8_0
+0.04% EN / +0.05% FR—14.9 GB in size. Indistinguishable from FP16 on normal text.
Q6_K
+0.15% EN / +0.21% FR — 11.5 GB in size. Excellent, already very difficult to distinguish.
Q5_K_M
+0.48% EN / +0.61% FR — 9.9 GB size. The quality sweet spot.
Q4_K_M
+1.12% EN / +1.48% FR — size 8.4 GB. The classic default: visible but moderate degradation.
Q3_K_M
+3.85% EN / +5.12% FR — 6.6 GB size. First real step down in quality.
Q2_K
+11.4% EN / +16.7% FR — 5.3 GB size. Avoid except in a VRAM emergency.

On Llama 4 Scout (17B-A2B MoE), the losses are more pronounced at low quantization: Q4_K_M degrades by +1.9% in FR, Q3_K_M by +6.4%. MoE models handle aggressive quants less well because each expert sees only a fraction of the tokens and has less redundancy. Q5_K_M is clearly preferable on Scout.

On DeepSeek V3.2-Lite, typical behavior: Q4_K_M holds at +1.3% FR, Q5_K_M at +0.5%. Q6_K is nearly free in terms of quality.

!
Perplexity doesn't tell the whole story
A model can have slightly degraded PPL and drop an entire level in reasoning. Conversely, identical PPL can conceal subtle hallucinations. Always cross-check with qualitative tests using your use case.

#Quality degradation in French

This is the angle most often overlooked in English-language benchmarks. LLMs see fewer French tokens during training, so their French representations are more fragile under compression. The empirical rule we've confirmed: French degradation is ~30 to 40% more pronounced than English degradation at the same quantization level.

Qualitative French scores on 50 prompts (average score out of 10, three blind judges), Qwen3-14B:

FP16
8.4 / 10—the benchmark
Q8_0
8.4 / 10 — strictly identical in practice
Q6_K
8.3 / 10 — imperceptible difference
Q5_K_M
8.1 / 10 — a few turns of phrase are slightly less elegant, substance unchanged
Q4_K_M
7.7 / 10 — correct answers but noticeable simplifications on complex questions
Q3_K_M
6.9 / 10 — hallucinations that appear, impoverished French vocabulary
Q2_K
5.2 / 10 — agreement errors, mistranslations, and occasional unintended shifts into English
i
The psychological threshold
Between Q4_K_M and Q5_K_M, the difference in French quality is around 4–5%. In professional or legal writing, you can see it. In technical French-English chat, it is inaudible. The right approach is to test with 10 prompts representative of your use case before finalizing the choice.

#Speed: speed / quality tradeoff

On RTX 4090 with full GPU offload, Qwen3-14B context 4k:

Q4_K_M
78 tok/s — the fastest viable K-quant
Q5_K_M
65 tok/s (-17%) — a real but not dramatic penalty
Q6_K
54 tok/s (-31%) — you start to notice it
Q8_0
42 tok/s (-46%) — VRAM usage barely changes
FP16
26 tok/s (-67%)—not competitive on this card

On a Mac M4 Pro with 48 GB of unified memory, the profile changes: Q4_K_M and Q5_K_M are nearly tied (32 vs 30 tok/s) because the bottleneck is memory bandwidth, not compute. On Mac, the “penalty” for Q5_K_M is negligible—might as well use it.

On RTX 3060 12 GB, VRAM becomes the dominant factor: Q5_K_M for Qwen3-14B just fits with 8k context, while Q4_K_M leaves room for 16k. Here, the choice is dictated by the context you want, not by quality.

#imatrix vs static: the difference is real

“Static” quantization (without imatrix) assigns precision uniformly. Quantization with imatrix uses a calibration corpus to identify important weights and preserve more precision for them. This has become the standard at Bartowski, mradermacher, and most serious releases on Hugging Face in 2026.

Measured gain on Qwen3-14B Q4_K_M, FR:

Static (without imatrix)
French PPL +2.1% vs FP16, qualitative score 7.4/10
imatrix EN-only
French PPL +1.6%, score 7.6/10—better but suboptimal for French
Mixed FR/EN/code imatrix
PPL FR +1.48%, score 7.7/10—the best practice
→
How to tell whether a GGUF is imatrix
The filename often contains imat or i1 (from mradermacher). On the HF page, the model card mentions the calibration. For French-language use, prefer GGUFs whose imatrix includes French—otherwise, recreate it yourself in 10 minutes with llama-imatrix.

The gap between imatrix and static is larger at lower quants (Q2, Q3) than at higher ones (Q6, Q8). With Q4_K_M, imatrix gives you roughly 0.5–1% better quality; with Q3_K_M, it’s 2–3%. With Q2_K, it’s the only way to get a usable result.

#Recommendation by use case

There is no universal answer. Here are the choices we stand behind, ranked by profile:

Comfortable VRAM headroom (model uses < 60% of the card)
Choose Q6_K. The difference from FP16 is imperceptible, and the file is 40% smaller. There is no reason to go lower.
Just enough VRAM for the model
Q4_K_M remains the reasonable default. Save VRAM for the KV cache and long context.
Demanding French-language use (writing, legal, medical)
Move up to Q5_K_M at a minimum. The Q4 penalty for French is noticeable. If possible, use Q6_K.
MoE model (Llama 4, DeepSeek, Qwen3-A3B)
Q5_K_M rather than Q4_K_M. MoE models suffer more from aggressive quantization.
Small model (1B–3B) for edge / CPU
Always use Q8_0. Small models have little redundancy, so losing bits hurts them badly.
Code model (Qwen3-Coder, Devstral)
Q5_K_M or Q6_K. The code is intolerant of subtle hallucinations; Q4’s VRAM savings are not worth it.
Very limited VRAM (8 GB), large model desired
IQ3_M or IQ3_XS with imatrix. Better than Q3_K_M at the same size.
No GPU, CPU only
Q4_K_M. IQ quants slow down on pure CPU, Q5 uses too much RAM, and Q4_K_M is the best throughput/quality compromise.

#Common pitfalls

Compare perplexity across model families
Not useful. PPL is comparable only with the model held constant, on the same corpus. Compare Qwen3 Q4 with Qwen3 Q5, not Qwen3 Q4 with Llama 4 Q4.
Q4_0 or Q4_1
These old non-K schemes have no reason to exist anymore. If you come across a recent GGUF Q4_0, it is probably a lazy upload—move on.
KV cache quantization
That is a separate topic (the --cache-type-k q8_0 parameter in llama.cpp). It is often confused with model quantization. Doing both at once cuts VRAM usage but compounds the quality loss.
Critical tensors at low precision
Some tools let you keep the embed and output layers at Q8 even in a Q4 GGUF. That's what Q4_K_M does by default. If you see Q4_K_S labeled “pure Q4” by an uploader, be careful: the actual quality is lower.
Believing that Q5_K_M uses 25% more VRAM than Q4_K_M
False. The actual file-size difference is ~18%. And once in VRAM, the KV cache (context) takes up the same amount of space in both cases.
!
Verify the SHA256
GGUF files circulate on Hugging Face and are sometimes repackaged elsewhere. A modified file (intentionally or not) can produce subtly biased outputs without any visible crash. Always verify the hash against the one published by the original uploader.

#Go further

This comparison assumes you already know how to work with GGUFs and run llama.cpp. If anything remains unclear, here are the related guides:

Choose your quantization (Q4, Q5, Q8, FP16)
Our introductory guide if you want the fundamentals before this advanced comparison.
TurboQuant: quantizing large models
The next step toward fitting a frontier model (>100B) on consumer hardware.
Compile llama.cpp with CUDA
Essential if you want to quantize freshly published weights yourself with imatrix.
Did this guide help you?

Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.