Methodology · updated 2026
Recommender for quantization
A guide that matches your exact VRAM budget to the optimal GGUF, GPTQ, or AWQ quantization — no guesswork, no wasted memory.
Basic formula: VRAM_Go = (Paramètres_Md × Bits / 8) × 1,2. The 1.2 multiplier absorbs the KV cache and activation overhead at an 8K context.
Which quantization should I use for my machine?
Enter your memory: the tool lists the catalog models that fit, with the best possible quantization for each.
Loading catalog…
The 30-second choice
Find your available VRAM below, then see the largest model and quantization that fit with an 8K context. Common GGUF file sizes (llama.cpp/Ollama default settings, FP16 KV cache); the last column shows what remains to move to a 16K context.
| VRAM | Typical GPU | Recommended model + quantization | 16K headroom |
|---|---|---|---|
| 6 GB | RTX 3060 Laptop (6 GB), RTX 4050 Laptop | Qwen 3 4B Q5_K_M, or Qwen 3 8B Q3_K_M | Exactly—switch back to Q4_K_S |
| 8 GB | RTX 3050 8 GB, RTX 4060, RTX 5060 | Qwen 3 8B Q4_K_M, Llama 3.1 8B Q4_K_M | Fine in Q4_K_M |
| 12 GB | RTX 3060 12 GB, RTX 4070, RTX 5070 | Qwen 3 14B Q4_K_M, Qwen 2.5 Coder 14B Q4_K_M, Gemma 3 12B Q5_K_M | Just for a 14B |
| 16 GB | RTX 4060 Ti 16 GB, RTX 5060 Ti 16 GB, RTX 5070 Ti, RTX 5080 | Qwen 3 14B Q6_K, Mistral Small 3.2 24B Q4_K_M (tight) | Good for a 14B |
| 24 GB | RTX 3090, RTX 4090, RX 7900 XTX | Qwen 3 32B Q4_K_M, Qwen3-Coder 30B-A3B Q4_K_M, Gemma 3 27B Q4_K_M | Just for a dense 32B; 27B with room to spare |
| 32 GB | RTX 5090 | Qwen 3 32B Q6_K, Qwen 3 30B-A3B Q6_K | Only for the 32B in Q6_K |
| 48 GB | RTX 6000 Ada, 2× RTX 3090/4090 | Llama 3.3 70B Q4_K_M, Qwen 2.5 72B Q4_K_S | Short context (8K) only |
| 80 GB | H100 80 GB, A100 80 GB | gpt-oss 120B (MXFP4), Llama 3.3 70B Q8_0 (short context) | Comfortable for gpt-oss 120B |
The calculation behind the recommendation
Quantization compresses each FP16 weight (16 bits) down to 2–8 bits. A 7B model in Q4_K_M uses an average of ~4.85 bits/weight, or 7 000 000 000 × 4.85 / 8 ≈ 4.2 GB. Adding the KV cache (which grows with the context), this reaches around 5.5 GB at 8K context for a grouped-query attention model (Mistral 7B, Qwen, Llama 3). For a long context, the KV cache dominates by a wide margin over the ×1.2 multiplier.
Bits per weight reference
| Quant | Effective bits/weights | VRAM (8B) | VRAM (32B) | VRAM (70B) |
|---|---|---|---|---|
| FP16 | 16,0 | 16.0 GB | 64.0 GB | 140.0 GB |
| Q8_0 | 8,5 | 8.5 GB | 34.0 GB | 74.4 GB |
| Q6_K | 6,6 | 6.6 GB | 26.4 GB | 57.8 GB |
| Q5_K_M | 5,7 | 5.7 GB | 22.8 GB | 49.9 GB |
| Q4_K_M | 4,83 | 4.83 GB | 19.3 GB | 42.3 GB |
| Q4_K_S | 4,58 | 4.58 GB | 18.3 GB | 40.1 GB |
| Q3_K_M | 3,9 | 3.9 GB | 15.6 GB | 34.1 GB |
| Q2_K | 3,35 | 3.35 GB | 13.4 GB | 29.3 GB |
Quality loss: what the numbers really say
The reference measurement remains the perplexity published with llama.cpp's k-quants (PR #1684, LLaMA 7B, FP16 = 5,9066). The smaller the gap, the more the quantized model behaves like the original. The cliff between Q4_K_M and Q3_K_M is clear, while the gain from Q5 to Q6 is minimal.
| Quant | Perplexity (LLaMA 7B) | Difference vs FP16 | Verdict |
|---|---|---|---|
| Q8_0 | — | negligible | Indistinguishable |
| Q6_K | 5,9110 | +0,07 % | Nearly lossless |
| Q5_K_M | 5,9208 | +0,24 % | Excellent |
| Q4_K_M | 5,9601 | +0,91 % | Sweet spot |
| Q4_K_S | 6,0215 | +1,95 % | Acceptable |
| Q3_K_M | 6,1503 | +4,13 % | Noticeable |
| Q2_K | 6,7764 | +14,7 % | Last resort |
Measurements from 2023 on LLaMA 7B: newer models, trained on much more data, are often slightly more sensitive to quantization. The order of magnitude remains valid.
With the same memory, a larger model in Q4_K_M generally performs better than a smaller model in Q8_0. When in doubt, increase the parameter count, not the bit count.
GGUF, GPTQ, or AWQ: the right format, not just the right number of bits
GGUF (llama.cpp / Ollama / LM Studio): the default for single-user inference, CPU offloading, and Apple Silicon. GPTQ : 4-bit, pure GPU inference with vLLM/TGI. AWQ : 4-bit, GPU-only—the fastest option in 2026 for batch inference on RTX 4090, RTX 5090, H100/H200.
| Use cases | Best format | Why |
|---|---|---|
| Single-user, desktop, ChatGPT style | GGUF Q4_K_M | CPU offload, partial GPU layers, runs everywhere |
| Apple Silicon (M1-M5) | GGUF Q4_K_M or 4-bit MLX | Metal kernels, unified memory |
| Batch API, >4 simultaneous users | 4-bit AWQ on vLLM | Best throughput at high batch sizes |
| Older Ampere (A100, RTX 3090) | 4-bit GPTQ or AWQ | Both work; benchmark them on your use case |
Frequently asked questions
Is Q4_K_M really the sweet spot, or just popular?
Both. In the perplexity measurements published with llama.cpp’s k-quants (LLaMA 7B), Q4_K_M deviates from FP16 by only about 0.9% while using roughly 30% of its memory. Q5_K_M reduces the gap to 0.2% for about 17% more memory—rarely worth it unless you have room to spare.
How much VRAM does a 70B model need?
In Q4_K_M, about 42 GB for the weights, plus 2 to 3 GB of KV cache at 8K for Llama 3.3 70B (thanks to grouped attention). This fits on a 48 GB card (RTX 6000 Ada) or two 24 GB cards (2×RTX 3090/4090), with a short context. In Q2_K, the weights drop to ~26 GB, but quality declines sharply.
GGUF, GPTQ, or AWQ: which should you choose?
GGUF for single-user use on a workstation or Apple Silicon. AWQ for batch API serving on Ada, Hopper, or Blackwell GPUs. GPTQ remains a historical choice that still works on Ampere but is rarely the fastest in 2026.
Does quantizing the KV cache degrade quality?
The KV cache in Q8_0 cuts its size in half with a loss generally considered negligible. In Q4_0, the loss becomes measurable: reserve it for cases where it is the only way to fit a long context.
Are MoE models like DeepSeek V3 different?
Yes. Storage uses the total number of parameters, so a 671B MoE still needs hundreds of GB even in Q4. Speed, however, depends on the active parameters (37B for DeepSeek V3). Under Q4, quality loss increases quickly; dynamic quantizations (such as Unsloth's for DeepSeek R1) nevertheless show that a very large MoE can remain usable at 2 bits.
How can you verify that a quantization really fits in VRAM?
Run nvidia-smi -l 1 while generating 500 tokens. If VRAM usage rises and then stabilizes, you're good. If it hits a ceiling and tokens/s collapses, you're spilling into system RAM—drop one quantization level or reduce the context.
Want to go further? The VRAM calculator specifies the exact requirements for your model, and our Q4 vs Q5 vs Q8 guide details the quality tests.