◆ Local AI — Your private ChatGPT, free, on your own machine, in an hour · $24 · or all kits $49 →

Methodology · updated 2026

Recommender for quantization

A guide that matches your exact VRAM budget to the optimal GGUF, GPTQ, or AWQ quantization — no guesswork, no wasted memory.

Basic formula: VRAM_Go = (Paramètres_Md × Bits / 8) × 1,2. The 1.2 multiplier absorbs the KV cache and activation overhead at an 8K context.

Which quantization should I use for my machine?

Enter your memory: the tool lists the catalog models that fit, with the best possible quantization for each.

Loading catalog…

The 30-second choice

Find your available VRAM below, then see the largest model and quantization that fit with an 8K context. Common GGUF file sizes (llama.cpp/Ollama default settings, FP16 KV cache); the last column shows what remains to move to a 16K context.

VRAMTypical GPURecommended model + quantization16K headroom
6 GBRTX 3060 Laptop (6 GB), RTX 4050 LaptopQwen 3 4B Q5_K_M, or Qwen 3 8B Q3_K_MExactly—switch back to Q4_K_S
8 GBRTX 3050 8 GB, RTX 4060, RTX 5060Qwen 3 8B Q4_K_M, Llama 3.1 8B Q4_K_MFine in Q4_K_M
12 GBRTX 3060 12 GB, RTX 4070, RTX 5070Qwen 3 14B Q4_K_M, Qwen 2.5 Coder 14B Q4_K_M, Gemma 3 12B Q5_K_MJust for a 14B
16 GBRTX 4060 Ti 16 GB, RTX 5060 Ti 16 GB, RTX 5070 Ti, RTX 5080Qwen 3 14B Q6_K, Mistral Small 3.2 24B Q4_K_M (tight)Good for a 14B
24 GBRTX 3090, RTX 4090, RX 7900 XTXQwen 3 32B Q4_K_M, Qwen3-Coder 30B-A3B Q4_K_M, Gemma 3 27B Q4_K_MJust for a dense 32B; 27B with room to spare
32 GBRTX 5090Qwen 3 32B Q6_K, Qwen 3 30B-A3B Q6_KOnly for the 32B in Q6_K
48 GBRTX 6000 Ada, 2× RTX 3090/4090Llama 3.3 70B Q4_K_M, Qwen 2.5 72B Q4_K_SShort context (8K) only
80 GBH100 80 GB, A100 80 GBgpt-oss 120B (MXFP4), Llama 3.3 70B Q8_0 (short context)Comfortable for gpt-oss 120B

The calculation behind the recommendation

Quantization compresses each FP16 weight (16 bits) down to 2–8 bits. A 7B model in Q4_K_M uses an average of ~4.85 bits/weight, or 7 000 000 000 × 4.85 / 8 ≈ 4.2 GB. Adding the KV cache (which grows with the context), this reaches around 5.5 GB at 8K context for a grouped-query attention model (Mistral 7B, Qwen, Llama 3). For a long context, the KV cache dominates by a wide margin over the ×1.2 multiplier.

Bits per weight reference

QuantEffective bits/weightsVRAM (8B)VRAM (32B)VRAM (70B)
FP1616,016.0 GB64.0 GB140.0 GB
Q8_08,58.5 GB34.0 GB74.4 GB
Q6_K6,66.6 GB26.4 GB57.8 GB
Q5_K_M5,75.7 GB22.8 GB49.9 GB
Q4_K_M4,834.83 GB19.3 GB42.3 GB
Q4_K_S4,584.58 GB18.3 GB40.1 GB
Q3_K_M3,93.9 GB15.6 GB34.1 GB
Q2_K3,353.35 GB13.4 GB29.3 GB

Quality loss: what the numbers really say

The reference measurement remains the perplexity published with llama.cpp's k-quants (PR #1684, LLaMA 7B, FP16 = 5,9066). The smaller the gap, the more the quantized model behaves like the original. The cliff between Q4_K_M and Q3_K_M is clear, while the gain from Q5 to Q6 is minimal.

QuantPerplexity (LLaMA 7B)Difference vs FP16Verdict
Q8_0—negligibleIndistinguishable
Q6_K5,9110+0,07 %Nearly lossless
Q5_K_M5,9208+0,24 %Excellent
Q4_K_M5,9601+0,91 %Sweet spot
Q4_K_S6,0215+1,95 %Acceptable
Q3_K_M6,1503+4,13 %Noticeable
Q2_K6,7764+14,7 %Last resort

Measurements from 2023 on LLaMA 7B: newer models, trained on much more data, are often slightly more sensitive to quantization. The order of magnitude remains valid.

With the same memory, a larger model in Q4_K_M generally performs better than a smaller model in Q8_0. When in doubt, increase the parameter count, not the bit count.

GGUF, GPTQ, or AWQ: the right format, not just the right number of bits

GGUF (llama.cpp / Ollama / LM Studio): the default for single-user inference, CPU offloading, and Apple Silicon. GPTQ : 4-bit, pure GPU inference with vLLM/TGI. AWQ : 4-bit, GPU-only—the fastest option in 2026 for batch inference on RTX 4090, RTX 5090, H100/H200.

Use casesBest formatWhy
Single-user, desktop, ChatGPT styleGGUF Q4_K_MCPU offload, partial GPU layers, runs everywhere
Apple Silicon (M1-M5)GGUF Q4_K_M or 4-bit MLXMetal kernels, unified memory
Batch API, >4 simultaneous users4-bit AWQ on vLLMBest throughput at high batch sizes
Older Ampere (A100, RTX 3090)4-bit GPTQ or AWQBoth work; benchmark them on your use case

Frequently asked questions

Is Q4_K_M really the sweet spot, or just popular?

Both. In the perplexity measurements published with llama.cpp’s k-quants (LLaMA 7B), Q4_K_M deviates from FP16 by only about 0.9% while using roughly 30% of its memory. Q5_K_M reduces the gap to 0.2% for about 17% more memory—rarely worth it unless you have room to spare.

How much VRAM does a 70B model need?

In Q4_K_M, about 42 GB for the weights, plus 2 to 3 GB of KV cache at 8K for Llama 3.3 70B (thanks to grouped attention). This fits on a 48 GB card (RTX 6000 Ada) or two 24 GB cards (2×RTX 3090/4090), with a short context. In Q2_K, the weights drop to ~26 GB, but quality declines sharply.

GGUF, GPTQ, or AWQ: which should you choose?

GGUF for single-user use on a workstation or Apple Silicon. AWQ for batch API serving on Ada, Hopper, or Blackwell GPUs. GPTQ remains a historical choice that still works on Ampere but is rarely the fastest in 2026.

Does quantizing the KV cache degrade quality?

The KV cache in Q8_0 cuts its size in half with a loss generally considered negligible. In Q4_0, the loss becomes measurable: reserve it for cases where it is the only way to fit a long context.

Are MoE models like DeepSeek V3 different?

Yes. Storage uses the total number of parameters, so a 671B MoE still needs hundreds of GB even in Q4. Speed, however, depends on the active parameters (37B for DeepSeek V3). Under Q4, quality loss increases quickly; dynamic quantizations (such as Unsloth's for DeepSeek R1) nevertheless show that a very large MoE can remain usable at 2 bits.

How can you verify that a quantization really fits in VRAM?

Run nvidia-smi -l 1 while generating 500 tokens. If VRAM usage rises and then stabilizes, you're good. If it hits a ceiling and tokens/s collapses, you're spilling into system RAM—drop one quantization level or reduce the context.

Want to go further? The VRAM calculator specifies the exact requirements for your model, and our Q4 vs Q5 vs Q8 guide details the quality tests.

More free tools