Advanced 11 minOptimization

Quantize the KV cache: save VRAM (long contexte)

You have enough VRAM to load the model, but as soon as you increase the context to 16k or 32k tokens, it overflows. The culprit is the KV cache: hidden memory that grows linearly with the context and, on long prompts, can weigh as much as the model itself. KV cache quantization compresses it to Q8 or Q4, doubling the usable context on the same card. Here’s how to enable it in Ollama and llama.cpp, along with the measured gains and its real impact on quality.

By Mohamed Meguedmi·Update 2026-08-27·Tested on Windows, macOS, and Linux

#Why the KV cache consumes your VRAM

When an LLM generates text, it doesn't recalculate attention over the entire prompt for every new token: it keeps the key (K) and value (V) vectors for each token it has already seen in memory. That's the KV cache, and it's what makes generation fast. The problem: this memory grows linearly with context length. Double the context, and you double the KV cache.

With a short prompt of a few hundred tokens, it is negligible. But as soon as you use RAG with large documents, summarize transcripts, or run agents with long-term memory, the context explodes—and so does the KV cache. On a 70B model with a 32k context, it can exceed 10 GB on its own, in addition to the model's ~40 GB. It is often the cache, not the model, that causes out-of-memory errors on long prompts.

i
Model vs. KV cache: two separate memory requirements
VRAM is divided into three parts: model weights (fixed, depending on size and quantization), the KV cache (variable, depending on context), and a bit of overhead. Quantizing the model (Q4_K_M) reduces the first part. Quantizing the KV cache reduces the second. These are two independent levers.

#How large is your KV cache?

The Local AI Kit

Your private ChatGPT, free, on your own machine in an hour — LM Studio, Ollama, Open WebUI, your documents, no cloud.

  • Lifetime online access
  • PDF + files
  • Lifetime updates

KV cache size depends on four factors: the number of model layers, the attention dimension, the context length, and the storage precision. In FP16 (the default), the approximate formula is: 2 (K and V) × layers × dim_kv × context × 2 bytes. In practice, use these rough figures for FP16 with a 32k-token context:

7-9B (e.g., Qwen 3.5 9B, Granite 4.2 8B)
≈ 2 to 4 GB of KV cache at 32k, depending on the architecture (GQA helps a lot).
14B
≈ 4 to 6 GB at 32k tokens.
32B
≈ 8 to 10 GB at 32k tokens.
70B
≈ 10 to 16 GB at 32k tokens — often the limiting factor.
i
GQA changes the game
Recent models use Grouped-Query Attention (GQA), which shares K/V heads across multiple query heads. As a result, their KV cache is already much lighter than that of older Multi-Head models. Qwen 3.5, Gemma 4, and Granite 4.2 benefit from this. It doesn't eliminate the need for quantization with very long contexts, but it raises the threshold.

#What KV cache quantization changes

The idea is the same as with the model weights: instead of storing each cache value in 16 bits (FP16), you store it in 8 bits (Q8_0) or 4 bits (Q4_0). This mechanically divides cache memory by 2 (Q8) or by 4 (Q4). Because the KV cache can account for a large share of VRAM with long contexts, the benefit is direct: with constant VRAM, you can approximately double the context length that can be handled in Q8.

FP16
Reference-level accuracy, no loss, but the most demanding. The default.
Q8_0
Half the memory, with an almost imperceptible quality loss on most models. The best compromise.
Q4_0
Four times less memory, but a measurable quality loss that varies by model. Reserve it for cases where VRAM is truly the limiting factor.
!
Prerequisite: Flash Attention
KV-cache quantization requires Flash Attention to be enabled. Without it, llama.cpp and Ollama refuse to apply any cache type other than FP16 or fail. That makes sense: Flash Attention and a quantized cache work together to reduce the attention memory footprint.

#Enable KV quantization in Ollama

Ollama exposes KV-cache quantization through two daemon environment variables. You must enable Flash Attention, then choose the cache type. Set these variables on the Ollama service, not when running ollama run.

  1. 01
    Enable Flash Attention
    Set OLLAMA_FLASH_ATTENTION=1 in the daemon environment. This is required for any non-FP16 cache.
  2. 02
    Choosing the cache type
    Set OLLAMA_KV_CACHE_TYPE to the desired value: f16 (default), q8_0 (recommended), or q4_0 (aggressive).
  3. 03
    Restart the daemon
    Environment variables are read only when the service starts. Restart Ollama for them to take effect.
  4. 04
    Check the gain
    Load a model with a large context and monitor VRAM with nvidia-smi or ollama ps. You should be able to raise num_ctx higher than before.
Terminal — Linux (systemd)
# Éditer l'unité systemd du service
sudo systemctl edit ollama

# Ajouter dans la section [Service] :
[Service]
Environment="OLLAMA_FLASH_ATTENTION=1"
Environment="OLLAMA_KV_CACHE_TYPE=q8_0"

# Recharger et redémarrer
sudo systemctl daemon-reload
sudo systemctl restart ollama
Terminal — manual launch
# Sur macOS ou pour un lancement direct du serveur
export OLLAMA_FLASH_ATTENTION=1
export OLLAMA_KV_CACHE_TYPE=q8_0
ollama serve
PowerShell — Windows
# Poser les variables au niveau utilisateur puis redémarrer Ollama
setx OLLAMA_FLASH_ATTENTION 1
setx OLLAMA_KV_CACHE_TYPE q8_0

# Quitter Ollama depuis la barre des tâches et le relancer
→
Increase the context to take advantage of it
Quantifying the cache is pointless if your num_ctx remains low. Once quantization is enabled, increase the model's context window (the num_ctx parameter in the Modelfile or API call) to convert the saved VRAM into usable context.

#Enable in llama.cpp

On the command line with llama.cpp, KV-cache quantization is controlled with two separate flags for keys (K) and values (V), plus the Flash Attention flag. K and V can be quantized independently, but in practice they are set to the same level.

Terminal — llama-cli / llama-server
# -fa active Flash Attention (obligatoire)
# -ctk = type du cache des clés, -ctv = type du cache des valeurs
llama-server \
  -m modele.gguf \
  -c 32768 \
  -fa \
  -ctk q8_0 \
  -ctv q8_0 \
  -ngl 99
-fa
Enable Flash Attention. Without this flag, -ctk/-ctv below f16 fails.
-ctk q8_0
Quantifies the key cache in 8-bit format. Possible values: f16, q8_0, q4_0, q4_1, q5_0, q5_1.
-ctv q8_0
Quantifies the value cache. Same set of values as -ctk.
-c 32768
The context size you are targeting. Quantization is what lets you increase it without overflowing.
→
Possible K/V asymmetry
The value (V) cache handles aggressive quantization better than the key (K) cache, which is more sensitive. An interesting compromise when VRAM is tight: -ctk q8_0 -ctv q4_0. You save space on the V side without damaging key precision.

#How much context you gain: real-world figures

The practical benefit depends on how much of your VRAM budget the KV cache represents. On a model that fits comfortably, quantizing the cache merely frees up a little headroom. On a model that is already saturated, it can make the difference between 8k and 24k of context. Here are observed ballpark figures at constant VRAM:

FP16 → Q8_0
Cache divided by 2. In practice, the sustainable context is roughly doubled when the cache dominated the budget.
FP16 → Q4_0
Cache divided by 4. Manageable context up to ~3-4× longer, at the cost of measurable quality loss.
14B example on RTX 4080 16GB
From ~16k context in FP16 to ~32k+ in Q8_0, with the model and everything else still fitting in the same 16 GB.
32B example on RTX 4090 24GB
Switching the cache to Q8_0 often lets you handle long RAG documents without CPU offloading.
i
The benefit is not just memory
A smaller KV cache also means less memory bandwidth to scan for each token. With very long contexts, you sometimes see a slight generation-speed gain in Q8, in addition to the VRAM savings. Don't count on it consistently, but it's a frequent bonus.

#Quality impact by model

That’s the real question. Quantizing the cache introduces noise into attention, and not all models respond to it the same way. The rule of thumb emerging from community testing:

Q8_0 in the cache
Nearly indistinguishable on the vast majority of models. Perplexity and perceived quality are nearly identical to FP16. This is the setting to enable by default, almost without thinking.
Q4_0 in the cache
Visible and variable loss. Some models handle it very well; others start rambling on very long contexts, lose the thread, or hallucinate more. Test it on YOUR use case.
Models with GQA
Generally more robust to cache quantization, since their cache is already compact and well structured.
Sensitive tasks (code, calculations, strict extraction)
More exposed to Q4 degradation. Stay at Q8 for anything requiring factual precision.
!
Test before adopting Q4
Never use Q4 for the production cache without comparing outputs on your own long prompts. The VRAM savings are tempting, but losing reliability in a RAG system or agent costs more than the VRAM saved. Q8 is the safe default; Q4 is an optimization that must be validated empirically.

#Troubleshooting

“flash attention required” or cache ignored
You forgot -fa (llama.cpp) or OLLAMA_FLASH_ATTENTION=1 (Ollama). The quantized cache is silently reverted to FP16, or an error is raised.
No visible VRAM improvement
Your context is too short for the cache to matter. The benefit only appears with long prompts. Increase num_ctx / -c to see it.
Quality degrades on long prompts
You are probably using Q4 on a model that handles it poorly. Switch K back to q8_0 (or even everything to q8_0) and test again.
Ollama variables have no effect
They are read only when the daemon starts. Restart the service after setting them, and verify that they are visible to the ollama process.
Still running out of memory despite quantization
It's the model itself that's overflowing, not the cache. Quantize the weights too (Q4_K_M), or move down to a smaller model.
Quick diagnostics
# Vérifier ce qui tourne sur GPU et la VRAM consommée
ollama ps
nvidia-smi

# Confirmer que les variables sont bien vues par le daemon (Linux)
systemctl show ollama --property=Environment

#Go further

KV-cache quantization is one of several levers for fitting large contexts locally. These guides complete the picture:

Flash Attention 2 on a local LLM: enable it and benchmark the gain
The prerequisite for KV quantization, and a memory/speed gain in its own right. Read this first if Flash Attention isn't active on your system yet.
Choose your quantization (Q4, Q5, Q8, FP16)
To quantize the model weights—the other major VRAM allocation, complementary to cache quantization.
Install Ollama: Windows, macOS, and Linux
If your stack isn't in place yet, including the GPU requirements for proper sizing.
Did this guide help you?

Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.