Flash Attention 2: enable it (llama.cpp, Ollama, vLLM)
Flash Attention 2 (FA2) is the optimization that takes a long-context local LLM from unusable to comfortable. Enabled via an environment variable in Ollama, a flag in llama.cpp, and by default in vLLM—but your GPU still has to genuinely support it. This guide shows how to enable it in the three main runtimes, measure the actual gains (KV-cache memory and tokens/sec) by context size, and identify cases where it provides no benefit.
#Why Flash Attention 2
A transformer’s standard attention has quadratic memory complexity as a function of context length. Doubling the context quadruples the VRAM consumed by attention. At 32k tokens, on an 8B model in FP16, the attention matrix alone can weigh more than the model’s weights.
Flash Attention, introduced by Tri Dao in 2022 and refined into FA2 in 2023, does not change the mathematical result: it changes how the result is computed. The idea is to process attention in blocks (tiling) directly in the GPU's SRAM instead of materializing the full matrix in HBM. The result is strictly identical to naïve attention, apart from precision differences.
In practice, on a local LLM during inference, FA2 delivers two cumulative gains: memory dedicated to attention becomes nearly linear in context size instead of quadratic (the KV cache is greatly reduced), and throughput increases by 2 to 4× on long contexts because expensive memory accesses are eliminated.
#GPU and software requirements
Your private ChatGPT, free, on your own machine in an hour — LM Studio, Ollama, Open WebUI, your documents, no cloud.
- Lifetime online access
- PDF + files
- Lifetime updates
Flash Attention 2 uses specific matrix instructions (tensor cores) that exist only on certain GPU generations. This is the first criterion to check.
- NVIDIA Ampere (RTX 3000, A100) and newer
- Full, high-performance FA2 support. This is the sweet spot: compute capability 8.0+.
- NVIDIA Ada Lovelace (RTX 4000) and Blackwell (RTX 5000)
- Optimal native support. Maximum gains on RTX 4090 / 5090 thanks to Hopper-class tensor cores.
- NVIDIA Turing (RTX 2000, T4)
- Partial and slower support. FA2 works, but without benefiting from the BF16 instructions. Enable it and measure: sometimes neutral, sometimes positive.
- NVIDIA Pascal (GTX 1080 Ti, P40) and earlier
- Unsupported. Tensor cores do not exist. The flag is silently ignored or produces a runtime error.
- AMD RDNA 3/4 (RX 7000/9000)
- Support via ROCm with the official Flash Attention port (composable_kernel). Near-native performance on RX 7900 XTX and above.
- Apple Silicon (M1 to M4)
- Not Flash Attention 2 in the strict sense. Metal has its own fused-attention kernels. The runtimes (llama.cpp Metal, MLX) use the equivalent automatically; you do not need to enable anything.
#Enable Flash Attention 2 in Ollama
Since version 0.3, Ollama has supported Flash Attention through an environment variable. It has historically not been enabled by default because older cards may not support it. Turn it on yourself.
To make it persistent through systemd (the common Linux case):
On Windows, add the user environment variable via Control Panel > System > Environment Variables, then restart Ollama. On macOS, use launchctl setenv or add the variable to your shell rc file.
To verify that FA2 is active, run ollama with OLLAMA_DEBUG=1 and look for the line flash_attention=true in the logs when the model loads.
#Enable Flash Attention 2 in llama.cpp
llama.cpp exposes an explicit flag that is more controllable than Ollama: -fa (or --flash-attn in its long form). It works with both llama-cli and llama-server.
The --cache-type-k and --cache-type-v flags quantize the KV cache's keys and values, respectively. q8_0 is the safe choice (virtually no loss); q4_0 is the aggressive option. Note that KV-cache quantization requires FA2 to be enabled: without -fa, these flags are rejected.
#Enable Flash Attention 2 in vLLM
vLLM has used Flash Attention 2 by default since version 0.2, and switches to FlashAttention-3 on Hopper (H100) and Blackwell when possible. You generally do not need to enable anything. The only time you need to intervene is to force a backend, for example to compare them or work around a bug on an exotic GPU.
Available backends are FLASH_ATTN (FA2), FLASHINFER (even faster at some sizes, requires installing flashinfer), XFORMERS (fallback for Ampere and earlier), and TORCH_SDPA (generic fallback). On a recent GPU, FLASH_ATTN or FLASHINFER are the right choices.
#Measured gain by context size
The benefit of Flash Attention 2 depends heavily on context size. With a short context, it is marginal or even negative (tiling overhead). With a long context, it is the difference between running and not running. Here are representative figures on RTX 4090 24 GB with Mistral Small 24B Q4_K_M.
- 2k-token context
- Without FA2: 40 tok/s, 14.6 GB VRAM. With FA2: 41 tok/s, 14.4 GB. Gain: ~2%. Not worth it at this scale.
- 8k-token context
- Without FA2: 36 tok/s, 16.2 GB. With FA2: 39 tok/s, 15.0 GB. Gain: +8% speed, –1.2 GB.
- 16k-token context
- Without FA2: 29 tok/s, 18.8 GB. With FA2: 37 tok/s, 16.0 GB. Gain: +28% speed, –2.8 GB.
- 32k-token context
- Without FA2: OOM (>24 GB). With FA2: 32 tok/s, 18.6 GB. FA2 simply makes 32k context accessible.
- 64k-token context (FA2 + KV q8_0)
- 30 tok/s, 21.5 GB. Without FA2 + without KV quantization: impossible on 24 GB.
On Qwen 3.5 9B (smaller model, different attention head dimension), the gain at 32k is more modest (+15% speed) because attention accounts for less of the total cost. On Qwen 3.8 27B, a heavier reasoning model, the gain at 16k is massive (+45%) because the attention/feedforward ratio favors FA2.
#Cases where it doesn't help (or makes things worse)
- Pascal / Maxwell GPU
- GTX 1080 Ti, P40, Tesla M40, Titan X Maxwell. No compatible tensor cores. The flag is ignored on Ollama and rejected by recent llama.cpp. No gain is possible—use xformers or nothing.
- Very short inference (500-token chat)
- If you consistently generate short responses from short prompts, FA2 adds a few percent of overhead without providing any benefit. Most noticeable on small 1B-3B models.
- CPU only
- FA2 is a GPU optimization. With llama.cpp CPU, the flag is ignored. CPU optimizations use other paths (ARM SVE, AVX-512, etc.).
- Models with non-standard sliding-window attention
- Gemma 4 (alternating global/local attention) and the Mistral variants with a sliding window: depending on the runtime version, FA2 may fall back. Check the logs; sometimes the backend still doesn't support it.
- Apple Silicon
- On M1–M4, llama.cpp accepts the -fa flag, but the implementation uses Metal, which already performs its own optimizations. The measured gain is marginal because Metal attention is already natively fused.
#FA2 vs xformers vs PyTorch SDPA
You will encounter these three names when reading runtime documentation. They do not do exactly the same thing and are not the same age.
- xformers (Meta, 2021)
- Efficient kernel library including memory_efficient_attention. Precursor to Flash Attention. Still used for fine-tuning (Unsloth, Axolotl) because it supports more attention variants. Slower than FA2 for inference on Ampere+.
- Flash Attention 2 (Tri Dao, 2023)
- Successor to Flash Attention 1. Specialized for inference and training, with hand-written CUDA kernels, and the fastest on Ampere and Hopper. The de facto standard in 2026.
- PyTorch SDPA
- torch.nn.functional.scaled_dot_product_attention. Since PyTorch 2.0, it automatically dispatches to Flash Attention 2 when possible, otherwise to xformers, and otherwise to the naive implementation. This is what vLLM and many runtimes use internally.
- FlashAttention-3 (Tri Dao + NVIDIA, 2024)
- Specific to Hopper (H100) and Blackwell GPUs. Uses FP8 and warp-specialized asymmetry. If you're running on H100, FA3 beats FA2. On consumer RTX, FA2 remains the target.
#Troubleshooting
- The flag appears to be ignored (no improvement)
- Check the runtime version (Ollama ≥ 0.3, recent llama.cpp commit). Enable verbose logs: OLLAMA_DEBUG=1 or llama-cli --verbose. Look for flash_attention=true or flash_attn=enabled.
- Error: unsupported head dimension
- Certain head dimensions (96, 192) are not supported by all FA2 versions. Update the runtime or fall back to the default backend. This mainly affects exotic models.
- Visible quality regression
- Very rare with FA2 alone (mathematically equivalent result). If you see degradation, it is probably KV-cache quantization, not FA2. Disable --cache-type-k/v and see whether the problem persists.
- OOM despite FA2 being enabled
- FA2 reduces attention memory, not weight memory. If your 30B Q5 model doesn't fit in 16 GB of VRAM, FA2 won't save you. Drop down one quantization level (Q4_K_M) or use a smaller model.
- vLLM falls back to xformers instead of FA2
- Check the installation: pip install flash-attn --no-build-isolation. On Turing, vLLM prefers xformers because FA2 is less optimized there. This is expected.
#Go further
Flash Attention 2 is the first thing to try when you want to extend the usable context for a given amount of VRAM. If you want to optimize further, these related guides are useful complements:
- GGUF quantization in 2026: Q4_K_M vs Q5_K_M vs Q6_K
- The obvious companion: reduce the model size to free up VRAM, combined with FA2 to extend the context further.
- Compile llama.cpp with CUDA
- Essential if you want the latest FA2 version in llama.cpp and need to benchmark cleanly.
- Deploy vLLM in production
- FA2 really shines on a batched, multi-request server: that’s where the gain explodes.
How do you enable Flash Attention in llama.cpp and LM Studio in 2026?+
Does Flash Attention 3 change anything for a local LLM?+
Why is Flash Attention slower on my AMD card?+
Can I combine Flash Attention and a quantized KV cache?+
Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.