Advanced 12 minOptimization

Flash Attention 2: enable it (llama.cpp, Ollama, vLLM)

Flash Attention 2 (FA2) is the optimization that takes a long-context local LLM from unusable to comfortable. Enabled via an environment variable in Ollama, a flag in llama.cpp, and by default in vLLM—but your GPU still has to genuinely support it. This guide shows how to enable it in the three main runtimes, measure the actual gains (KV-cache memory and tokens/sec) by context size, and identify cases where it provides no benefit.

By Mohamed Meguedmi·Update 2026-08-31·Tested on Windows, macOS, and Linux
i
In brief
Flash Attention 2 recomputes attention block by block in the GPU's SRAM instead of materializing the entire matrix, without changing the result. · It is enabled with OLLAMA_FLASH_ATTENTION=1 under Ollama, the -fa flag in llama.cpp, and by default since version 0.2 in vLLM. · Requires a recent GPU (NVIDIA Ampere or newer, AMD RDNA 3/4 via ROCm); on Apple Silicon, Metal already handles the equivalent natively. · Measured gain on RTX 4090 (Mistral Small 24B Q4_K_M, 16k tokens): +28% speed and -2.8 GB of VRAM.

#Why Flash Attention 2

A transformer’s standard attention has quadratic memory complexity as a function of context length. Doubling the context quadruples the VRAM consumed by attention. At 32k tokens, on an 8B model in FP16, the attention matrix alone can weigh more than the model’s weights.

Flash Attention, introduced by Tri Dao in 2022 and refined into FA2 in 2023, does not change the mathematical result: it changes how the result is computed. The idea is to process attention in blocks (tiling) directly in the GPU's SRAM instead of materializing the full matrix in HBM. The result is strictly identical to naïve attention, apart from precision differences.

In practice, on a local LLM during inference, FA2 delivers two cumulative gains: memory dedicated to attention becomes nearly linear in context size instead of quadratic (the KV cache is greatly reduced), and throughput increases by 2 to 4× on long contexts because expensive memory accesses are eliminated.

i
What FA2 is not
This is not compression. Generation quality is identical bit-for-bit (apart from the order of floating-point additions). It is not quantization either: FA2 and KV-cache quantization (Q8_0, for example) are two distinct optimizations that you can combine.

#GPU and software requirements

The Local AI Kit

Your private ChatGPT, free, on your own machine in an hour — LM Studio, Ollama, Open WebUI, your documents, no cloud.

  • Lifetime online access
  • PDF + files
  • Lifetime updates

Flash Attention 2 uses specific matrix instructions (tensor cores) that exist only on certain GPU generations. This is the first criterion to check.

NVIDIA Ampere (RTX 3000, A100) and newer
Full, high-performance FA2 support. This is the sweet spot: compute capability 8.0+.
NVIDIA Ada Lovelace (RTX 4000) and Blackwell (RTX 5000)
Optimal native support. Maximum gains on RTX 4090 / 5090 thanks to Hopper-class tensor cores.
NVIDIA Turing (RTX 2000, T4)
Partial and slower support. FA2 works, but without benefiting from the BF16 instructions. Enable it and measure: sometimes neutral, sometimes positive.
NVIDIA Pascal (GTX 1080 Ti, P40) and earlier
Unsupported. Tensor cores do not exist. The flag is silently ignored or produces a runtime error.
AMD RDNA 3/4 (RX 7000/9000)
Support via ROCm with the official Flash Attention port (composable_kernel). Near-native performance on RX 7900 XTX and above.
Apple Silicon (M1 to M4)
Not Flash Attention 2 in the strict sense. Metal has its own fused-attention kernels. The runtimes (llama.cpp Metal, MLX) use the equivalent automatically; you do not need to enable anything.
!
Required precision: FP16 or BF16
FA2 doesn’t work in FP32. GGUF models (Q4_K_M, Q5_K_M, etc.) are dequantized on the fly to FP16 for attention, so it works. But if you load an FP32 model (rare in inference), disable FA2 or the runtime will refuse to run.

#Enable Flash Attention 2 in Ollama

Since version 0.3, Ollama has supported Flash Attention through an environment variable. It has historically not been enabled by default because older cards may not support it. Turn it on yourself.

One-time activation (Linux/macOS)
OLLAMA_FLASH_ATTENTION=1 ollama serve

To make it persistent through systemd (the common Linux case):

systemd drop-in
sudo systemctl edit ollama.service

# Dans l'éditeur, ajoutez :
[Service]
Environment="OLLAMA_FLASH_ATTENTION=1"
Environment="OLLAMA_KV_CACHE_TYPE=q8_0"

# Puis :
sudo systemctl daemon-reload
sudo systemctl restart ollama

On Windows, add the user environment variable via Control Panel > System > Environment Variables, then restart Ollama. On macOS, use launchctl setenv or add the variable to your shell rc file.

→
Combine with KV-cache quantization
OLLAMA_KV_CACHE_TYPE=q8_0 quantizes the KV cache to 8 bits, halving its size again with virtually no loss in quality. Combined with FA2, you can triple the context size your VRAM can hold. q4_0 is even more aggressive, but its impact starts to show on reasoning models.

To verify that FA2 is active, run ollama with OLLAMA_DEBUG=1 and look for the line flash_attention=true in the logs when the model loads.

#Enable Flash Attention 2 in llama.cpp

llama.cpp exposes an explicit flag that is more controllable than Ollama: -fa (or --flash-attn in its long form). It works with both llama-cli and llama-server.

llama-cli with FA2
./llama-cli -m models/mistral-small-24b-instruct-q4_k_m.gguf \
  --flash-attn \
  --ctx-size 32768 \
  -p "Résume ce document..."
llama-server with FA2
./llama-server -m models/mistral-small-24b-instruct-q4_k_m.gguf \
  -fa \
  --ctx-size 32768 \
  --cache-type-k q8_0 \
  --cache-type-v q8_0 \
  --host 0.0.0.0 --port 8080

The --cache-type-k and --cache-type-v flags quantize the KV cache's keys and values, respectively. q8_0 is the safe choice (virtually no loss); q4_0 is the aggressive option. Note that KV-cache quantization requires FA2 to be enabled: without -fa, these flags are rejected.

i
Compilation required
If you compiled llama.cpp with -DGGML_CUDA=ON, FA2 is included automatically. On Vulkan builds, FA2 support is newer and less mature—test it before pushing to production. On Metal (macOS), FA2 is integrated natively; the -fa flag is accepted, but the implementation uses the Apple kernels.

#Enable Flash Attention 2 in vLLM

vLLM has used Flash Attention 2 by default since version 0.2, and switches to FlashAttention-3 on Hopper (H100) and Blackwell when possible. You generally do not need to enable anything. The only time you need to intervene is to force a backend, for example to compare them or work around a bug on an exotic GPU.

Force the Flash Attention backend
VLLM_ATTENTION_BACKEND=FLASH_ATTN \
  vllm serve mistralai/Mistral-Small-24B-Instruct-2501 \
  --max-model-len 32768 \
  --gpu-memory-utilization 0.90

Available backends are FLASH_ATTN (FA2), FLASHINFER (even faster at some sizes, requires installing flashinfer), XFORMERS (fallback for Ampere and earlier), and TORCH_SDPA (generic fallback). On a recent GPU, FLASH_ATTN or FLASHINFER are the right choices.

→
FlashInfer on Hopper / Blackwell
If you serve Qwen 3.6 35B or larger on H100/B200 with many concurrent requests, install flashinfer and switch to VLLM_ATTENTION_BACKEND=FLASHINFER. The gain in batched throughput is measurable (10–20% depending on the load) because FlashInfer optimizes vLLM’s paged KV cache more effectively.

#Measured gain by context size

The benefit of Flash Attention 2 depends heavily on context size. With a short context, it is marginal or even negative (tiling overhead). With a long context, it is the difference between running and not running. Here are representative figures on RTX 4090 24 GB with Mistral Small 24B Q4_K_M.

2k-token context
Without FA2: 40 tok/s, 14.6 GB VRAM. With FA2: 41 tok/s, 14.4 GB. Gain: ~2%. Not worth it at this scale.
8k-token context
Without FA2: 36 tok/s, 16.2 GB. With FA2: 39 tok/s, 15.0 GB. Gain: +8% speed, –1.2 GB.
16k-token context
Without FA2: 29 tok/s, 18.8 GB. With FA2: 37 tok/s, 16.0 GB. Gain: +28% speed, –2.8 GB.
32k-token context
Without FA2: OOM (>24 GB). With FA2: 32 tok/s, 18.6 GB. FA2 simply makes 32k context accessible.
64k-token context (FA2 + KV q8_0)
30 tok/s, 21.5 GB. Without FA2 + without KV quantization: impossible on 24 GB.

On Qwen 3.5 9B (smaller model, different attention head dimension), the gain at 32k is more modest (+15% speed) because attention accounts for less of the total cost. On Qwen 3.8 27B, a heavier reasoning model, the gain at 16k is massive (+45%) because the attention/feedforward ratio favors FA2.

i
Measure it yourself
These figures provide an order of magnitude. Attention-head size, model quantization, GPU frequency, and runtime version can vary the gain by ±20%. Run your typical prompt with and without -fa to decide.

#Cases where it doesn't help (or makes things worse)

Pascal / Maxwell GPU
GTX 1080 Ti, P40, Tesla M40, Titan X Maxwell. No compatible tensor cores. The flag is ignored on Ollama and rejected by recent llama.cpp. No gain is possible—use xformers or nothing.
Very short inference (500-token chat)
If you consistently generate short responses from short prompts, FA2 adds a few percent of overhead without providing any benefit. Most noticeable on small 1B-3B models.
CPU only
FA2 is a GPU optimization. With llama.cpp CPU, the flag is ignored. CPU optimizations use other paths (ARM SVE, AVX-512, etc.).
Models with non-standard sliding-window attention
Gemma 4 (alternating global/local attention) and the Mistral variants with a sliding window: depending on the runtime version, FA2 may fall back. Check the logs; sometimes the backend still doesn't support it.
Apple Silicon
On M1–M4, llama.cpp accepts the -fa flag, but the implementation uses Metal, which already performs its own optimizations. The measured gain is marginal because Metal attention is already natively fused.

#FA2 vs xformers vs PyTorch SDPA

You will encounter these three names when reading runtime documentation. They do not do exactly the same thing and are not the same age.

xformers (Meta, 2021)
Efficient kernel library including memory_efficient_attention. Precursor to Flash Attention. Still used for fine-tuning (Unsloth, Axolotl) because it supports more attention variants. Slower than FA2 for inference on Ampere+.
Flash Attention 2 (Tri Dao, 2023)
Successor to Flash Attention 1. Specialized for inference and training, with hand-written CUDA kernels, and the fastest on Ampere and Hopper. The de facto standard in 2026.
PyTorch SDPA
torch.nn.functional.scaled_dot_product_attention. Since PyTorch 2.0, it automatically dispatches to Flash Attention 2 when possible, otherwise to xformers, and otherwise to the naive implementation. This is what vLLM and many runtimes use internally.
FlashAttention-3 (Tri Dao + NVIDIA, 2024)
Specific to Hopper (H100) and Blackwell GPUs. Uses FP8 and warp-specialized asymmetry. If you're running on H100, FA3 beats FA2. On consumer RTX, FA2 remains the target.
→
In practice, you don't have to choose
The runtime decides for you: vLLM chooses the best backend based on the detected GPU, llama.cpp uses its own FA2 implementation behind the -fa flag, and Ollama does the same. The distinctions above matter when you write PyTorch code or fine-tune.

#Troubleshooting

The flag appears to be ignored (no improvement)
Check the runtime version (Ollama ≥ 0.3, recent llama.cpp commit). Enable verbose logs: OLLAMA_DEBUG=1 or llama-cli --verbose. Look for flash_attention=true or flash_attn=enabled.
Error: unsupported head dimension
Certain head dimensions (96, 192) are not supported by all FA2 versions. Update the runtime or fall back to the default backend. This mainly affects exotic models.
Visible quality regression
Very rare with FA2 alone (mathematically equivalent result). If you see degradation, it is probably KV-cache quantization, not FA2. Disable --cache-type-k/v and see whether the problem persists.
OOM despite FA2 being enabled
FA2 reduces attention memory, not weight memory. If your 30B Q5 model doesn't fit in 16 GB of VRAM, FA2 won't save you. Drop down one quantization level (Q4_K_M) or use a smaller model.
vLLM falls back to xformers instead of FA2
Check the installation: pip install flash-attn --no-build-isolation. On Turing, vLLM prefers xformers because FA2 is less optimized there. This is expected.

#Go further

Flash Attention 2 is the first thing to try when you want to extend the usable context for a given amount of VRAM. If you want to optimize further, these related guides are useful complements:

GGUF quantization in 2026: Q4_K_M vs Q5_K_M vs Q6_K
The obvious companion: reduce the model size to free up VRAM, combined with FA2 to extend the context further.
Compile llama.cpp with CUDA
Essential if you want the latest FA2 version in llama.cpp and need to benchmark cleanly.
Deploy vLLM in production
FA2 really shines on a batched, multi-request server: that’s where the gain explodes.
Frequently asked questions about Flash Attention
How do you enable Flash Attention in llama.cpp and LM Studio in 2026?+
In llama.cpp, the -fa flag is enough on the command line (or --flash-attn on for the server). LM Studio now enables it by default on CUDA, Metal, and Vulkan: if your version is older, check the option in the inference engine settings. In Ollama, the OLLAMA_FLASH_ATTENTION=1 variable remains the method described above.
Does Flash Attention 3 change anything for a local LLM?+
Not for the general public: FA3 targets Hopper datacenter GPUs (H100) and their specific instructions. On a consumer card (RTX 30/40/50, Mac, Radeon), the gains come from the FA2 implementation built into llama.cpp—there are no additional flags to add.
Why is Flash Attention slower on my AMD card?+
Known ROCm/HIP trap: the fused kernel is used only when the two KV cache types are symmetric (for example, q4_0 for K and V). A combination such as q4_0 + f16 silently falls back to a slower non-fused path, with no warning. Align the two types and rerun the benchmark.
Can I combine Flash Attention and a quantized KV cache?+
Yes on NVIDIA, provided your llama.cpp build was compiled with FA_ALL_QUANTS—otherwise attention falls back to the CPU and performance collapses. On Mac, LM Studio handles the combination natively through Metal. When in doubt, compare tok/s with and without cache quantization: the difference is immediately apparent.
Did this guide help you?

Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.