What Is Flash Attention? The Real Speed and Memory Gains, Explained
Flash Attention computes exactly the same attention as before, without ever building the giant matrix that made long contexts impossible. What it changes, and how to turn it on locally.
Key takeaways
- Flash Attention is an algorithm that computes standard attention faster and with far less memory, with no approximation. The output is mathematically the same.
- It works by never materializing the full attention matrix. It processes it in small tiles that fit in the GPU's fast on-chip memory.
- Attention memory goes from growing with the square of the context length to growing linearly. At 32K tokens that is the difference between tens of gigabytes and well under one.
- The main benefit for local users is on long prompts: faster prompt processing and lower peak VRAM. It barely changes token-by-token generation speed.
- In llama.cpp and Ollama it is a single switch, and it is a prerequisite for quantizing the KV cache.
Flash Attention in one paragraph
Your private ChatGPT, free, on your own machine in an hour — LM Studio, Ollama, Open WebUI, your documents, no cloud.
- Lifetime online access
- PDF + files
- 30-day refund
Attention is the operation that lets each token in a transformer look at every other token. Computed the textbook way, it builds a table with one row and one column per token, so its size grows with the square of the sequence length. Flash Attention, introduced by Tri Dao and colleagues in 2022 (arXiv 2205.14135), reorganizes the same computation so that the table is never stored. It is an IO-aware algorithm: it is designed around how data moves inside a GPU rather than around the arithmetic, because the data movement was the real bottleneck.
The problem: a matrix that grows with the square of the context
For a sequence of N tokens, standard attention stores an N × N score matrix for every attention head in every layer, reads it back to apply softmax, and reads it again to weight the values. Here is what that intermediate matrix weighs for one layer of a model with 32 attention heads at FP16:
| Context length | Scores per head (N²) | Per head, FP16 | Per layer (32 heads) |
|---|---|---|---|
| 2,048 | 4.2 million | 8 MB | 0.27 GB |
| 8,192 | 67 million | 134 MB | 4.3 GB |
| 32,768 | 1.07 billion | 2.1 GB | 69 GB |
| 131,072 | 17.2 billion | 34 GB | 1,100 GB |
Computed as N² × 2 bytes. This is the transient matrix that naive attention would materialize while processing a full prompt of that length; it is separate from the KV cache.
Frameworks never literally allocated a terabyte; they chunked, recomputed or simply failed. But the table shows why long-context inference was impractical on a single GPU before 2022, and why the problem is memory traffic rather than arithmetic.
How Flash Attention works
A GPU has two kinds of memory. The large pool, VRAM (technically HBM or GDDR), holds gigabytes but is comparatively slow to reach. A tiny on-chip pool, SRAM, holds megabytes and is an order of magnitude faster. Standard attention shuttles the giant score matrix between the two several times. Flash Attention uses two ideas to avoid that:
- Tiling. Split queries, keys and values into blocks small enough to fit in SRAM. Compute attention block by block, and combine the partial results with a running, numerically exact softmax. The full matrix never exists anywhere.
- Recomputation. For training, instead of saving the matrix for the backward pass, recompute the needed tiles on the fly. Redoing the arithmetic is cheaper than moving the data.
| Standard attention | Flash Attention | |
|---|---|---|
| Result | Exact | Exact (identical up to floating-point rounding) |
| Extra memory vs sequence length | Quadratic | Linear |
| Limiting factor | Reads and writes to VRAM | Arithmetic in on-chip SRAM |
| Speed of the attention operation | Baseline | Typically 2–4× faster, more at long sequences |
This distinguishes it from the "efficient attention" methods that came before, such as sparse or low-rank attention, which traded accuracy for speed. Flash Attention changes nothing about what the model computes.
FlashAttention 1, 2 and 3
| Version | Year | What changed |
|---|---|---|
| FlashAttention | 2022 | The original tiling and recomputation algorithm |
| FlashAttention-2 | 2023 | Better parallelism and work partitioning; roughly 2× faster than v1 |
| FlashAttention-3 | 2024 | Tuned for NVIDIA Hopper data-center GPUs, with FP8 support |
The reference implementation targets NVIDIA data-center and recent consumer GPUs. Local runtimes such as llama.cpp ship their own implementations of the same idea for CUDA, Metal, Vulkan and ROCm, so you do not install the research library to benefit from it.
What you actually gain on a local machine
| Phase | Bottleneck | Effect of Flash Attention |
|---|---|---|
| Prompt processing (prefill) | Compute and attention memory traffic | Large. Faster time to first token and lower peak VRAM, growing with prompt length |
| Token generation (decode) | Reading model weights from VRAM | Small. A few percent at most on short contexts; more noticeable once the context is very long |
| KV cache size | Not affected directly | None by itself, but it unlocks KV-cache quantization in llama.cpp and Ollama, which does shrink it |
So if you chat in short exchanges with an 8B model, you will struggle to notice it. If you paste in a 40-page document, run a coding agent with a large context, or do retrieval with long passages, it is the difference between waiting and working. The memory side of long contexts is covered in context window vs VRAM cost.
How to turn it on
| Tool | Setting |
|---|---|
llama.cpp (llama-server, llama-cli) | -fa / --flash-attn; recent builds enable it automatically when the backend supports it |
| Ollama | Environment variable OLLAMA_FLASH_ATTENTION=1, then restart the server |
| LM Studio | "Flash Attention" toggle in the model load settings |
| vLLM | Used by default on supported GPUs; nothing to set |
| Hugging Face transformers | attn_implementation="flash_attention_2" when loading the model |
# Ollama on Linux (systemd): enable Flash Attention and an 8-bit KV cache
sudo systemctl edit ollama
# [Service]
# Environment="OLLAMA_FLASH_ATTENTION=1"
# Environment="OLLAMA_KV_CACHE_TYPE=q8_0"
sudo systemctl restart ollama
The second variable is the reason many people enable the first: Ollama and llama.cpp require Flash Attention before they will quantize the KV cache, and an 8-bit cache halves context memory with negligible quality loss.
When it does not help, or is not available
- Short contexts. Below a couple of thousand tokens the attention matrix is small and the gain disappears into noise.
- CPU-only inference. The algorithm is about GPU memory hierarchy; the benefit on CPU is limited.
- Older or unusual hardware. Some backends and older GPUs fall back to standard attention silently. If enabling the flag changes nothing, that is usually why.
- A few model architectures. Models with unconventional attention variants may not be covered by a given runtime's kernel yet.
It is safe to leave on: when it is unsupported, runtimes fall back rather than fail, and when it is supported the output is unchanged. For which engines implement what, see ExLlama vs vLLM vs llama.cpp and what is vLLM. Model and hardware figures across BestLLMfor are open through our public API (CC BY 4.0) and MCP server.
Frequently asked questions
Does Flash Attention reduce model quality?
No. It is an exact algorithm: it computes the same attention as the standard method, up to normal floating-point rounding. It is not an approximation like sparse or low-rank attention.
Does Flash Attention make token generation faster?
Only slightly. Generation speed is limited by reading model weights from memory, which Flash Attention does not change. Its main effect is on prompt processing and on peak memory with long contexts.
Should I enable Flash Attention in Ollama?
Yes, on a supported GPU. Set OLLAMA_FLASH_ATTENTION=1 and restart the server. It lowers memory use on long prompts and is required if you want to quantize the KV cache with OLLAMA_KV_CACHE_TYPE.
Does Flash Attention reduce VRAM usage?
It removes the large temporary attention matrix, which lowers peak VRAM during prompt processing. It does not shrink the model weights or the KV cache itself, though it enables KV-cache quantization, which does.
Does Flash Attention work on AMD GPUs and Macs?
In llama.cpp-based tools, yes: there are implementations for ROCm, Vulkan and Apple's Metal. The original research library targets NVIDIA GPUs, but local runtimes ship their own kernels.
What is the difference between Flash Attention and PagedAttention?
They solve different problems. Flash Attention speeds up the attention computation and removes its quadratic temporary memory. PagedAttention, from vLLM, manages how the KV cache is stored across many concurrent requests. Servers commonly use both.
A current option for local AI: GMKtec EVO-X2 64GB / 1TB (Ryzen AI Max+ 395). Match memory to your model and software. A mini PC is a complete PC alternative; Mac/MLX and CUDA instructions require compatible hardware.
Amazon Check GMKtec EVO-X2 64GB / 1TB (Ryzen AI Max+ 395) price →As an Amazon Associate, BestLLMfor earns from qualifying purchases, at no extra cost to you. It does not influence our independent rankings.
Found an error or have feedback? Let us know — it helps everyone who reads this guide.