Tool · instant calculation
Calculator for VRAM for LLMs
How much GPU memory does a local LLM really require? Weights + KV cache + overhead, in one honest estimate.
The rule: VRAM ≈ parameters (B) × bytes per weight + KV cache + ~20% overhead. In Q4, that's approximately 0.55–0.6 GB per billion parameters ; in FP16, 2 GB per billion. A longer context makes the KV cache grow well beyond 20%.
Estimate = weights × (1 + 0.2 × context/8K). Heuristic for a dense model; MoE (mixture of experts) models activate fewer weights per token and may perform better than their size suggests. Always leave 1–2 GB of headroom for the OS and display.
VRAM required by model size (8K context)
| Size | Q4_K_M | Q8_0 | FP16 | Fits on (Q4) |
|---|---|---|---|---|
| 7B | 4.9 GB | 9.0 GB | 16.8 GB | 8 GB cards (RTX 3050, 4060) |
| 9B | 6.3 GB | 11.6 GB | 21.6 GB | 8–12 GB cards |
| 14B | 9.7 GB | 18.0 GB | 33.6 GB | 12–16 GB cards |
| 27B | 18.8 GB | 34.7 GB | 64.8 GB | 24 GB cards (RTX 3090/4090) |
| 32B | 22.3 GB | 41.1 GB | 76.8 GB | 24 GB cards, short context |
| 70B | 48.7 GB | 89.9 GB | 168.0 GB | 48 GB+ or dual 24 GB GPUs |
Figures include a 20% margin for KV cache/overhead at an 8K context, rounded to the nearest tenth of a GB.
Now, choosing the model
Knowing your budget is only half the journey. The hardware configurator matches real models to your exact GPU or Mac; the local LLM leaderboard ranks all models that can run locally; and the tier rankings get straight to the point: best LLM for 8 GB of VRAM, RTX 4090, 16 GB Mac. Can't decide between cloud and local? Run the numbers with the cost calculator.
Frequently asked questions
How much VRAM per billion parameters?
About 2 GB per billion parameters in FP16, 1.1 GB per billion in Q8, and 0.55–0.6 GB per billion in Q4 quantization. Add ~20% for the KV cache and overhead at a standard 8K context, and more for a longer context.
Is 8 GB of VRAM enough for a local LLM?
Yes. In Q4, an 8–9B model (about 5–6 GB of weights) fits entirely on an 8 GB GPU, with room for context. See our ranking of the best LLMs for 8 GB of VRAM.
Can a 24 GB GPU (RTX 3090/4090) run a 32B model?
Yes, in Q4 a dense 32B model requires about 20-22 GB, which fits on a 24 GB card with a short context. For a long context, a 27B model is safer.
Does a Mac’s unified memory count as VRAM?
Apple Silicon Macs share a single unified memory pool between the CPU and GPU: a 32 GB Mac can therefore load larger models than a dedicated 16 GB GPU—but reserve 8–10 GB for macOS itself.