Which LLM on RTX 4080 / 4080 Super (16 GB) ?
The RTX 4080 and the 4080 Super both have 16 GB of GDDR6X (716.8 and 736 GB/s): they load all models up to about 14 GB into VRAM, including Gemma 4 12B, gpt-oss 20B, and Mistral Small 24B. A public third-party measurement records 106 tokens/s on an 8B model in Q4 with a RTX 4080. The bandwidth difference between the two variants is about 3%; it does not justify a significant premium.
The RTX 4080 and the 4080 Super are the high-end 16 GB cards of the Ada generation. For a local LLM, they differ less from each other than from their same-capacity competitors: RTX 4070 Ti Super, RTX 5070 Ti, RTX 5080. This guide quantifies the real gap between the two variants, what 16 GB can handle, and the only published throughput measurement available for this card.
Choosing a machine? Our picks by budget →
Buying alternative for this guide: RTX 5080 16GB (GIGABYTE Gaming OC).
Compare all options by budget, from €800 to €3,500 →
On the go: which laptop for local AI →
Affiliate links — possible commission at no extra cost to you. As an Amazon Associate, BestLLMfor earns from qualifying purchases.
#RTX 4080 for a local LLM: what 16 GB and 716 GB/s enable
A RTX 4080 or 4080 Super loads any 4B to 24B model in Q4 into VRAM: Gemma 4 12B (7.7 to 8 GB), gpt-oss 20B (14 GB), and Mistral Small 24B (14 GB) fit, with the latter leaving about 2 GB of headroom for context. The third-party measurement cited below records 106 tokens/s generating on an 8B model in Q4 with a RTX 4080, putting the card in the fast range for interactive use. The limitation is capacity: Qwen 3.5 27B weighs 17 GB according to Ollama and doesn't fit. Models of this size require 24 GB of VRAM. The 4080 is therefore a speed card for medium-sized models, not a capacity card.
- Architecture
- Ada Lovelace, AD103 chip.
- Memory
- 16 GB of GDDR6X on a 256-bit bus: 716.8 GB/s (4080) and 736 GB/s (4080 Super), according to Wikipedia.
- CUDA cores
- 9,728 (4080) and 10,240 (4080 Super), according to NVIDIA.
- Power
- 320 W of total graphics power across both variants; NVIDIA specifies a 750 W minimum power supply for the PC.
- Launch price
- $1,199 in November 2022 for the 4080, then $999 for the 4080 Super in January 2024.
#4080 or 4080 Super: the real difference for LLMs
The Super adds 5% more CUDA cores (10,240 versus 9,728) and 2.7% more bandwidth (736 versus 716.8 GB/s). For text generation, only bandwidth really matters, so the theoretical gain is under 3%—invisible in practice. Both cards have the same 16 GB capacity and the same 320 W power consumption, so they support the same loaded models. A 4080 Super is worth paying extra for only if it’s sold at the same price or nearly so. Used prices change every week: check the tracker instead of relying on a fixed figure.
| Criterion | RTX 4080 | RTX 4080 Super |
|---|---|---|
| CUDA cores | 9 728 | 10 240 |
| Memory | 16 GB GDDR6X, 256-bit | 16 GB GDDR6X, 256-bit |
| Bandwidth | 716.8 GB/s | 736 GB/s |
| Graphics power | 320 W | 320 W |
| Minimum power supply | 750 W | 750 W |
| Bandwidth gap | reference | about 2.7% more |
#Which models fit in 16 GB
| Model | Size | Headroom on 16 GB | Verdict |
|---|---|---|---|
| Granite 4.2 8B | 5.3 GB | 10.7 GB | Very long context possible |
| Gemma 4 12B | 7.7 to 8.0 GB | 8 GB | Comfortable |
| gpt-oss 20B | 14 GB | 2 GB | Fits, moderate context |
| Mistral Small 24B | 14 GB | 2 GB | Fits, moderate context |
| Qwen 3.5 27B | 17 GB | Negative | Does not fit in VRAM |
Ollama uses a context of 4,096 tokens by default. With 2 GB of headroom, a 14 GB model supports a longer context provided you quantize the K/V cache: Ollama offers the q8_0 format, which uses about half the memory of f16, with Flash Attention. These two settings are the key to using gpt-oss 20B or Mistral Small 24B with a 16,000-token context.
#The only published measurement: what it says about the 4080
A public GitHub repository compares GPUs with llama.cpp on LLaMA 3 (snapshot from May 2024, Ubuntu system, RunPod). For the RTX 4080 16 GB, it records 106.22 tokens/s during generation on 1,024 tokens with an 8B model in Q4_K_M, and 40.29 tokens/s with the same model in F16, using 14.96 GB—nearly the entire card. These are third-party measurements on an older model than those featured on the site, using an earlier version of llama.cpp: treat them as an order of magnitude, not a promise.
| Step | 512 tokens | 1,024 tokens | 8,192 tokens |
|---|---|---|---|
| Generation (tokens/s) | 108,15 | 106,22 | 93,71 |
| Prompt processing (tokens/s) | 5 389,74 | 5 064,99 | 2 882,03 |
Two takeaways. First, generation slows with context: going from 512 to 8,192 context tokens costs about 13% in this measurement because the K/V cache is added to the weights read for each token. Second, we can calculate the ratio between the measurement and the theoretical ceiling: 716.8 GB/s divided by the 4.58 GB used by the 8B model in Q4 gives 156 tokens/s; the measurement reaches 68% of that. In F16, the ceiling is 48 tokens/s and the measurement reaches 84%. The ratio therefore ranges from 68% to 84%, depending on the format.
| Model (Q4) | Weights | Cap | Estimated at 70% |
|---|---|---|---|
| Granite 4.2 8B | 5.3 GB | 135 tokens/s | about 95 tokens/s |
| Gemma 4 12B | 7.7 GB | 93 tokens/s | approximately 65 tokens/s |
| Mistral Small 24B | 14 GB | 51 tokens/s | about 36 tokens/s |
These estimates apply to a short context. The gpt-oss 20B model reads only its active experts for each token: it exceeds a 14 GB dense model, but no reliable measured value is reported here.
#Reading the prompt is much faster than generating
The same measurement also covers prompt processing—that is, reading your question and documents before the first response: 5,065 tokens/s for 1,024 tokens and 2,882 tokens/s for 8,192 tokens. A document of about 8,000 tokens, roughly fifteen pages, is therefore read in about three seconds before generation starts. This step is limited by compute rather than bandwidth, and this is where the 4080's extra cores show compared with a 288 GB/s card such as the 4060 Ti.
#Use cases that benefit from a 4080: agents, RAG, code
A card at this speed is suitable for workloads that chain many generations: agents, code-correction loops, and batch document processing. At 100 tokens/s with 8B, an agent producing 500 tokens per step responds in a few seconds. The 16 GB, however, limits context size and the number of loaded models: keep a single 12B model at 14 GB, plus a small embedding model for RAG. A reasoning model such as gpt-oss 20B produces lengthy thoughts before its answer, consuming context.
| Usage | Configuration | Key consideration |
|---|---|---|
| Code assistant | 8B to 12B model, 16,000-token context, q8_0 cache | Code models of 15 GB and up: virtually no context |
| RAG on documents | Gemma 4 12B and lightweight embeddings | Load only one generation model |
| Quality chat | Mistral Small 24B alone | Moderate context, 2 GB of headroom |
#The model slows down: check that it fits properly in VRAM
On 16 GB, a 14 GB model leaves little headroom, and a growing context can make it spill over during the conversation. The symptom is generation that suddenly becomes slower: part of the model has moved into system RAM. This check takes one minute and prevents you from wrongly concluding that the card is slow when the context setting is actually to blame.
- 01Observe placementRun ollama ps in a second terminal: the PROCESSOR column should show 100% GPU. A CPU/GPU split indicates overflow.
- 02Reduce the contextLower OLLAMA_CONTEXT_LENGTH to 8192, or to its default value of 4096, then restart the server.
- 03Quantize the cacheEnable Flash Attention and OLLAMA_KV_CACHE_TYPE=q8_0: the K/V cache uses about half the memory of f16.
- 04Free up VRAMClose applications using the GPU—games and browsers with hardware acceleration—and check with nvidia-smi.
- 05Change modelsIf nothing is enough, switch to a smaller model, for example Gemma 4 12B instead of Mistral Small 24B.
#4080 versus the 4070 Ti Super, 5070 Ti, and 5080
| Card | VRAM | Bandwidth | Difference from the 4080 |
|---|---|---|---|
| RTX 4070 Ti Super | 16 GB GDDR6X | 672 GB/s | about 6% less |
| RTX 4080 | 16 GB GDDR6X | 716.8 GB/s | reference |
| RTX 4080 Super | 16 GB GDDR6X | 736 GB/s | about 3% more |
| RTX 5070 Ti | 16 GB GDDR7 | 896 GB/s | about 25% more |
| RTX 5080 | 16 GB GDDR7 | 960 GB/s | about 34% more |
All these cards load the same models because their capacity is identical. Only throughput changes, proportionally to bandwidth. The 4080 is only 6% faster than the 4070 Ti Super: if the price difference is significant, the Ti Super is the better buy. Compared with 50-series cards, the 4080 is 20 to 25% slower but remains unmatched if the used price is low. This tradeoff comes down to price, not the spec sheet.
- Which LLM on RTX 4070 Ti Super
- Which LLM on RTX 5070 Ti
- Which LLM on RTX 5080
- Local AI graphics card prices
#2026 verdict: keep, buy, or upgrade
- Keep your 4080 if
- Your models fit in 16 GB. It remains fast, and there is no reason to replace it before you need 24 GB.
- Buy it used if
- Its price is significantly lower than that of a new 5070 Ti, whose bandwidth is about 25% higher.
- Aim for 24 GB if
- You want 27B or larger models: no setting can fit 17 GB into 16 GB with context.
A 2022 or 2023 card may have been used for mining or intensive gaming: ask for the original invoice if available, and favor a seller that accepts returns. Before buying used, use nvidia-smi to verify that the card reports 16 GB, run a 14 GB model, and monitor the temperature during a long generation. The card draws up to 320 W: check the power supply; NVIDIA recommends at least 750 W.
- Source: official NVIDIA specifications (RTX 4080 and 4080 Super)
- Source: third-party llama.cpp throughput measurements on GPU (GitHub)
- Source: Ollama's official FAQ (Flash Attention, K/V cache)
#Frequently asked questions
What’s the difference between RTX 4080 and the 4080 Super for an LLM?+
What generation speed can you expect on a RTX 4080?+
Can the RTX 4080 run Qwen 3.5 27B?+
RTX 4080 Super or RTX 4070 Ti Great for an LLM?+
What power supply do you need for a RTX 4080?+
Is a used RTX 4080 worth it for an LLM?+
Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.