Intermediate 11 minRTX 40

Which LLM on RTX 4080 / 4080 Super (16 GB) ?

Direct response

The RTX 4080 and the 4080 Super both have 16 GB of GDDR6X (716.8 and 736 GB/s): they load all models up to about 14 GB into VRAM, including Gemma 4 12B, gpt-oss 20B, and Mistral Small 24B. A public third-party measurement records 106 tokens/s on an 8B model in Q4 with a RTX 4080. The bandwidth difference between the two variants is about 3%; it does not justify a significant premium.

The RTX 4080 and the 4080 Super are the high-end 16 GB cards of the Ada generation. For a local LLM, they differ less from each other than from their same-capacity competitors: RTX 4070 Ti Super, RTX 5070 Ti, RTX 5080. This guide quantifies the real gap between the two variants, what 16 GB can handle, and the only published throughput measurement available for this card.

Choosing a machine? Our picks by budget →

By Mohamed Meguedmi·Update 2026-09-30·Tested on Windows, macOS, and Linux
Recommended hardware

Buying alternative for this guide: RTX 5080 16GB (GIGABYTE Gaming OC).

Compare all options by budget, from €800 to €3,500 →

On the go: which laptop for local AI →

Affiliate links — possible commission at no extra cost to you. As an Amazon Associate, BestLLMfor earns from qualifying purchases.

#RTX 4080 for a local LLM: what 16 GB and 716 GB/s enable

A RTX 4080 or 4080 Super loads any 4B to 24B model in Q4 into VRAM: Gemma 4 12B (7.7 to 8 GB), gpt-oss 20B (14 GB), and Mistral Small 24B (14 GB) fit, with the latter leaving about 2 GB of headroom for context. The third-party measurement cited below records 106 tokens/s generating on an 8B model in Q4 with a RTX 4080, putting the card in the fast range for interactive use. The limitation is capacity: Qwen 3.5 27B weighs 17 GB according to Ollama and doesn't fit. Models of this size require 24 GB of VRAM. The 4080 is therefore a speed card for medium-sized models, not a capacity card.

Architecture
Ada Lovelace, AD103 chip.
Memory
16 GB of GDDR6X on a 256-bit bus: 716.8 GB/s (4080) and 736 GB/s (4080 Super), according to Wikipedia.
CUDA cores
9,728 (4080) and 10,240 (4080 Super), according to NVIDIA.
Power
320 W of total graphics power across both variants; NVIDIA specifies a 750 W minimum power supply for the PC.
Launch price
$1,199 in November 2022 for the 4080, then $999 for the 4080 Super in January 2024.

#4080 or 4080 Super: the real difference for LLMs

The Super adds 5% more CUDA cores (10,240 versus 9,728) and 2.7% more bandwidth (736 versus 716.8 GB/s). For text generation, only bandwidth really matters, so the theoretical gain is under 3%—invisible in practice. Both cards have the same 16 GB capacity and the same 320 W power consumption, so they support the same loaded models. A 4080 Super is worth paying extra for only if it’s sold at the same price or nearly so. Used prices change every week: check the tracker instead of relying on a fixed figure.

RTX 4080 and 4080 Super: the two profiles
CriterionRTX 4080RTX 4080 Super
CUDA cores9 72810 240
Memory16 GB GDDR6X, 256-bit16 GB GDDR6X, 256-bit
Bandwidth716.8 GB/s736 GB/s
Graphics power320 W320 W
Minimum power supply750 W750 W
Bandwidth gapreferenceabout 2.7% more

#Which models fit in 16 GB

Models on RTX 4080 (sizes Ollama, weights only)
ModelSizeHeadroom on 16 GBVerdict
Granite 4.2 8B5.3 GB10.7 GBVery long context possible
Gemma 4 12B7.7 to 8.0 GB8 GBComfortable
gpt-oss 20B14 GB2 GBFits, moderate context
Mistral Small 24B14 GB2 GBFits, moderate context
Qwen 3.5 27B17 GBNegativeDoes not fit in VRAM

Ollama uses a context of 4,096 tokens by default. With 2 GB of headroom, a 14 GB model supports a longer context provided you quantize the K/V cache: Ollama offers the q8_0 format, which uses about half the memory of f16, with Flash Attention. These two settings are the key to using gpt-oss 20B or Mistral Small 24B with a 16,000-token context.

Ollama server settings (Linux, macOS)
export OLLAMA_FLASH_ATTENTION=1
export OLLAMA_KV_CACHE_TYPE=q8_0
export OLLAMA_CONTEXT_LENGTH=16384
ollama serve

#The only published measurement: what it says about the 4080

A public GitHub repository compares GPUs with llama.cpp on LLaMA 3 (snapshot from May 2024, Ubuntu system, RunPod). For the RTX 4080 16 GB, it records 106.22 tokens/s during generation on 1,024 tokens with an 8B model in Q4_K_M, and 40.29 tokens/s with the same model in F16, using 14.96 GB—nearly the entire card. These are third-party measurements on an older model than those featured on the site, using an earlier version of llama.cpp: treat them as an order of magnitude, not a promise.

RTX 4080 16 GB: measurements from the third-party benchmark (LLaMA 3 8B in Q4_K_M, llama.cpp)
Step512 tokens1,024 tokens8,192 tokens
Generation (tokens/s)108,15106,2293,71
Prompt processing (tokens/s)5 389,745 064,992 882,03

Two takeaways. First, generation slows with context: going from 512 to 8,192 context tokens costs about 13% in this measurement because the K/V cache is added to the weights read for each token. Second, we can calculate the ratio between the measurement and the theoretical ceiling: 716.8 GB/s divided by the 4.58 GB used by the 8B model in Q4 gives 156 tokens/s; the measurement reaches 68% of that. In F16, the ceiling is 48 tokens/s and the measurement reaches 84%. The ratio therefore ranges from 68% to 84%, depending on the format.

Theoretical ceiling and estimate on RTX 4080 (716.8 GB/s)
Model (Q4)WeightsCapEstimated at 70%
Granite 4.2 8B5.3 GB135 tokens/sabout 95 tokens/s
Gemma 4 12B7.7 GB93 tokens/sapproximately 65 tokens/s
Mistral Small 24B14 GB51 tokens/sabout 36 tokens/s

These estimates apply to a short context. The gpt-oss 20B model reads only its active experts for each token: it exceeds a 14 GB dense model, but no reliable measured value is reported here.

#Reading the prompt is much faster than generating

The same measurement also covers prompt processing—that is, reading your question and documents before the first response: 5,065 tokens/s for 1,024 tokens and 2,882 tokens/s for 8,192 tokens. A document of about 8,000 tokens, roughly fifteen pages, is therefore read in about three seconds before generation starts. This step is limited by compute rather than bandwidth, and this is where the 4080's extra cores show compared with a 288 GB/s card such as the 4060 Ti.

#Use cases that benefit from a 4080: agents, RAG, code

A card at this speed is suitable for workloads that chain many generations: agents, code-correction loops, and batch document processing. At 100 tokens/s with 8B, an agent producing 500 tokens per step responds in a few seconds. The 16 GB, however, limits context size and the number of loaded models: keep a single 12B model at 14 GB, plus a small embedding model for RAG. A reasoning model such as gpt-oss 20B produces lengthy thoughts before its answer, consuming context.

Three typical configurations on RTX 4080
UsageConfigurationKey consideration
Code assistant8B to 12B model, 16,000-token context, q8_0 cacheCode models of 15 GB and up: virtually no context
RAG on documentsGemma 4 12B and lightweight embeddingsLoad only one generation model
Quality chatMistral Small 24B aloneModerate context, 2 GB of headroom

#The model slows down: check that it fits properly in VRAM

On 16 GB, a 14 GB model leaves little headroom, and a growing context can make it spill over during the conversation. The symptom is generation that suddenly becomes slower: part of the model has moved into system RAM. This check takes one minute and prevents you from wrongly concluding that the card is slow when the context setting is actually to blame.

  1. 01
    Observe placement
    Run ollama ps in a second terminal: the PROCESSOR column should show 100% GPU. A CPU/GPU split indicates overflow.
  2. 02
    Reduce the context
    Lower OLLAMA_CONTEXT_LENGTH to 8192, or to its default value of 4096, then restart the server.
  3. 03
    Quantize the cache
    Enable Flash Attention and OLLAMA_KV_CACHE_TYPE=q8_0: the K/V cache uses about half the memory of f16.
  4. 04
    Free up VRAM
    Close applications using the GPU—games and browsers with hardware acceleration—and check with nvidia-smi.
  5. 05
    Change models
    If nothing is enough, switch to a smaller model, for example Gemma 4 12B instead of Mistral Small 24B.

#4080 versus the 4070 Ti Super, 5070 Ti, and 5080

16 GB cards: the same capacity, different bandwidths
CardVRAMBandwidthDifference from the 4080
RTX 4070 Ti Super16 GB GDDR6X672 GB/sabout 6% less
RTX 408016 GB GDDR6X716.8 GB/sreference
RTX 4080 Super16 GB GDDR6X736 GB/sabout 3% more
RTX 5070 Ti16 GB GDDR7896 GB/sabout 25% more
RTX 508016 GB GDDR7960 GB/sabout 34% more

All these cards load the same models because their capacity is identical. Only throughput changes, proportionally to bandwidth. The 4080 is only 6% faster than the 4070 Ti Super: if the price difference is significant, the Ti Super is the better buy. Compared with 50-series cards, the 4080 is 20 to 25% slower but remains unmatched if the used price is low. This tradeoff comes down to price, not the spec sheet.

#2026 verdict: keep, buy, or upgrade

Keep your 4080 if
Your models fit in 16 GB. It remains fast, and there is no reason to replace it before you need 24 GB.
Buy it used if
Its price is significantly lower than that of a new 5070 Ti, whose bandwidth is about 25% higher.
Aim for 24 GB if
You want 27B or larger models: no setting can fit 17 GB into 16 GB with context.

A 2022 or 2023 card may have been used for mining or intensive gaming: ask for the original invoice if available, and favor a seller that accepts returns. Before buying used, use nvidia-smi to verify that the card reports 16 GB, run a 14 GB model, and monitor the temperature during a long generation. The card draws up to 320 W: check the power supply; NVIDIA recommends at least 750 W.

#Frequently asked questions

FAQ
What’s the difference between RTX 4080 and the 4080 Super for an LLM?+
Almost none. The Super has 5% more CUDA cores and 2.7% more bandwidth (736 versus 716.8 GB/s), with the same 16 GB of memory and the same 320 W power draw. The loaded models are the same, and the speed difference, under 3%, is not noticeable in practice.
What generation speed can you expect on a RTX 4080?+
A public third-party benchmark measured 106.22 tokens/s with LLaMA 3 8B in Q4_K_M under llama.cpp in May 2024. It was not measured here, and the model is older than those on the site. The theoretical ceiling for a 5.3 GB model is approximately 135 tokens/s.
Can the RTX 4080 run Qwen 3.5 27B?+
Not entirely in VRAM: the model weighs 17 GB according to Ollama, while the card has 16 GB. Some of the weights would remain in system RAM, which significantly slows generation. This model requires 24 GB of VRAM. On 16 GB, choose Mistral Small 24B or gpt-oss 20B, which weigh 14 GB.
RTX 4080 Super or RTX 4070 Ti Great for an LLM?+
Both have 16 GB, so they support the same models. The 4080 Super reads its memory at 736 GB/s versus 672 GB/s, about 10% more: that is the maximum speed gain during generation. Choose the 4070 Ti Super if its used price is significantly lower.
What power supply do you need for a RTX 4080?+
NVIDIA specifies 320 W of total graphics power and a 750 W minimum power supply for the entire PC. Check your card model's connectors. A quality 750 W power supply is suitable; higher wattage leaves headroom for a powerful processor.
Is a used RTX 4080 worth it for an LLM?+
Yes, if its price is significantly lower than that of a new RTX 5070 Ti, which has the same capacity and about 25% higher bandwidth. Use nvidia-smi to verify that the card reports 16 GB, test a long generation, and check the temperature before buying.
Did this guide help you?

Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.