Which LLM on RTX 5080 (16 GB) ?
A RTX 5080 offers 16 GB of GDDR7 VRAM and 960 GB/s of bandwidth: it keeps models of up to 14 billion parameters in VRAM (Qwen3 14B, 9.3 GB) and gpt-oss 20B (14 GB), at 64 tokens per second on a 14B with a 16k context, according to Hardware Corner. A dense 27B (17 GB) and a 70B do not fit.
The RTX 5080 generates quickly, but with 16 GB it shares the memory limit of less expensive cards. This guide lists the models that actually fit, with their exact weights, throughput measured by Hardware Corner, context budget, and the real gap versus the 5070 Ti, 4080 Super, and 3090. You’ll also know when to move to a 24 GB card.
Choosing a machine? Our picks by budget →
For this setup: RTX 5080 16GB (GIGABYTE Gaming OC).
Compare all options by budget, from €800 to €3,500 →
On the go: which laptop for local AI →
Affiliate links — possible commission at no extra cost to you. As an Amazon Associate, BestLLMfor earns from qualifying purchases.
#RTX 5080 for a local LLM: what 16 GB enables
For a local LLM, a RTX 5080 earns its keep with 16 GB of GDDR7 and 960 GB/s of bandwidth: it generates quickly, but its capacity limits the models. According to the Ollama library, Qwen3 14B (9.3 GB) and gpt-oss 20B (14 GB) fit entirely in VRAM, while Qwen 3.5 27B (17 GB) and Qwen3 30B (19 GB) exceed the card’s capacity. Hardware Corner measures 64 tokens per second on Qwen3 14B with a 16,000-token context, and 105.6 on gpt-oss 20B with 128,000 tokens. It is therefore very well suited to models with 8 to 20 billion parameters, but not to a dense 27B.
- Memory
- 16 GB of GDDR7 on a 256-bit bus, according to NVIDIA; 960 GB/s of bandwidth at 30 Gbit/s, according to Hardware Corner.
- Compute
- 10,752 CUDA cores and 5th-generation Tensor Cores, according to NVIDIA.
- Power consumption
- 360 W of graphics power, 850 W of required system power, and 88 °C maximum temperature, according to NVIDIA.
#Models that fit in 16 GB
The weights are those of the Ollama library files, checked on September 29, 2026. Headroom is what remains on 16 GB before accounting for the context, cache, and engine buffers.
| Model | File Ollama | Gross margin | Verdict |
|---|---|---|---|
| Granite 4.2 8B | 5.3 GB | 10.7 GB | Very comfortable |
| Qwen 3.5 9B | 6.6 GB | 9.4 GB | Very comfortable, text and image |
| Gemma 4 12B | 7.6 GB | 8.4 GB | Comfortable |
| Qwen3 14B | 9.3 GB | 6.7 GB | Comfortable, with a possible 32k context |
| gpt-oss 20B | 14 GB | 2 GB | Fits, with long context possible in MXFP4 |
| Devstral Small 2 24B / Mistral Small 3.2 24B | 15 GB | 1 GB | Too little headroom for a useful context |
| Qwen 3.5 27B | 17 GB | Short by 1 GB | Doesn't fit |
| Qwen3 30B | 19 GB | Short by 3 GB | Doesn't fit |
The 14-billion model in Q4 is the largest comfortable dense model: it leaves nearly 7 GB for context. The 24-billion models leave only one GB and overflow as soon as the context grows: the card works with a 24B model as long as the conversation stays short, but not for an agent. gpt-oss 20B is a separate case: according to Ollama, the expert weights are quantized in MXFP4 at 4,25 bits per parameter, allowing it to run within 16 GB of memory, and Hardware Corner measures it at up to 128k context on this card.
A local RAG adds an embedding model: bge-m3 weighs 1.2 GB in Ollama, for a total of 7.8 GB with Qwen 3.5 9B and 8.8 GB with Gemma 4 12B. More than 7 GB remains for the context and the injected fragment database. A coding agent, on the other hand, requires a context of at least 64,000 tokens according to the Ollama documentation: on 16 GB, that means dropping to an 8- to 14-billion-parameter model and quantizing the cache instead of targeting a 24B.
#Measured throughput: speed and bandwidth ceiling
The figures come from Hardware Corner, which tests on llama.cpp with contexts ranging from 4k to 128k. These are not QuelLLM measurements.
| Model (Q4_K, or MXFP4 for gpt-oss) | 4k context | 32k context | 128k context | 4k prompt processing |
|---|---|---|---|---|
| Qwen3 8B | 129,1 | 72,5 | not measured | 6 410,1 |
| Qwen3 14B | 80,6 | 51,9 | not measured | 3 820,5 |
| gpt-oss 20B (MXFP4) | 172,4 | 149,3 | 105,6 | 9 146,1 |
Generation is bandwidth-bound: the theoretical ceiling is approximately bandwidth divided by model size. For Qwen3 14B (9.3 GB), 960 ÷ 9.3 gives 103 tokens per second, and the 4k measurement, 80.6, reaches 78% of that. gpt-oss 20B exceeds its apparent ceiling (960 ÷ 14 = 69) at 172 tokens per second because a mixture-of-experts model reads only some of its weights for each token: the ceiling is calculated from the active weights, not the entire file.
Context affects speed. Qwen3 14B drops from 80.6 to 51.9 tokens per second between 4k and 32k (-36%), while gpt-oss 20B loses only 13% between the same points (172.4 to 149.3). If you work with long documents, the model choice matters more than the card.
#Context and KV cache: the real ceiling of 16 GB
Ollama sets the default context based on video memory: its documentation specifies 4k tokens with less than 24 GiB of VRAM and recommends at least 64,000 tokens for agents, web research, and coding tools. A RTX 5080 therefore receives 4k by default: an agent launched with the default settings quickly loses its history. Increasing the context consumes VRAM as a KV cache, and that cost comes out of the headroom shown in the table above.
The main lever is cache quantization: according to the Ollama FAQ, q8_0 uses about half the memory of f16, with very little loss of precision, and requires Flash Attention, which Ollama enables automatically when the card and model support it. On 16 GB, the priority order is simple: choose a model that leaves at least 5 to 6 GB of headroom, then quantize the cache, then increase the context in stages while monitoring ollama ps.
#Install Ollama on a RTX 5080
- 01Check the driverRun nvidia-smi in a terminal: the driver must be version 550 or later (551.61 on Windows), and total memory should show approximately 16 GB.
- 02Install OllamaInstall Ollama from its official website. The local API listens on http://localhost:11434.
- 03Run a modelRun ollama run qwen3:14b: the 9.3 GB download happens once. For a faster model, try ollama run gpt-oss:20b (14 GB).
- 04Control the distributionIn a second terminal, run ollama ps: the Processor column should show 100% GPU. A CPU/GPU split means the model or context overflows.
- 05Adjust the contextSet the context with OLLAMA_CONTEXT_LENGTH, then run ollama ps again. If there isn't enough headroom, set OLLAMA_FLASH_ATTENTION to 1 and OLLAMA_KV_CACHE_TYPE to q8_0 at startup.
#Blackwell: MXFP4, NVFP4, and power limits
The 5080 provides hardware acceleration for floating-point 4-bit formats: Hardware Corner reports that it supports MXFP4 and NVFP4, and the benefit is visible with gpt-oss 20B, whose weights are distributed in MXFP4. NVIDIA describes NVFP4 as a 4-bit floating-point format introduced with Blackwell. NVFP4 support in tools (llama.cpp, Ollama, vLLM) is evolving quickly: check the installed version and the model card, and keep Q4_K_M as the safe choice.
Flash Attention 3 is limited to Hopper (H100), according to the official Flash Attention repository: on a 5080, version 2 or the implementation from Ollama and llama.cpp applies, and Ollama enables it automatically when available. To reduce power consumption, nvidia-smi -pl sets a power limit in watts; because the LLM primarily stresses memory, the performance loss is modest, but measure it on your model.
#5080, 5070 Ti, 4080 Super, 3090: the measured gap
| Card | Memory | Relative generation |
|---|---|---|
| RTX 4090 | 24 GB | 123 % |
| RTX 5080 | 16 GB | 100 % |
| RTX 5070 Ti | 16 GB | 91 % |
| RTX 3090 Ti | 24 GB | 89 % |
| RTX 4080 Super | 16 GB | 82 % |
| RTX 3090 | 24 GB | 81 % |
This table tells three stories. The 5070 Ti, also with 16 GB, reaches 91% of the 5080's speed: in pure throughput, the gap is narrow. The 3090, at 81%, is almost as fast with 24 GB, which changes the catalog of accessible models (Qwen 3.5 27B fits on it, but not on the 5080). Finally, the 4090 is 23% faster and also offers 24 GB. For pure LLM use, the question is therefore less about speed than capacity.
#Verdict: when the 5080 is worth choosing
| Situation | Decision | Quantified rationale |
|---|---|---|
| 8- to 20-billion-parameter models, fast use | Suitable 5080 | 64 t/s on a 14B, 105.6 t/s on gpt-oss 20B at 128k |
| You want a 27B dense model | Choose a 24 GB card | Qwen 3.5 27B weighs 17 GB, while the 5080 has 16 GB |
| Tight budget for 16 GB | Compare the 5070 Ti and 4080 Super | 91% and 82% of the 5080’s speed |
| You already have a 4080 or 4080 Super | Keep it | Same 16 GB capacity |
| You already have a 5070 Ti | Keep it | 91% of the 5080's speed, same capacity |
The 5080 is a good 16 GB card whose drawback for LLMs is that it has no more memory than its cheaper peers. Before buying, ask yourself one question: what is the largest dense model you want to run? Up to 14 billion, 16 GB is enough and the 5080's speed is noticeable; beyond that, capacity matters more than speed, and a used 24 GB card offers more models at similar speed. For current prices, check our tracker, which records the lowest price for each card twice a week.
- Source: NVIDIA, RTX 5080 specifications
- Source: Hardware Corner, LLM benchmarks on RTX 5080
- Source: Ollama, gpt-oss model page
- Source: Ollama documentation, context length
- Source: official Flash Attention repository
#Frequently asked questions
Which LLM should you install on a RTX 5080?+
Can the RTX 5080 run a 70B model?+
RTX 5080 or RTX 5070 Ti for a local LLM?+
Is 16 GB enough for an LLM in 2026?+
What power supply for a RTX 5080?+
Does the RTX 5080 support FP4?+
Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.