Intermediate 11 minRTX 50

Which LLM on RTX 5080 (16 GB) ?

Direct response

A RTX 5080 offers 16 GB of GDDR7 VRAM and 960 GB/s of bandwidth: it keeps models of up to 14 billion parameters in VRAM (Qwen3 14B, 9.3 GB) and gpt-oss 20B (14 GB), at 64 tokens per second on a 14B with a 16k context, according to Hardware Corner. A dense 27B (17 GB) and a 70B do not fit.

The RTX 5080 generates quickly, but with 16 GB it shares the memory limit of less expensive cards. This guide lists the models that actually fit, with their exact weights, throughput measured by Hardware Corner, context budget, and the real gap versus the 5070 Ti, 4080 Super, and 3090. You’ll also know when to move to a 24 GB card.

Choosing a machine? Our picks by budget →

By Mohamed Meguedmi·Update 2026-09-29·Tested on Windows, macOS, and Linux
Recommended hardware

For this setup: RTX 5080 16GB (GIGABYTE Gaming OC).

Compare all options by budget, from €800 to €3,500 →

On the go: which laptop for local AI →

Affiliate links — possible commission at no extra cost to you. As an Amazon Associate, BestLLMfor earns from qualifying purchases.

#RTX 5080 for a local LLM: what 16 GB enables

For a local LLM, a RTX 5080 earns its keep with 16 GB of GDDR7 and 960 GB/s of bandwidth: it generates quickly, but its capacity limits the models. According to the Ollama library, Qwen3 14B (9.3 GB) and gpt-oss 20B (14 GB) fit entirely in VRAM, while Qwen 3.5 27B (17 GB) and Qwen3 30B (19 GB) exceed the card’s capacity. Hardware Corner measures 64 tokens per second on Qwen3 14B with a 16,000-token context, and 105.6 on gpt-oss 20B with 128,000 tokens. It is therefore very well suited to models with 8 to 20 billion parameters, but not to a dense 27B.

Memory
16 GB of GDDR7 on a 256-bit bus, according to NVIDIA; 960 GB/s of bandwidth at 30 Gbit/s, according to Hardware Corner.
Compute
10,752 CUDA cores and 5th-generation Tensor Cores, according to NVIDIA.
Power consumption
360 W of graphics power, 850 W of required system power, and 88 °C maximum temperature, according to NVIDIA.
i
Hardware Corner's take
As of March 2026, the site considers the 5080 to offer poor price-to-capacity value for local LLMs: its 16 GB limits models to about 28 billion parameters with aggressive quantization, and it recommends the RTX 4080 instead at this memory capacity. This opinion is based on prices at the time: check the price tracking before deciding.

#Models that fit in 16 GB

The weights are those of the Ollama library files, checked on September 29, 2026. Headroom is what remains on 16 GB before accounting for the context, cache, and engine buffers.

Models for RTX 5080 (Ollama files, Q4 unless noted)
ModelFile OllamaGross marginVerdict
Granite 4.2 8B5.3 GB10.7 GBVery comfortable
Qwen 3.5 9B6.6 GB9.4 GBVery comfortable, text and image
Gemma 4 12B7.6 GB8.4 GBComfortable
Qwen3 14B9.3 GB6.7 GBComfortable, with a possible 32k context
gpt-oss 20B14 GB2 GBFits, with long context possible in MXFP4
Devstral Small 2 24B / Mistral Small 3.2 24B15 GB1 GBToo little headroom for a useful context
Qwen 3.5 27B17 GBShort by 1 GBDoesn't fit
Qwen3 30B19 GBShort by 3 GBDoesn't fit

The 14-billion model in Q4 is the largest comfortable dense model: it leaves nearly 7 GB for context. The 24-billion models leave only one GB and overflow as soon as the context grows: the card works with a 24B model as long as the conversation stays short, but not for an agent. gpt-oss 20B is a separate case: according to Ollama, the expert weights are quantized in MXFP4 at 4,25 bits per parameter, allowing it to run within 16 GB of memory, and Hardware Corner measures it at up to 128k context on this card.

A local RAG adds an embedding model: bge-m3 weighs 1.2 GB in Ollama, for a total of 7.8 GB with Qwen 3.5 9B and 8.8 GB with Gemma 4 12B. More than 7 GB remains for the context and the injected fragment database. A coding agent, on the other hand, requires a context of at least 64,000 tokens according to the Ollama documentation: on 16 GB, that means dropping to an 8- to 14-billion-parameter model and quantizing the cache instead of targeting a 24B.

#Measured throughput: speed and bandwidth ceiling

The figures come from Hardware Corner, which tests on llama.cpp with contexts ranging from 4k to 128k. These are not QuelLLM measurements.

Generation on RTX 5080, tokens per second (Hardware Corner)
Model (Q4_K, or MXFP4 for gpt-oss)4k context32k context128k context4k prompt processing
Qwen3 8B129,172,5not measured6 410,1
Qwen3 14B80,651,9not measured3 820,5
gpt-oss 20B (MXFP4)172,4149,3105,69 146,1

Generation is bandwidth-bound: the theoretical ceiling is approximately bandwidth divided by model size. For Qwen3 14B (9.3 GB), 960 ÷ 9.3 gives 103 tokens per second, and the 4k measurement, 80.6, reaches 78% of that. gpt-oss 20B exceeds its apparent ceiling (960 ÷ 14 = 69) at 172 tokens per second because a mixture-of-experts model reads only some of its weights for each token: the ceiling is calculated from the active weights, not the entire file.

Context affects speed. Qwen3 14B drops from 80.6 to 51.9 tokens per second between 4k and 32k (-36%), while gpt-oss 20B loses only 13% between the same points (172.4 to 149.3). If you work with long documents, the model choice matters more than the card.

#Context and KV cache: the real ceiling of 16 GB

Ollama sets the default context based on video memory: its documentation specifies 4k tokens with less than 24 GiB of VRAM and recommends at least 64,000 tokens for agents, web research, and coding tools. A RTX 5080 therefore receives 4k by default: an agent launched with the default settings quickly loses its history. Increasing the context consumes VRAM as a KV cache, and that cost comes out of the headroom shown in the table above.

The main lever is cache quantization: according to the Ollama FAQ, q8_0 uses about half the memory of f16, with very little loss of precision, and requires Flash Attention, which Ollama enables automatically when the card and model support it. On 16 GB, the priority order is simple: choose a model that leaves at least 5 to 6 GB of headroom, then quantize the cache, then increase the context in stages while monitoring ollama ps.

#Install Ollama on a RTX 5080

  1. 01
    Check the driver
    Run nvidia-smi in a terminal: the driver must be version 550 or later (551.61 on Windows), and total memory should show approximately 16 GB.
  2. 02
    Install Ollama
    Install Ollama from its official website. The local API listens on http://localhost:11434.
  3. 03
    Run a model
    Run ollama run qwen3:14b: the 9.3 GB download happens once. For a faster model, try ollama run gpt-oss:20b (14 GB).
  4. 04
    Control the distribution
    In a second terminal, run ollama ps: the Processor column should show 100% GPU. A CPU/GPU split means the model or context overflows.
  5. 05
    Adjust the context
    Set the context with OLLAMA_CONTEXT_LENGTH, then run ollama ps again. If there isn't enough headroom, set OLLAMA_FLASH_ATTENTION to 1 and OLLAMA_KV_CACHE_TYPE to q8_0 at startup.
Basic commands
ollama run qwen3:14b
ollama run gpt-oss:20b
ollama ps

# Exemple de la documentation Ollama : contexte de 64 000 tokens
OLLAMA_CONTEXT_LENGTH=64000 ollama serve

#Blackwell: MXFP4, NVFP4, and power limits

The 5080 provides hardware acceleration for floating-point 4-bit formats: Hardware Corner reports that it supports MXFP4 and NVFP4, and the benefit is visible with gpt-oss 20B, whose weights are distributed in MXFP4. NVIDIA describes NVFP4 as a 4-bit floating-point format introduced with Blackwell. NVFP4 support in tools (llama.cpp, Ollama, vLLM) is evolving quickly: check the installed version and the model card, and keep Q4_K_M as the safe choice.

Flash Attention 3 is limited to Hopper (H100), according to the official Flash Attention repository: on a 5080, version 2 or the implementation from Ollama and llama.cpp applies, and Ollama enables it automatically when available. To reduce power consumption, nvidia-smi -pl sets a power limit in watts; because the LLM primarily stresses memory, the performance loss is modest, but measure it on your model.

Limit power (administrator rights required)
sudo nvidia-smi -pl 300

#5080, 5070 Ti, 4080 Super, 3090: the measured gap

Relative generation by card (Hardware Corner, 5080 = 100%)
CardMemoryRelative generation
RTX 409024 GB123 %
RTX 508016 GB100 %
RTX 5070 Ti16 GB91 %
RTX 3090 Ti24 GB89 %
RTX 4080 Super16 GB82 %
RTX 309024 GB81 %

This table tells three stories. The 5070 Ti, also with 16 GB, reaches 91% of the 5080's speed: in pure throughput, the gap is narrow. The 3090, at 81%, is almost as fast with 24 GB, which changes the catalog of accessible models (Qwen 3.5 27B fits on it, but not on the 5080). Finally, the 4090 is 23% faster and also offers 24 GB. For pure LLM use, the question is therefore less about speed than capacity.

#Verdict: when the 5080 is worth choosing

Which decision fits your situation
SituationDecisionQuantified rationale
8- to 20-billion-parameter models, fast useSuitable 508064 t/s on a 14B, 105.6 t/s on gpt-oss 20B at 128k
You want a 27B dense modelChoose a 24 GB cardQwen 3.5 27B weighs 17 GB, while the 5080 has 16 GB
Tight budget for 16 GBCompare the 5070 Ti and 4080 Super91% and 82% of the 5080’s speed
You already have a 4080 or 4080 SuperKeep itSame 16 GB capacity
You already have a 5070 TiKeep it91% of the 5080's speed, same capacity

The 5080 is a good 16 GB card whose drawback for LLMs is that it has no more memory than its cheaper peers. Before buying, ask yourself one question: what is the largest dense model you want to run? Up to 14 billion, 16 GB is enough and the 5080's speed is noticeable; beyond that, capacity matters more than speed, and a used 24 GB card offers more models at similar speed. For current prices, check our tracker, which records the lowest price for each card twice a week.

#Frequently asked questions

FAQ
Which LLM should you install on a RTX 5080?+
Start with Qwen3 14B: its 9.3 GB Ollama file leaves nearly 7 GB of headroom and generates 80.6 tokens per second at 4k according to Hardware Corner. For more speed, gpt-oss 20B (14 GB) exceeds 170 tokens per second at 4k. For everyday use, Qwen 3.5 9B (6.6 GB) is the lightest.
Can the RTX 5080 run a 70B model?+
No. Llama 3.3 70B weighs 43 GB in Q4, nearly three times the card’s 16 GB. Splitting it across system memory causes throughput to collapse because a dense model rereads all its weights for every token. For a 70B, you need two 24 GB cards, a 5090 with more memory, or a unified-memory machine.
RTX 5080 or RTX 5070 Ti for a local LLM?+
Both have 16 GB, so they support the same models. According to Hardware Corner, the 5070 Ti reaches 91% of the 5080’s generation throughput. Compare the price difference: at equal capacity, a 9% speed difference rarely justifies a significant premium. The price tracker shows current prices.
Is 16 GB enough for an LLM in 2026?+
Yes for models with 8 to 14 billion parameters and for gpt-oss 20B, which cover chat, light coding, and RAG. No for the 27 to 32 billion models in Q4, which weigh 17 to 20 GB. The 24-billion-parameter models (15 GB) fit, but leave no room for a long context.
What power supply for a RTX 5080?+
NVIDIA specifies a required system power supply of 850 W for 360 W of graphics power. These are the manufacturer's values for a complete PC. A power limit reduces the card's power consumption at the cost of a speed reduction that must be measured on your model, without changing the recommended power supply.
Does the RTX 5080 support FP4?+
Yes: according to Hardware Corner, it supports MXFP4 and NVFP4 in hardware. The gain is visible on gpt-oss 20B, shipped in MXFP4 and measured at 172 tokens per second at 4k. For other models, Q4_K_M remains the most common format: FP4 depends on the availability of the converted file.
Did this guide help you?

Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.