Which LLM on RTX 3080 / 3080 Ti (10–12 GB) ?
For a local LLM, a RTX 3080 or 3080 Ti with 12 GB can run Qwen 3.5 9B in Q8 (10 GB) and Gemma 4 12B in Q4 (7 GB) with room to spare, at around 140 tokens/s on a 7B: the 3080 10 GB scores 139.7 on the llama.cpp test. The 10 GB version is limited to 8–9B models in Q4. None of the three can load a 20B model in Q4 (13 GB) without spilling into RAM.
Three cards carry the name RTX 3080: the 10 GB version from September 2020, the 12 GB version from January 2022, and the 12 GB 3080 Ti from June 2021. They use the same chip, but not the same memory, and memory determines what you can run. This guide relies on a public speed measurement and official specifications for capacity.
Choosing a machine? Our picks by budget →
Buying alternative for this guide: RTX 5080 16GB (GIGABYTE Gaming OC).
Compare all options by budget, from €800 to €3,500 →
On the go: which laptop for local AI →
Affiliate links — possible commission at no extra cost to you. As an Amazon Associate, BestLLMfor earns from qualifying purchases.
#RTX 3080 and 3080 Ti for a local LLM: the three variants
A RTX 3080 or a 3080 Ti with 12 GB comfortably runs a 12-billion-parameter model in Q4 or a 9-billion-parameter model in Q8; the 10 GB version is limited to 7- to 9-billion-parameter models in Q4. According to the QuelLLM catalog, Gemma 4 12B weighs 7 GB in Q4 and 9 GB in Q5; Qwen 3.5 9B weighs 6 GB in Q4 and 10 GB in Q8, requiring at least 12 GB once context is added. In terms of speed, these cards are among the fastest of previous generations: the 10 GB 3080 generates 139.7 tokens per second on a 7-billion-parameter model in Q4_0 according to the public llama.cpp table, nearly twice the RTX 3060 12 GB (75.6). Choosing a RTX 3080 for an LLM therefore comes down to memory, never speed.
- RTX 3080 10 GB
- September 2020, 8,704 CUDA cores, 10 GB of GDDR6X on a 320-bit bus, for 760 GB/s. Board power: 320 W.
- RTX 3080 12 GB
- January 2022, 8,960 cores, 12 GB of GDDR6X on a 384-bit bus, or 912 GB/s, 350 W. About 20% more bandwidth.
- RTX 3080 Ti
- June 2021, 10,240 cores, 12 GB of GDDR6X on a 384-bit bus like the 3080 12 GB, so the expected bandwidth is identical, 350 W.
The ceiling calculation helps explain it: text generation reads the weights for every token, so bandwidth determines the maximum speed. The 3080 12 GB and Ti gain 20% more bandwidth than the 10 GB model, but the 3080 Ti mainly adds compute cores, which speed up prompt processing more than generation.
#3080 10 GB, 12 GB, or 3080 Ti: which one to choose
| Criterion | 3080 10 GB | 3080 12 GB | 3080 Ti 12 GB |
|---|---|---|---|
| CUDA cores | 8 704 | 8 960 | 10 240 |
| Memory | 10 GB GDDR6X | 12 GB GDDR6X | 12 GB GDDR6X |
| Bus / bandwidth | 320 bits / 760 GB/s | 384 bits / 912 GB/s | 384 bits, same expected bandwidth |
| Card power | 320 W | 350 W | 350 W |
| Generation on Llama 2 7B Q4_0 | 139.7 tok/s (measured) | estimate: up to approximately 167 | estimate: close to 12 GB |
For local AI, the 3080 Ti and the 3080 12 GB are interchangeable: the same 12 GB, the same bus, and nearly identical generation speed. The estimate for these two cards is an upper bound that applies the measured efficiency of the 10 GB model to 912 GB/s; in practice, the 3090, which also has a 384-bit bus, measures 158.2 tokens per second. So choose the cheaper one, without being guided by the core count. The 10 GB version is one step down in capacity: its 10 GB prevents Q8 use with 9-billion-parameter models.
#Compatible models by memory capacity
| Model and quantization | Weights | 3080 10 GB | 3080 12 GB / Ti |
|---|---|---|---|
| Qwen 3.5 9B in Q4 | 6 GB | Yes, with context | Yes, with a lot of context |
| Gemma 4 12B in Q4 | 7 GB | Yes, moderate context | Yes, comfortable |
| Gemma 4 12B in Q5 | 9 GB | Exactly right | Yes, moderate context |
| Qwen 3.5 9B in Q8 | 10 GB | No | Yes, but with limited context |
| gpt-oss 20B in Q4 | 13 GB | No | No: exceeds by 1 GB |
Two points in this table deserve attention. First, the 3080 10 GB handles Gemma 4 12B in Q4 (7 GB), despite what you might think: 3 GB remains for context, which is enough for average conversations. Second, the gap between 12 and 16 GB is clear: a 20-billion-class model in Q4, such as gpt-oss 20B, weighs 13 GB and fits on none of the three cards. If that is your target, you need a 16 GB card or larger, and the guide to 12 GB of VRAM explains exactly what you can do with 12 GB.
#Installation, power, and settings
The RTX 3080 and 3080 Ti are among the compute-capability 8.6 cards listed by Ollama, which requires driver 550 or later. A specific point to watch with these cards: the NVIDIA specification lists 350 W for the card, 320 W for the 10 GB version, and recommends a 750 W system power supply. Check your power supply before installing the card, and connect its power connectors to separate cables.
- 01Check the power supplyMake sure it reaches 750 W and has the connectors required by your card model.
- 02Install the driver and OllamaInstall driver 550 or later, then Ollama, and launch a model of your choice.
- 03Adjust memoryOn the 10 GB version, start the server with OLLAMA_FLASH_ATTENTION=1 and OLLAMA_KV_CACHE_TYPE=q8_0: the documentation specifies that the K/V cache can be quantized when Flash Attention is enabled.
- 04ControlUse nvidia-smi to monitor memory and temperature under load: these cards run hotter than newer models.
#Speed: the public benchmark and its extrapolations
The collaborative llama.cpp table, which measures Llama 2 7B in Q4_0 (3.56 GiB file, or 3.82 GB), reports 139.65 tokens per second during generation and 5,014 during prompt processing for the RTX 3080 10 GB. With Flash Attention, generation remains at 139.95, but prompt processing rises to 5,570 tokens per second, about 11% better. The 760 GB/s bandwidth would theoretically allow nearly 199 tokens per second: the measurement reaches 70% of that ceiling. Throughput also depends on the driver, system, and card manufacturer, even with the same chip.
| Card | Memory | Generation (tok/s) |
|---|---|---|
| RTX 3060 12 GB | 12 GB | 75,6 |
| RTX 3080 10 GB | 10 GB | 139,7 |
| RTX 3090 | 24 GB | 158,2 |
Using a rule of three based on file sizes, the 3080 10 GB would deliver at most about 89 tokens per second on Qwen 3.5 9B in Q4 (6 GB) and about 76 on Gemma 4 12B in Q4 (7 GB). These are theoretical upper bounds, not measurements. On a 3080 12 GB, Qwen 3.5 9B's Q8 (10 GB) would theoretically remain below 65 tokens per second, which is comfortable. In all cases, these cards are well beyond the read speed; capacity, not speed, is the limit.
Prompt processing is the second advantage of these cards. At 5,014 tokens per second, the 3080 10 GB reads a 10,000-token document in about two seconds, and at 5,570 with Flash Attention, in slightly less time—two theoretical calculations that assume a constant throughput. That's about twice as fast as a 3060 12 GB (2,138 tokens per second). Prompt processing matters even more when several people share the same machine, since each request must first reread its own context before generating any response, and that's where the gap with a more modest card widens. For RAG over large documents or a coding assistant that rereads large files for every request, this is a real usability gain, much more noticeable than the generation gap, which is already well above reading speed.
#Buying in 2026: where these cards stand
Compared with a new 16 GB card, the argument for a used RTX 3080 is speed at a moderate price, but it lacks memory. The 3090 measures 158.2 tokens per second with 24 GB, making it the used-market anchor for anyone targeting 20- to 30-billion-parameter models. Used-card prices vary every week: the site’s tracker records the lowest price for each card and the price per GB of VRAM, which is more useful than a figure written here.
| Target model | Required memory | Suitable card |
|---|---|---|
| 8 to 9B in Q4 | 6 GB of weights | Any 8 GB card |
| 12B in Q4 or 9B in Q8 | 7 to 10 GB of weights | 3080 12 GB or Ti, or 3080 10 GB for Q4 |
| 20B in Q4 (gpt-oss 20B) | 13 GB of weights | 16 GB or more card |
- Local AI graphics card prices (weekly tracking)
- Which LLM on RTX 3090 / 3090 Ti (24 GB)?
- Which LLM on RTX 4070 Ti Super (16 GB)?
#2026 verdict
- Used 3080 12 GB or 3080 Ti
- A very good choice if the price remains well below that of a new 12 to 16 GB card: 12 GB, high speed, and enough headroom for a 9B Q8 or a 12B Q4.
- 3080 10 GB
- Useful for 8- to 12-billion-parameter models in Q4. Without the 9B's Q8 or any context headroom, a 3060 12 GB may be a better fit if you are looking for capacity rather than speed.
- At purchase
- Check the power supply, the 10 or 12 GB labeling, and the fans’ condition; these cards have run hot and have several years of use.
#Frequently asked questions
RTX 3080 10 GB or 12 GB for an LLM?+
RTX 3080 Ti or RTX 3090 for an LLM?+
Can the RTX 3080 12 GB run gpt-oss 20B?+
How many tokens per second on a RTX 3080?+
What power supply do you need for a RTX 3080 Ti?+
Does Ollama work on a RTX 3080?+
Is RTX 3080 still worth it in 2026?+
- Source: llama.cpp performance table on CUDA
- Source: NVIDIA spec sheet for the RTX 3080 and 3080 Ti
- Source: maps NVIDIA supported by Ollama
- Source: Ollama FAQ (Flash Attention, K/V cache)
Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.