Intermediate 11 minRTX 30

Which LLM on RTX 3080 / 3080 Ti (10–12 GB) ?

Direct response

For a local LLM, a RTX 3080 or 3080 Ti with 12 GB can run Qwen 3.5 9B in Q8 (10 GB) and Gemma 4 12B in Q4 (7 GB) with room to spare, at around 140 tokens/s on a 7B: the 3080 10 GB scores 139.7 on the llama.cpp test. The 10 GB version is limited to 8–9B models in Q4. None of the three can load a 20B model in Q4 (13 GB) without spilling into RAM.

Three cards carry the name RTX 3080: the 10 GB version from September 2020, the 12 GB version from January 2022, and the 12 GB 3080 Ti from June 2021. They use the same chip, but not the same memory, and memory determines what you can run. This guide relies on a public speed measurement and official specifications for capacity.

Choosing a machine? Our picks by budget →

By Mohamed Meguedmi·Update 2026-09-30·Tested on Windows, macOS, and Linux
Recommended hardware

Buying alternative for this guide: RTX 5080 16GB (GIGABYTE Gaming OC).

Compare all options by budget, from €800 to €3,500 →

On the go: which laptop for local AI →

Affiliate links — possible commission at no extra cost to you. As an Amazon Associate, BestLLMfor earns from qualifying purchases.

#RTX 3080 and 3080 Ti for a local LLM: the three variants

A RTX 3080 or a 3080 Ti with 12 GB comfortably runs a 12-billion-parameter model in Q4 or a 9-billion-parameter model in Q8; the 10 GB version is limited to 7- to 9-billion-parameter models in Q4. According to the QuelLLM catalog, Gemma 4 12B weighs 7 GB in Q4 and 9 GB in Q5; Qwen 3.5 9B weighs 6 GB in Q4 and 10 GB in Q8, requiring at least 12 GB once context is added. In terms of speed, these cards are among the fastest of previous generations: the 10 GB 3080 generates 139.7 tokens per second on a 7-billion-parameter model in Q4_0 according to the public llama.cpp table, nearly twice the RTX 3060 12 GB (75.6). Choosing a RTX 3080 for an LLM therefore comes down to memory, never speed.

RTX 3080 10 GB
September 2020, 8,704 CUDA cores, 10 GB of GDDR6X on a 320-bit bus, for 760 GB/s. Board power: 320 W.
RTX 3080 12 GB
January 2022, 8,960 cores, 12 GB of GDDR6X on a 384-bit bus, or 912 GB/s, 350 W. About 20% more bandwidth.
RTX 3080 Ti
June 2021, 10,240 cores, 12 GB of GDDR6X on a 384-bit bus like the 3080 12 GB, so the expected bandwidth is identical, 350 W.

The ceiling calculation helps explain it: text generation reads the weights for every token, so bandwidth determines the maximum speed. The 3080 12 GB and Ti gain 20% more bandwidth than the 10 GB model, but the 3080 Ti mainly adds compute cores, which speed up prompt processing more than generation.

#3080 10 GB, 12 GB, or 3080 Ti: which one to choose

The three RTX 3080
Criterion3080 10 GB3080 12 GB3080 Ti 12 GB
CUDA cores8 7048 96010 240
Memory10 GB GDDR6X12 GB GDDR6X12 GB GDDR6X
Bus / bandwidth320 bits / 760 GB/s384 bits / 912 GB/s384 bits, same expected bandwidth
Card power320 W350 W350 W
Generation on Llama 2 7B Q4_0139.7 tok/s (measured)estimate: up to approximately 167estimate: close to 12 GB

For local AI, the 3080 Ti and the 3080 12 GB are interchangeable: the same 12 GB, the same bus, and nearly identical generation speed. The estimate for these two cards is an upper bound that applies the measured efficiency of the 10 GB model to 912 GB/s; in practice, the 3090, which also has a 384-bit bus, measures 158.2 tokens per second. So choose the cheaper one, without being guided by the core count. The 10 GB version is one step down in capacity: its 10 GB prevents Q8 use with 9-billion-parameter models.

#Compatible models by memory capacity

What fits on 10 GB and 12 GB (weights according to the QuelLLM catalog)
Model and quantizationWeights3080 10 GB3080 12 GB / Ti
Qwen 3.5 9B in Q46 GBYes, with contextYes, with a lot of context
Gemma 4 12B in Q47 GBYes, moderate contextYes, comfortable
Gemma 4 12B in Q59 GBExactly rightYes, moderate context
Qwen 3.5 9B in Q810 GBNoYes, but with limited context
gpt-oss 20B in Q413 GBNoNo: exceeds by 1 GB

Two points in this table deserve attention. First, the 3080 10 GB handles Gemma 4 12B in Q4 (7 GB), despite what you might think: 3 GB remains for context, which is enough for average conversations. Second, the gap between 12 and 16 GB is clear: a 20-billion-class model in Q4, such as gpt-oss 20B, weighs 13 GB and fits on none of the three cards. If that is your target, you need a 16 GB card or larger, and the guide to 12 GB of VRAM explains exactly what you can do with 12 GB.

#Installation, power, and settings

The RTX 3080 and 3080 Ti are among the compute-capability 8.6 cards listed by Ollama, which requires driver 550 or later. A specific point to watch with these cards: the NVIDIA specification lists 350 W for the card, 320 W for the 10 GB version, and recommends a 750 W system power supply. Check your power supply before installing the card, and connect its power connectors to separate cables.

  1. 01
    Check the power supply
    Make sure it reaches 750 W and has the connectors required by your card model.
  2. 02
    Install the driver and Ollama
    Install driver 550 or later, then Ollama, and launch a model of your choice.
  3. 03
    Adjust memory
    On the 10 GB version, start the server with OLLAMA_FLASH_ATTENTION=1 and OLLAMA_KV_CACHE_TYPE=q8_0: the documentation specifies that the K/V cache can be quantized when Flash Attention is enabled.
  4. 04
    Control
    Use nvidia-smi to monitor memory and temperature under load: these cards run hotter than newer models.
Linux: launch Ollama with the memory settings
OLLAMA_FLASH_ATTENTION=1 OLLAMA_KV_CACHE_TYPE=q8_0 ollama serve

#Speed: the public benchmark and its extrapolations

The collaborative llama.cpp table, which measures Llama 2 7B in Q4_0 (3.56 GiB file, or 3.82 GB), reports 139.65 tokens per second during generation and 5,014 during prompt processing for the RTX 3080 10 GB. With Flash Attention, generation remains at 139.95, but prompt processing rises to 5,570 tokens per second, about 11% better. The 760 GB/s bandwidth would theoretically allow nearly 199 tokens per second: the measurement reaches 70% of that ceiling. Throughput also depends on the driver, system, and card manufacturer, even with the same chip.

Comparison with other cards (Llama 2 7B Q4_0, generation)
CardMemoryGeneration (tok/s)
RTX 3060 12 GB12 GB75,6
RTX 3080 10 GB10 GB139,7
RTX 309024 GB158,2

Using a rule of three based on file sizes, the 3080 10 GB would deliver at most about 89 tokens per second on Qwen 3.5 9B in Q4 (6 GB) and about 76 on Gemma 4 12B in Q4 (7 GB). These are theoretical upper bounds, not measurements. On a 3080 12 GB, Qwen 3.5 9B's Q8 (10 GB) would theoretically remain below 65 tokens per second, which is comfortable. In all cases, these cards are well beyond the read speed; capacity, not speed, is the limit.

Prompt processing is the second advantage of these cards. At 5,014 tokens per second, the 3080 10 GB reads a 10,000-token document in about two seconds, and at 5,570 with Flash Attention, in slightly less time—two theoretical calculations that assume a constant throughput. That's about twice as fast as a 3060 12 GB (2,138 tokens per second). Prompt processing matters even more when several people share the same machine, since each request must first reread its own context before generating any response, and that's where the gap with a more modest card widens. For RAG over large documents or a coding assistant that rereads large files for every request, this is a real usability gain, much more noticeable than the generation gap, which is already well above reading speed.

#Buying in 2026: where these cards stand

Compared with a new 16 GB card, the argument for a used RTX 3080 is speed at a moderate price, but it lacks memory. The 3090 measures 158.2 tokens per second with 24 GB, making it the used-market anchor for anyone targeting 20- to 30-billion-parameter models. Used-card prices vary every week: the site’s tracker records the lowest price for each card and the price per GB of VRAM, which is more useful than a figure written here.

Which card for your target model
Target modelRequired memorySuitable card
8 to 9B in Q46 GB of weightsAny 8 GB card
12B in Q4 or 9B in Q87 to 10 GB of weights3080 12 GB or Ti, or 3080 10 GB for Q4
20B in Q4 (gpt-oss 20B)13 GB of weights16 GB or more card

#2026 verdict

Used 3080 12 GB or 3080 Ti
A very good choice if the price remains well below that of a new 12 to 16 GB card: 12 GB, high speed, and enough headroom for a 9B Q8 or a 12B Q4.
3080 10 GB
Useful for 8- to 12-billion-parameter models in Q4. Without the 9B's Q8 or any context headroom, a 3060 12 GB may be a better fit if you are looking for capacity rather than speed.
At purchase
Check the power supply, the 10 or 12 GB labeling, and the fans’ condition; these cards have run hot and have several years of use.

#Frequently asked questions

Frequently asked questions
RTX 3080 10 GB or 12 GB for an LLM?+
Choose the 12 GB model if the price is close: its extra 2 GB opens up Qwen 3.5 9B's Q8 (10 GB of weights) and leaves room for Gemma 4 12B, whereas the 10 GB model is limited to 8–9B models in Q4 and a 12B model in Q4 with moderate context. Its wider bus also provides about 20% more bandwidth.
RTX 3080 Ti or RTX 3090 for an LLM?+
The 3090 has 24 GB versus 12 GB and measures 158.2 tokens per second on the llama.cpp test, opening up 20- to 30-billion-parameter models in Q4. The 3080 Ti is sufficient if you stay at 12 billion parameters or fewer. For these models, generation speeds remain close: capacity is the deciding factor.
Can the RTX 3080 12 GB run gpt-oss 20B?+
Not entirely on the card. The catalog lists 13 GB in Q4 for this model, which is 1 GB more than a 3080 12 GB's memory: the model would be split with RAM and speed would drop. You need a 16 GB card to run it at full speed.
How many tokens per second on a RTX 3080?+
The public llama.cpp benchmark table gives 139.7 tokens per second for the 3080 10 GB on Llama 2 7B in Q4_0. A heavier model will be slower: about 89 for a 9B in Q4 and 76 for a 12B in Q4, as a theoretical estimate. The 12 GB and the Ti should be slightly faster, with no public measurement available.
What power supply do you need for a RTX 3080 Ti?+
NVIDIA lists 350 W for the card and recommends a 750 W system power supply for both the 3080 and the 3080 Ti. A 650 W supply is below that recommendation. Also plan for the appropriate connectors and good case ventilation, because these cards dissipate a lot of heat under load.
Does Ollama work on a RTX 3080?+
Yes. Ollama lists the RTX 3080 and 3080 Ti among the cards with compute capability 8.6 and requires driver 550 or newer. On the 10 GB version, enable Flash Attention and K/V cache quantization to save memory on long contexts.
Is RTX 3080 still worth it in 2026?+
Yes, if its used price remains significantly lower than that of an equivalent new card with the same memory. Its speed is excellent, but its 10 or 12 GB capacity is what matters: check the site's price tracking and compare it with 16 GB cards before deciding.
Did this guide help you?

Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.