Intermediate 11 minRTX 20

Which LLM on RTX 2080 / 2080 Super (8 GB) ?

Direct response

A RTX 2080 or 2080 Super (8 GB of GDDR6) runs 7- to 9-billion-parameter models in Q4 without difficulty, such as Granite 4.2 8B (4.6 GB) or Qwen 3.5 9B (6 GB). No public measurements exist for these two cards, but their 256-bit bus puts them at the level of the 2070 Super, at about 88 tokens/s on a 7B. The 8 GB capacity remains the ceiling.

The RTX 2080 (September 2018) and 2080 Super (July 2019) were high-end cards; in 2026, their advantage for local AI is modest because they have only 8 GB. This guide separates documented facts from estimates, compares these cards with 12 GB alternatives, and explains how to verify that a model actually fits in memory.

Choosing a machine? Our picks by budget →

By Mohamed Meguedmi·Update 2026-09-30·Tested on Windows, macOS, and Linux
Recommended hardware

Buying alternative for this guide: RTX 5060 Ti 16GB (ASUS Prime).

Compare all options by budget, from €800 to €3,500 →

On the go: which laptop for local AI →

Affiliate links — possible commission at no extra cost to you. As an Amazon Associate, BestLLMfor earns from qualifying purchases.

#RTX 2080 and 2080 Super: what they can run

On a RTX 2080 or a 2080 Super, a 7- to 9-billion-parameter model in Q4 fits entirely in video memory and runs quickly. According to the QuelLLM catalog, Granite 4.2 8B requires 4.6 GB in Q4 and Qwen 3.5 9B 6 GB, leaving room for a context of a few thousand tokens. Gemma 4 12B (7 GB in Q4) reaches the 8 GB ceiling and spills over as soon as the conversation gets longer. Neither card appears in llama.cpp's public performance table; they can nevertheless be placed approximately because they share the same 256-bit bus as the 2070 Super, measured at 88 tokens per second on a 7-billion-parameter model in Q4_0. The 2080 Super, with faster memory, should perform slightly better. In short: fast, but limited by its 8 GB.

RTX 2080
September 2018, 2,944 CUDA cores, 8 GB of GDDR6, 256-bit bus, delivering 448 GB/s at 14 Gbit/s.
RTX 2080 Super
July 2019, 3,072 CUDA cores, 8 GB of GDDR6 at 15.5 Gbit/s, or 496 GB/s—about 11% more than the 2080.
For an LLM
Generation depends mainly on bandwidth: the Super is slightly faster, but capacity does not change.

#Compatible models: what fits in 8 GB

Models for RTX 2080 / 2080 Super (weights according to the QuelLLM catalog)
ModelWeights in Q4What to expect on 8 GB
Granite 4.2 8B4.6 GBComfortable, with context of several thousand tokens
Qwen 3.5 9B6 GBFits in Q4; moderate context, adjust the K/V cache if needed
Gemma 4 12B7 GBToo tight: the cache and system no longer fit
Qwen 3.5 9B in Q810 GBDoes not fit: exceeds 8 GB

The reasoning is the same for all models: weights, plus the context cache, plus a margin of about one gigabyte for the driver and display. If your display is connected to this card, some of its memory is already reserved. On 8 GB, targeting weights of no more than 6 GB is a safe rule. For a longer context, quantizing the K/V cache is the tool that yields the greatest benefit.

#Verify that a model actually fits on the card

Much of the performance loss comes from a model that does not fit entirely in VRAM: Ollama then places part of it on the processor, and throughput drops without an error message. The Ollama documentation shows how to see this: the ollama ps command displays, in the Processor column, where the model was loaded. 100% GPU means the model is entirely on the card, while a ratio such as 48%/52% CPU/GPU indicates that loading is split between RAM and VRAM.

  1. 01
    Load the model
    Launch the model with ollama run followed by its name, then send an initial prompt.
  2. 02
    Read the distribution
    In another terminal, type ollama ps. Note the Processor column and the reported size.
  3. 03
    Correct if necessary
    If the model is not at 100% on the GPU, reduce the context, choose a lighter quantization, or switch to a smaller model. Quantizing the K/V cache also helps; the documentation specifies that this requires Flash Attention.
Control GPU / CPU allocation
ollama ps

#Speed: what is measured and what is estimated

The collaborative llama.cpp table (Llama 2 7B in Q4_0, 3.56 GiB file) includes the 2070 Super, 2080 Ti, 3060 12 GB, and 4060 Ti 8 GB, but neither the 2080 nor the 2080 Super. The table notes that results vary by driver, system, and card, even with the same chip. Here is what was measured, followed by what it is reasonable to infer from it.

llama.cpp test, Llama 2 7B Q4_0, generation (tokens/s)
CardMemoryBusGenerationStatus
RTX 2070 Super8 GB256-bit88,1Published measurement
RTX 2080 / 2080 Super8 GB256-bitapproximately 88 to 97Estimate: same bus, faster memory on the Super
RTX 2080 Ti11 GB352 bits107,5Published measurement
RTX 3060 12 GB12 GB192 bits75,6Published measurement
RTX 4060 Ti 8 GB8 GB128 bits63,9Published measurement

The 2080 estimate follows a simple rule: at 448 GB/s, the 2080 should produce throughput very close to the 2070 Super, and at 496 GB/s, the 2080 Super can gain up to an additional 10%, or roughly 95 tokens per second. This is a projection, not a measurement: do not cite it as a result.

The table offers a useful lesson for anyone considering replacing their card: the more recent RTX 4060 Ti with 8 GB generates more slowly (63.9 tokens per second) than the 2070 Super because its 128-bit bus limits bandwidth. In text generation, a newer GPU generation does not mean higher speed. What it mainly provides is energy efficiency and newer features, not memory capacity.

#2080 Super or RTX 3060 12 GB: capacity wins

The 3060 12 GB generates more slowly than the 2070 Super (75.6 versus 88.1 tokens per second in the test), but it offers 4 GB more. That is what matters for local AI: on 12 GB, Gemma 4 12B (7 GB in Q4) fits with room for context, which 8 GB cards cannot handle. A 2080 Super does not make up for this limitation with its speed, since the increase to 95 tokens per second does not change the fact that a 12B model will not fit.

2080 Super or 3060 12 GB: trade-off
CriterionRTX 2080 SuperRTX 3060 12 GB
Memory8 GB12 GB
Generation (7B Q4_0)Estimated at around 95 tokens/s75.6 tokens/s (measured)
12-billion-parameter modelsNo (7 GB of weights, no headroom)Yes, with context
7- to 9-billion-parameter modelsYesYes

The conclusion therefore depends on your target. For an assistant with 8 or 9 billion parameters, your 2080 is sufficient, and moving to a 3060 12 GB adds nothing in speed. If you are targeting a 12B model or long contexts, the upgrade is justified. Used prices change every week: check the tracking page instead of relying on a fixed figure.

#Long context on 8 GB: the K/V cache decides

On an 8 GB card, the problem isn't the model; it's the context cache. Each conversation token adds data to memory, and an 8B model in Q4 that fit with an empty context can overflow when you paste in a long document. There are two levers in Ollama. Flash Attention, which the software enables automatically when the card supports it, reduces the memory required as the context grows. Once enabled, it allows K/V cache quantization; according to the documentation, the q8_0 format uses about half the memory of the default format. The q4_0 format saves more, but degrades accuracy more noticeably at large contexts.

Linux: force Flash Attention and quantize the cache to q8_0
OLLAMA_FLASH_ATTENTION=1 OLLAMA_KV_CACHE_TYPE=q8_0 ollama serve

This setting is global: it applies to every model served by the instance, which is convenient on a dedicated machine but worth keeping in mind if you change models often. The best practice is to measure before tuning: load your model, fill a context representative of your usage, then check with ollama ps that the model remains at 100% on the GPU. If so, you have headroom; otherwise, quantize the cache or reduce the context.

#When 8 GB is no longer enough: the next tiers

What each memory tier unlocks (Q4 weights, QuelLLM catalog)
Card memoryModels that become feasibleWeights in Q4
8 GB (your card)7- to 9-billion-parameter models4.6 to 6 GB
12 GBGemma 4 12B with context7 GB
16 GBgpt-oss 20B (13 GB) or Qwen 3.5 27B (16 GB, very tight)13 to 16 GB

The table is easy to read: each memory increase opens up a category of models, while raw speed barely changes the experience beyond 30 or 40 tokens per second. If you are considering an investment, first ask yourself which model you want to run, then choose the card with the corresponding memory. A 27-billion-parameter model on 16 GB is already at the limit; if that is your target, look toward 24 GB instead.

#Drivers and support in 2026

The RTX 2080 and the 2080 Ti are among the compute capability 7.5 cards listed by Ollama, which requires driver 550 or later. This covers current versions of Ollama; if installation fails, the first thing to check is the driver version with nvidia-smi. There is therefore no reason to fear a software block when installing Ollama or llama.cpp on these cards. The only point to watch is the driver version: update it before looking for a problem elsewhere.

#2026 verdict

Keep it if
It is already in your PC, and you are targeting 7- to 9-billion-parameter models in Q4. More than enough speed, at zero cost.
Don’t buy for AI
Unless the price is low and the warranty is clear. At a similar price, a 12 GB card offers more than a few extra tokens per second.
Move on if
You want a 12B, a 14B or larger, or a context of more than ten thousand tokens on an 8B. You then need 12 to 16 GB.

#Frequently asked questions

Frequently asked questions
Can the RTX 2080 run an 8- to 9-billion-parameter model?+
Yes. In Q4, Granite 4.2 8B weighs 4.6 GB and Qwen 3.5 9B 6 GB according to the QuelLLM catalog, fitting within the card's 8 GB with a moderate context. Its speed is comparable to that of a 2070 Super, measured at 88 tokens per second on a 7B in Q4_0.
RTX 2080 Super or RTX 2080 Ti for an LLM?+
The 2080 Ti has 11 GB and 616 GB/s of bandwidth; the 2080 Super has 8 GB and 496 GB/s. For an LLM, the extra 3 GB matters more than speed: it opens up 12-billion-parameter models. If your budget allows, the 2080 Ti is the better choice of the two.
Can Gemma 4 12B run on a RTX 2080 Super?+
Not comfortably. The catalog lists 7 GB of weights in Q4, leaving barely one gigabyte for the context cache and system on 8 GB. The model loads, but Ollama quickly offloads part of it to RAM, as ollama ps will show you. Prefer an 8B or 9B model.
How can you tell whether your model is fully on the GPU?+
Use the ollama ps command: the Processor column shows 100% GPU when the model is fully loaded into the graphics card, and a ratio such as 48%/52% CPU/GPU when it is split between system memory and the GPU. In the latter case, speed drops significantly, so it is better to reduce the context or the model.
Is a RTX 2080 faster than an RTX 4060 Ti 8 GB for generating text?+
On the public llama.cpp table, the 8 GB 4060 Ti generates 63.9 tokens per second versus 88.1 for the 2070 Super, whose bus is identical to the 2080’s. The 4060 Ti’s 128-bit bus limits its bandwidth. So expect higher throughput from the 2080 at the same capacity.
Does Ollama still work on a RTX 2080 in 2026?+
Yes. Ollama lists the RTX 2080 among cards with compute capability 7.5 and requires only driver 550 or later. No end of support has been announced for this family in the documentation reviewed; simply monitor driver and CUDA release notes during major updates.
Did this guide help you?

Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.