Which LLM on RTX 2080 / 2080 Super (8 GB) ?
A RTX 2080 or 2080 Super (8 GB of GDDR6) runs 7- to 9-billion-parameter models in Q4 without difficulty, such as Granite 4.2 8B (4.6 GB) or Qwen 3.5 9B (6 GB). No public measurements exist for these two cards, but their 256-bit bus puts them at the level of the 2070 Super, at about 88 tokens/s on a 7B. The 8 GB capacity remains the ceiling.
The RTX 2080 (September 2018) and 2080 Super (July 2019) were high-end cards; in 2026, their advantage for local AI is modest because they have only 8 GB. This guide separates documented facts from estimates, compares these cards with 12 GB alternatives, and explains how to verify that a model actually fits in memory.
Choosing a machine? Our picks by budget →
Buying alternative for this guide: RTX 5060 Ti 16GB (ASUS Prime).
Compare all options by budget, from €800 to €3,500 →
On the go: which laptop for local AI →
Affiliate links — possible commission at no extra cost to you. As an Amazon Associate, BestLLMfor earns from qualifying purchases.
#RTX 2080 and 2080 Super: what they can run
On a RTX 2080 or a 2080 Super, a 7- to 9-billion-parameter model in Q4 fits entirely in video memory and runs quickly. According to the QuelLLM catalog, Granite 4.2 8B requires 4.6 GB in Q4 and Qwen 3.5 9B 6 GB, leaving room for a context of a few thousand tokens. Gemma 4 12B (7 GB in Q4) reaches the 8 GB ceiling and spills over as soon as the conversation gets longer. Neither card appears in llama.cpp's public performance table; they can nevertheless be placed approximately because they share the same 256-bit bus as the 2070 Super, measured at 88 tokens per second on a 7-billion-parameter model in Q4_0. The 2080 Super, with faster memory, should perform slightly better. In short: fast, but limited by its 8 GB.
- RTX 2080
- September 2018, 2,944 CUDA cores, 8 GB of GDDR6, 256-bit bus, delivering 448 GB/s at 14 Gbit/s.
- RTX 2080 Super
- July 2019, 3,072 CUDA cores, 8 GB of GDDR6 at 15.5 Gbit/s, or 496 GB/s—about 11% more than the 2080.
- For an LLM
- Generation depends mainly on bandwidth: the Super is slightly faster, but capacity does not change.
#Compatible models: what fits in 8 GB
| Model | Weights in Q4 | What to expect on 8 GB |
|---|---|---|
| Granite 4.2 8B | 4.6 GB | Comfortable, with context of several thousand tokens |
| Qwen 3.5 9B | 6 GB | Fits in Q4; moderate context, adjust the K/V cache if needed |
| Gemma 4 12B | 7 GB | Too tight: the cache and system no longer fit |
| Qwen 3.5 9B in Q8 | 10 GB | Does not fit: exceeds 8 GB |
The reasoning is the same for all models: weights, plus the context cache, plus a margin of about one gigabyte for the driver and display. If your display is connected to this card, some of its memory is already reserved. On 8 GB, targeting weights of no more than 6 GB is a safe rule. For a longer context, quantizing the K/V cache is the tool that yields the greatest benefit.
#Verify that a model actually fits on the card
Much of the performance loss comes from a model that does not fit entirely in VRAM: Ollama then places part of it on the processor, and throughput drops without an error message. The Ollama documentation shows how to see this: the ollama ps command displays, in the Processor column, where the model was loaded. 100% GPU means the model is entirely on the card, while a ratio such as 48%/52% CPU/GPU indicates that loading is split between RAM and VRAM.
- 01Load the modelLaunch the model with ollama run followed by its name, then send an initial prompt.
- 02Read the distributionIn another terminal, type ollama ps. Note the Processor column and the reported size.
- 03Correct if necessaryIf the model is not at 100% on the GPU, reduce the context, choose a lighter quantization, or switch to a smaller model. Quantizing the K/V cache also helps; the documentation specifies that this requires Flash Attention.
- Quantize the KV cache: save VRAM (long context)
- Troubleshoot Ollama: GPU not detected, slowdowns, out-of-memory errors
#Speed: what is measured and what is estimated
The collaborative llama.cpp table (Llama 2 7B in Q4_0, 3.56 GiB file) includes the 2070 Super, 2080 Ti, 3060 12 GB, and 4060 Ti 8 GB, but neither the 2080 nor the 2080 Super. The table notes that results vary by driver, system, and card, even with the same chip. Here is what was measured, followed by what it is reasonable to infer from it.
| Card | Memory | Bus | Generation | Status |
|---|---|---|---|---|
| RTX 2070 Super | 8 GB | 256-bit | 88,1 | Published measurement |
| RTX 2080 / 2080 Super | 8 GB | 256-bit | approximately 88 to 97 | Estimate: same bus, faster memory on the Super |
| RTX 2080 Ti | 11 GB | 352 bits | 107,5 | Published measurement |
| RTX 3060 12 GB | 12 GB | 192 bits | 75,6 | Published measurement |
| RTX 4060 Ti 8 GB | 8 GB | 128 bits | 63,9 | Published measurement |
The 2080 estimate follows a simple rule: at 448 GB/s, the 2080 should produce throughput very close to the 2070 Super, and at 496 GB/s, the 2080 Super can gain up to an additional 10%, or roughly 95 tokens per second. This is a projection, not a measurement: do not cite it as a result.
The table offers a useful lesson for anyone considering replacing their card: the more recent RTX 4060 Ti with 8 GB generates more slowly (63.9 tokens per second) than the 2070 Super because its 128-bit bus limits bandwidth. In text generation, a newer GPU generation does not mean higher speed. What it mainly provides is energy efficiency and newer features, not memory capacity.
#2080 Super or RTX 3060 12 GB: capacity wins
The 3060 12 GB generates more slowly than the 2070 Super (75.6 versus 88.1 tokens per second in the test), but it offers 4 GB more. That is what matters for local AI: on 12 GB, Gemma 4 12B (7 GB in Q4) fits with room for context, which 8 GB cards cannot handle. A 2080 Super does not make up for this limitation with its speed, since the increase to 95 tokens per second does not change the fact that a 12B model will not fit.
| Criterion | RTX 2080 Super | RTX 3060 12 GB |
|---|---|---|
| Memory | 8 GB | 12 GB |
| Generation (7B Q4_0) | Estimated at around 95 tokens/s | 75.6 tokens/s (measured) |
| 12-billion-parameter models | No (7 GB of weights, no headroom) | Yes, with context |
| 7- to 9-billion-parameter models | Yes | Yes |
The conclusion therefore depends on your target. For an assistant with 8 or 9 billion parameters, your 2080 is sufficient, and moving to a 3060 12 GB adds nothing in speed. If you are targeting a 12B model or long contexts, the upgrade is justified. Used prices change every week: check the tracking page instead of relying on a fixed figure.
- Local AI graphics card prices (weekly tracking)
- Which LLM for RTX 3060 with 12 GB?
- Which LLM on RTX 2080 Ti (11 GB)?
#Long context on 8 GB: the K/V cache decides
On an 8 GB card, the problem isn't the model; it's the context cache. Each conversation token adds data to memory, and an 8B model in Q4 that fit with an empty context can overflow when you paste in a long document. There are two levers in Ollama. Flash Attention, which the software enables automatically when the card supports it, reduces the memory required as the context grows. Once enabled, it allows K/V cache quantization; according to the documentation, the q8_0 format uses about half the memory of the default format. The q4_0 format saves more, but degrades accuracy more noticeably at large contexts.
This setting is global: it applies to every model served by the instance, which is convenient on a dedicated machine but worth keeping in mind if you change models often. The best practice is to measure before tuning: load your model, fill a context representative of your usage, then check with ollama ps that the model remains at 100% on the GPU. If so, you have headroom; otherwise, quantize the cache or reduce the context.
#When 8 GB is no longer enough: the next tiers
| Card memory | Models that become feasible | Weights in Q4 |
|---|---|---|
| 8 GB (your card) | 7- to 9-billion-parameter models | 4.6 to 6 GB |
| 12 GB | Gemma 4 12B with context | 7 GB |
| 16 GB | gpt-oss 20B (13 GB) or Qwen 3.5 27B (16 GB, very tight) | 13 to 16 GB |
The table is easy to read: each memory increase opens up a category of models, while raw speed barely changes the experience beyond 30 or 40 tokens per second. If you are considering an investment, first ask yourself which model you want to run, then choose the card with the corresponding memory. A 27-billion-parameter model on 16 GB is already at the limit; if that is your target, look toward 24 GB instead.
#Drivers and support in 2026
The RTX 2080 and the 2080 Ti are among the compute capability 7.5 cards listed by Ollama, which requires driver 550 or later. This covers current versions of Ollama; if installation fails, the first thing to check is the driver version with nvidia-smi. There is therefore no reason to fear a software block when installing Ollama or llama.cpp on these cards. The only point to watch is the driver version: update it before looking for a problem elsewhere.
#2026 verdict
- Keep it if
- It is already in your PC, and you are targeting 7- to 9-billion-parameter models in Q4. More than enough speed, at zero cost.
- Don’t buy for AI
- Unless the price is low and the warranty is clear. At a similar price, a 12 GB card offers more than a few extra tokens per second.
- Move on if
- You want a 12B, a 14B or larger, or a context of more than ten thousand tokens on an 8B. You then need 12 to 16 GB.
#Frequently asked questions
Can the RTX 2080 run an 8- to 9-billion-parameter model?+
RTX 2080 Super or RTX 2080 Ti for an LLM?+
Can Gemma 4 12B run on a RTX 2080 Super?+
How can you tell whether your model is fully on the GPU?+
Is a RTX 2080 faster than an RTX 4060 Ti 8 GB for generating text?+
Does Ollama still work on a RTX 2080 in 2026?+
- Source: llama.cpp performance table on CUDA
- Source: maps NVIDIA supported by Ollama
- Source: Ollama FAQ (ollama ps, K/V cache)
- Source: GeForce 20-series specifications
Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.