Which LLM on RTX 2080 Ti (11 GB) ?
With 11 GB of GDDR6 and 616 GB/s of bandwidth, the RTX 2080 Ti can run a 12-billion-parameter model in Q4 without running out of memory, such as Gemma 4 12B (7 GB), or an 8B model in Q8, such as Granite 4.2 8B (9 GB). On the llama.cpp benchmark, it generates 107 tokens/s on a 7B model in Q4_0. It is a rare example of a 2018 graphics card whose capacity remains useful.
The RTX 2080 Ti (September 2018) is the only card in the Turing generation to exceed 8 GB: it has 11. This guide connects its specifications to public benchmarks, identifies which models fit on it, compares its capacity-to-speed ratio with newer cards, and explains how to install it and verify that nothing spills over.
Choosing a machine? Our picks by budget →
Buying alternative for this guide: RTX 5060 Ti 16GB (ASUS Prime).
Compare all options by budget, from €800 to €3,500 →
On the go: which laptop for local AI →
Affiliate links — possible commission at no extra cost to you. As an Amazon Associate, BestLLMfor earns from qualifying purchases.
#RTX 2080 Ti in 2026: the capacity that matters
The RTX 2080 Ti runs a 12-billion-parameter model in Q4 without spilling over, something no 8 GB card does cleanly. According to the QuelLLM catalog, Gemma 4 12B weighs 7 GB in Q4, Granite 4.2 8B 4.6 GB, and Qwen 3.5 9B 6 GB; in Q8, Granite 4.2 8B reaches 9 GB, which also fits within 11 GB. The card uses a TU102 chip, 4,352 CUDA cores, and 11 GB of GDDR6 memory on a 352-bit bus, delivering 616 GB/s. On llama.cpp's public leaderboard, it reaches 107.5 tokens per second when generating with a 7-billion-parameter model in Q4_0. This combination of speed and capacity explains why the 2080 Ti remains relevant while newer 8 GB cards outperform it in gaming but not in AI.
- Architecture
- Turing, TU102 chip, released in September 2018.
- Memory
- 11 GB of GDDR6, a 352-bit bus, and 616 GB/s of bandwidth.
- Cœurs
- 4,352 CUDA cores and second-generation Tensor Cores.
- Power supply
- Check the specifications for your exact model: factory versions differ by manufacturer.
#Models compatible with 11 GB
| Model | Q4 | Q5 | Q8 | On 11 GB |
|---|---|---|---|---|
| Granite 4.2 8B | 4.6 GB | 6 GB | 9 GB | Q4 and Q5 handle comfortably, Q8 possible with a moderate context |
| Qwen 3.5 9B | 6 GB | 7 GB | 10 GB | Q4 and Q5 are comfortable; Q8 leaves barely 1 GB for context |
| Gemma 4 12B | 7 GB | 9 GB | 13 GB | Q4 with 3 to 4 GB of headroom; Q5 is tight; Q8 won't fit |
The Q5 column shows the value of an 11 GB card: with an 8B or 9B model, you can choose a more faithful quantization than Q4 without sacrificing anything. Q5 for a 12B model (9 GB) leaves less than 2 GB for the cache, limiting the context. A Q4 model with a large context is often more useful than a Q5 with a cramped window: choose based on your workload, whether summarizing long documents or having short conversations.
Two practical benchmarks. First, an 8B's Q8 (9 GB) fits on the card and delivers quality very close to full precision—something 8 GB cards can't do. Second, the 3 to 4 GB of headroom on a 12B in Q4 gets consumed by the context: beyond a few thousand tokens, quantize the K/V cache or reduce the context window. For choosing the quantization level, the dedicated guide details the quality losses.
#Installation and verification
Turing is supported by Ollama: the documentation lists the RTX 2080 Ti, 2080, 2070, and 2060 among the compute capability 7.5 cards and requires driver 550 or later. The procedure is the same as for any NVIDIA card, with one final check that matters: make sure the model fits entirely on the card.
- 01Update the driverInstall an NVIDIA driver version 550 or later, then check with nvidia-smi that the card is detected with its 11 GB.
- 02Install OllamaFollow the Ollama installation guide for your system, then launch a model with ollama run.
- 03Control the distributionRun ollama ps: the Processor column should show 100% GPU. If a CPU percentage appears, the model is spilling into RAM: reduce the context or the model.
- 04Adjust the K/V cache if neededStart Ollama with OLLAMA_FLASH_ATTENTION=1 and OLLAMA_KV_CACHE_TYPE=q8_0 to save space on long contexts.
#Speed: what llama.cpp measures
The llama.cpp repository's collaborative table reports 107.5 tokens per second for generation and 2,891 for prompt processing on the RTX 2080 Ti, using Llama 2 7B at Q4_0, with a 3.56 GiB file. With Flash Attention, these figures rise to 109.2 and 3,108. One point stands out: 616 GB/s of bandwidth would theoretically allow about 161 tokens per second on this 3.82 GB file, yet the measurement reaches 67% of that. More modest cards come closer. Actual throughput therefore depends on factors beyond memory alone, such as clock speeds and drivers, and the table itself notes that results vary by card and system.
To estimate the speed of a current model, you can apply the rule of three to the file size. A 4.6 GB model such as Granite 4.2 8B in Q4 would yield at most 107.5 × 3.82 ÷ 4.6, or about 89 tokens per second. For a 7 GB Gemma 4 12B, that drops to about 59 tokens per second. These are theoretical upper bounds, not measurements; newer architectures may differ. In all cases, these speeds exceed the reading rate.
#Long context and documents: prompt processing and K/V cache
Generation speed is only half the experience: as soon as you paste a document or search your files, prompt processing is what makes you wait. llama.cpp’s table measures it separately (pp512). At 2 891 tokens per second, the 2080 Ti reads a 10 000-token document in a little over three and a half seconds, compared with about four and a half seconds for the 3060 12 GB (2 138) and two seconds for the 3080 10 GB (5 014). This calculation is theoretical and assumes a constant throughput, but it gives the right order of magnitude: the 2080 Ti remains pleasant to use for RAG.
The second constraint is memory. The K/V cache grows with the context, and on a 12B model in Q4 with 3 to 4 GB of headroom, a long document quickly uses up that reserve. The Ollama documentation specifies that the K/V cache can be quantized when Flash Attention is enabled, and that the q8_0 format uses about half as much memory as the default f16 format. The setting is global to the Ollama instance: it applies to all served models.
Proceed through measured tests: load the model, fill a context comparable to your real-world usage, then check with ollama ps that the card remains at 100%. If the model overflows, the least expensive solution is almost always to reduce the context window before changing the model or card. The guide to the context window explains how to configure it.
#2080 Ti, 3060 12 GB, 5060 Ti 16 GB: which trade-off?
| Card | Memory | Generation (tok/s) | Prompt processing (tok/s) |
|---|---|---|---|
| RTX 3060 12 GB | 12 GB | 75,6 | 2 138 |
| RTX 5060 Ti 16 GB | 16 GB | 90,9 | 3 737 |
| RTX 2080 Ti | 11 GB | 107,5 | 2 891 |
| RTX 3080 10 GB | 10 GB | 139,7 | 5 014 |
The 2080 Ti generates approximately 42% faster than the 3060 12 GB, and even faster than the RTX 5060 Ti 16 GB, whose 128-bit bus is a bottleneck. Its memory, however, puts the 2080 Ti behind both: 11 GB versus 12 and then 16 GB. The choice therefore depends on the target model. For a 12B in Q4, the 2080 Ti is sufficient and remains the fastest of the three. For a model with 20 billion parameters or more, or a very long context, you need 16 GB, and the newer-generation card becomes the better choice. Prices change every week; the site’s tracking gives the price per GB of VRAM.
- Question 1: which model do you want?
- For an 8 to 12B model: keep or buy the 2080 Ti. For a 20B model or larger: aim for 16 GB.
- Question 2: how long should the context be?
- A few thousand tokens: 11 GB is enough. Tens of thousands: there will not be enough headroom.
- Question 3: what is your electricity budget?
- Compare the power consumption listed on your card's specification sheet with that of newer cards, which are more efficient at comparable speeds.
- Local AI graphics card prices (weekly tracking)
- Which LLM for RTX 3060 with 12 GB?
- Which LLM on RTX 5060 Ti (8 / 16 GB)?
#Turing support: what's certain and what isn't
Turing is now the minimum CUDA 13 baseline: the release notes state that support for earlier architectures (Maxwell, Volta, Pascal) was dropped from its libraries. This means the 2080 Ti is supported, but it is also one of the oldest cards still included. No source consulted announces end of support for Turing; however, newer generations' hardware features (more efficient compute formats, for example) will never be available on it. For llama.cpp and Ollama with standard GGUF models, this is not an obstacle.
#2026 verdict
- Keep it if
- You have it: 11 GB and 107 tokens per second on a 7B make it a more useful AI card than many recent 8 GB cards.
- Buy used if
- The price is significantly lower than that of a 12 to 16 GB card, and your target is an 8B or 12B model. Check the fans and warranty; these cards have several years of use.
- Move to 16 GB if
- You are targeting a model with 20 billion parameters or more, or a context of several tens of thousands of tokens.
#Frequently asked questions
Is the RTX 2080 Ti still relevant for an LLM in 2026?+
RTX 2080 Ti or RTX 3060 12 GB for an LLM?+
How many tokens per second on a RTX 2080 Ti?+
Does the RTX 2080 Ti support Flash Attention?+
How can I tell whether my model fits entirely on the card?+
Is the RTX 2080 Ti at risk of no longer being supported?+
- Source: llama.cpp performance table on CUDA
- Source: maps NVIDIA supported by Ollama
- Source: Ollama FAQ (ollama ps, K/V cache)
- Source: GeForce 20-series specifications
Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.