Intermediate 11 minRTX 20

Which LLM on RTX 2080 Ti (11 GB) ?

Direct response

With 11 GB of GDDR6 and 616 GB/s of bandwidth, the RTX 2080 Ti can run a 12-billion-parameter model in Q4 without running out of memory, such as Gemma 4 12B (7 GB), or an 8B model in Q8, such as Granite 4.2 8B (9 GB). On the llama.cpp benchmark, it generates 107 tokens/s on a 7B model in Q4_0. It is a rare example of a 2018 graphics card whose capacity remains useful.

The RTX 2080 Ti (September 2018) is the only card in the Turing generation to exceed 8 GB: it has 11. This guide connects its specifications to public benchmarks, identifies which models fit on it, compares its capacity-to-speed ratio with newer cards, and explains how to install it and verify that nothing spills over.

Choosing a machine? Our picks by budget →

By Mohamed Meguedmi·Update 2026-09-30·Tested on Windows, macOS, and Linux
Recommended hardware

Buying alternative for this guide: RTX 5060 Ti 16GB (ASUS Prime).

Compare all options by budget, from €800 to €3,500 →

On the go: which laptop for local AI →

Affiliate links — possible commission at no extra cost to you. As an Amazon Associate, BestLLMfor earns from qualifying purchases.

#RTX 2080 Ti in 2026: the capacity that matters

The RTX 2080 Ti runs a 12-billion-parameter model in Q4 without spilling over, something no 8 GB card does cleanly. According to the QuelLLM catalog, Gemma 4 12B weighs 7 GB in Q4, Granite 4.2 8B 4.6 GB, and Qwen 3.5 9B 6 GB; in Q8, Granite 4.2 8B reaches 9 GB, which also fits within 11 GB. The card uses a TU102 chip, 4,352 CUDA cores, and 11 GB of GDDR6 memory on a 352-bit bus, delivering 616 GB/s. On llama.cpp's public leaderboard, it reaches 107.5 tokens per second when generating with a 7-billion-parameter model in Q4_0. This combination of speed and capacity explains why the 2080 Ti remains relevant while newer 8 GB cards outperform it in gaming but not in AI.

Architecture
Turing, TU102 chip, released in September 2018.
Memory
11 GB of GDDR6, a 352-bit bus, and 616 GB/s of bandwidth.
Cœurs
4,352 CUDA cores and second-generation Tensor Cores.
Power supply
Check the specifications for your exact model: factory versions differ by manufacturer.

#Models compatible with 11 GB

Models for a RTX 2080 Ti (weights according to the QuelLLM catalog)
ModelQ4Q5Q8On 11 GB
Granite 4.2 8B4.6 GB6 GB9 GBQ4 and Q5 handle comfortably, Q8 possible with a moderate context
Qwen 3.5 9B6 GB7 GB10 GBQ4 and Q5 are comfortable; Q8 leaves barely 1 GB for context
Gemma 4 12B7 GB9 GB13 GBQ4 with 3 to 4 GB of headroom; Q5 is tight; Q8 won't fit

The Q5 column shows the value of an 11 GB card: with an 8B or 9B model, you can choose a more faithful quantization than Q4 without sacrificing anything. Q5 for a 12B model (9 GB) leaves less than 2 GB for the cache, limiting the context. A Q4 model with a large context is often more useful than a Q5 with a cramped window: choose based on your workload, whether summarizing long documents or having short conversations.

Two practical benchmarks. First, an 8B's Q8 (9 GB) fits on the card and delivers quality very close to full precision—something 8 GB cards can't do. Second, the 3 to 4 GB of headroom on a 12B in Q4 gets consumed by the context: beyond a few thousand tokens, quantize the K/V cache or reduce the context window. For choosing the quantization level, the dedicated guide details the quality losses.

#Installation and verification

Turing is supported by Ollama: the documentation lists the RTX 2080 Ti, 2080, 2070, and 2060 among the compute capability 7.5 cards and requires driver 550 or later. The procedure is the same as for any NVIDIA card, with one final check that matters: make sure the model fits entirely on the card.

  1. 01
    Update the driver
    Install an NVIDIA driver version 550 or later, then check with nvidia-smi that the card is detected with its 11 GB.
  2. 02
    Install Ollama
    Follow the Ollama installation guide for your system, then launch a model with ollama run.
  3. 03
    Control the distribution
    Run ollama ps: the Processor column should show 100% GPU. If a CPU percentage appears, the model is spilling into RAM: reduce the context or the model.
  4. 04
    Adjust the K/V cache if needed
    Start Ollama with OLLAMA_FLASH_ATTENTION=1 and OLLAMA_KV_CACHE_TYPE=q8_0 to save space on long contexts.

#Speed: what llama.cpp measures

The llama.cpp repository's collaborative table reports 107.5 tokens per second for generation and 2,891 for prompt processing on the RTX 2080 Ti, using Llama 2 7B at Q4_0, with a 3.56 GiB file. With Flash Attention, these figures rise to 109.2 and 3,108. One point stands out: 616 GB/s of bandwidth would theoretically allow about 161 tokens per second on this 3.82 GB file, yet the measurement reaches 67% of that. More modest cards come closer. Actual throughput therefore depends on factors beyond memory alone, such as clock speeds and drivers, and the table itself notes that results vary by card and system.

To estimate the speed of a current model, you can apply the rule of three to the file size. A 4.6 GB model such as Granite 4.2 8B in Q4 would yield at most 107.5 × 3.82 ÷ 4.6, or about 89 tokens per second. For a 7 GB Gemma 4 12B, that drops to about 59 tokens per second. These are theoretical upper bounds, not measurements; newer architectures may differ. In all cases, these speeds exceed the reading rate.

#Long context and documents: prompt processing and K/V cache

Generation speed is only half the experience: as soon as you paste a document or search your files, prompt processing is what makes you wait. llama.cpp’s table measures it separately (pp512). At 2 891 tokens per second, the 2080 Ti reads a 10 000-token document in a little over three and a half seconds, compared with about four and a half seconds for the 3060 12 GB (2 138) and two seconds for the 3080 10 GB (5 014). This calculation is theoretical and assumes a constant throughput, but it gives the right order of magnitude: the 2080 Ti remains pleasant to use for RAG.

The second constraint is memory. The K/V cache grows with the context, and on a 12B model in Q4 with 3 to 4 GB of headroom, a long document quickly uses up that reserve. The Ollama documentation specifies that the K/V cache can be quantized when Flash Attention is enabled, and that the q8_0 format uses about half as much memory as the default f16 format. The setting is global to the Ollama instance: it applies to all served models.

Linux: Flash Attention and K/V cache in q8_0
OLLAMA_FLASH_ATTENTION=1 OLLAMA_KV_CACHE_TYPE=q8_0 ollama serve

Proceed through measured tests: load the model, fill a context comparable to your real-world usage, then check with ollama ps that the card remains at 100%. If the model overflows, the least expensive solution is almost always to reduce the context window before changing the model or card. The guide to the context window explains how to configure it.

#2080 Ti, 3060 12 GB, 5060 Ti 16 GB: which trade-off?

Generation on Llama 2 7B Q4_0 (public llama.cpp table)
CardMemoryGeneration (tok/s)Prompt processing (tok/s)
RTX 3060 12 GB12 GB75,62 138
RTX 5060 Ti 16 GB16 GB90,93 737
RTX 2080 Ti11 GB107,52 891
RTX 3080 10 GB10 GB139,75 014

The 2080 Ti generates approximately 42% faster than the 3060 12 GB, and even faster than the RTX 5060 Ti 16 GB, whose 128-bit bus is a bottleneck. Its memory, however, puts the 2080 Ti behind both: 11 GB versus 12 and then 16 GB. The choice therefore depends on the target model. For a 12B in Q4, the 2080 Ti is sufficient and remains the fastest of the three. For a model with 20 billion parameters or more, or a very long context, you need 16 GB, and the newer-generation card becomes the better choice. Prices change every week; the site’s tracking gives the price per GB of VRAM.

Question 1: which model do you want?
For an 8 to 12B model: keep or buy the 2080 Ti. For a 20B model or larger: aim for 16 GB.
Question 2: how long should the context be?
A few thousand tokens: 11 GB is enough. Tens of thousands: there will not be enough headroom.
Question 3: what is your electricity budget?
Compare the power consumption listed on your card's specification sheet with that of newer cards, which are more efficient at comparable speeds.

#Turing support: what's certain and what isn't

Turing is now the minimum CUDA 13 baseline: the release notes state that support for earlier architectures (Maxwell, Volta, Pascal) was dropped from its libraries. This means the 2080 Ti is supported, but it is also one of the oldest cards still included. No source consulted announces end of support for Turing; however, newer generations' hardware features (more efficient compute formats, for example) will never be available on it. For llama.cpp and Ollama with standard GGUF models, this is not an obstacle.

#2026 verdict

Keep it if
You have it: 11 GB and 107 tokens per second on a 7B make it a more useful AI card than many recent 8 GB cards.
Buy used if
The price is significantly lower than that of a 12 to 16 GB card, and your target is an 8B or 12B model. Check the fans and warranty; these cards have several years of use.
Move to 16 GB if
You are targeting a model with 20 billion parameters or more, or a context of several tens of thousands of tokens.

#Frequently asked questions

Frequently asked questions
Is the RTX 2080 Ti still relevant for an LLM in 2026?+
Yes, especially because of its 11 GB of memory. It runs a 12B in Q4 like Gemma 4 12B (7 GB) and an 8B in Q8, which 8 GB cards cannot do. In the public llama.cpp test, it generates 107 tokens per second on a 7B in Q4_0, a very comfortable speed.
RTX 2080 Ti or RTX 3060 12 GB for an LLM?+
The 2080 Ti generates about 42% faster (107.5 versus 75.6 tokens per second), while the 3060 offers 1 GB more memory and uses less power. For a 12B in Q4, both are sufficient and the 2080 Ti is faster. The 3060 12 GB is mainly justified by its lower price and power consumption.
How many tokens per second on a RTX 2080 Ti?+
The public llama.cpp table lists 107.5 tokens per second for Llama 2 7B in Q4_0. A larger model will be slower, roughly proportionally: an estimated 89 tokens per second for an 8B model at 4.6 GB and 59 for a 12B model at 7 GB, based on theoretical estimates. Long context reduces these values further.
Does the RTX 2080 Ti support Flash Attention?+
Yes: the llama.cpp table includes a result with Flash Attention for the 2080 Ti, at 109.2 tokens per second instead of 107.5. Ollama enables it automatically when the card supports it, and you can force it with OLLAMA_FLASH_ATTENTION=1. The main benefit is reduced memory use with long contexts.
How can I tell whether my model fits entirely on the card?+
Run ollama ps: the Processor column shows 100% GPU when the model is fully loaded into the GPU. A ratio such as 48%/52% CPU/GPU indicates loading split between the CPU and RAM, resulting in lower speed. In that case, reduce the context, quantize the K/V cache, or choose a lighter model.
Is the RTX 2080 Ti at risk of no longer being supported?+
No source consulted announces an end of support. Ollama lists the card among those with compute capability 7.5 with a 550 or later driver, and CUDA 13 makes Turing its baseline. Monitor the release notes with every major CUDA update: that is where a retirement would be announced.
Did this guide help you?

Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.