Beginner 11 minGTX 10

Which LLM on GTX 1080 Ti (11 GB) ?

Direct response

Yes: the GTX 1080 Ti (11 GB, 484 GB/s) remains one of the best Pascal cards for a local LLM. The public llama.cpp benchmark measures 62.49 tokens/s in generation on a Llama 2 7B in Q4_0, on par with a RTX 2060 Super, with 3 GB more memory. It handles an 8B or 9B in Q4 with room to spare, and a 12B with a modest context. Its software future is limited, however: Pascal has been removed from CUDA 13.

The GTX 1080 Ti was released in March 2017 with 11 GB of GDDR5X on a 352-bit bus, a capacity that was rare for a long time and still gives it an advantage over 8 GB cards. This guide quantifies what that memory enables today, what public benchmarks show, and corrects a common misconception: FP16 on Pascal is not “twice as slow”; it is nearly unusable, but that matters little for quantized models.

Choosing a machine? Our picks by budget →

By Mohamed Meguedmi·Update 2026-09-30·Tested on Windows, macOS, and Linux
Recommended hardware

Good value for local AI: a GMKtec EVO-X2 64GB / 1TB (Ryzen AI Max+ 395).

A mini PC is a complete machine: check the available memory and engine compatibility. It does not replace macOS/MLX or CUDA.

Why this choice? Our complete guide on GMKtec EVO-X2 64GB / 1TB (Ryzen AI Max+ 395) →

Compare all options by budget, from €800 to €3,500 →

Small budget: RTX 5060 · Large models: RTX 5090 · Mac Studio.

On the go: which laptop for local AI →

Affiliate links — possible commission at no extra cost to you. As an Amazon Associate, BestLLMfor earns from qualifying purchases.

#The GTX 1080 Ti in 2026: what it has and what it doesn't

The GTX 1080 Ti is a Pascal card (GP102 chip) with 3,584 CUDA cores, 11 GB of GDDR5X, and 484 GB/s of bandwidth, with a 250 W power envelope. On the public llama.cpp benchmark, it produces 62.49 tokens/s during generation and 1,084.41 tokens/s during prompt processing on a Q4_0 Llama 2 7B. These results put it on par with a RTX 2060 Super for generation, with 3 GB more: it is this capacity, more than speed, that still makes it worth considering.

Specifications (Wikipedia, GeForce 10 series)
ItemValue
ChipGP102-350, Pascal
CUDA cores3 584
Memory11 GB GDDR5X, 352-bit bus
Bandwidth484 GB/s
FP32 compute power10,609 GFLOPS
FP16 compute power166 GFLOPS
Thermal design power (TDP)250 W

The table already contains the most important information: FP16 is worth about 1/64 of FP32 on this card. The next paragraph explains why this does not hurt performance as much as you might think.

#Pascal's limitations: what matters and what doesn't

Three limitations are real. The first is the lack of Tensor Cores, the matrix-computation units introduced with Volta and Turing: they accelerate prompt processing and training, not generation, which depends on memory. The second is FP16, with its ridiculously low throughput on consumer Pascal (about 166 GFLOPS versus 10,609 in FP32): engines such as llama.cpp work around the problem by computing quantized models with integers. The third is the end of software support, discussed below.

Two frequently repeated claims are false. FP16 isn't simply twice as slow; it's sixty times slower. And Flash Attention isn't limited to Tensor Cores: the llama.cpp benchmark runs the 1080 Ti with Flash Attention enabled, at 1,138.14 tokens/s for prompt processing and 61.38 tokens/s for generation, versus 1,084.41 and 62.49 without it. The gain is small for throughput, but the feature is available and reduces context memory usage.

!
No FP8 or FP4
Recent hardware quantization formats (FP8, FP4) require newer cards. On the 1080 Ti, stick with standard GGUF quantizations (Q4_K_M, Q5_K_M, Q8_0), which use integer computation.

#What 11 GB enables: the decision table

The site's reference point puts an 8B in Q4 at around 5 GB and a 14B at around 9 GB, excluding the KV cache. With 11 GB, a 12B leaves 4 GB of headroom and a 14B only 2 GB. The table applies the performance observed on the benchmark (49% of the theoretical 484 GB/s ceiling, or 62.49 tokens/s for 3.82 GB of weights) to common models: these are projections, not measurements.

Models at 11 GB, theoretical ceiling and projection at 49% efficiency
Model (Q4)WeightsHeadroom on 11 GBCapped at 484 GB/sProjection
Granite 4.2 8B4.6 GB6.4 GB105 tokens/sabout 51 tokens/s
Qwen3.5 9B6 GB5 GB81 tokens/sabout 40 tokens/s
Gemma 4 12B7 GB4 GB69 tokens/sapproximately 34 tokens/s
Phi-4 14B9 GB2 GB54 tokens/sabout 26 tokens/s, short context
Gemma 4 26B-A4B (MoE)16 GBnegativespills over from the GPUout of reach

The most interesting range is 12B: it is the largest model that remains comfortable, with a context of several thousand tokens. On an 8 GB card, this same model leaves barely 1 GB and forces you to reduce the context drastically. Beyond that, a 14B model in Q4 fits but leaves no headroom, and a 30-billion-parameter model does not fit.

#What public benchmarks say

QuelLLM does not test this card. The benchmark in the “Performance of llama.cpp on Nvidia CUDA” discussion compares all cards using the same model, making it possible to place the 1080 Ti relative to its peers.

llama.cpp CUDA benchmark, Llama 2 7B Q4_0, without Flash Attention
CardMemoryPrompt (pp512)Generation (tg128)
GTX 1080 Ti11 GB GDDR5X, 352-bit1,084.41 tokens/s62.49 tokens/s
Tesla P40 (Pascal)24 GB GDDR5, 384-bit1,007.42 tokens/s54.74 tokens/s
RTX 2060 Super8 GB GDDR6, 256-bit1,420.24 tokens/s60.04 tokens/s
RTX 306012 GB GDDR6, 192-bit2,137.50 tokens/s75.57 tokens/s
RTX 2080 Ti11 GB GDDR6, 352-bit2,890.66 tokens/s107.51 tokens/s

Three patterns emerge. In generation, the 1080 Ti matches RTX 2060 Super and remains 17% behind a 12 GB RTX 3060. In prompt processing, the gap widens: RTX 3060 nearly doubles its speed, while RTX 2080 Ti, with the same capacity, is 2.7 times faster. Finally, the Tesla P40, a 24 GB Pascal card, reaches 54.74 tokens/s, showing that 24 GB is achievable on this architecture, albeit at a lower throughput than the 1080 Ti.

i
Generation versus prompt processing
For chat with short messages, generation dominates and the 1080 Ti holds its own. For summarizing long documents or using RAG, prompt processing matters: Turing and Ampere cards gain a factor of 2 to 3.

#Make the Most of 11 GB: Use a Lower Quantization

An advantage specific to 11 GB is that you can increase quality rather than size. The QuelLLM catalog lists, for a Granite 4.2 8B: 4.6 GB in Q4, 6 GB in Q5, and 9 GB in Q8. On 8 GB, only Q4 leaves room to spare. On 11 GB, Q5 (6 GB) remains very comfortable, while Q8 (9 GB) works with a short context.

Granite 4.2 8B depending on quantization at 11 GB
QuantizationWeightsMarginCapped at 484 GB/s
Q44.6 GB6.4 GB105 tokens/s
Q56 GB5 GB81 tokens/s
Q89 GB2 GB54 tokens/s

The trade-off is clear: each quantization level reduces model errors but slows generation, since more weights are reread for each token. For everyday chat, Q5 is a good balance; Q8 makes sense when precision matters (extraction, code) and throughput matters less. The quantization guide details the quality difference.

#Two cards for 22 GB: possible, but verify first

The llama.cpp benchmark notes in its instructions that contributors with multiple GPUs must force a single GPU with -sm none -mg, unless the model is too large for VRAM: the engine therefore does distribute a model across multiple cards. Two GTX 1080 Ti would provide 22 GB for the weights, enough to host a 32-billion-parameter model in Q4 (19 to 20 GB according to the site's benchmark), with no context headroom.

The split does not add the speeds together: layers are divided among cards and transferred over the PCIe bus, which limits the gain. This setup also requires a suitable case, power supply, and motherboard, and doubles your exposure to Pascal’s end of support. For a 24 GB project, first compare the cost of a single newer card.

#Install and verify in four steps

  1. 01
    Check the driver
    Run nvidia-smi. For compute capability 5.0 to 6.2 cards, including GTX 1080 Ti, Ollama's documentation requires driver 570 or newer.
  2. 02
    Install Ollama
    Follow your system's installation guide. Ollama recognizes the card as GTX 1080 Ti (compute capability 6.1) in its official list.
  3. 03
    Choose a suitable model
    Start with an 8B in Q4 to validate the pipeline, then test a 12B with a 4,096-token context.
  4. 04
    Check placement
    Run ollama ps: the PROCESSOR column should show 100% GPU. If not, reduce the context or the model.

If loading fails, do not look for an exotic workaround variable: Ollama enables Flash Attention automatically when the engine and card support it, and the documented variable is OLLAMA_FLASH_ATTENTION (1 to force it, 0 to disable it). The troubleshooting guide details common errors.

#Pascal end of life: what to plan for

The CUDA 13 release notes indicate that Maxwell, Pascal, and Volta are no longer supported. Wikipedia reports that NVIDIA ended Game Ready support for these architectures in December 2025, with security updates through October 2028. Today, the GTX 1080 Ti works with Ollama and driver 570 or later; the risk is that future inference engines will target only CUDA 13. Record your driver and Ollama versions before any update.

#Verdict: keep it, yes; buy it, provided that

Decision based on your situation
Your situationRecommendation
You already have a 1080 TiKeep it: 8–12B comfortably, 62 tokens/s in the benchmark
You want a 12B model with context without spendingIt fits: 11 GB, with 4 GB of headroom
You work with long documents or RAGAn Ampere or Turing card reads the prompt 2 to 3 times faster
You’re looking for a purchase that will lastTarget Turing or newer: CUDA 13 no longer targets Pascal
You want 16 GB or moreMove up a tier: the 1080 Ti tops out at 11 GB

To compare current options without listing prices that change every week, the site's pages show the models suited to each memory capacity and link to the graphics card pricing page.

FAQ
Can the GTX 1080 Ti still run an LLM in 2026?+
Yes. The public llama.cpp benchmark measures 62.49 tokens/s on a 7B in Q4_0, and its 11 GB can accommodate an 8B or 9B with room to spare, and a 12B with a modest context. It remains relevant as long as the tools support Pascal.
GTX 1080 Ti or RTX 2060 Super for an LLM?+
For generation, they perform similarly: 62.49 versus 60.04 tokens/s in benchmarks. The 1080 Ti has 3 GB more memory, which matters for a 12B model; the 2060 Super processes prompts faster (1,420 versus 1,084 tokens/s) and remains supported by CUDA 13.
Is the GTX 1080 Ti slower than a RTX 3060 for an LLM?+
Yes, by 17% in generation (62.49 versus 75.57 tokens/s) and by almost half in prompt processing (1,084 versus 2,137 tokens/s). The 3060 has only 12 GB versus 11: the capacity advantage is minimal. The 3060 wins mainly because of its longer software support.
Does the 1080 Ti support Flash Attention?+
Yes. The llama.cpp benchmark measures it with Flash Attention at 1,138.14 tokens/s for prompt processing and 61.38 for generation. The throughput gain is small, but the feature reduces context memory usage. Ollama enables it automatically when the card supports it.
Which driver do you need for the GTX 1080 Ti with Ollama?+
Ollama's documentation requires an NVIDIA 570 or newer driver for compute capability 5.0 to 6.2 cards, including the GTX 1080 Ti. Check your version with nvidia-smi before installing. NVIDIA plans security updates for the 580 branch through October 2028, but no new optimizations.
When should you replace the GTX 1080 Ti?+
When a tool you need stops supporting Pascal, or when your workloads require more than 11 GB or fast prompt processing, such as long documents and RAG. CUDA 13 has already dropped Pascal: that’s the signal to watch with every Ollama or llama.cpp update.
Did this guide help you?

Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.

Prices in euros (€) are French market prices including VAT, checked by QuelLLM. US prices differ: the Amazon buttons show the current US price.