Which LLM on GTX 1080 Ti (11 GB) ?
Yes: the GTX 1080 Ti (11 GB, 484 GB/s) remains one of the best Pascal cards for a local LLM. The public llama.cpp benchmark measures 62.49 tokens/s in generation on a Llama 2 7B in Q4_0, on par with a RTX 2060 Super, with 3 GB more memory. It handles an 8B or 9B in Q4 with room to spare, and a 12B with a modest context. Its software future is limited, however: Pascal has been removed from CUDA 13.
The GTX 1080 Ti was released in March 2017 with 11 GB of GDDR5X on a 352-bit bus, a capacity that was rare for a long time and still gives it an advantage over 8 GB cards. This guide quantifies what that memory enables today, what public benchmarks show, and corrects a common misconception: FP16 on Pascal is not “twice as slow”; it is nearly unusable, but that matters little for quantized models.
Choosing a machine? Our picks by budget →
Good value for local AI: a GMKtec EVO-X2 64GB / 1TB (Ryzen AI Max+ 395).
A mini PC is a complete machine: check the available memory and engine compatibility. It does not replace macOS/MLX or CUDA.
Why this choice? Our complete guide on GMKtec EVO-X2 64GB / 1TB (Ryzen AI Max+ 395) →
Compare all options by budget, from €800 to €3,500 →
Small budget: RTX 5060 · Large models: RTX 5090 · Mac Studio.
On the go: which laptop for local AI →
Affiliate links — possible commission at no extra cost to you. As an Amazon Associate, BestLLMfor earns from qualifying purchases.
#The GTX 1080 Ti in 2026: what it has and what it doesn't
The GTX 1080 Ti is a Pascal card (GP102 chip) with 3,584 CUDA cores, 11 GB of GDDR5X, and 484 GB/s of bandwidth, with a 250 W power envelope. On the public llama.cpp benchmark, it produces 62.49 tokens/s during generation and 1,084.41 tokens/s during prompt processing on a Q4_0 Llama 2 7B. These results put it on par with a RTX 2060 Super for generation, with 3 GB more: it is this capacity, more than speed, that still makes it worth considering.
| Item | Value |
|---|---|
| Chip | GP102-350, Pascal |
| CUDA cores | 3 584 |
| Memory | 11 GB GDDR5X, 352-bit bus |
| Bandwidth | 484 GB/s |
| FP32 compute power | 10,609 GFLOPS |
| FP16 compute power | 166 GFLOPS |
| Thermal design power (TDP) | 250 W |
The table already contains the most important information: FP16 is worth about 1/64 of FP32 on this card. The next paragraph explains why this does not hurt performance as much as you might think.
#Pascal's limitations: what matters and what doesn't
Three limitations are real. The first is the lack of Tensor Cores, the matrix-computation units introduced with Volta and Turing: they accelerate prompt processing and training, not generation, which depends on memory. The second is FP16, with its ridiculously low throughput on consumer Pascal (about 166 GFLOPS versus 10,609 in FP32): engines such as llama.cpp work around the problem by computing quantized models with integers. The third is the end of software support, discussed below.
Two frequently repeated claims are false. FP16 isn't simply twice as slow; it's sixty times slower. And Flash Attention isn't limited to Tensor Cores: the llama.cpp benchmark runs the 1080 Ti with Flash Attention enabled, at 1,138.14 tokens/s for prompt processing and 61.38 tokens/s for generation, versus 1,084.41 and 62.49 without it. The gain is small for throughput, but the feature is available and reduces context memory usage.
#What 11 GB enables: the decision table
The site's reference point puts an 8B in Q4 at around 5 GB and a 14B at around 9 GB, excluding the KV cache. With 11 GB, a 12B leaves 4 GB of headroom and a 14B only 2 GB. The table applies the performance observed on the benchmark (49% of the theoretical 484 GB/s ceiling, or 62.49 tokens/s for 3.82 GB of weights) to common models: these are projections, not measurements.
| Model (Q4) | Weights | Headroom on 11 GB | Capped at 484 GB/s | Projection |
|---|---|---|---|---|
| Granite 4.2 8B | 4.6 GB | 6.4 GB | 105 tokens/s | about 51 tokens/s |
| Qwen3.5 9B | 6 GB | 5 GB | 81 tokens/s | about 40 tokens/s |
| Gemma 4 12B | 7 GB | 4 GB | 69 tokens/s | approximately 34 tokens/s |
| Phi-4 14B | 9 GB | 2 GB | 54 tokens/s | about 26 tokens/s, short context |
| Gemma 4 26B-A4B (MoE) | 16 GB | negative | spills over from the GPU | out of reach |
The most interesting range is 12B: it is the largest model that remains comfortable, with a context of several thousand tokens. On an 8 GB card, this same model leaves barely 1 GB and forces you to reduce the context drastically. Beyond that, a 14B model in Q4 fits but leaves no headroom, and a 30-billion-parameter model does not fit.
#What public benchmarks say
QuelLLM does not test this card. The benchmark in the “Performance of llama.cpp on Nvidia CUDA” discussion compares all cards using the same model, making it possible to place the 1080 Ti relative to its peers.
| Card | Memory | Prompt (pp512) | Generation (tg128) |
|---|---|---|---|
| GTX 1080 Ti | 11 GB GDDR5X, 352-bit | 1,084.41 tokens/s | 62.49 tokens/s |
| Tesla P40 (Pascal) | 24 GB GDDR5, 384-bit | 1,007.42 tokens/s | 54.74 tokens/s |
| RTX 2060 Super | 8 GB GDDR6, 256-bit | 1,420.24 tokens/s | 60.04 tokens/s |
| RTX 3060 | 12 GB GDDR6, 192-bit | 2,137.50 tokens/s | 75.57 tokens/s |
| RTX 2080 Ti | 11 GB GDDR6, 352-bit | 2,890.66 tokens/s | 107.51 tokens/s |
Three patterns emerge. In generation, the 1080 Ti matches RTX 2060 Super and remains 17% behind a 12 GB RTX 3060. In prompt processing, the gap widens: RTX 3060 nearly doubles its speed, while RTX 2080 Ti, with the same capacity, is 2.7 times faster. Finally, the Tesla P40, a 24 GB Pascal card, reaches 54.74 tokens/s, showing that 24 GB is achievable on this architecture, albeit at a lower throughput than the 1080 Ti.
#Make the Most of 11 GB: Use a Lower Quantization
An advantage specific to 11 GB is that you can increase quality rather than size. The QuelLLM catalog lists, for a Granite 4.2 8B: 4.6 GB in Q4, 6 GB in Q5, and 9 GB in Q8. On 8 GB, only Q4 leaves room to spare. On 11 GB, Q5 (6 GB) remains very comfortable, while Q8 (9 GB) works with a short context.
| Quantization | Weights | Margin | Capped at 484 GB/s |
|---|---|---|---|
| Q4 | 4.6 GB | 6.4 GB | 105 tokens/s |
| Q5 | 6 GB | 5 GB | 81 tokens/s |
| Q8 | 9 GB | 2 GB | 54 tokens/s |
The trade-off is clear: each quantization level reduces model errors but slows generation, since more weights are reread for each token. For everyday chat, Q5 is a good balance; Q8 makes sense when precision matters (extraction, code) and throughput matters less. The quantization guide details the quality difference.
#Two cards for 22 GB: possible, but verify first
The llama.cpp benchmark notes in its instructions that contributors with multiple GPUs must force a single GPU with -sm none -mg, unless the model is too large for VRAM: the engine therefore does distribute a model across multiple cards. Two GTX 1080 Ti would provide 22 GB for the weights, enough to host a 32-billion-parameter model in Q4 (19 to 20 GB according to the site's benchmark), with no context headroom.
The split does not add the speeds together: layers are divided among cards and transferred over the PCIe bus, which limits the gain. This setup also requires a suitable case, power supply, and motherboard, and doubles your exposure to Pascal’s end of support. For a 24 GB project, first compare the cost of a single newer card.
#Install and verify in four steps
- 01Check the driverRun nvidia-smi. For compute capability 5.0 to 6.2 cards, including GTX 1080 Ti, Ollama's documentation requires driver 570 or newer.
- 02Install OllamaFollow your system's installation guide. Ollama recognizes the card as GTX 1080 Ti (compute capability 6.1) in its official list.
- 03Choose a suitable modelStart with an 8B in Q4 to validate the pipeline, then test a 12B with a 4,096-token context.
- 04Check placementRun ollama ps: the PROCESSOR column should show 100% GPU. If not, reduce the context or the model.
If loading fails, do not look for an exotic workaround variable: Ollama enables Flash Attention automatically when the engine and card support it, and the documented variable is OLLAMA_FLASH_ATTENTION (1 to force it, 0 to disable it). The troubleshooting guide details common errors.
#Pascal end of life: what to plan for
The CUDA 13 release notes indicate that Maxwell, Pascal, and Volta are no longer supported. Wikipedia reports that NVIDIA ended Game Ready support for these architectures in December 2025, with security updates through October 2028. Today, the GTX 1080 Ti works with Ollama and driver 570 or later; the risk is that future inference engines will target only CUDA 13. Record your driver and Ollama versions before any update.
#Verdict: keep it, yes; buy it, provided that
| Your situation | Recommendation |
|---|---|
| You already have a 1080 Ti | Keep it: 8–12B comfortably, 62 tokens/s in the benchmark |
| You want a 12B model with context without spending | It fits: 11 GB, with 4 GB of headroom |
| You work with long documents or RAG | An Ampere or Turing card reads the prompt 2 to 3 times faster |
| You’re looking for a purchase that will last | Target Turing or newer: CUDA 13 no longer targets Pascal |
| You want 16 GB or more | Move up a tier: the 1080 Ti tops out at 11 GB |
To compare current options without listing prices that change every week, the site's pages show the models suited to each memory capacity and link to the graphics card pricing page.
- Which LLM on RTX 3060 12 GB
- Which LLM on RTX 2080 Ti (11 GB)
- Which models for 12 GB of VRAM, all cards
- Compare GPU prices for AI
- Source: public llama.cpp benchmark on CUDA
- Source: Ollama documentation, supported NVIDIA hardware
- Source: GeForce 10 series, specs
- Source: CUDA Toolkit release notes
Can the GTX 1080 Ti still run an LLM in 2026?+
GTX 1080 Ti or RTX 2060 Super for an LLM?+
Is the GTX 1080 Ti slower than a RTX 3060 for an LLM?+
Does the 1080 Ti support Flash Attention?+
Which driver do you need for the GTX 1080 Ti with Ollama?+
When should you replace the GTX 1080 Ti?+
Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.