Which LLM on GTX 1070 / 1080 (8 GB) ?
Yes: a GTX 1070, 1070 Ti, or 1080 with 8 GB can run a 7- to 9-billion-parameter model in Q4 without difficulty, with a reasonable context. The public llama.cpp benchmark measures 37.82 tokens/s for a GTX 1070 Ti on a 7B Q4_0; the 1080, with faster memory, should do better, though there is no direct measurement. Beyond 9 billion parameters, 8 GB becomes a hard limit, and the Pascal architecture has no software future.
The GTX 1070 (June 2016), 1070 Ti (November 2017), and 1080 (May 2016) share the same GP104 chip and 8 GB of memory. They have become the most tempting used cards for trying a local LLM without spending much. This guide quantifies what their memory allows, what public measurements say about their speed, and why Pascal's end-of-support date matters more than the price.
Choosing a machine? Our picks by budget →
Good value for local AI: a GMKtec EVO-X2 64GB / 1TB (Ryzen AI Max+ 395).
A mini PC is a complete machine: check the available memory and engine compatibility. It does not replace macOS/MLX or CUDA.
Why this choice? Our complete guide on GMKtec EVO-X2 64GB / 1TB (Ryzen AI Max+ 395) →
Compare all options by budget, from €800 to €3,500 →
Small budget: RTX 5060 · Large models: RTX 5090 · Mac Studio.
On the go: which laptop for local AI →
Affiliate links — possible commission at no extra cost to you. As an Amazon Associate, BestLLMfor earns from qualifying purchases.
#GTX 1070, 1070 Ti, 1080: what sets them apart for an LLM
For an LLM, these three cards have the same capacity, 8 GB, and differ in memory bandwidth, which determines generation speed. The 1070 and 1070 Ti use GDDR5 on a 256-bit bus, for 256 GB/s; the 1080 reaches 320 GB/s with GDDR5X. Because an LLM rereads all its weights for every token, this 25% difference translates almost directly into speed, far more than the 1080's 2,560 cores versus the 1070's 1,920. None has Tensor Cores: Pascal remains limited to standard CUDA computation.
| Card | CUDA cores | Memory | Bandwidth | Measured throughput (7B Q4_0) |
|---|---|---|---|---|
| GTX 1070 | 1 920 | 8 GB GDDR5 | 256 GB/s | not measured in the benchmark |
| GTX 1070 Ti | 2 432 | 8 GB GDDR5 | 256 GB/s | 37.82 tokens/s |
| GTX 1080 | 2 560 | 8 GB GDDR5X | 320 GB/s | not measured in the benchmark |
A GTX 1080 Ti, with 11 GB and 484 GB/s, belongs to another category: it has its own page. If your card is a 1080 Ti, read the guide dedicated to it instead.
#What 8 GB really enable
The key criterion is the ratio between the weights and the available memory, since the KV cache is added to them. The site's benchmark puts a 7–8B model in Q4 at around 5 GB of weights: on 8 GB, about 3 GB remains for context and the system, providing real headroom. A 9B model in Q4, at 6 GB, leaves 2 GB: it is possible with a short context. At 12 billion parameters, the weights reach 7 GB: almost nothing remains.
| Model (Q4) | Weights | Headroom on 8 GB | Capped at 256 GB/s | Ceiling at 320 GB/s |
|---|---|---|---|---|
| Granite 4.1 3B | 2 GB | 6 GB | 128 tokens/s | 160 tokens/s |
| Granite 4.2 8B | 4.6 GB | 3.4 GB | 55 tokens/s | 70 tokens/s |
| Qwen3.5 9B | 6 GB | 2 GB | 43 tokens/s | 53 tokens/s |
| Gemma 4 12B | 7 GB | 1 GB | not recommended | not recommended |
| Phi-4 14B | 9 GB | negative | spills over from the GPU | spills over from the GPU |
A theoretical ceiling is never reached. In the public benchmark, the GTX 1070 Ti reaches 37.82 tokens/s on a 3.56 GiB model, or about 56% of its 256 GB/s ceiling. Applying the same efficiency to the models above, a Granite 4.2 8B would run at around 30 tokens/s on a 1070 Ti: this is a proportional extrapolation, not a measurement.
#What public benchmarks say
QuelLLM does not test these cards. The figures come from the “Performance of llama.cpp on Nvidia CUDA” discussion in the llama.cpp repository, where each contributor runs llama-bench on Llama 2 7B in Q4_0 with all layers on the GPU. The GTX 1070 Ti reports 714.44 tokens/s for prompt processing and 37.82 tokens/s for generation.
For the cards on the same page that have a measurement, the comparison is useful: the RTX 2060 Super, an 8 GB, 256-bit Turing card, produces 60.04 tokens/s on the same test, versus 37.82 for the 1070 Ti. With similar memory bandwidth, the gain is 59%. The benchmark therefore explains the gap better than the core count or price: Turing makes better use of its memory.
| Card | Memory | Prompt (pp512) | Generation (tg128) |
|---|---|---|---|
| GTX 1070 Ti | 8 GB GDDR5, 256-bit | 714.44 tokens/s | 37.82 tokens/s |
| RTX 2060 Super | 8 GB GDDR6, 256-bit | 1,420.24 tokens/s | 60.04 tokens/s |
| GTX 1080 Ti | 11 GB GDDR5X, 352-bit | 1,084.41 tokens/s | 62.49 tokens/s |
These figures should be read with three caveats. Each benchmark row is an individual contribution: the driver, system, cooling, and card manufacturer vary from one contributor to another, explaining a few points of difference for the same chip. The test model, a Llama 2 7B, is older and smaller than current models: it serves as a common benchmark, not a prediction for any specific model. Finally, generation depends mainly on memory, while prompt processing depends on compute power: the two columns are measuring different things.
#Estimate GTX 1080 without measuring: method and limitations
The 1080 does not appear in the benchmark. To get an idea, we can apply the memory bandwidth ratio: 320 ÷ 256 = 1.25, or approximately 47 tokens/s (37.82 × 1.25) if efficiency remains identical to that of the 1070 Ti. This is a plausible assumption, not a result: drivers, the llama.cpp version, temperature, or card cooling can significantly affect throughput.
If you buy a used card for this purpose, ask the seller for a screenshot of a real test, or run llama-bench yourself on the benchmark model as soon as it arrives: the result can be compared directly with the public table.
#The context: where the 3 GB of headroom goes
On 8 GB, a Granite 4.2 8B in Q4 leaves about 3,4 GB. This reserve is shared by the KV cache, which grows with every conversation token, the computation buffers, and the space used by the driver and display if the card also drives your screen. Ollama starts with a 4 096-token window, adjustable through the OLLAMA_CONTEXT_LENGTH variable; doubling the window doubles the KV cache.
A simple rule of thumb: if you want to work with long documents (16,000 tokens or more), choose a model with 3 to 4 billion parameters, or an 8B model with a quantized KV cache. If you chat in short messages, the 8B in Q4 is the right choice. In both cases, use ollama ps to verify that the model is loaded 100% on the GPU: sharing with the processor indicates that you have run out of headroom.
#Flash Attention on Pascal: useful or not?
A common misconception is that Flash Attention is limited to cards with Tensor Cores. The llama.cpp benchmark disproves it: it runs every card with and without Flash Attention, and the GTX 1080 Ti, from the Pascal generation, appears in both tables. With Flash Attention, it goes from 1,084.41 to 1,138.14 tokens/s for prompt processing, about 5% better, and from 62.49 to 61.38 tokens/s for generation, 2% worse. Flash Attention's main benefit is not speed but memory: it reduces the context footprint, which matters on 8 GB.
Ollama automatically enables Flash Attention when the engine and GPU support it; the OLLAMA_FLASH_ATTENTION=1 variable forces it on, and 0 disables it. KV-cache quantization, which further reduces context memory usage, requires Flash Attention to be enabled.
#What can you use an 8 GB Pascal for?
| Usage | Suitability | Recommended setting |
|---|---|---|
| Chat, rewriting, translation | Good | Granite 4.2 8B in Q4, 4,096-token context |
| Code assistance in the editor | Correct | 7–8B model, medium context, a single model loaded |
| Summarizing long documents | Limited | 3–4B model or quantized KV cache |
| Search your documents (RAG) | Correct | 3–8B model, short excerpts, no 12B |
| Agents with chained tool calls | Low | Too much context and too many calls for 8 GB |
#Pascal: software support is the real limitation
The main risk comes not from speed but from the software. Ollama requires a NVIDIA 570 driver or newer for compute capability 5.0 to 6.2 cards, and the GTX 1070 through 1080 are included. But the CUDA 13 release notes state that Pascal has been removed, and Wikipedia notes that NVIDIA ended Game Ready support for this generation in December 2025, with security updates through October 2028.
- Aujourd'hui
- Ollama and llama.cpp work with driver 570 or later and CUDA 12 builds.
- Tomorrow
- New features and new engines will target CUDA 13: each update may drop Pascal.
- Security
- Driver security fixes are available through October 2028, but no new optimizations.
#Is a 1070 Ti better, or should I get something else?
For the same secondhand budget, the comparison comes down to three criteria: memory (8 GB minimum), architecture generation (Turing or newer to benefit from CUDA 13), and bandwidth. Without quoting prices, which change every week, here is how the options compare.
| Option | Strengths | Weaknesses |
|---|---|---|
| GTX 1070 / 1070 Ti / 1080 (already owned) | Free; comfortable with 7–9B | Pascal at end of support, slower at equal bandwidth |
| RTX 2060 Super (8 GB) | 60.04 tokens/s in benchmarks, Turing, Tensor Cores | Still 8 GB |
| RTX 3050 (6 or 8 GB) | Ampere, long software support | Bandwidth limited depending on the version |
| 12 GB card (RTX 3060 12 GB) | Headroom for 12–14B and long contexts | Higher budget |
- Which LLM on RTX 2060 / 2060 Super
- Which LLM on RTX 3060 12 GB
- Which models for 8 GB of VRAM, all cards
- Compare GPU prices for AI
#Verdict: worth exploring, not worth investing in
If you already have a GTX 1070, 1070 Ti, or 1080, it’s an excellent entry point: 8 GB is enough for an 8B model in Q4, and throughput remains comfortable for chat. If you’re looking to buy, prioritize an 8 GB card with a Turing or newer architecture: at the same memory capacity, throughput is significantly better and software support lasts longer.
One final practical criterion: power consumption. These cards are 150 to 180 W models while gaming; during inference, they run cooler than in games, but a poorly ventilated case can reduce throughput. Check temperatures during a long generation before concluding that the card is slow.
- Source: public llama.cpp benchmark on CUDA
- Source: Ollama documentation, supported NVIDIA hardware
- Source: GeForce 10 series, specs
- Source: CUDA Toolkit release notes
Can a GTX 1070 or 1080 run an 8B model?+
GTX 1070 or GTX 1080 for a local LLM?+
How many tokens per second on a GTX 1080?+
Can Qwen3.5 9B run on a GTX 1070?+
Should you buy a used GTX 1080 in 2026?+
Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.