Beginner 11 minGTX 10

Which LLM on GTX 1070 / 1080 (8 GB) ?

Direct response

Yes: a GTX 1070, 1070 Ti, or 1080 with 8 GB can run a 7- to 9-billion-parameter model in Q4 without difficulty, with a reasonable context. The public llama.cpp benchmark measures 37.82 tokens/s for a GTX 1070 Ti on a 7B Q4_0; the 1080, with faster memory, should do better, though there is no direct measurement. Beyond 9 billion parameters, 8 GB becomes a hard limit, and the Pascal architecture has no software future.

The GTX 1070 (June 2016), 1070 Ti (November 2017), and 1080 (May 2016) share the same GP104 chip and 8 GB of memory. They have become the most tempting used cards for trying a local LLM without spending much. This guide quantifies what their memory allows, what public measurements say about their speed, and why Pascal's end-of-support date matters more than the price.

Choosing a machine? Our picks by budget →

By Mohamed Meguedmi·Update 2026-09-30·Tested on Windows, macOS, and Linux
Recommended hardware

Good value for local AI: a GMKtec EVO-X2 64GB / 1TB (Ryzen AI Max+ 395).

A mini PC is a complete machine: check the available memory and engine compatibility. It does not replace macOS/MLX or CUDA.

Why this choice? Our complete guide on GMKtec EVO-X2 64GB / 1TB (Ryzen AI Max+ 395) →

Compare all options by budget, from €800 to €3,500 →

Small budget: RTX 5060 · Large models: RTX 5090 · Mac Studio.

On the go: which laptop for local AI →

Affiliate links — possible commission at no extra cost to you. As an Amazon Associate, BestLLMfor earns from qualifying purchases.

#GTX 1070, 1070 Ti, 1080: what sets them apart for an LLM

For an LLM, these three cards have the same capacity, 8 GB, and differ in memory bandwidth, which determines generation speed. The 1070 and 1070 Ti use GDDR5 on a 256-bit bus, for 256 GB/s; the 1080 reaches 320 GB/s with GDDR5X. Because an LLM rereads all its weights for every token, this 25% difference translates almost directly into speed, far more than the 1080's 2,560 cores versus the 1070's 1,920. None has Tensor Cores: Pascal remains limited to standard CUDA computation.

The three cards at a glance (Wikipedia, GeForce 10 series; measurement: llama.cpp benchmark)
CardCUDA coresMemoryBandwidthMeasured throughput (7B Q4_0)
GTX 10701 9208 GB GDDR5256 GB/snot measured in the benchmark
GTX 1070 Ti2 4328 GB GDDR5256 GB/s37.82 tokens/s
GTX 10802 5608 GB GDDR5X320 GB/snot measured in the benchmark

A GTX 1080 Ti, with 11 GB and 484 GB/s, belongs to another category: it has its own page. If your card is a 1080 Ti, read the guide dedicated to it instead.

#What 8 GB really enable

The key criterion is the ratio between the weights and the available memory, since the KV cache is added to them. The site's benchmark puts a 7–8B model in Q4 at around 5 GB of weights: on 8 GB, about 3 GB remains for context and the system, providing real headroom. A 9B model in Q4, at 6 GB, leaves 2 GB: it is possible with a short context. At 12 billion parameters, the weights reach 7 GB: almost nothing remains.

Models on 8 GB, theoretical ceilings (bandwidth ÷ weights; these are not measurements)
Model (Q4)WeightsHeadroom on 8 GBCapped at 256 GB/sCeiling at 320 GB/s
Granite 4.1 3B2 GB6 GB128 tokens/s160 tokens/s
Granite 4.2 8B4.6 GB3.4 GB55 tokens/s70 tokens/s
Qwen3.5 9B6 GB2 GB43 tokens/s53 tokens/s
Gemma 4 12B7 GB1 GBnot recommendednot recommended
Phi-4 14B9 GBnegativespills over from the GPUspills over from the GPU

A theoretical ceiling is never reached. In the public benchmark, the GTX 1070 Ti reaches 37.82 tokens/s on a 3.56 GiB model, or about 56% of its 256 GB/s ceiling. Applying the same efficiency to the models above, a Granite 4.2 8B would run at around 30 tokens/s on a 1070 Ti: this is a proportional extrapolation, not a measurement.

#What public benchmarks say

QuelLLM does not test these cards. The figures come from the “Performance of llama.cpp on Nvidia CUDA” discussion in the llama.cpp repository, where each contributor runs llama-bench on Llama 2 7B in Q4_0 with all layers on the GPU. The GTX 1070 Ti reports 714.44 tokens/s for prompt processing and 37.82 tokens/s for generation.

For the cards on the same page that have a measurement, the comparison is useful: the RTX 2060 Super, an 8 GB, 256-bit Turing card, produces 60.04 tokens/s on the same test, versus 37.82 for the 1070 Ti. With similar memory bandwidth, the gain is 59%. The benchmark therefore explains the gap better than the core count or price: Turing makes better use of its memory.

llama.cpp CUDA benchmark, Llama 2 7B Q4_0, without Flash Attention
CardMemoryPrompt (pp512)Generation (tg128)
GTX 1070 Ti8 GB GDDR5, 256-bit714.44 tokens/s37.82 tokens/s
RTX 2060 Super8 GB GDDR6, 256-bit1,420.24 tokens/s60.04 tokens/s
GTX 1080 Ti11 GB GDDR5X, 352-bit1,084.41 tokens/s62.49 tokens/s

These figures should be read with three caveats. Each benchmark row is an individual contribution: the driver, system, cooling, and card manufacturer vary from one contributor to another, explaining a few points of difference for the same chip. The test model, a Llama 2 7B, is older and smaller than current models: it serves as a common benchmark, not a prediction for any specific model. Finally, generation depends mainly on memory, while prompt processing depends on compute power: the two columns are measuring different things.

i
Throughput differences also come from the prompt
Prompt processing is about twice as slow on the 1070 Ti as on the 2060 Super. With a long document pasted into the conversation, the wait before the first response is therefore significantly longer on Pascal.

#Estimate GTX 1080 without measuring: method and limitations

The 1080 does not appear in the benchmark. To get an idea, we can apply the memory bandwidth ratio: 320 ÷ 256 = 1.25, or approximately 47 tokens/s (37.82 × 1.25) if efficiency remains identical to that of the 1070 Ti. This is a plausible assumption, not a result: drivers, the llama.cpp version, temperature, or card cooling can significantly affect throughput.

If you buy a used card for this purpose, ask the seller for a screenshot of a real test, or run llama-bench yourself on the benchmark model as soon as it arrives: the result can be compared directly with the public table.

#The context: where the 3 GB of headroom goes

On 8 GB, a Granite 4.2 8B in Q4 leaves about 3,4 GB. This reserve is shared by the KV cache, which grows with every conversation token, the computation buffers, and the space used by the driver and display if the card also drives your screen. Ollama starts with a 4 096-token window, adjustable through the OLLAMA_CONTEXT_LENGTH variable; doubling the window doubles the KV cache.

A simple rule of thumb: if you want to work with long documents (16,000 tokens or more), choose a model with 3 to 4 billion parameters, or an 8B model with a quantized KV cache. If you chat in short messages, the 8B in Q4 is the right choice. In both cases, use ollama ps to verify that the model is loaded 100% on the GPU: sharing with the processor indicates that you have run out of headroom.

#Flash Attention on Pascal: useful or not?

A common misconception is that Flash Attention is limited to cards with Tensor Cores. The llama.cpp benchmark disproves it: it runs every card with and without Flash Attention, and the GTX 1080 Ti, from the Pascal generation, appears in both tables. With Flash Attention, it goes from 1,084.41 to 1,138.14 tokens/s for prompt processing, about 5% better, and from 62.49 to 61.38 tokens/s for generation, 2% worse. Flash Attention's main benefit is not speed but memory: it reduces the context footprint, which matters on 8 GB.

Ollama automatically enables Flash Attention when the engine and GPU support it; the OLLAMA_FLASH_ATTENTION=1 variable forces it on, and 0 disables it. KV-cache quantization, which further reduces context memory usage, requires Flash Attention to be enabled.

Force Flash Attention when starting the server
OLLAMA_FLASH_ATTENTION=1 ollama serve

#What can you use an 8 GB Pascal for?

Common workloads and suitability
UsageSuitabilityRecommended setting
Chat, rewriting, translationGoodGranite 4.2 8B in Q4, 4,096-token context
Code assistance in the editorCorrect7–8B model, medium context, a single model loaded
Summarizing long documentsLimited3–4B model or quantized KV cache
Search your documents (RAG)Correct3–8B model, short excerpts, no 12B
Agents with chained tool callsLowToo much context and too many calls for 8 GB

#Pascal: software support is the real limitation

The main risk comes not from speed but from the software. Ollama requires a NVIDIA 570 driver or newer for compute capability 5.0 to 6.2 cards, and the GTX 1070 through 1080 are included. But the CUDA 13 release notes state that Pascal has been removed, and Wikipedia notes that NVIDIA ended Game Ready support for this generation in December 2025, with security updates through October 2028.

Aujourd'hui
Ollama and llama.cpp work with driver 570 or later and CUDA 12 builds.
Tomorrow
New features and new engines will target CUDA 13: each update may drop Pascal.
Security
Driver security fixes are available through October 2028, but no new optimizations.
!
Don’t be misled by the date 2028
This date concerns driver security, not AI tool compatibility. An inference engine may stop supporting Pascal much earlier: plan for a replacement card.

#Is a 1070 Ti better, or should I get something else?

For the same secondhand budget, the comparison comes down to three criteria: memory (8 GB minimum), architecture generation (Turing or newer to benefit from CUDA 13), and bandwidth. Without quoting prices, which change every week, here is how the options compare.

Entry-level options for a local LLM
OptionStrengthsWeaknesses
GTX 1070 / 1070 Ti / 1080 (already owned)Free; comfortable with 7–9BPascal at end of support, slower at equal bandwidth
RTX 2060 Super (8 GB)60.04 tokens/s in benchmarks, Turing, Tensor CoresStill 8 GB
RTX 3050 (6 or 8 GB)Ampere, long software supportBandwidth limited depending on the version
12 GB card (RTX 3060 12 GB)Headroom for 12–14B and long contextsHigher budget

#Verdict: worth exploring, not worth investing in

If you already have a GTX 1070, 1070 Ti, or 1080, it’s an excellent entry point: 8 GB is enough for an 8B model in Q4, and throughput remains comfortable for chat. If you’re looking to buy, prioritize an 8 GB card with a Turing or newer architecture: at the same memory capacity, throughput is significantly better and software support lasts longer.

One final practical criterion: power consumption. These cards are 150 to 180 W models while gaming; during inference, they run cooler than in games, but a poorly ventilated case can reduce throughput. Check temperatures during a long generation before concluding that the card is slow.

FAQ
Can a GTX 1070 or 1080 run an 8B model?+
Yes. A Granite 4.2 8B in Q4 weighs about 4.6 GB, leaving more than 3 GB on 8 GB for the context and system. The llama.cpp benchmark reports 37.82 tokens/s for a 1070 Ti on a 7B Q4_0, a comfortable rate for chat.
GTX 1070 or GTX 1080 for a local LLM?+
The 1080. Its GDDR5X memory at 320 GB/s versus 256 GB/s for the 1070 increases generation throughput by about 25% with the same model, since speed is memory-bound. Capacity remains the same, 8 GB. Check the condition and cooling of a used card.
How many tokens per second on a GTX 1080?+
No measurement of GTX 1080 exists on the public llama.cpp benchmark. GTX 1070 Ti, at 256 GB/s, reaches 37,82 tokens/s on a 7B Q4_0. In proportion to memory bandwidth, the 1080 should exceed 45 tokens/s—an estimate to verify with llama-bench.
Can Qwen3.5 9B run on a GTX 1070?+
Yes, but with little headroom: the Q4 weights take up about 6 GB of 8, leaving 2 GB for the KV cache and system. Keep the context short, and use ollama ps to check that the model stays on the GPU.
Should you buy a used GTX 1080 in 2026?+
Only at a very low price and for experimentation. Pascal is removed from CUDA 13, and Game Ready support ended in December 2025; an RTX 2060 Super 8 GB produces 60.04 tokens/s in the benchmark versus 37.82 for a 1070 Ti. At a similar budget, the Turing card is the better buy.
Did this guide help you?

Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.