Which LLM for RTX 2070 / 2070 Super (8 GB) ?
A RTX 2070 or 2070 Super (8 GB of GDDR6) comfortably runs 7- to 9-billion-parameter models in Q4: Granite 4.2 8B (4.6 GB of weights) or Qwen 3.5 9B (6 GB) fit entirely on the card. In llama.cpp's reference benchmark, the 2070 Super generates 88 tokens/s. Beyond 9B, it must spill into RAM and throughput collapses.
The RTX 2070 (October 2018) and the RTX 2070 Super (July 2019) are Turing cards with 8 GB of memory. In 2026, they are at the right level for an 8B assistant in Q4, and at the limit for everything else. This guide connects their specifications to a public speed test, explains where the 8 GB goes, and indicates when it is worth changing cards.
Choosing a machine? Our picks by budget →
Buying alternative for this guide: RTX 5060 Ti 16GB (ASUS Prime).
Compare all options by budget, from €800 to €3,500 →
On the go: which laptop for local AI →
Affiliate links — commission possible at no extra cost to you. As an Amazon Associate, BestLLMfor earns from qualifying purchases.
#RTX 2070 and 2070 Super: what 8 GB can do
On a RTX 2070 or a 2070 Super, 7- to 9-billion-parameter models in Q4 quantization fit entirely in VRAM: this is the comfort zone for these cards. According to the QuelLLM catalog, Granite 4.2 8B requires 4.6 GB of weights in Q4 and Qwen 3.5 9B about 6 GB. That leaves room for context, provided you keep its length reasonable. A 12-billion-parameter model such as Gemma 4 12B (7 GB in Q4) leaves barely one gigabyte for the cache and system: it works on paper, but poorly in practice. In terms of speed, the 2070 Super reaches 88 tokens per second on llama.cpp's reference benchmark, using a 7-billion-parameter model in Q4_0—well above human reading speed. The real limit of these cards is therefore not speed but memory capacity.
- RTX 2070
- October 2018, 2,304 CUDA cores, 8 GB of GDDR6 on a 256-bit bus, or 448 GB/s (256 bits at 14 Gbit/s).
- RTX 2070 Super
- July 2019, 2,560 CUDA cores, the same 8 GB and the same 448 GB/s bandwidth, 215 W power consumption.
- What distinguishes them for an LLM
- Almost nothing: bandwidth and capacity are identical. The Super has more cores, which mainly helps with prompt processing, not generation.
This last point is counterintuitive: to generate text, the GPU rereads all the model’s weights for every token, so memory bandwidth matters more than the number of cores. Two cards at 448 GB/s will produce similar generation rates, even if one is significantly faster in games.
#Which models fit in 8 GB
Weights are only part of the bill: you also need the context cache (KV cache), which grows with conversation length, plus some headroom for the driver and display. On an 8 GB card, the rule of thumb is to target weights of no more than 6 GB if you want a context of several thousand tokens. As a reference point, the llama.cpp test tool sees 7,838 MiB of free memory on an 8 GB RTX 3060 Ti, slightly less than the theoretical 8,192 MiB: a few hundred megabytes are already occupied before a model is even loaded.
| Model | Weights in Q4 | Verdict on 8 GB |
|---|---|---|
| Qwen 3.5 4B | 2.3 GB (Q8: 4.3 GB) | Very capable, even in Q8 and with a long context |
| Granite 4.2 8B | 4.6 GB (Q8: 9 GB) | Comfortable range in Q4, with a context of several thousand tokens |
| Qwen 3.5 9B | 6 GB (Q8: 10 GB) | Switch to Q4 with a moderate context; use the quantized K/V cache beyond that |
| Gemma 4 12B | 7 GB (Q8: 13 GB) | Limit: no room for context, likely RAM overflow |
Q8 for an 8B model (9 GB) won’t fit: with 8 GB, Q4 quantization is the rule, and Q5 remains possible for smaller models. To compare quantization levels, the guide to choosing Q4, Q5, or Q8 details the expected quality loss. For a specific project, the site’s VRAM calculator gives you the exact footprint based on the model and context.
#Speed: what the benchmark measures
The only usable public benchmark is the collaborative table in the llama.cpp repository, where users publish the result of the same command: llama-bench on Llama 2 7B in Q4_0, a 3.56 GiB file, with all layers on the GPU. The table notes that results vary by driver, operating system, and card manufacturer, even with the same chip. These are not measurements from the site: the figures below come from that table, and the “reading” column is prompt-processing throughput (pp512), while “generation” is response throughput (tg128).
| Card | Memory | Generation (tok/s) | Prompt processing (tok/s) |
|---|---|---|---|
| RTX 2060 Super | 8 GB | 60,0 | 1 420 |
| RTX 2070 Super | 8 GB | 88,1 | 2 088 |
| RTX 3060 12 GB | 12 GB | 75,6 | 2 138 |
| RTX 2080 Ti | 11 GB | 107,5 | 2 891 |
Two takeaways. First, the 2070 Super generates faster than a RTX 3060 12 GB (88.1 versus 75.6 tokens per second, or about a 17% difference): speed is not what justifies replacing it. Second, the RTX 2060 Super has the same 448 GB/s bandwidth as the 2070 Super but delivers 60 tokens per second, 32% less. A bandwidth ceiling is therefore not a prediction: driver state, clock speeds, and the card's specifications also matter. Treat these values as rough estimates.
#From a test model to a current model: the rule of three
The test model weighs 3.82 GB (3.56 GiB). Since generation is limited by reading the weights, a heavier model is proportionally slower, all else being equal. For Granite 4.2 8B in Q4 (4.6 GB), the best possible result is 88 × 3.82 ÷ 4.6, or about 73 tokens per second; for Qwen 3.5 9B in Q4 (6 GB), about 56 tokens per second. These are theoretical upper bounds, not measurements: newer architectures (hybrid models, mixtures of experts) and context length shift the result in either direction.
This rule has a practical use. If a figure reported for your configuration is far below these orders of magnitude—for example, less than 15 tokens per second on an 8B in Q4—that indicates part of the model is being offloaded to system RAM: check VRAM usage before blaming the card. The Ollama troubleshooting guide describes what to do.
#Context and KV cache: where 8 GB makes a difference
Throughput drops as the conversation gets longer because the model also rereads the cache for each token. On an 8 GB card, the drop remains moderate—on the order of a few percent in the public measurements available for comparable cards. Memory is the second effect: the K/V cache uses VRAM in proportion to the context. Two Ollama settings ease the load on an 8 GB card. Flash Attention reduces memory usage as the context grows, and the Ollama documentation specifies that the K/V cache can be quantized when it is enabled. The q8_0 type uses about half the memory of the default f16 format, with very little loss of precision according to the documentation.
- 01Enable Flash AttentionStart the server with the OLLAMA_FLASH_ATTENTION=1 variable. Ollama already enables it automatically when the card and model support it; the variable is used to force it on.
- 02Quantize the K/V cacheAdd OLLAMA_KV_CACHE_TYPE=q8_0. Switch to q4_0 only as a last resort: the documentation warns of possible degradation with large contexts.
- 03Check usageRun a long prompt and watch the card’s memory with nvidia-smi: if VRAM fills up, reduce the context instead of letting Ollama spill into RAM.
On the 2070 Super, the llama.cpp test with Flash Attention delivers 87.7 tokens per second during generation versus 88.1 without it, so do not expect any speedup here. Prompt processing, however, rises from 2,088 to 2,293 tokens per second, which matters for long documents and RAG. For more on this topic, a complete guide is dedicated to cache quantization.
#Turing in 2026: drivers, support, and limitations
Turing cards remain supported. The Ollama documentation requires a compute capability of at least 5.0 and driver 550 or later; the RTX 2070, 2080, and 2080 Ti appear in its list of 7.5-capability cards. Turing has even become the baseline: the CUDA 13 release notes state that support for architectures earlier than Turing (Maxwell, Volta, and Pascal) has been dropped. Nothing indicates when Turing support will end, and announcing an end date would be unwise; the prudent approach is to monitor these release notes with every major CUDA update.
One practical limitation remains: recent models often launch first in formats designed for newer hardware, and the community then provides GGUF conversions in Q4 for cards like yours. This does not block Ollama or llama.cpp, but it can delay a model's availability in a readable format by a day or two.
#2070 Super, 3060 12 GB, or 3060 Ti: what trade-off
| Card | Memory | Generation (tok/s) | What it offers |
|---|---|---|---|
| RTX 2070 Super | 8 GB | 88,1 | Decent speed; 9B ceiling in Q4 |
| RTX 3060 Ti | 8 GB | 92,2 | Slightly faster; same capacity ceiling |
| RTX 3060 12 GB | 12 GB | 75,6 | Slower, but with 4 GB more: Gemma 4 12B in Q4 with headroom |
At comparable speeds, capacity is decisive. Moving from a 2070 Super to a 3060 Ti does nothing to the 9B limit; moving to a 3060 12 GB unlocks 12-billion-parameter models and long contexts, at the cost of somewhat lower throughput. Used prices change every week: check the site's tracking page before deciding.
- Local AI graphics card prices (weekly tracking)
- Which LLM for RTX 3060 with 12 GB?
- Which LLM for 8 GB of VRAM? (all cards)
#2026 verdict: keep it, buy it, or move on
- Keep it if
- You're using a 7 to 9B assistant in Q4 for short code, summaries, or moderate RAG. Throughput is very comfortable, and nothing requires a change.
- Move on if
- You want a 12B model with a real context window, a 14B model or larger, or very long documents. In that case, you need at least 12 GB, and 16 GB to be comfortable.
- When buying used
- A 2070 Super is worthwhile only if its price remains well below that of a 12 GB card. Check the warranty and fan condition; these cards have been in service for several years.
#Frequently asked questions
Can the RTX 2070 run an LLM in 2026?+
RTX 2070 or RTX 2070 Super for an LLM?+
Can you run Gemma 4 12B on a RTX 2070?+
How many tokens per second on a RTX 2070 Super?+
Do current drivers still support RTX 2070?+
Should you enable Flash Attention on a RTX 2070?+
- Source: llama.cpp performance table on CUDA
- Source: maps NVIDIA supported by Ollama
- Source: Ollama FAQ (Flash Attention, K/V cache)
- Source: GeForce 20-series specifications
Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.