Beginner 11 minRTX 20

Which LLM for RTX 2070 / 2070 Super (8 GB) ?

Direct response

A RTX 2070 or 2070 Super (8 GB of GDDR6) comfortably runs 7- to 9-billion-parameter models in Q4: Granite 4.2 8B (4.6 GB of weights) or Qwen 3.5 9B (6 GB) fit entirely on the card. In llama.cpp's reference benchmark, the 2070 Super generates 88 tokens/s. Beyond 9B, it must spill into RAM and throughput collapses.

The RTX 2070 (October 2018) and the RTX 2070 Super (July 2019) are Turing cards with 8 GB of memory. In 2026, they are at the right level for an 8B assistant in Q4, and at the limit for everything else. This guide connects their specifications to a public speed test, explains where the 8 GB goes, and indicates when it is worth changing cards.

Choosing a machine? Our picks by budget →

By Mohamed Meguedmi·Update 2026-09-30·Tested on Windows, macOS, and Linux
Recommended hardware

Buying alternative for this guide: RTX 5060 Ti 16GB (ASUS Prime).

Compare all options by budget, from €800 to €3,500 →

On the go: which laptop for local AI →

Affiliate links — commission possible at no extra cost to you. As an Amazon Associate, BestLLMfor earns from qualifying purchases.

#RTX 2070 and 2070 Super: what 8 GB can do

On a RTX 2070 or a 2070 Super, 7- to 9-billion-parameter models in Q4 quantization fit entirely in VRAM: this is the comfort zone for these cards. According to the QuelLLM catalog, Granite 4.2 8B requires 4.6 GB of weights in Q4 and Qwen 3.5 9B about 6 GB. That leaves room for context, provided you keep its length reasonable. A 12-billion-parameter model such as Gemma 4 12B (7 GB in Q4) leaves barely one gigabyte for the cache and system: it works on paper, but poorly in practice. In terms of speed, the 2070 Super reaches 88 tokens per second on llama.cpp's reference benchmark, using a 7-billion-parameter model in Q4_0—well above human reading speed. The real limit of these cards is therefore not speed but memory capacity.

RTX 2070
October 2018, 2,304 CUDA cores, 8 GB of GDDR6 on a 256-bit bus, or 448 GB/s (256 bits at 14 Gbit/s).
RTX 2070 Super
July 2019, 2,560 CUDA cores, the same 8 GB and the same 448 GB/s bandwidth, 215 W power consumption.
What distinguishes them for an LLM
Almost nothing: bandwidth and capacity are identical. The Super has more cores, which mainly helps with prompt processing, not generation.

This last point is counterintuitive: to generate text, the GPU rereads all the model’s weights for every token, so memory bandwidth matters more than the number of cores. Two cards at 448 GB/s will produce similar generation rates, even if one is significantly faster in games.

#Which models fit in 8 GB

Weights are only part of the bill: you also need the context cache (KV cache), which grows with conversation length, plus some headroom for the driver and display. On an 8 GB card, the rule of thumb is to target weights of no more than 6 GB if you want a context of several thousand tokens. As a reference point, the llama.cpp test tool sees 7,838 MiB of free memory on an 8 GB RTX 3060 Ti, slightly less than the theoretical 8,192 MiB: a few hundred megabytes are already occupied before a model is even loaded.

Models for a RTX 2070 / 2070 Super (weights according to the QuelLLM catalog)
ModelWeights in Q4Verdict on 8 GB
Qwen 3.5 4B2.3 GB (Q8: 4.3 GB)Very capable, even in Q8 and with a long context
Granite 4.2 8B4.6 GB (Q8: 9 GB)Comfortable range in Q4, with a context of several thousand tokens
Qwen 3.5 9B6 GB (Q8: 10 GB)Switch to Q4 with a moderate context; use the quantized K/V cache beyond that
Gemma 4 12B7 GB (Q8: 13 GB)Limit: no room for context, likely RAM overflow

Q8 for an 8B model (9 GB) won’t fit: with 8 GB, Q4 quantization is the rule, and Q5 remains possible for smaller models. To compare quantization levels, the guide to choosing Q4, Q5, or Q8 details the expected quality loss. For a specific project, the site’s VRAM calculator gives you the exact footprint based on the model and context.

#Speed: what the benchmark measures

The only usable public benchmark is the collaborative table in the llama.cpp repository, where users publish the result of the same command: llama-bench on Llama 2 7B in Q4_0, a 3.56 GiB file, with all layers on the GPU. The table notes that results vary by driver, operating system, and card manufacturer, even with the same chip. These are not measurements from the site: the figures below come from that table, and the “reading” column is prompt-processing throughput (pp512), while “generation” is response throughput (tg128).

llama.cpp test, Llama 2 7B Q4_0, without Flash Attention
CardMemoryGeneration (tok/s)Prompt processing (tok/s)
RTX 2060 Super8 GB60,01 420
RTX 2070 Super8 GB88,12 088
RTX 3060 12 GB12 GB75,62 138
RTX 2080 Ti11 GB107,52 891

Two takeaways. First, the 2070 Super generates faster than a RTX 3060 12 GB (88.1 versus 75.6 tokens per second, or about a 17% difference): speed is not what justifies replacing it. Second, the RTX 2060 Super has the same 448 GB/s bandwidth as the 2070 Super but delivers 60 tokens per second, 32% less. A bandwidth ceiling is therefore not a prediction: driver state, clock speeds, and the card's specifications also matter. Treat these values as rough estimates.

#From a test model to a current model: the rule of three

The test model weighs 3.82 GB (3.56 GiB). Since generation is limited by reading the weights, a heavier model is proportionally slower, all else being equal. For Granite 4.2 8B in Q4 (4.6 GB), the best possible result is 88 × 3.82 ÷ 4.6, or about 73 tokens per second; for Qwen 3.5 9B in Q4 (6 GB), about 56 tokens per second. These are theoretical upper bounds, not measurements: newer architectures (hybrid models, mixtures of experts) and context length shift the result in either direction.

This rule has a practical use. If a figure reported for your configuration is far below these orders of magnitude—for example, less than 15 tokens per second on an 8B in Q4—that indicates part of the model is being offloaded to system RAM: check VRAM usage before blaming the card. The Ollama troubleshooting guide describes what to do.

#Context and KV cache: where 8 GB makes a difference

Throughput drops as the conversation gets longer because the model also rereads the cache for each token. On an 8 GB card, the drop remains moderate—on the order of a few percent in the public measurements available for comparable cards. Memory is the second effect: the K/V cache uses VRAM in proportion to the context. Two Ollama settings ease the load on an 8 GB card. Flash Attention reduces memory usage as the context grows, and the Ollama documentation specifies that the K/V cache can be quantized when it is enabled. The q8_0 type uses about half the memory of the default f16 format, with very little loss of precision according to the documentation.

  1. 01
    Enable Flash Attention
    Start the server with the OLLAMA_FLASH_ATTENTION=1 variable. Ollama already enables it automatically when the card and model support it; the variable is used to force it on.
  2. 02
    Quantize the K/V cache
    Add OLLAMA_KV_CACHE_TYPE=q8_0. Switch to q4_0 only as a last resort: the documentation warns of possible degradation with large contexts.
  3. 03
    Check usage
    Run a long prompt and watch the card’s memory with nvidia-smi: if VRAM fills up, reduce the context instead of letting Ollama spill into RAM.
Linux: run Ollama with the two settings
OLLAMA_FLASH_ATTENTION=1 OLLAMA_KV_CACHE_TYPE=q8_0 ollama serve

On the 2070 Super, the llama.cpp test with Flash Attention delivers 87.7 tokens per second during generation versus 88.1 without it, so do not expect any speedup here. Prompt processing, however, rises from 2,088 to 2,293 tokens per second, which matters for long documents and RAG. For more on this topic, a complete guide is dedicated to cache quantization.

#Turing in 2026: drivers, support, and limitations

Turing cards remain supported. The Ollama documentation requires a compute capability of at least 5.0 and driver 550 or later; the RTX 2070, 2080, and 2080 Ti appear in its list of 7.5-capability cards. Turing has even become the baseline: the CUDA 13 release notes state that support for architectures earlier than Turing (Maxwell, Volta, and Pascal) has been dropped. Nothing indicates when Turing support will end, and announcing an end date would be unwise; the prudent approach is to monitor these release notes with every major CUDA update.

One practical limitation remains: recent models often launch first in formats designed for newer hardware, and the community then provides GGUF conversions in Q4 for cards like yours. This does not block Ollama or llama.cpp, but it can delay a model's availability in a readable format by a day or two.

#2070 Super, 3060 12 GB, or 3060 Ti: what trade-off

Three 8 to 12 GB cards for LLM use
CardMemoryGeneration (tok/s)What it offers
RTX 2070 Super8 GB88,1Decent speed; 9B ceiling in Q4
RTX 3060 Ti8 GB92,2Slightly faster; same capacity ceiling
RTX 3060 12 GB12 GB75,6Slower, but with 4 GB more: Gemma 4 12B in Q4 with headroom

At comparable speeds, capacity is decisive. Moving from a 2070 Super to a 3060 Ti does nothing to the 9B limit; moving to a 3060 12 GB unlocks 12-billion-parameter models and long contexts, at the cost of somewhat lower throughput. Used prices change every week: check the site's tracking page before deciding.

#2026 verdict: keep it, buy it, or move on

Keep it if
You're using a 7 to 9B assistant in Q4 for short code, summaries, or moderate RAG. Throughput is very comfortable, and nothing requires a change.
Move on if
You want a 12B model with a real context window, a 14B model or larger, or very long documents. In that case, you need at least 12 GB, and 16 GB to be comfortable.
When buying used
A 2070 Super is worthwhile only if its price remains well below that of a 12 GB card. Check the warranty and fan condition; these cards have been in service for several years.

#Frequently asked questions

Frequently asked questions
Can the RTX 2070 run an LLM in 2026?+
Yes, within its 8 GB limit. A 7- to 9-billion-parameter model in Q4, such as Granite 4.2 8B or Qwen 3.5 9B, fits entirely on the card. In llama.cpp's reference test, the 2070 Super generates 88 tokens per second, which is more than enough for chat or coding assistance.
RTX 2070 or RTX 2070 Super for an LLM?+
Both have 8 GB of GDDR6 and the same 448 GB/s bandwidth, while generation depends primarily on bandwidth. The Super, with more cores, mainly helps with reading long prompts. If the price difference is significant, the standard 2070 delivers a very similar generation experience.
Can you run Gemma 4 12B on a RTX 2070?+
It's borderline. The catalog lists 7 GB of weights in Q4, leaving almost nothing for the context cache and system on 8 GB. The model loads but quickly spills into RAM as the conversation gets longer. Prefer Granite 4.2 8B or Qwen 3.5 9B, or move to 12 GB of VRAM.
How many tokens per second on a RTX 2070 Super?+
The public llama.cpp table gives 88 tokens per second for Llama 2 7B in Q4_0. A larger model will be slower, roughly in proportion to its size: estimate around 55 to 75 tokens per second on an 8 to 9B model in Q4, as a theoretical estimate. A long context lowers this figure.
Do current drivers still support RTX 2070?+
Yes. Ollama lists the RTX 2070, 2080, and 2080 Ti among the cards with compute capability 7.5 and simply requires driver 550 or later. CUDA 13 even made Turing its baseline by dropping earlier architectures. No end of support has been announced in the sources consulted.
Should you enable Flash Attention on a RTX 2070?+
Ollama enables it automatically when the hardware allows it; you can force it with OLLAMA_FLASH_ATTENTION=1. On the 2070 Super, it does not change generation speed but accelerates prompt processing by about 10% and reduces memory usage with large contexts, which matters on 8 GB.
Did this guide help you?

Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.

Prices in euros (€) are French market prices including VAT, as checked by BestLLMfor. US prices differ: the Amazon buttons show the current US price.