Which LLM on RTX 2060 / 2060 Super (6–8 GB) ?
Yes, a RTX 2060 can run a local LLM: with 6 GB, a 3- to 4-billion-parameter model in Q4 fits comfortably; with 8 GB (2060 Super), an 8B model in Q4 (4.6 to 5.2 GB) fits with a moderate context; the 12 GB variant from late 2021 can host a 14B model. Launched in January 2019, the 2060 remains supported by Ollama and CUDA. Its real limitation is memory, not age.
There are three cards under this name: the RTX 2060 from January 2019 (6 GB), the July 2019 2060 Super (8 GB), and a December 2021 2060 with 12 GB. For local AI, they are far from equivalent. This page gives the specifications for each, what it can run today, what the speed calculations say, and when to move on to something else.
Choosing a machine? Our picks by budget →
Buying alternative for this guide: RTX 5060 Ti 16GB (ASUS Prime).
Compare all options by budget, from €800 to €3,500 →
On the go: which laptop for local AI →
Affiliate links — commission possible at no extra cost to you. As an Amazon Associate, BestLLMfor earns from qualifying purchases.
#RTX 2060, 2060 Super, 2060 12 GB: three cards, three use cases
The RTX 2060 was released on January 15, 2019, the 2060 Super on July 9, 2019, and a 12 GB variant of the 2060 on December 7, 2021, according to Wikipedia's GeForce 20-series table. All are based on the Turing TU106 chip. For an LLM, what matters is memory (capacity, bus, bandwidth), not CUDA cores.
| Card | Output | CUDA cores | Memory | Bandwidth | Power |
|---|---|---|---|---|---|
| RTX 2060 | January 15, 2019 | 1 920 | 6 GB, 192-bit bus | 336 GB/s | 160 W |
| RTX 2060 Super | July 9, 2019 | 2 176 | 8 GB, 256-bit bus | 448 GB/s | 175 W |
| RTX 2060 12 GB | December 7, 2021 | 2 176 | 12 GB | not specified in this table | 185 W |
The 2060 Super offers 448 GB/s of bandwidth versus 336 GB/s for the original 2060; with the same model, it generates about one-third faster. The 12 GB variant is rare and arrived late, but it's by far the most interesting for local AI: it's the only one of the three that lets a 14B model in Q4 (about 9 GB) fit with headroom. So check the card's exact capacity before buying: the name alone doesn't tell you.
#What each variant can run
The rule is the same as for any card: Q4 weights (the site's reference point: 3B ≈ 2 GB, 7-8B ≈ 5 GB, 14B ≈ 9 GB), plus the context cache and about 0.5 GB reserved by the system and engine. The table also gives the theoretical generation ceiling, which is bandwidth divided by weight size. We have no throughput measurements for these cards and do not cite any.
| Model (Q4) | Weights | 2060 6 GB (336 GB/s) | 2060 Super 8 GB (448 GB/s) | 2060 12 GB |
|---|---|---|---|---|
| Gemma 4 2B | 1.2 GB | Yes, ceiling of about 280 t/s | Yes, around 370 t/s | Yes |
| Qwen 3.5 4B | 2.3 GB | Yes, about 146 t/s | Yes, approximately 195 t/s | Yes |
| Granite 4.2 8B | 4.6 GB | Fair, very short context, about 73 t/s | Yes, about 97 t/s | Yes |
| Qwen3 8B (Ollama) | 5.2 GB | Not reliably | Yes, moderate context, approximately 86 t/s | Yes |
| Qwen 3.5 9B | 6 GB | No | Just about 75 t/s | Yes |
| Qwen3 14B (Ollama) | 9.3 GB | No | No | Yes, short context |
These ceilings are never reached in practice; they are used to compare cards with one another and determine whether a model will be interactive. A model that weighs twice as much generates at roughly half the speed. For conversational use, anything above around ten tokens per second in real-world performance is comfortable to read.
#Turing in 2026: still supported
The card’s age is less concerning than you might think. Ollama's documentation states that it supports NVIDIA GPUs with compute capability 5.0 and higher with driver 550 or newer, and lists the RTX 2060 in the 7.5 capability family. CUDA 13 release notes, meanwhile, state that support for Maxwell, Pascal, and Volta was removed from some libraries: Turing is the oldest generation that remains on this path. This is an argument for the 2060 over GTX 10 cards in the long term, even though Ollama still supports capabilities 5.0 through 6.2 with driver 570 or newer: the risk is that future features will target only newer architectures.
Text generation on this card is limited not by compute power but by memory bandwidth, so do not pay a premium for Tensor Cores alone. The driver is the real prerequisite: an up-to-date installation avoids most GPU detection failures.
#The memory calculation for your card
The calculation takes three lines: card memory, minus about 0.5 GB reserved by the system and engine, minus the model size, equals the space remaining for the context cache. What the cache consumes per token depends on the model architecture; the site's calculator gives the figure for a specific model, and q8_0 cache quantization cuts it roughly in half.
| Card | Usable (memory − 0.5 GB) | With Granite 4.2 8B (4.6 GB) | With Qwen3 8B (5.2 GB) | With Qwen3 14B (9.3 GB) |
|---|---|---|---|---|
| RTX 2060 6 GB | 5.5 GB | 0.9 GB | 0.3 GB | Doesn't fit |
| RTX 2060 Super 8 GB | 7.5 GB | 2.9 GB | 2.3 GB | Doesn't fit |
| RTX 2060 12 GB | 11,5 GB | 6.9 GB | 6.3 GB | 2.2 GB |
This table explains why the same model family is pleasant on the Super and painful on the 6 GB version: with less than a gigabyte of headroom, a few thousand context tokens are enough to push the model onto the processor. It also shows that the 14B on 12 GB is real but tight: 2.2 GB for context, which requires a short window.
#Tips for the 6 GB and 8 GB versions
For the 6 GB card, the guide “Which LLM for 6 GB of VRAM” breaks down the complete memory budget; in short, target a 4B model or an 8B model with a short context. On the 8 GB Super, context becomes the first setting to monitor.
- 01Install a recent driverUpdate the NVIDIA driver to version 550 or later, as required by Ollama, then install Ollama for your system.
- 02Choose the model size6 GB: a 4B, or an 8B with a short context. 8 GB: an 8B in Q4 with a context of 4 000 to 8 000 tokens. 12 GB: a 14B in Q4, with a short context.
- 03Reduce the context cacheEnable Flash Attention and q8_0 cache quantization: according to the Ollama FAQ, it cuts cache memory usage by about half compared with f16.
- 04Check with ollama psThe Processor column should show 100% GPU. A split between CPU and GPU indicates offloading, which means lower speed.
- 05Close resource-hungry applicationsA browser with hardware acceleration or a game uses VRAM that the model cannot use.
The Ollama FAQ also offers a more aggressive q4_0 cache quantization that further reduces memory at the possible cost of quality; test it with your workload before keeping it.
#What you can really do with a 2060
- Chat and writing
- A 4B or 8B responds correctly for everyday use; this is where the card delivers on its promise best.
- RAG and documents
- A small 4B model and an embedding model fit in memory; limit the number of excerpts you send.
- Code
- Autocompletion with a small code model; no reliable agent at this size.
- What exceeds the card's capacity
- Large contexts, heavy vision workloads, 14B models on 6 or 8 GB, and any agent that chains tools with long prompts.
The 2060 is also a 160 to 185 W card: under sustained load, it gets hot and loud like any card in its class. The thermal guide explains how to cap power for a server that runs all day.
#Verdict and when to move on
An already-installed RTX 2060 does not need replacing if you use it for chat, text summarization, and light RAG. If you are buying, memory takes priority over everything else: the 12 GB variant or the 2060 Super 8 GB are preferable to the 6 GB version. Prices change every week and are not listed on this page; the site’s tracker records the lowest price for each card.
| Card | What it offers | The tradeoffs it entails |
|---|---|---|
| RTX 2060 6 GB | The entry point: 3–4B models | Little headroom; 8B only with a very short context |
| RTX 2060 Super 8 GB | Comfortable 8B in Q4 | 448 GB/s, faster than the 3060 12 GB |
| RTX 3060 12 GB | 12 GB: 14B in Q4 and long context | 360 GB/s: slower than the 2060 Super on the same model that fits on both |
| RTX 3050 6 GB | Energy efficiency | 168 GB/s, half as fast as the 2060 6 GB |
This comparison corrects a common misconception: the RTX 3060 12 GB is not faster than the 2060 Super for text generation. Its bandwidth is 360 GB/s according to Wikipedia, versus 448 GB/s for the 2060 Super. It wins on capacity: 12 GB can load models that the Super cannot accommodate. Choose based on the model you want to run, not on a general speed percentage.
If you buy a used card, check three things before keeping it. Use nvidia-smi to verify the advertised total memory: the card name alone is not enough to distinguish the 6, 8, and 12 GB variants. Run a continuous generation for about twenty minutes and monitor the temperature and core frequency: a card that slows down significantly needs cleaning or new thermal paste. Finally, test a real model with the intended context and verify with ollama ps that it remains entirely on the card.
- Which LLM on RTX 3060 12 GB
- Which LLM on RTX 2070 / 2070 Super
- Which LLM on RTX 3050
- Local AI graphics card price tracking
- Source: Wikipedia, GeForce 20 Series
- Source: Ollama, supported hardware
- Source: official Ollama FAQ
#Frequently asked questions
Can the RTX 2060 run an LLM?+
What is the release date for RTX 2060?+
RTX 2060 or RTX 2060 Super for an LLM?+
Can a modern 8B LLM fit on a RTX 2060 Super?+
Is the RTX 2060 6 GB still worth it for local AI?+
Is the RTX 3060 12 GB faster than the 2060 Super?+
Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.