Which LLM on GTX 1650 / 1660 / Super / Ti ?
A GTX 1660 (6 GB) runs a local LLM with 7 to 8 billion parameters in Q4: the public llama.cpp benchmark measures 41.35 tokens/s for generation on a 7B Q4_0, but only 148.91 tokens/s for prompt processing. The 1660 Super and Ti, with faster memory, should be quicker. The GTX 1650, limited to 4 GB, is restricted to models with 2 to 3 billion parameters. All remain supported by CUDA 13.
The GTX 16 series, launched in 2019, uses the Turing architecture without Tensor Cores or ray-tracing cores. Unlike the GTX 10 series, it is not affected by the end of Pascal support. This guide distinguishes the four cards based on what matters for an LLM—memory and memory bandwidth—relies on public measurements, and flags a pitfall: correct generation does not imply fast processing of long prompts.
Choosing a machine? Our picks by budget →
Good value for local AI: a GMKtec EVO-X2 64GB / 1TB (Ryzen AI Max+ 395).
A mini PC is a complete machine: check the available memory and engine compatibility. It does not replace macOS/MLX or CUDA.
Why this choice? Our complete guide on GMKtec EVO-X2 64GB / 1TB (Ryzen AI Max+ 395) →
Compare all options by budget, from €800 to €3,500 →
Small budget: RTX 5060 · Large models: RTX 5090 · Mac Studio.
On the go: which laptop for local AI →
Affiliate links — possible commission at no extra cost to you. As an Amazon Associate, BestLLMfor earns from qualifying purchases.
#The GTX 16 series for an LLM: key takeaways
For an LLM, the GTX 16 series is best understood by memory: 4 GB for the GTX 1650, and 6 GB for the GTX 1660, 1660 Super, and 1660 Ti. Generation speed follows bandwidth, which ranges from one to three times as much: 128 GB/s for the 1650 with GDDR5, 192 GB/s for the 1660, 288 GB/s for the 1660 Ti, and 336 GB/s for the 1660 Super. At the same 6 GB capacity, the 1660 Super is therefore the fastest in the series, and the original 1660 the slowest.
| Card | Memory | Bandwidth | CUDA cores |
|---|---|---|---|
| GTX 1650 (GDDR5) | 4 GB, GDDR5, 128 bits | 128 GB/s | 896 |
| GTX 1650 (GDDR6, April 2020) | GDDR6 | 192 GB/s | 896 |
| GTX 1660 | 6 GB, GDDR5, 192-bit | 192 GB/s | 1 408 |
| GTX 1660 Super | 6 GB, GDDR6 | 336 GB/s | not detailed |
| GTX 1660 Ti | 6 GB, GDDR6 | 288 GB/s | 1 536 |
The series has neither Tensor Cores nor RT Cores: Wikipedia states that it is based on the Turing architecture of the RTX 20 series, omitting these two types of units, and that it retains the dedicated integer cores. For models quantized to integers, the lack of Tensor Cores has little impact on generation.
#Turing without Tensor Cores: what changes compared with Pascal
The series’ strength is FP16. On a GTX 1660, Wikipedia reports 8 617 GFLOPS in half precision versus 4 309 in single precision: FP16 runs twice as fast as FP32, whereas it runs much more slowly than FP32 on a consumer GTX 10, as the GTX 1080 Ti guide shows. This explains why the 1660, with the same 192 GB/s bandwidth, clearly outperforms a GTX 1060 in generation, and why CUDA 13 has not dropped support for it.
The counterexample appears when reading the prompt. The llama.cpp benchmark gives 148,91 tokens/s on the 1660 for prompt processing, versus 416,85 for the GTX 1060: a newer card can read a long text almost three times more slowly. The benchmark does not detail the cause; the result comes from a single contributor and should be read as an order of magnitude. Even so, it is enough to temper enthusiasm for long documents.
| Card | Bandwidth | Prompt (pp512) | Generation (tg128) |
|---|---|---|---|
| GTX 1060 6 GB (Pascal) | 192 GB/s | 416.85 tokens/s | 27.79 tokens/s |
| GTX 1660 6 GB (Turing) | 192 GB/s | 148.91 tokens/s | 41.35 tokens/s |
| Quadro T1000 4 GB | 128 GB/s | 79.44 tokens/s | 27,82 tokens/s |
#The four variants, one by one
- GTX 1650 (4 GB)
- The test bench does not have a 1650 but contains a 4 GB, 128-bit Quadro T1000 with the same bandwidth: it handles a 7B in Q4_0 at 27.82 tokens/s with a minimal context. That is the limit: on 4 GB, a 2- to 3-billion-parameter model is preferable.
- GTX 1660 (6 GB GDDR5)
- The test-bench card: 41.35 tokens/s during generation. The slowest of the 6 GB cards, but sufficient for Q4 chat.
- GTX 1660 Ti (6 GB GDDR6)
- 288 GB/s: 50% more bandwidth than the 1660, so a throughput gain of roughly the same order is expected, pending measurement.
- GTX 1660 Super (6 GB GDDR6)
- 336 GB/s: the best in the series for an LLM, with 75% more bandwidth than the 1660.
These extrapolations rely solely on bandwidth. Using the 1660's efficiency (82% of its theoretical ceiling), a 1660 Ti would produce around 62 tokens/s and a 1660 Super about 72 on the same 7B Q4_0. These are projections to be confirmed by measuring your card, never results.
Buying pitfall: there are two GTX 1650. The original 2019 version uses GDDR5 on a 128-bit bus, for 128 GB/s; an April 2020 revision switched to GDDR6 at 12 Gbit/s and 192 GB/s. For an LLM, that is a 50% potential throughput gain with the same model. Check whether the listing says GDDR5 or GDDR6, or read the specification in GPU-Z, before buying a used 1650.
#The 4 and 6 GB context: the real ceiling
Ollama starts with a 4,096-token context window, adjustable through OLLAMA_CONTEXT_LENGTH. The KV cache grows with every conversation token: doubling the window doubles the cache. On 6 GB, with a Granite 4.2 8B in Q4, the roughly 1.4 GB of headroom is quickly consumed; on 4 GB, a 3B model leaves 2 GB, providing more comfortable headroom than an 8B model on 6 GB.
The rule of thumb: rather than using more aggressive quantization, reduce the model size or context. A 3B model with a long context performs better on these cards than an 8B model that spills over to the processor. Always check with ollama ps that the model remains at 100% on the GPU: spilling over sharply reduces throughput because system memory is much slower than GPU memory.
| Usage | GTX 1650 (4 GB) | GTX 1660 / Super / Ti (6 GB) |
|---|---|---|
| Short chats, translation, rewriting | 2B to 3B model | 8B in Q4, short context |
| Code completion in the editor | 2 to 3B model, limited quality | 7–8B model possible |
| Summarizing long documents | Not recommended | Slow: prompt reading at 148,91 tokens/s on the 1660 |
| Search its documents (RAG) | 3B model, short excerpts | 3-8B model, short excerpts |
#What fits: 4 GB versus 6 GB
The site's reference point puts a 3B model in Q4 at around 2 GB and a 7–8B model at around 5 GB, plus the KV cache. On 6 GB, about 1.4 GB remains with a Granite 4.2 8B in Q4: enough for a context of a few thousand tokens, but no more. On 4 GB, a 3B model is the maximum reasonable choice.
| Model (Q4) | Weights | On 4 GB | Over 6 GB |
|---|---|---|---|
| Gemma 4 2B | 1.2 GB | Comfortable | Comfortable |
| Granite 4.1 3B | 2 GB | Comfortable | Comfortable |
| Phi-4 Mini 3.8B | 3 GB | Tight | Comfortable |
| Granite 4.2 8B | 4.6 GB | Doesn't fit | 1.4 GB of headroom |
| Qwen3.5 9B | 6 GB | Doesn't fit | Overflows with the context |
On the 1650, stick to models with 2 to 4 billion parameters; that is also the right range for drafting, translation, or classification. On the 1660, the 8B in Q4 is the limit: it remains usable for chat but leaves little room for a long context.
#Software support: a clear advantage over GTX 10 cards
Ollama requires driver NVIDIA 550 or later for newer cards, versus 570 for cards with compute capability 5.0 to 6.2 (including Pascal). The official Ollama list names the GTX 1650 Ti, as well as the RTX 20 series, with compute capability 7.5. The other GTX 16 cards share the same Turing architecture without being named individually in the list: verify this at installation time.
On the CUDA side, the release notes for version 13 state that the dropped support applies to architectures older than Turing (Maxwell, Volta, and Pascal): Turing, including the GTX 16 series, remains supported. For a used purchase, this is the main argument for choosing a GTX 16 over a similarly priced GTX 10.
| Point | GTX 16 (Turing) | GTX 10 (Pascal) |
|---|---|---|
| Driver required by Ollama | 550 or newer | 570 or newer |
| CUDA 13 | Supported | Removed |
| Flash Attention in llama.cpp | Available | Available |
#Install and configure
- 01Check the driverRun nvidia-smi: the version must be higher than 550.
- 02Install OllamaFollow your system’s installation guide, then download a model suited to your available memory.
- 03Control placementRun ollama ps: the PROCESSOR column should show 100% GPU. Sharing with the processor indicates an overflow.
- 04Limit contextOllama starts at 4,096 tokens. On 4 or 6 GB, increase it only after confirming that there is enough room.
Flash Attention is not limited to cards with Tensor Cores: llama.cpp runs it on the GTX 1660 as well as on a Pascal card, and the benchmark reports 154.45 tokens/s for prompt reading and 41.43 for generation with it, versus 148.91 and 41.35 without it. The gain is small, but context memory usage decreases, which is useful on 6 GB. Ollama enables it automatically when the card supports it.
- Install Ollama step by step
- Troubleshoot Ollama: GPU not detected, slow performance
- Quantize the KV cache to save VRAM
#Verdict: which GTX 16 and when to move on
| Your situation | Recommendation |
|---|---|
| You have a 1660 Super or 1660 Ti | Best of the series: 8B in Q4, smooth chat |
| You have a 1660 | Sufficient for chat; slow prompt processing |
| You have a 1650 (4 GB) | Only 2- to 4-billion-parameter models |
| You want to read long documents | An Ampere or newer card reads the prompt much faster |
| You want a 12B model or larger | Move up a tier: target 12 GB or more |
One final reference point: these cards are not designed for heavy workloads, but their thermal envelope remains modest (120 W for the GTX 1660, according to Wikipedia). For a local assistant that answers a few questions per day, that is a real advantage over a 250 W card.
At a comparable used price, the 1660 Super is a better option than a GTX 1060 or 1070: same capacity, faster memory, and longer software support. For a new purchase or regular use, a RTX 3050 or RTX 2060 adds Tensor Cores and more memory, depending on the version. Prices change every week: check the site's pricing page.
- Which LLM on RTX 3050
- Which LLM on RTX 2060 / 2060 Super
- Which models for 6 GB of VRAM, across all GPUs
- Compare GPU prices for AI
- Source: public llama.cpp benchmark on CUDA
- Source: Ollama documentation, supported NVIDIA hardware
- Source: GeForce 16 series specifications
- Source: CUDA Toolkit release notes
Can you run an LLM on a GTX 1660?+
Which GTX 16 should you choose for an LLM?+
Can the GTX 1650 4 GB run a 7B model?+
Do GTX 16 cards support Flash Attention?+
Is a GTX 1660 Super better than a GTX 1060 for AI?+
Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.