Beginner 11 minGTX 16

Which LLM on GTX 1650 / 1660 / Super / Ti ?

Direct response

A GTX 1660 (6 GB) runs a local LLM with 7 to 8 billion parameters in Q4: the public llama.cpp benchmark measures 41.35 tokens/s for generation on a 7B Q4_0, but only 148.91 tokens/s for prompt processing. The 1660 Super and Ti, with faster memory, should be quicker. The GTX 1650, limited to 4 GB, is restricted to models with 2 to 3 billion parameters. All remain supported by CUDA 13.

The GTX 16 series, launched in 2019, uses the Turing architecture without Tensor Cores or ray-tracing cores. Unlike the GTX 10 series, it is not affected by the end of Pascal support. This guide distinguishes the four cards based on what matters for an LLM—memory and memory bandwidth—relies on public measurements, and flags a pitfall: correct generation does not imply fast processing of long prompts.

Choosing a machine? Our picks by budget →

By Mohamed Meguedmi·Update 2026-09-30·Tested on Windows, macOS, and Linux
Recommended hardware

Good value for local AI: a GMKtec EVO-X2 64GB / 1TB (Ryzen AI Max+ 395).

A mini PC is a complete machine: check the available memory and engine compatibility. It does not replace macOS/MLX or CUDA.

Why this choice? Our complete guide on GMKtec EVO-X2 64GB / 1TB (Ryzen AI Max+ 395) →

Compare all options by budget, from €800 to €3,500 →

Small budget: RTX 5060 · Large models: RTX 5090 · Mac Studio.

On the go: which laptop for local AI →

Affiliate links — possible commission at no extra cost to you. As an Amazon Associate, BestLLMfor earns from qualifying purchases.

#The GTX 16 series for an LLM: key takeaways

For an LLM, the GTX 16 series is best understood by memory: 4 GB for the GTX 1650, and 6 GB for the GTX 1660, 1660 Super, and 1660 Ti. Generation speed follows bandwidth, which ranges from one to three times as much: 128 GB/s for the 1650 with GDDR5, 192 GB/s for the 1660, 288 GB/s for the 1660 Ti, and 336 GB/s for the 1660 Super. At the same 6 GB capacity, the 1660 Super is therefore the fastest in the series, and the original 1660 the slowest.

Bandwidth and memory of the GTX 16 series (Wikipedia, GeForce 16 series)
CardMemoryBandwidthCUDA cores
GTX 1650 (GDDR5)4 GB, GDDR5, 128 bits128 GB/s896
GTX 1650 (GDDR6, April 2020)GDDR6192 GB/s896
GTX 16606 GB, GDDR5, 192-bit192 GB/s1 408
GTX 1660 Super6 GB, GDDR6336 GB/snot detailed
GTX 1660 Ti6 GB, GDDR6288 GB/s1 536

The series has neither Tensor Cores nor RT Cores: Wikipedia states that it is based on the Turing architecture of the RTX 20 series, omitting these two types of units, and that it retains the dedicated integer cores. For models quantized to integers, the lack of Tensor Cores has little impact on generation.

#Turing without Tensor Cores: what changes compared with Pascal

The series’ strength is FP16. On a GTX 1660, Wikipedia reports 8 617 GFLOPS in half precision versus 4 309 in single precision: FP16 runs twice as fast as FP32, whereas it runs much more slowly than FP32 on a consumer GTX 10, as the GTX 1080 Ti guide shows. This explains why the 1660, with the same 192 GB/s bandwidth, clearly outperforms a GTX 1060 in generation, and why CUDA 13 has not dropped support for it.

The counterexample appears when reading the prompt. The llama.cpp benchmark gives 148,91 tokens/s on the 1660 for prompt processing, versus 416,85 for the GTX 1060: a newer card can read a long text almost three times more slowly. The benchmark does not detail the cause; the result comes from a single contributor and should be read as an order of magnitude. Even so, it is enough to temper enthusiasm for long documents.

llama.cpp CUDA benchmark, Llama 2 7B Q4_0: a 1660 versus its peers
CardBandwidthPrompt (pp512)Generation (tg128)
GTX 1060 6 GB (Pascal)192 GB/s416.85 tokens/s27.79 tokens/s
GTX 1660 6 GB (Turing)192 GB/s148.91 tokens/s41.35 tokens/s
Quadro T1000 4 GB128 GB/s79.44 tokens/s27,82 tokens/s
!
One contribution, one card
Each benchmark row comes from a contributor, along with their driver and system. Comparing a 1660 with a 1060 on this benchmark gives you an order of magnitude, not a rule. If prompt processing matters to you, test your card with llama-bench.

#The four variants, one by one

GTX 1650 (4 GB)
The test bench does not have a 1650 but contains a 4 GB, 128-bit Quadro T1000 with the same bandwidth: it handles a 7B in Q4_0 at 27.82 tokens/s with a minimal context. That is the limit: on 4 GB, a 2- to 3-billion-parameter model is preferable.
GTX 1660 (6 GB GDDR5)
The test-bench card: 41.35 tokens/s during generation. The slowest of the 6 GB cards, but sufficient for Q4 chat.
GTX 1660 Ti (6 GB GDDR6)
288 GB/s: 50% more bandwidth than the 1660, so a throughput gain of roughly the same order is expected, pending measurement.
GTX 1660 Super (6 GB GDDR6)
336 GB/s: the best in the series for an LLM, with 75% more bandwidth than the 1660.

These extrapolations rely solely on bandwidth. Using the 1660's efficiency (82% of its theoretical ceiling), a 1660 Ti would produce around 62 tokens/s and a 1660 Super about 72 on the same 7B Q4_0. These are projections to be confirmed by measuring your card, never results.

Buying pitfall: there are two GTX 1650. The original 2019 version uses GDDR5 on a 128-bit bus, for 128 GB/s; an April 2020 revision switched to GDDR6 at 12 Gbit/s and 192 GB/s. For an LLM, that is a 50% potential throughput gain with the same model. Check whether the listing says GDDR5 or GDDR6, or read the specification in GPU-Z, before buying a used 1650.

#The 4 and 6 GB context: the real ceiling

Ollama starts with a 4,096-token context window, adjustable through OLLAMA_CONTEXT_LENGTH. The KV cache grows with every conversation token: doubling the window doubles the cache. On 6 GB, with a Granite 4.2 8B in Q4, the roughly 1.4 GB of headroom is quickly consumed; on 4 GB, a 3B model leaves 2 GB, providing more comfortable headroom than an 8B model on 6 GB.

The rule of thumb: rather than using more aggressive quantization, reduce the model size or context. A 3B model with a long context performs better on these cards than an 8B model that spills over to the processor. Always check with ollama ps that the model remains at 100% on the GPU: spilling over sharply reduces throughput because system memory is much slower than GPU memory.

Which use case for which card?
UsageGTX 1650 (4 GB)GTX 1660 / Super / Ti (6 GB)
Short chats, translation, rewriting2B to 3B model8B in Q4, short context
Code completion in the editor2 to 3B model, limited quality7–8B model possible
Summarizing long documentsNot recommendedSlow: prompt reading at 148,91 tokens/s on the 1660
Search its documents (RAG)3B model, short excerpts3-8B model, short excerpts

#What fits: 4 GB versus 6 GB

The site's reference point puts a 3B model in Q4 at around 2 GB and a 7–8B model at around 5 GB, plus the KV cache. On 6 GB, about 1.4 GB remains with a Granite 4.2 8B in Q4: enough for a context of a few thousand tokens, but no more. On 4 GB, a 3B model is the maximum reasonable choice.

What fits on 4 GB and 6 GB (Q4 weights, QuelLLM catalog)
Model (Q4)WeightsOn 4 GBOver 6 GB
Gemma 4 2B1.2 GBComfortableComfortable
Granite 4.1 3B2 GBComfortableComfortable
Phi-4 Mini 3.8B3 GBTightComfortable
Granite 4.2 8B4.6 GBDoesn't fit1.4 GB of headroom
Qwen3.5 9B6 GBDoesn't fitOverflows with the context

On the 1650, stick to models with 2 to 4 billion parameters; that is also the right range for drafting, translation, or classification. On the 1660, the 8B in Q4 is the limit: it remains usable for chat but leaves little room for a long context.

#Software support: a clear advantage over GTX 10 cards

Ollama requires driver NVIDIA 550 or later for newer cards, versus 570 for cards with compute capability 5.0 to 6.2 (including Pascal). The official Ollama list names the GTX 1650 Ti, as well as the RTX 20 series, with compute capability 7.5. The other GTX 16 cards share the same Turing architecture without being named individually in the list: verify this at installation time.

On the CUDA side, the release notes for version 13 state that the dropped support applies to architectures older than Turing (Maxwell, Volta, and Pascal): Turing, including the GTX 16 series, remains supported. For a used purchase, this is the main argument for choosing a GTX 16 over a similarly priced GTX 10.

Software support compared
PointGTX 16 (Turing)GTX 10 (Pascal)
Driver required by Ollama550 or newer570 or newer
CUDA 13SupportedRemoved
Flash Attention in llama.cppAvailableAvailable

#Install and configure

  1. 01
    Check the driver
    Run nvidia-smi: the version must be higher than 550.
  2. 02
    Install Ollama
    Follow your system’s installation guide, then download a model suited to your available memory.
  3. 03
    Control placement
    Run ollama ps: the PROCESSOR column should show 100% GPU. Sharing with the processor indicates an overflow.
  4. 04
    Limit context
    Ollama starts at 4,096 tokens. On 4 or 6 GB, increase it only after confirming that there is enough room.

Flash Attention is not limited to cards with Tensor Cores: llama.cpp runs it on the GTX 1660 as well as on a Pascal card, and the benchmark reports 154.45 tokens/s for prompt reading and 41.43 for generation with it, versus 148.91 and 41.35 without it. The gain is small, but context memory usage decreases, which is useful on 6 GB. Ollama enables it automatically when the card supports it.

#Verdict: which GTX 16 and when to move on

Decision based on your situation
Your situationRecommendation
You have a 1660 Super or 1660 TiBest of the series: 8B in Q4, smooth chat
You have a 1660Sufficient for chat; slow prompt processing
You have a 1650 (4 GB)Only 2- to 4-billion-parameter models
You want to read long documentsAn Ampere or newer card reads the prompt much faster
You want a 12B model or largerMove up a tier: target 12 GB or more

One final reference point: these cards are not designed for heavy workloads, but their thermal envelope remains modest (120 W for the GTX 1660, according to Wikipedia). For a local assistant that answers a few questions per day, that is a real advantage over a 250 W card.

At a comparable used price, the 1660 Super is a better option than a GTX 1060 or 1070: same capacity, faster memory, and longer software support. For a new purchase or regular use, a RTX 3050 or RTX 2060 adds Tensor Cores and more memory, depending on the version. Prices change every week: check the site's pricing page.

FAQ
Can you run an LLM on a GTX 1660?+
Yes. The public llama.cpp benchmark measures 41,35 tokens/s for generation with a 7B in Q4_0 on a GTX 1660 with 6 GB, which is sufficient for chat. Prompt processing is slower (148,91 tokens/s), so long documents take time. An 8B in Q4 is the reasonable limit.
Which GTX 16 should you choose for an LLM?+
The GTX 1660 Super: 6 GB and 336 GB/s of bandwidth, the best in the series for generation. The 1660 Ti (288 GB/s) comes next, followed by the 1660 (192 GB/s). Choose the 1650 only if you are targeting models with 2 to 3 billion parameters.
Can the GTX 1650 4 GB run a 7B model?+
Barely. The benchmark includes a 4 GB Quadro T1000 running a 7B in Q4_0 at 27.82 tokens/s, with minimal context. In practice, a 1650 is more comfortable with 2- to 3-billion-parameter models, which leave room for context.
Do GTX 16 cards support Flash Attention?+
Yes. The llama.cpp benchmark measures it on a GTX 1660: 154.45 tokens/s for prompt processing and 41.43 for generation with Flash Attention, versus 148.91 and 41.35 without it. The gain is small, but it reduces context memory usage. The lack of Tensor Cores therefore does not prevent it from working.
Is a GTX 1660 Super better than a GTX 1060 for AI?+
Yes. At the same 6 GB capacity, the 1660 has 336 GB/s versus 192 GB/s, and Turing uses its memory more effectively: a 1660 measures 41.35 tokens/s versus 27.79 for a 1060. Turing is also still supported by CUDA 13, unlike Pascal.
Did this guide help you?

Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.

Prices in euros (€) are French market prices including VAT, checked by QuelLLM. US prices differ: the Amazon buttons show the current US price.