Beginner 11 minRTX 30

Which LLM on RTX 3060 Ti (8 GB) ?

Direct response

The RTX 3060 Ti (8 GB, 448 GB/s) runs 7B to 9B models in Q4 without spilling over, such as Granite 4.2 8B (4.6 GB) or Qwen 3.5 9B (6 GB). It generates 92 tokens/s on llama.cpp's reference benchmark, faster than a RTX 3060 12 GB (75.6), but its 8 GB rule it out for 12B models. For local AI, its speed does not make up for the missing 4 GB.

The RTX 3060 Ti (December 2020) is significantly faster in games than the RTX 3060, but it has only 8 GB of memory. This guide draws on recent public measurements to explain what it can run, what Flash Attention adds, how it differs from the 3060 12 GB, and when it is better to move on.

Choosing a machine? Our picks by budget →

By Mohamed Meguedmi·Update 2026-09-30·Tested on Windows, macOS, and Linux
Recommended hardware

Buying alternative for this guide: RTX 5060 Ti 16GB (ASUS Prime).

Compare all options by budget, from €800 to €3,500 →

On the go: which laptop for local AI →

Affiliate links — commission possible at no extra cost to you. As an Amazon Associate, BestLLMfor earns from qualifying purchases.

#RTX 3060 Ti vs RTX 3060: the summary for local AI

The RTX 3060 Ti is faster than the 3060 12 GB for generating text, but it has 4 GB less, and for local AI, capacity is the limiting factor. The llama.cpp collaborative table gives 92.2 tokens per second for the 3060 Ti, versus 75.6 for the 3060 12 GB, on a 7-billion-parameter model in Q4_0: about 22% more. This advantage comes from its 256-bit bus, which provides 448 GB/s of bandwidth versus 360 GB/s. However, with 8 GB it runs 7- to 9-billion-parameter models (Granite 4.2 8B at 4.6 GB, Qwen 3.5 9B at 6 GB) but not Gemma 4 12B (7 GB of weights), which leaves no room for context. If your target is an 8B, the Ti is the faster of the two; if you are aiming for a 12B, it is out of the running.

RTX 3060 Ti
GA104 chip, December 2020, 4,864 CUDA cores, 8 GB of GDDR6, 256-bit bus, 448 GB/s.
RTX 3060
GA106 chip, 3,584 cores, 12 GB of GDDR6 on a 192-bit bus, 360 GB/s.
For an LLM
More bandwidth means more speed; less memory means fewer models. The two effects offset each other, but capacity carries more weight.

#Models compatible with 8 GB

Models for a RTX 3060 Ti (weights according to the QuelLLM catalog)
ModelWeights in Q4Q8 sizeOn 8 GB
Granite 4.2 8B4.6 GB9 GBQ4 is very comfortable; Q8 does not fit
Qwen 3.5 9B6 GB10 GBQ4 with moderate context
Gemma 4 12B7 GB13 GBDoesn't hold up with context

The card’s generation throughput is not what will limit you; space is. The llama.cpp test sees 7,838 MiB of memory on the 3060 Ti, slightly less than the nominal 8,192 MiB, and that budget is shared between the weights, context cache, and display if a monitor is connected. Allow at most 6 GB for weights to keep a context of several thousand tokens.

#Installation and settings

The card has compute capability 8.6 in the Ollama list, which requires driver 550 or newer. For power, the NVIDIA specification lists 200 W for the card and recommends a 600 W system power supply, versus 550 W for the 3060: a 500 W power supply, sometimes cited, is below that recommendation.

  1. 01
    Install the driver
    Install an NVIDIA 550 or newer driver and verify with nvidia-smi that the card is recognized with approximately 8 GB of memory.
  2. 02
    Install Ollama
    Follow the installation guide, then launch an 8-billion-parameter model in Q4 with ollama run.
  3. 03
    Enable memory settings
    Start the server with OLLAMA_FLASH_ATTENTION=1 and OLLAMA_KV_CACHE_TYPE=q8_0. The documentation specifies that the K/V cache can be quantized when Flash Attention is enabled.
  4. 04
    Monitor memory
    During a long conversation, check nvidia-smi: if memory approaches 7,800 MiB, reduce the context.
Linux: launch Ollama with the memory settings
OLLAMA_FLASH_ATTENTION=1 OLLAMA_KV_CACHE_TYPE=q8_0 ollama serve

#Speed: generation, prompt processing, and Flash Attention

A September 2026 measurement published on the llama.cpp dashboard provides a detailed look at the 3060 Ti’s behavior, using a 7-billion-parameter model in Q4_0 at 3.56 GiB. The 448 GB/s bandwidth would allow about 117 tokens per second at maximum; the measurement reaches 92 tokens, or about 79% of the ceiling. The most interesting result concerns Flash Attention.

RTX 3060 Ti, Llama 2 7B Q4_0, Flash Attention effect
MetricWithout Flash AttentionWith Flash AttentionGap
Generation (128 tokens)92.2 tok/s96.5 tok/s+5 %
Reading a prompt of 4,096 tokens1,897 tok/s2,598 tok/s+37 %

Generation improves little, but reading a long prompt gets more than one-third faster, which matters for RAG and document summaries. At this rate, a prompt of 10,000 tokens is read in about four seconds, a theoretical calculation that assumes a constant throughput. Ollama automatically enables Flash Attention when the card supports it; you can force it with the OLLAMA_FLASH_ATTENTION=1 variable.

For a model larger than the test file, the rule of three gives rough figures: 92 × 3,82 ÷ 4,6, or about 76 tokens per second, for Granite 4.2 8B in Q4, and about 59 for Qwen 3.5 9B (6 GB). These are high theoretical upper bounds, not measurements. If you want to compare your own card with these values, the llama-bench tool from llama.cpp reproduces the test: same file Llama 2 7B in Q4_0, with all layers placed on the GPU, with and without Flash Attention. Results vary by driver, system, and card manufacturer, even with the same chip, so a difference of a few percent is nothing unusual. These speeds exceed reading speed: the card is never the limiting factor in a conversation.

#Documents and RAG: where the time goes

When you query your documents, the card first spends time reading the prompt, then generating the response. On the 3060 Ti, a 4,096-token prompt—about ten pages—takes roughly 2.2 seconds to read without Flash Attention and 1.6 seconds with it, according to speeds measured by llama.cpp. The 400-token response then takes about four to five seconds at 92 tokens per second on the test model, and longer on a current 8B. These are theoretical calculations that assume constant throughput, but they show that the wait remains a few seconds.

The RAG limit on 8 GB is therefore not time but memory: a context of several tens of thousands of tokens fills the context cache of an 8-billion-parameter model. Two ways to stay within the budget: reduce the number of excerpts injected into the prompt, and quantize the K/V cache to q8_0, which uses approximately half the memory of the default format according to Ollama's documentation.

#When to move on

What each need requires (weights according to the QuelLLM catalog)
Your needsWeight memorySuitable card
An 8B in Q4 (Granite 4.2 8B)4.6 GB8 GB is enough
A 9B in Q4 with medium context (Qwen 3.5 9B)6 GB8 GB, at the edge of comfortable
An 8B in Q8 (Granite 4.2 8B)9 GB12 GB
A 12B in Q4 (Gemma 4 12B)7 GB plus context12 GB
A 12B in Q8 (Gemma 4 12B)13 GB16 GB

Read this table from top to bottom: the 3060 Ti covers the first two rows, while the next three require more memory. If your needs fall in the lower half, changing cards is justified; otherwise, the 3060 Ti gets the job done. A higher-quality but larger model can sometimes be better than a faster card: at 92 tokens per second, speed is not your bottleneck.

#3060 Ti and 3060 12 GB: the false twin

3060 Ti or 3060 12 GB for an LLM
CriterionRTX 3060 TiRTX 3060 12 GB
Memory8 GB12 GB
Bus / bandwidth256 bits / 448 GB/s192 bits / 360 GB/s
Generation (7B Q4_0)92.2 tok/s75.6 tok/s
Gemma 4 12B in Q4 (7 GB)Doesn't fit with contextFits comfortably
Recommended system power supply600 W550 W

The first instinct is to choose the faster one. However, for the 8- to 9-billion-parameter models that both cards can load, the extra 22% speed does not change the experience: in both cases, text appears faster than you can read it. By contrast, the 3060 12 GB opens up an entire category of models. It is therefore the better choice for local AI, unless your use is permanently limited to 8- to 9B models. One buying precaution: NVIDIA also sold a 3060 Ti with GDDR6X memory, distinguished by its specification sheet; the 8 GB capacity remains the same.

#3060 Ti and RTX 4060 Ti 8 GB

The newer RTX 4060 Ti 8 GB measures 63.9 tokens per second on the same test, significantly less than the 3060 Ti. The reason is its 128-bit bus, which limits its bandwidth despite a more efficient architecture. For text generation, the 3060 Ti therefore remains faster than this newer card at the same capacity. Its strengths lie elsewhere: newer architectural features and, according to the manufacturers’ specifications, lower power consumption, which is useful in a compact case or on a machine that stays on to serve a model. This is not a text-generation speed improvement, nor is it a memory increase. If you are choosing between the two, look at the 16 GB version of the 4060 Ti, which offers twice the memory.

#2026 verdict

Keep it if
It works, and you use 7- to 9-billion-parameter models: speed is more than sufficient, and you should not expect any gain from switching at the same capacity.
Don't buy for AI if 12 GB is available
A 3060 12 GB opens up 12-billion-parameter models. The used-price gap changes every week: check the site's tracker.
Move to 12 or 16 GB if
You want a 12B model, a long context, or RAG over large documents.

#Frequently asked questions

Frequently asked questions
RTX 3060 Ti or RTX 3060 12 GB for an LLM?+
The 12 GB 3060 for local AI, despite its lower speed: 75.6 tokens per second versus 92.2 for the Ti on the llama.cpp test, but with 4 GB more memory. Those 4 GB support 12-billion-parameter models such as Gemma 4 12B, which the 3060 Ti cannot load with context.
Can the RTX 3060 Ti run Gemma 4 12B?+
Not comfortably. The catalog lists 7 GB of Q4 weights, while llama.cpp sees only 7,838 MiB on the card, even before the context cache and display. The model spills over as soon as the conversation gets longer. Stick with Granite 4.2 8B (4.6 GB) or Qwen 3.5 9B (6 GB).
How many tokens per second on a RTX 3060 Ti?+
92.2 tokens per second during generation, and 96.5 with Flash Attention, on Llama 2 7B in Q4_0 according to llama.cpp’s public table. A larger model will be slower, roughly in proportion to its size: expect around 60 to 75 tokens per second for an 8B or 9B model in Q4, as a theoretical estimate.
Is RTX 3060 Ti better than RTX 4060 for an LLM?+
With the same 8 GB of memory, the 3060 Ti generates faster: 92.2 tokens per second versus 63.9 for the RTX 4060 Ti 8 GB, whose bus is only 128 bits. The 4060 uses less power and offers newer features, but it is not faster for text generation.
What power supply do you need for a RTX 3060 Ti?+
NVIDIA lists 200 W for the card and recommends a 600 W system power supply, versus 550 W for the RTX 3060. A 500 W power supply is therefore below the manufacturer's recommendation: choose a good-quality 600 W unit, especially if your processor is power-hungry.
Is Flash Attention worth it on a RTX 3060 Ti?+
Yes, especially for long prompts: the llama.cpp table measures reading 4,096 tokens at 2,598 tokens per second with Flash Attention versus 1,897 without it, or 37% faster. Generation improves by about 5%. Ollama enables it automatically when the card supports it.
Did this guide help you?

Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.

Prices in euros (€) are French market prices including VAT, as checked by BestLLMfor. US prices differ: the Amazon buttons show the current US price.