Beginner 11 minRTX 30

Which LLM on RTX 3060 8 GB ?

Direct response

The RTX 3060 8 GB runs 7B to 9B models in Q4, such as Granite 4.2 8B (4.6 GB) or Qwen 3.5 9B (6 GB), but not 12B models. Its distinguishing feature is a 128-bit bus and 240 GB/s, versus 192-bit and 360 GB/s for the 3060 12 GB, making it significantly slower for generation. The public llama.cpp table lists 75.6 tokens/s for the 12 GB model on a 7B.

Under the name RTX 3060, NVIDIA sold two cards: the original 12 GB model and an 8 GB version released in late 2022, with a narrower memory bus. This guide explains what that changes for local AI, which models fit in 8 GB, how to avoid confusing the two versions when buying, and what speeds to expect, distinguishing measured results from unmeasured ones.

Choosing a machine? Our picks by budget →

By Mohamed Meguedmi·Update 2026-09-30·Tested on Windows, macOS, and Linux
Recommended hardware

Buying alternative for this guide: RTX 5060 Ti 16GB (ASUS Prime).

Compare all options by budget, from €800 to €3,500 →

On the go: which laptop for local AI →

Affiliate links — possible commission at no extra cost to you. As an Amazon Associate, BestLLMfor earns from qualifying purchases.

#RTX 3060 8 GB: what it changes compared with the 12 GB version

The RTX 3060 8 GB runs 7- to 9-billion-parameter models in Q4: Granite 4.2 8B requires 4.6 GB and Qwen 3.5 9B 6 GB according to the QuelLLM catalog. Gemma 4 12B (7 GB in Q4) does not fit comfortably. Its differences from the 3060 12 GB are twofold: 4 GB less memory and a 128-bit bus instead of 192-bit. According to the specialist press that recorded the specifications when it launched, the narrower bus reduces bandwidth to 240 GB/s, or 33% less than the 360 GB/s of the 12 GB version. For local AI, this matters: text generation depends primarily on memory bandwidth. No public benchmark exists for the 8 GB version, but we can estimate that it generates about 50 tokens per second on a 7-billion-parameter model in Q4_0, compared with 75.6 measured on the 12 GB version.

Chip
The same GA106 as the 3060 12 GB and the same 170 W card power.
Memory
8 GB of GDDR6 on a 128-bit bus, or 240 GB/s, versus 12 GB, 192 bits, and 360 GB/s.
Consequence
The same number of compute cores, but less capacity and bandwidth: memory, not compute, is what limits local AI.

#Models compatible with 8 GB

Models for an RTX 3060 8 GB (weights according to the QuelLLM catalog)
ModelWeights in Q4Q8 sizeOn 8 GB
Qwen 3.5 4B2.3 GB4.3 GBQ8 possible, comfortable context
Granite 4.2 8B4.6 GB9 GBComfortable Q4; Q8 doesn’t fit
Qwen 3.5 9B6 GB10 GBQ4 with a moderate context
Gemma 4 12B7 GB13 GBDoesn't hold up with context

Its capacity is identical to that of a 3070 or 3060 Ti, two faster 8 GB cards. The same rules apply: target a maximum of 6 GB for weights, keep one gigabyte for the driver and context, and verify that the model fits entirely on the card. To do this, the ollama ps command displays 100% GPU when everything is loaded on the card. If a CPU percentage appears, the model spills over and speed collapses.

#Speed: why the memory bus dominates

The llama.cpp comparison table (Llama 2 7B in Q4_0, 3.56 GiB file) doesn't include the 3060 8 GB, but it helps explain what the memory bus changes. It reports 75.6 tokens per second for the 3060 12 GB (192-bit bus), 92.2 for the 3060 Ti 8 GB (256-bit bus), and 63.9 for an RTX 4060 Ti 8 GB card whose bus is only 128 bits wide. At equal capacity, the ranking follows bus width, with a few nuances: a newer card with a narrow bus can end up behind an older card with a wide bus.

Generation on Llama 2 7B Q4_0: the bus makes the difference
CardMemoryBusGeneration (tok/s)Status
RTX 3060 Ti8 GB256-bit92,2Published measurement
RTX 3060 12 GB12 GB192 bits75,6Published measurement
RTX 4060 Ti 8 GB8 GB128 bits63,9Published measurement
RTX 3060 8 GB8 GB128 bitsapproximately 47 to 50Estimate, not measured

The final-line estimate applies the measured efficiency of the 12 GB card—about 80% of what its bandwidth allows—to 240 GB/s. This is a projection and should be presented as such. Scaling by file size gives roughly 40 tokens per second for Granite 4.2 8B (4.6 GB) and about thirty for Qwen 3.5 9B (6 GB), theoretical upper bounds. This throughput remains readable in a conversation, but it puts the 8 GB card about one-third below the 12 GB card and more than 40% below the 3060 Ti.

#Pitfalls to avoid

Two cards under one name
NVIDIA lists the RTX 3060 with 12 GB or 8 GB, and a 192- or 128-bit bus. The name alone tells you nothing: require the capacity to be stated in the listing.
Installation verification
Run nvidia-smi: the 12 GB card exposes about 12 287 MiB according to llama.cpp measurements, while the 8 GB card exposes about 7 800 MiB, like its neighbor, the 3060 Ti.
Priced in line with the 12 GB model
Some vendors list the same price for both versions. Compare against the site's weekly tracking rather than a fixed price.
Unrealistic expectations for 12B
Gemma 4 12B and models of this size do not fit with context; the ceiling is 9B in Q4.
Read the card's name and total memory
nvidia-smi --query-gpu=name,memory.total --format=csv

The second check also applies to a refurbished PC: the manufacturer's description or the card name in Windows do not always distinguish between versions, whereas the total memory reported by nvidia-smi does not lie.

#The calculation behind the estimate

The principle is simple: to produce a token, the card rereads most of the model’s weights. The maximum number of tokens per second is therefore roughly the bandwidth divided by the file size. The test file weighs 3.56 GiB, or about 3.82 GB. For the 3060 12 GB, 360 GB/s divided by 3.82 gives a ceiling of 94 tokens per second, and the published measurement reaches 80% of that. For the 8 GB, 240 GB/s gives a ceiling of about 63 tokens per second; applying the same efficiency yields around fifty.

From the theoretical ceiling to the estimate
CardBandwidthCeiling (GB/s ÷ 3.82 GB)Result
RTX 3060 12 GB360 GB/sapproximately 94 tok/s75.6 tok/s measured (about 80%)
RTX 3060 8 GB240 GB/sabout 63 tok/sapproximately 50 tok/s estimated

This method has one limitation: performance varies from one card to another, as shown by other cards in the same table whose measured-to-theoretical ratio ranges from 67% to 85%. So use a range rather than a precise figure. Also note that published results depend on the driver, system, and card manufacturer, even with the same chip, and measure on your own hardware if speed affects a purchase decision.

#What it’s like in practice

Translate throughput into wait time. At around 40 tokens per second, a 500-token response—about a page of explanation—takes roughly a dozen seconds; on a 3060 12 GB, the same response would take about eight, still as a theoretical estimate. For a conversation, the difference is tolerable: the text appears faster than you can read it. It becomes burdensome in two cases: agents that make many successive model calls, and batch processing of many documents, where the gap accumulates.

The 8 GB capacity is a more significant limitation than speed. It forces you to choose between a long context and a larger model, whereas a 12 GB card gives you both. If your use case is limited to asking short questions to an 8B model, 8 GB is sufficient. If you want to connect the model to your documents, more memory is the better choice.

  1. 01
    Control the announcement
    Look for the mention 8 GB or 12 GB in the title, description, or box photos. If it is not specified, ask the seller.
  2. 02
    Request a screenshot
    Require a capture from nvidia-smi or GPU-Z showing the total memory and bus width.
  3. 03
    Test after removal
    On the system, run nvidia-smi: approximately 12,287 MiB for the 12 GB card and around 7,800 MiB for the 8 GB card.
  4. 04
    Run a model
    Load an 8B model in Q4 with Ollama and verify with ollama ps that it is 100% on the GPU before deciding to buy.

#Context: making the most of 8 GB

On 8 GB, the context cache is the real limit. Ollama's documentation indicates that the K/V cache can be quantized when Flash Attention is enabled: set OLLAMA_FLASH_ATTENTION=1 and then OLLAMA_KV_CACHE_TYPE=q8_0 before starting the server. With a 128-bit bus, the card is already struggling for speed, so it's better not to add an oversized context. A context of a few thousand tokens with an 8B model in Q4 is a good compromise.

Linux: launch Ollama with quantized K/V cache
OLLAMA_FLASH_ATTENTION=1 OLLAMA_KV_CACHE_TYPE=q8_0 ollama serve

#What to use instead

Alternatives at a comparable budget
CardMemoryGeneration (tok/s)What it offers
RTX 3060 12 GB12 GB75.6 (measured)Opens the 12B, faster than the 8 GB
RTX 3060 Ti8 GB92.2 (measurement)Same capacity, but much faster
RTX 4060 Ti 8 GB8 GB63.9 (measurement)Newer, but with a 128-bit bus like the 3060 8 GB

The choice depends on your needs: for capacity, the 3060 12 GB; for speed on 8-billion-parameter models, the 3060 Ti. Used prices change every week: the site's tracker reports the lowest price for cards used in local AI.

#2026 verdict

Keep it if
It is already in your PC, and you are targeting an 8B or 9B model in Q4 for chat or coding assistance, without demanding high speed.
Don’t buy for AI
A 3060 12 GB offers more, often for a similar budget. Only pay for the 8 GB model if it is significantly cheaper.
Always check the version
Require the memory capacity to be listed in the advertisement and verify it with nvidia-smi upon delivery.

#Frequently asked questions

Frequently asked questions
How do you tell the RTX 3060 8 GB from the 12 GB model?+
Check the memory specification on the packaging or listing, then run nvidia-smi: total memory is close to 12 000 MiB for the 12 GB version (12 287 MiB measured by llama.cpp), and 8 000 MiB for the 8 GB version. NVIDIA also distinguishes the buses: 192 bits for the 12 GB version, 128 bits for the 8 GB version.
How many tokens per second on an 8 GB RTX 3060?+
No public measurement exists. The 12 GB generates 75.6 tokens per second on a 7B in Q4_0 according to the llama.cpp table; with 240 GB/s instead of 360, the 8 GB should run at around 47 to 50 tokens per second on the same test, as a theoretical estimate.
RTX 3060 8 GB or RTX 4060 8 GB for an LLM?+
Both have 8 GB and a 128-bit bus. The RTX 4060 Ti 8 GB, another card with a 128-bit bus, measures 63.9 tokens per second on the public llama.cpp leaderboard. Choose based on price and power consumption: with equal capacity and bus width, the speed will be in the same range.
Can the RTX 3060 8 GB run Gemma 4 12B?+
With difficulty. The catalog lists 7 GB of Q4 weights, leaving barely one gigabyte for context and the system. The model is then split between the card and RAM, as ollama ps indicates, and speed drops. Prefer Qwen 3.5 9B (6 GB) or Granite 4.2 8B (4.6 GB).
Does Ollama work on a RTX 3060 8 GB?+
Yes. Ollama lists the RTX 3060 among cards with compute capability 8.6 and requires driver 550 or later. After the first load, check with ollama ps that 100% GPU appears: this indicates that the model fits entirely within 8 GB.
Would a RTX 3060 12 GB be better?+
For local AI, yes in almost all cases: it reaches 75.6 tokens per second on the public llama.cpp benchmark, has 4 GB more, and opens 12-billion-parameter models. The 8 GB version is worthwhile only if it is already installed or significantly cheaper.
Did this guide help you?

Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.