Which LLM on RTX 3060 8 GB ?
The RTX 3060 8 GB runs 7B to 9B models in Q4, such as Granite 4.2 8B (4.6 GB) or Qwen 3.5 9B (6 GB), but not 12B models. Its distinguishing feature is a 128-bit bus and 240 GB/s, versus 192-bit and 360 GB/s for the 3060 12 GB, making it significantly slower for generation. The public llama.cpp table lists 75.6 tokens/s for the 12 GB model on a 7B.
Under the name RTX 3060, NVIDIA sold two cards: the original 12 GB model and an 8 GB version released in late 2022, with a narrower memory bus. This guide explains what that changes for local AI, which models fit in 8 GB, how to avoid confusing the two versions when buying, and what speeds to expect, distinguishing measured results from unmeasured ones.
Choosing a machine? Our picks by budget →
Buying alternative for this guide: RTX 5060 Ti 16GB (ASUS Prime).
Compare all options by budget, from €800 to €3,500 →
On the go: which laptop for local AI →
Affiliate links — possible commission at no extra cost to you. As an Amazon Associate, BestLLMfor earns from qualifying purchases.
#RTX 3060 8 GB: what it changes compared with the 12 GB version
The RTX 3060 8 GB runs 7- to 9-billion-parameter models in Q4: Granite 4.2 8B requires 4.6 GB and Qwen 3.5 9B 6 GB according to the QuelLLM catalog. Gemma 4 12B (7 GB in Q4) does not fit comfortably. Its differences from the 3060 12 GB are twofold: 4 GB less memory and a 128-bit bus instead of 192-bit. According to the specialist press that recorded the specifications when it launched, the narrower bus reduces bandwidth to 240 GB/s, or 33% less than the 360 GB/s of the 12 GB version. For local AI, this matters: text generation depends primarily on memory bandwidth. No public benchmark exists for the 8 GB version, but we can estimate that it generates about 50 tokens per second on a 7-billion-parameter model in Q4_0, compared with 75.6 measured on the 12 GB version.
- Chip
- The same GA106 as the 3060 12 GB and the same 170 W card power.
- Memory
- 8 GB of GDDR6 on a 128-bit bus, or 240 GB/s, versus 12 GB, 192 bits, and 360 GB/s.
- Consequence
- The same number of compute cores, but less capacity and bandwidth: memory, not compute, is what limits local AI.
#Models compatible with 8 GB
| Model | Weights in Q4 | Q8 size | On 8 GB |
|---|---|---|---|
| Qwen 3.5 4B | 2.3 GB | 4.3 GB | Q8 possible, comfortable context |
| Granite 4.2 8B | 4.6 GB | 9 GB | Comfortable Q4; Q8 doesn’t fit |
| Qwen 3.5 9B | 6 GB | 10 GB | Q4 with a moderate context |
| Gemma 4 12B | 7 GB | 13 GB | Doesn't hold up with context |
Its capacity is identical to that of a 3070 or 3060 Ti, two faster 8 GB cards. The same rules apply: target a maximum of 6 GB for weights, keep one gigabyte for the driver and context, and verify that the model fits entirely on the card. To do this, the ollama ps command displays 100% GPU when everything is loaded on the card. If a CPU percentage appears, the model spills over and speed collapses.
#Speed: why the memory bus dominates
The llama.cpp comparison table (Llama 2 7B in Q4_0, 3.56 GiB file) doesn't include the 3060 8 GB, but it helps explain what the memory bus changes. It reports 75.6 tokens per second for the 3060 12 GB (192-bit bus), 92.2 for the 3060 Ti 8 GB (256-bit bus), and 63.9 for an RTX 4060 Ti 8 GB card whose bus is only 128 bits wide. At equal capacity, the ranking follows bus width, with a few nuances: a newer card with a narrow bus can end up behind an older card with a wide bus.
| Card | Memory | Bus | Generation (tok/s) | Status |
|---|---|---|---|---|
| RTX 3060 Ti | 8 GB | 256-bit | 92,2 | Published measurement |
| RTX 3060 12 GB | 12 GB | 192 bits | 75,6 | Published measurement |
| RTX 4060 Ti 8 GB | 8 GB | 128 bits | 63,9 | Published measurement |
| RTX 3060 8 GB | 8 GB | 128 bits | approximately 47 to 50 | Estimate, not measured |
The final-line estimate applies the measured efficiency of the 12 GB card—about 80% of what its bandwidth allows—to 240 GB/s. This is a projection and should be presented as such. Scaling by file size gives roughly 40 tokens per second for Granite 4.2 8B (4.6 GB) and about thirty for Qwen 3.5 9B (6 GB), theoretical upper bounds. This throughput remains readable in a conversation, but it puts the 8 GB card about one-third below the 12 GB card and more than 40% below the 3060 Ti.
#Pitfalls to avoid
- Two cards under one name
- NVIDIA lists the RTX 3060 with 12 GB or 8 GB, and a 192- or 128-bit bus. The name alone tells you nothing: require the capacity to be stated in the listing.
- Installation verification
- Run nvidia-smi: the 12 GB card exposes about 12 287 MiB according to llama.cpp measurements, while the 8 GB card exposes about 7 800 MiB, like its neighbor, the 3060 Ti.
- Priced in line with the 12 GB model
- Some vendors list the same price for both versions. Compare against the site's weekly tracking rather than a fixed price.
- Unrealistic expectations for 12B
- Gemma 4 12B and models of this size do not fit with context; the ceiling is 9B in Q4.
The second check also applies to a refurbished PC: the manufacturer's description or the card name in Windows do not always distinguish between versions, whereas the total memory reported by nvidia-smi does not lie.
#The calculation behind the estimate
The principle is simple: to produce a token, the card rereads most of the model’s weights. The maximum number of tokens per second is therefore roughly the bandwidth divided by the file size. The test file weighs 3.56 GiB, or about 3.82 GB. For the 3060 12 GB, 360 GB/s divided by 3.82 gives a ceiling of 94 tokens per second, and the published measurement reaches 80% of that. For the 8 GB, 240 GB/s gives a ceiling of about 63 tokens per second; applying the same efficiency yields around fifty.
| Card | Bandwidth | Ceiling (GB/s ÷ 3.82 GB) | Result |
|---|---|---|---|
| RTX 3060 12 GB | 360 GB/s | approximately 94 tok/s | 75.6 tok/s measured (about 80%) |
| RTX 3060 8 GB | 240 GB/s | about 63 tok/s | approximately 50 tok/s estimated |
This method has one limitation: performance varies from one card to another, as shown by other cards in the same table whose measured-to-theoretical ratio ranges from 67% to 85%. So use a range rather than a precise figure. Also note that published results depend on the driver, system, and card manufacturer, even with the same chip, and measure on your own hardware if speed affects a purchase decision.
#What it’s like in practice
Translate throughput into wait time. At around 40 tokens per second, a 500-token response—about a page of explanation—takes roughly a dozen seconds; on a 3060 12 GB, the same response would take about eight, still as a theoretical estimate. For a conversation, the difference is tolerable: the text appears faster than you can read it. It becomes burdensome in two cases: agents that make many successive model calls, and batch processing of many documents, where the gap accumulates.
The 8 GB capacity is a more significant limitation than speed. It forces you to choose between a long context and a larger model, whereas a 12 GB card gives you both. If your use case is limited to asking short questions to an 8B model, 8 GB is sufficient. If you want to connect the model to your documents, more memory is the better choice.
- 01Control the announcementLook for the mention 8 GB or 12 GB in the title, description, or box photos. If it is not specified, ask the seller.
- 02Request a screenshotRequire a capture from nvidia-smi or GPU-Z showing the total memory and bus width.
- 03Test after removalOn the system, run nvidia-smi: approximately 12,287 MiB for the 12 GB card and around 7,800 MiB for the 8 GB card.
- 04Run a modelLoad an 8B model in Q4 with Ollama and verify with ollama ps that it is 100% on the GPU before deciding to buy.
#Context: making the most of 8 GB
On 8 GB, the context cache is the real limit. Ollama's documentation indicates that the K/V cache can be quantized when Flash Attention is enabled: set OLLAMA_FLASH_ATTENTION=1 and then OLLAMA_KV_CACHE_TYPE=q8_0 before starting the server. With a 128-bit bus, the card is already struggling for speed, so it's better not to add an oversized context. A context of a few thousand tokens with an 8B model in Q4 is a good compromise.
#What to use instead
| Card | Memory | Generation (tok/s) | What it offers |
|---|---|---|---|
| RTX 3060 12 GB | 12 GB | 75.6 (measured) | Opens the 12B, faster than the 8 GB |
| RTX 3060 Ti | 8 GB | 92.2 (measurement) | Same capacity, but much faster |
| RTX 4060 Ti 8 GB | 8 GB | 63.9 (measurement) | Newer, but with a 128-bit bus like the 3060 8 GB |
The choice depends on your needs: for capacity, the 3060 12 GB; for speed on 8-billion-parameter models, the 3060 Ti. Used prices change every week: the site's tracker reports the lowest price for cards used in local AI.
- Local AI graphics card prices (weekly tracking)
- Which LLM for RTX 3060 with 12 GB?
- Which LLM on RTX 3060 Ti (8 GB)?
#2026 verdict
- Keep it if
- It is already in your PC, and you are targeting an 8B or 9B model in Q4 for chat or coding assistance, without demanding high speed.
- Don’t buy for AI
- A 3060 12 GB offers more, often for a similar budget. Only pay for the 8 GB model if it is significantly cheaper.
- Always check the version
- Require the memory capacity to be listed in the advertisement and verify it with nvidia-smi upon delivery.
#Frequently asked questions
How do you tell the RTX 3060 8 GB from the 12 GB model?+
How many tokens per second on an 8 GB RTX 3060?+
RTX 3060 8 GB or RTX 4060 8 GB for an LLM?+
Can the RTX 3060 8 GB run Gemma 4 12B?+
Does Ollama work on a RTX 3060 8 GB?+
Would a RTX 3060 12 GB be better?+
- Source: llama.cpp performance table on CUDA
- Source: NVIDIA documentation from the RTX 3060
- Source: Tom's Hardware, RTX 3060 8 GB
- Source: Ollama FAQ (ollama ps, K/V cache)
Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.