Which LLM on RTX 3060 12 GB ?
The RTX 3060 12 GB runs 7- to 14-billion-parameter models in Q4 without compromise: Granite 4.2 8B (5.3 GB), Qwen3.5 9B (6.6 GB), Gemma 4 12B (7.6 GB), and Qwen3 14B (9.3 GB). Its memory determines the model and context size, while its 360 GB/s bandwidth sets the speed ceiling: about 50 tokens/s on an 8-9B. Above 14B, you need a MoE model with CPU assistance or 16 to 24 GB of VRAM.
This guide starts with the card: what its 12 GB of VRAM and 360 GB/s enable, what third parties have measured, and where the real limits are (context, 27B to 32B models, and the second card). Every weight comes from the model's Ollama page, and every throughput figure is attributed to its source, because the numbers circulating about this card often contradict one another.
Choosing a machine? Our picks by budget →
Buying alternative for this guide: RTX 5060 Ti 16GB (ASUS Prime).
Compare all options by budget, from €800 to €3,500 →
On the go: which laptop for local AI →
Affiliate links — commission possible at no extra cost to you. As an Amazon Associate, BestLLMfor earns from qualifying purchases.
#The RTX 3060 12 GB in figures
The RTX 3060 12 GB loads all models with 7 to 14 billion parameters in Q4 completely: Granite 4.2 8B (5.3 GB in Ollama), Qwen3.5 9B (6.6 GB), Gemma 4 12B (7.6 GB), or Qwen3 14B (9.3 GB). Its compute power is not what determines this; its memory does. The 12 GB of GDDR6 sets the maximum model and context size. The 360 GB/s of bandwidth sets the speed, since the card rereads most of the weights for every token. Third-party measurements report 50 to 65 tokens per second on a model with 8 to 9 billion parameters and 26 to 33 on a 14B, enough for a chat or coding assistant. A dense model with 27 to 32 billion parameters overflows the capacity: you need an MoE, a second card, or more VRAM.
- Architecture
- Ampere, 3,584 CUDA cores, 3rd-generation Tensor Cores (NVIDIA spec sheet).
- Memory
- 12 GB of GDDR6 on a 192-bit bus. Tom's Hardware puts the peak bandwidth at 360 GB/s.
- Power supply
- 170 W of board power; NVIDIA specifies 550 W for the system's power supply.
- Compatibility Ollama
- Compute capability 8.6, NVIDIA driver 550 or later on Linux, 551.61 or later on Windows.
- Pricing
- No prices here; they change every week: see the site's AI GPU Price Tracker.
#1. Which models fit in 12 GB
The rule of thumb: file weights, plus the context cache, plus about 1 GB of headroom for the runtime. The table shows the weights listed by Ollama for models that the card loads entirely and for a few neighboring models that do not fit. The full list, across all cards, is in the guide dedicated to 12 GB of VRAM.
| Model (tag Ollama) | Weights | Advertised context | Fits in 12 GB? |
|---|---|---|---|
| Granite 4.2 8B (granite4.2:8b) | 5.3 GB | 128K | Yes, plenty of headroom |
| Qwen3.5 9B in Q4_K_M (qwen3.5:9b) | 6.6 GB | 256K | Yes, plenty of headroom |
| Gemma 4 12B (gemma4:12b) | 7.6 GB | 256K | Yes, about 4 GB of headroom |
| Qwen3 14B (qwen3:14b) | 9.3 GB | 40K | Yes, context to limit |
| Qwen3.5 9B in Q8_0 (qwen3.5:9b-q8_0) | 11 GB | 256K | Barely enough: about 1 GB of headroom |
| Qwen3.5 27B (qwen3.5:27b) | 17 GB | 256K | No |
| Gemma 4 26B (gemma4:26b) | 19 GB | 256K | No |
| Mistral Small 3.2 24B (mistral-small3.2) | 15 GB | 128K | No |
Two pitfalls. The number in the name does not indicate the size: gemma4:e4b weighs 9.6 GB and gemma4:e2b weighs 7.2 GB in Ollama, which is more than the 12B (7.6 GB). Check the size before downloading. Also, Qwen3.5 9B’s Q8_0 technically fits, but leaves only about 1 GB for context and buffers: prefer the 6.6 GB Q4_K_M, which frees up more than 5 GB.
#2. Install Ollama and verify that everything runs on the GPU
- 01Check the NVIDIA driverOllama requires driver version 550 or later on Linux and 551.61 or later on Windows. The nvidia-smi command displays the installed version and the card name.
- 02Install OllamaOn Windows and macOS, use the installer from ollama.com. On Linux, a single command is enough; the NVIDIA GPU is detected automatically.
- 03Run a first 12 GB modelollama run gemma4:12b downloads 7.6 GB and launches a model that also accepts images. For more speed, ollama run qwen3.5:9b weighs 6.6 GB.
- 04Check placementollama ps, in a second terminal, displays the PROCESSOR column. The expected value is 100% GPU.
#3. Published speeds and the theoretical ceiling
For each token, the card rereads the active weights from VRAM. The theoretical ceiling is therefore bandwidth divided by model weight: 360 GB/s divided by 5.3 GB gives about 68 tokens/s for Granite 4.2 8B, 55 for Qwen3.5 9B (6.6 GB), 47 for Gemma 4 12B (7.6 GB), and 39 for Qwen3 14B (9.3 GB). These are ceilings, never measurements: Qwen3.5 9B is published at 49.9 tokens/s (about 90% of the ceiling) and Qwen3 14B at 33.4 (about 86%), confirming that memory limits speed.
| Model and quantization | Generation (tokens/s) | Source and conditions |
|---|---|---|
| Llama 3.1 8B Instruct, Q4_K_M | 51,3 | LocalScore (Mozilla Builders project) |
| Llama 3.1 8B Instruct, Q4_K_M | 64,5 | Tyolab blog, llama.cpp, 8,192-token context, May 2026 |
| Llama 7B, Q4_0 (3.56 GiB) | 75,6 | llama.cpp discussion about CUDA GPUs (llama-bench, WSL2) |
| Qwen3.5 9B, Q4_K_M | 49,9 | Tyolab blog, same conditions |
| Gemma 4 12B, Q5_K_XL | about 33 | Gemma4All, based on a community test (llama.cpp, Flash Attention, Q8_0 KV cache) |
| Qwen3 14B, Q4_K_M | 33,4 | Tyolab blog, context reduced to 4,096 |
| Qwen2.5 14B Instruct, Q4_K_M | 26,6 | LocalScore |
Differences on the same model (51.3 versus 64.5 tokens/s for Llama 3.1 8B) come from the method: software, prompt length, and system. Compare only figures from the same source; none were measured by the site.
Generation is only half the experience. LocalScore measures prompt processing at 1,483 tokens/s on Llama 3.1 8B and 759 tokens/s on Qwen2.5 14B. A 10,000-token document therefore waits about 7 seconds before the first response on an 8B model and 13 seconds on a 14B model. This processing depends on compute, not bandwidth.
#4. A 35B model on 12 GB: the MoE case
An MoE (mixture of experts) model activates only a fraction of its parameters for each token. Qwen3.6 35B-A3B activates approximately 3 billion, but its 23 GB file in Ollama does not fit in 12 GB. The solution planned by llama.cpp is the --n-cpu-moe option: it keeps the expert weights from the first N layers in RAM and leaves attention on the GPU.
A user published their measurements on September 22, 2026: RTX 3060 12 GB, Ryzen 5 3600, 32 GB of DDR4-3200, Qwen3.6-35B-A3B in Q4_K_M (about 21 GB). They measured 47 tokens/s on a short question, 45 with 12,000 tokens of context, and 37 with 45,000. With a conventional layer split, the same machine dropped to 7.1 tokens/s on 15,000 tokens (measured with Qwen3-30B-A3B), versus 25.3 after tuning. Only one in four layers of the Qwen3.5 and 3.6 series uses standard attention, so their speed drops less as the context grows.
Three caveats: this is one person’s measurement, using llama.cpp rather than Ollama; it assumes 32 GB of RAM; and the RAM speed matters. Its DDR4-3200 was running at 2,400 by default, and the XMP or DOCP profile added 10 to 17% throughput.
The value 26 is the author's value for their model: the higher it goes, the more VRAM is freed, but the slower generation becomes. Look for the lowest value that leaves about 1 GB of VRAM free.
#5. Context and KV cache
#Context, the first limit of 12 GB
Ollama allocates 4,096 context tokens by default under 24 GiB of VRAM, while its documentation recommends at least 64,000 tokens for agents and coding tools. Between the two, each token costs memory. Example with Qwen3 14B, whose architecture is published: 40 layers, 8 KV heads, 128 dimensions per head. Per token, the KV cache in 16-bit occupies 2 (keys and values) × 40 × 8 × 128 × 2 bytes, or 160 KiB: about 2.4 GiB for 16,000 tokens and 3.7 GiB for 24,000.
The 9.3 GB model represents 8.7 GiB, leaving 12 GiB of VRAM. At 16,000 tokens, the total reaches 11.1 GiB: less than 1 GiB remains for runtime buffers, which is too tight. At 24,000 tokens, it reaches 12.3 GiB: loading fails or spills over to the CPU. With an 8-bit cache, 24,000 tokens use only 1.8 GiB, for 10.5 GiB total: it fits.
#8-bit KV cache and Flash Attention
Ollama automatically enables Flash Attention when the software and card support it; OLLAMA_FLASH_ATTENTION=1 forces it. This is a prerequisite for quantized caching: with OLLAMA_KV_CACHE_TYPE=q8_0, the cache uses about half the memory of f16 with very little precision loss, according to the documentation. The setting is global: all served models use the same cache type.
#6. Two RTX 3060 12 GB: 24 GB of VRAM, not twice as fast
Two cards offer 24 GB of VRAM, enough for a dense 27 to 32B model in Q4: Qwen3.5 27B weighs 17 GB in Ollama. Ollama automatically distributes the model across all cards when it does not fit on one. In llama.cpp, the default split is by layer, in a pipeline: the layers run one after another. The tensor mode, which parallelizes computation, is marked experimental.
The gain is therefore capacity, not speed: each token rereads all the weights at 360 GB/s per card. The estimated ceiling for a 32B model weighing 20 GB is 360 divided by 20, or about 18 tokens/s. A RTX 3090 offers 24 GB on a single card at 936 GB/s according to RunAIHome: the same model then tops out at around 47 tokens/s, or 2.6 times more. Two 170 W cards also mean 340 W for the GPU alone, whereas the 550 W of NVIDIA applies to a single card.
This setup makes sense if you already own a card and want to reach 24 GB. If you are starting from scratch, compare it first with a single 16 or 24 GB card.
#7. RTX 3060 8 GB, 3060 Ti, 4060, 5060: what 12 GB changes
| Card | VRAM | Memory and bus | 8B Q4_K_M, tokens/s | Gemma 4 12B (7.6 GB) |
|---|---|---|---|---|
| RTX 3060 12 GB | 12 GB | GDDR6, 192-bit | 51,3 | Yes, with about 4 GB of context |
| RTX 3060 8 GB | 8 GB | GDDR6, 128-bit | 37,0 | No in practice (0.4 GB of headroom) |
| RTX 3060 Ti | 8 GB | GDDR6 or GDDR6X, 256 bits | 60,2 | Not in practice |
| RTX 4060 | 8 GB | GDDR6, 128-bit | 38,1 | Not in practice |
| RTX 5060 | 8 GB | GDDR7, 128-bit | — | Not in practice |
| RTX 5060 Ti | 16 GB or 8 GB | GDDR7, 128-bit | 51.2 (16 GB) | Yes, in the 16 GB version |
The 3060 8 GB is not simply a 3060 12 GB with less memory: its bus drops from 192 to 128 bits, and Tom's Hardware puts its bandwidth at 240 GB/s versus 360 GB/s, or 33% less. LocalScore records it at 37.0 tokens/s versus 51.3, or 28% less on the same model. These speeds are isolated results, without controlled conditions: the 3060 Ti comes out ahead there, which its 256-bit bus makes plausible, but its 8 GB excludes it from 12- or 14-billion-parameter models. The 5060 offsets its narrow bus with GDDR7, without changing its capacity. Only the 5060 Ti 16 GB exceeds the 3060 in VRAM.
#Buying a 3060 12 GB: when it makes sense and when to move on
- Buy a 3060 12 GB if
- you want a first GPU for Ollama, 7B to 14B models, chat, document RAG, or coding assistance, and capacity matters more than speed.
- Aim for 16 GB if
- you want Gemma 4 12B with a long context, or more headroom for a 14B. The RTX 5060 Ti 16 GB is the first step; the comparison page quantifies the difference.
- Aim for 24 GB if
- Dense models from 27B to 32B are your target: a 24 GB card avoids splitting the model across two cards and provides much higher bandwidth.
- Avoid or verify if
- The announcement does not specify the memory: the 8 GB version offers neither the capacity nor the bandwidth of the 12 GB version.
- Source: NVIDIA spec sheet for RTX 3060 and 3060 Ti
- Source: Ollama documentation, context length
- Source: Ollama documentation, supported GPU
- Source: llama.cpp, server options (--n-cpu-moe, KV cache)
- Source: LocalScore, results from the RTX 3060 12 GB
#Frequently asked questions
Is RTX 3060 12 GB enough for a local LLM?+
Which model should you run first with Ollama on a RTX 3060 12 GB?+
How many tokens per second on a 12 GB RTX 3060?+
Can you run a 32B or 70B model on a RTX 3060 with 12 GB?+
Are two RTX 3060 12 GB cards as good as one RTX 3090 for local AI?+
RTX 3060 12 GB or RTX 4060 8 GB for an LLM?+
Prices change quickly: every Monday and Thursday, our tracker records the lowest price for local AI graphics cards, along with the price per GB of VRAM.
Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.