Beginner 14 minRTX 30

Which LLM on RTX 3060 12 GB ?

Direct response

The RTX 3060 12 GB runs 7- to 14-billion-parameter models in Q4 without compromise: Granite 4.2 8B (5.3 GB), Qwen3.5 9B (6.6 GB), Gemma 4 12B (7.6 GB), and Qwen3 14B (9.3 GB). Its memory determines the model and context size, while its 360 GB/s bandwidth sets the speed ceiling: about 50 tokens/s on an 8-9B. Above 14B, you need a MoE model with CPU assistance or 16 to 24 GB of VRAM.

This guide starts with the card: what its 12 GB of VRAM and 360 GB/s enable, what third parties have measured, and where the real limits are (context, 27B to 32B models, and the second card). Every weight comes from the model's Ollama page, and every throughput figure is attributed to its source, because the numbers circulating about this card often contradict one another.

Choosing a machine? Our picks by budget →

By Mohamed Meguedmi·Update 2026-09-29·Tested on Windows, macOS, and Linux
Recommended hardware

Buying alternative for this guide: RTX 5060 Ti 16GB (ASUS Prime).

Compare all options by budget, from €800 to €3,500 →

On the go: which laptop for local AI →

Affiliate links — commission possible at no extra cost to you. As an Amazon Associate, BestLLMfor earns from qualifying purchases.

#The RTX 3060 12 GB in figures

The RTX 3060 12 GB loads all models with 7 to 14 billion parameters in Q4 completely: Granite 4.2 8B (5.3 GB in Ollama), Qwen3.5 9B (6.6 GB), Gemma 4 12B (7.6 GB), or Qwen3 14B (9.3 GB). Its compute power is not what determines this; its memory does. The 12 GB of GDDR6 sets the maximum model and context size. The 360 GB/s of bandwidth sets the speed, since the card rereads most of the weights for every token. Third-party measurements report 50 to 65 tokens per second on a model with 8 to 9 billion parameters and 26 to 33 on a 14B, enough for a chat or coding assistant. A dense model with 27 to 32 billion parameters overflows the capacity: you need an MoE, a second card, or more VRAM.

Architecture
Ampere, 3,584 CUDA cores, 3rd-generation Tensor Cores (NVIDIA spec sheet).
Memory
12 GB of GDDR6 on a 192-bit bus. Tom's Hardware puts the peak bandwidth at 360 GB/s.
Power supply
170 W of board power; NVIDIA specifies 550 W for the system's power supply.
Compatibility Ollama
Compute capability 8.6, NVIDIA driver 550 or later on Linux, 551.61 or later on Windows.
Pricing
No prices here; they change every week: see the site's AI GPU Price Tracker.
→
Why 12 GB matters more than compute speed
The RTX 4060 and 5060 from NVIDIA come with 8 GB. Gemma 4 12B takes up 7.6 GB in Ollama: on 8 GB, only 0.4 GB remains for context and the runtime—that is, nothing. On 12 GB, about 4 GB remains. This capacity gap makes the 3060 12 GB a credible local AI card.

#1. Which models fit in 12 GB

The rule of thumb: file weights, plus the context cache, plus about 1 GB of headroom for the runtime. The table shows the weights listed by Ollama for models that the card loads entirely and for a few neighboring models that do not fit. The full list, across all cards, is in the guide dedicated to 12 GB of VRAM.

Ollama weights and 12 GB performance (default context for Ollama: 4,096 tokens)
Model (tag Ollama)WeightsAdvertised contextFits in 12 GB?
Granite 4.2 8B (granite4.2:8b)5.3 GB128KYes, plenty of headroom
Qwen3.5 9B in Q4_K_M (qwen3.5:9b)6.6 GB256KYes, plenty of headroom
Gemma 4 12B (gemma4:12b)7.6 GB256KYes, about 4 GB of headroom
Qwen3 14B (qwen3:14b)9.3 GB40KYes, context to limit
Qwen3.5 9B in Q8_0 (qwen3.5:9b-q8_0)11 GB256KBarely enough: about 1 GB of headroom
Qwen3.5 27B (qwen3.5:27b)17 GB256KNo
Gemma 4 26B (gemma4:26b)19 GB256KNo
Mistral Small 3.2 24B (mistral-small3.2)15 GB128KNo

Two pitfalls. The number in the name does not indicate the size: gemma4:e4b weighs 9.6 GB and gemma4:e2b weighs 7.2 GB in Ollama, which is more than the 12B (7.6 GB). Check the size before downloading. Also, Qwen3.5 9B’s Q8_0 technically fits, but leaves only about 1 GB for context and buffers: prefer the 6.6 GB Q4_K_M, which frees up more than 5 GB.

#2. Install Ollama and verify that everything runs on the GPU

  1. 01
    Check the NVIDIA driver
    Ollama requires driver version 550 or later on Linux and 551.61 or later on Windows. The nvidia-smi command displays the installed version and the card name.
  2. 02
    Install Ollama
    On Windows and macOS, use the installer from ollama.com. On Linux, a single command is enough; the NVIDIA GPU is detected automatically.
  3. 03
    Run a first 12 GB model
    ollama run gemma4:12b downloads 7.6 GB and launches a model that also accepts images. For more speed, ollama run qwen3.5:9b weighs 6.6 GB.
  4. 04
    Check placement
    ollama ps, in a second terminal, displays the PROCESSOR column. The expected value is 100% GPU.
Linux (Windows: installer from ollama.com)
nvidia-smi
curl -fsSL https://ollama.com/install.sh | sh
ollama run gemma4:12b
ollama ps
i
What a GPU value other than 100% means
If PROCESSOR shows anything other than 100% GPU, part of the model or cache is running on the processor, and throughput collapses. Before concluding that the card is too weak, reduce the context or switch the cache to 8-bit (section 5): these are the two quickest levers.

#3. Published speeds and the theoretical ceiling

For each token, the card rereads the active weights from VRAM. The theoretical ceiling is therefore bandwidth divided by model weight: 360 GB/s divided by 5.3 GB gives about 68 tokens/s for Granite 4.2 8B, 55 for Qwen3.5 9B (6.6 GB), 47 for Gemma 4 12B (7.6 GB), and 39 for Qwen3 14B (9.3 GB). These are ceilings, never measurements: Qwen3.5 9B is published at 49.9 tokens/s (about 90% of the ceiling) and Qwen3 14B at 33.4 (about 86%), confirming that memory limits speed.

Generation throughput published on RTX 3060 12 GB (third-party measurements, not from the site)
Model and quantizationGeneration (tokens/s)Source and conditions
Llama 3.1 8B Instruct, Q4_K_M51,3LocalScore (Mozilla Builders project)
Llama 3.1 8B Instruct, Q4_K_M64,5Tyolab blog, llama.cpp, 8,192-token context, May 2026
Llama 7B, Q4_0 (3.56 GiB)75,6llama.cpp discussion about CUDA GPUs (llama-bench, WSL2)
Qwen3.5 9B, Q4_K_M49,9Tyolab blog, same conditions
Gemma 4 12B, Q5_K_XLabout 33Gemma4All, based on a community test (llama.cpp, Flash Attention, Q8_0 KV cache)
Qwen3 14B, Q4_K_M33,4Tyolab blog, context reduced to 4,096
Qwen2.5 14B Instruct, Q4_K_M26,6LocalScore

Differences on the same model (51.3 versus 64.5 tokens/s for Llama 3.1 8B) come from the method: software, prompt length, and system. Compare only figures from the same source; none were measured by the site.

Generation is only half the experience. LocalScore measures prompt processing at 1,483 tokens/s on Llama 3.1 8B and 759 tokens/s on Qwen2.5 14B. A 10,000-token document therefore waits about 7 seconds before the first response on an 8B model and 13 seconds on a 14B model. This processing depends on compute, not bandwidth.

#4. A 35B model on 12 GB: the MoE case

An MoE (mixture of experts) model activates only a fraction of its parameters for each token. Qwen3.6 35B-A3B activates approximately 3 billion, but its 23 GB file in Ollama does not fit in 12 GB. The solution planned by llama.cpp is the --n-cpu-moe option: it keeps the expert weights from the first N layers in RAM and leaves attention on the GPU.

A user published their measurements on September 22, 2026: RTX 3060 12 GB, Ryzen 5 3600, 32 GB of DDR4-3200, Qwen3.6-35B-A3B in Q4_K_M (about 21 GB). They measured 47 tokens/s on a short question, 45 with 12,000 tokens of context, and 37 with 45,000. With a conventional layer split, the same machine dropped to 7.1 tokens/s on 15,000 tokens (measured with Qwen3-30B-A3B), versus 25.3 after tuning. Only one in four layers of the Qwen3.5 and 3.6 series uses standard attention, so their speed drops less as the context grows.

Three caveats: this is one person’s measurement, using llama.cpp rather than Ollama; it assumes 32 GB of RAM; and the RAM speed matters. Its DDR4-3200 was running at 2,400 by default, and the XMP or DOCP profile added 10 to 17% throughput.

llama.cpp: experts on the CPU, attention on the GPU
llama-server -m modele-moe-Q4_K_M.gguf -ngl 99 --n-cpu-moe 26 -c 32768 -ctk q8_0 -ctv q8_0

The value 26 is the author's value for their model: the higher it goes, the more VRAM is freed, but the slower generation becomes. Look for the lowest value that leaves about 1 GB of VRAM free.

#5. Context and KV cache

#Context, the first limit of 12 GB

Ollama allocates 4,096 context tokens by default under 24 GiB of VRAM, while its documentation recommends at least 64,000 tokens for agents and coding tools. Between the two, each token costs memory. Example with Qwen3 14B, whose architecture is published: 40 layers, 8 KV heads, 128 dimensions per head. Per token, the KV cache in 16-bit occupies 2 (keys and values) × 40 × 8 × 128 × 2 bytes, or 160 KiB: about 2.4 GiB for 16,000 tokens and 3.7 GiB for 24,000.

The 9.3 GB model represents 8.7 GiB, leaving 12 GiB of VRAM. At 16,000 tokens, the total reaches 11.1 GiB: less than 1 GiB remains for runtime buffers, which is too tight. At 24,000 tokens, it reaches 12.3 GiB: loading fails or spills over to the CPU. With an 8-bit cache, 24,000 tokens use only 1.8 GiB, for 10.5 GiB total: it fits.

#8-bit KV cache and Flash Attention

Ollama automatically enables Flash Attention when the software and card support it; OLLAMA_FLASH_ATTENTION=1 forces it. This is a prerequisite for quantized caching: with OLLAMA_KV_CACHE_TYPE=q8_0, the cache uses about half the memory of f16 with very little precision loss, according to the documentation. The setting is global: all served models use the same cache type.

Linux: Ollama server with a 16,000-token context and 8-bit KV cache
OLLAMA_FLASH_ATTENTION=1 OLLAMA_KV_CACHE_TYPE=q8_0 OLLAMA_CONTEXT_LENGTH=16000 ollama serve

#6. Two RTX 3060 12 GB: 24 GB of VRAM, not twice as fast

Two cards offer 24 GB of VRAM, enough for a dense 27 to 32B model in Q4: Qwen3.5 27B weighs 17 GB in Ollama. Ollama automatically distributes the model across all cards when it does not fit on one. In llama.cpp, the default split is by layer, in a pipeline: the layers run one after another. The tensor mode, which parallelizes computation, is marked experimental.

The gain is therefore capacity, not speed: each token rereads all the weights at 360 GB/s per card. The estimated ceiling for a 32B model weighing 20 GB is 360 divided by 20, or about 18 tokens/s. A RTX 3090 offers 24 GB on a single card at 936 GB/s according to RunAIHome: the same model then tops out at around 47 tokens/s, or 2.6 times more. Two 170 W cards also mean 340 W for the GPU alone, whereas the 550 W of NVIDIA applies to a single card.

This setup makes sense if you already own a card and want to reach 24 GB. If you are starting from scratch, compare it first with a single 16 or 24 GB card.

#7. RTX 3060 8 GB, 3060 Ti, 4060, 5060: what 12 GB changes

Entry-level NVIDIA cards: memory (NVIDIA specifications) and generation on Llama 3.1 8B Q4_K_M (LocalScore)
CardVRAMMemory and bus8B Q4_K_M, tokens/sGemma 4 12B (7.6 GB)
RTX 3060 12 GB12 GBGDDR6, 192-bit51,3Yes, with about 4 GB of context
RTX 3060 8 GB8 GBGDDR6, 128-bit37,0No in practice (0.4 GB of headroom)
RTX 3060 Ti8 GBGDDR6 or GDDR6X, 256 bits60,2Not in practice
RTX 40608 GBGDDR6, 128-bit38,1Not in practice
RTX 50608 GBGDDR7, 128-bit—Not in practice
RTX 5060 Ti16 GB or 8 GBGDDR7, 128-bit51.2 (16 GB)Yes, in the 16 GB version

The 3060 8 GB is not simply a 3060 12 GB with less memory: its bus drops from 192 to 128 bits, and Tom's Hardware puts its bandwidth at 240 GB/s versus 360 GB/s, or 33% less. LocalScore records it at 37.0 tokens/s versus 51.3, or 28% less on the same model. These speeds are isolated results, without controlled conditions: the 3060 Ti comes out ahead there, which its 256-bit bus makes plausible, but its 8 GB excludes it from 12- or 14-billion-parameter models. The 5060 offsets its narrow bus with GDDR7, without changing its capacity. Only the 5060 Ti 16 GB exceeds the 3060 in VRAM.

!
Used listings: “RTX 3060” is not enough
Two versions of the 3060 exist, 12 GB and 8 GB, and the 3060 Ti has only 8 GB. Require the 12 GB designation and a screenshot of nvidia-smi: the command below displays the card's name and total memory.
Check the card's capacity
nvidia-smi --query-gpu=name,memory.total --format=csv

#Buying a 3060 12 GB: when it makes sense and when to move on

Buy a 3060 12 GB if
you want a first GPU for Ollama, 7B to 14B models, chat, document RAG, or coding assistance, and capacity matters more than speed.
Aim for 16 GB if
you want Gemma 4 12B with a long context, or more headroom for a 14B. The RTX 5060 Ti 16 GB is the first step; the comparison page quantifies the difference.
Aim for 24 GB if
Dense models from 27B to 32B are your target: a 24 GB card avoids splitting the model across two cards and provides much higher bandwidth.
Avoid or verify if
The announcement does not specify the memory: the 8 GB version offers neither the capacity nor the bandwidth of the 12 GB version.

#Frequently asked questions

FAQ
Is RTX 3060 12 GB enough for a local LLM?+
Yes, for 7- to 14-billion-parameter models, which cover chat, document summarization, and coding assistance. Its 12 GB can load Gemma 4 12B (7.6 GB) or Qwen3 14B (9.3 GB) in Q4, at about 33 tokens per second for both models according to published measurements. It is not sufficient for a dense 27- to 32-billion-parameter model.
Which model should you run first with Ollama on a RTX 3060 12 GB?+
Start with ollama run gemma4:12b: 7,6 GB, advertised 256K context, image support, and approximately 4 GB of headroom on the card. To prioritize speed, ollama run qwen3.5:9b weighs 6,6 GB and reaches nearly 50 tokens/s according to a published measurement. Then check with ollama ps that the PROCESSOR column shows 100% GPU.
How many tokens per second on a 12 GB RTX 3060?+
Published measurements show 51 to 65 tokens/s on Llama 3.1 8B in Q4_K_M, about 50 on Qwen3.5 9B, and 26 to 33 on a 14B model. The theoretical ceiling is the 360 GB/s bandwidth divided by the model’s weight in GB: 47 tokens/s for Gemma 4 12B. These figures vary with the software, context, and processor.
Can you run a 32B or 70B model on a RTX 3060 with 12 GB?+
A dense 32B in Q4 (19 to 20 GB of weights) doesn't fit: it spills into RAM and speed collapses. A 70B (about 40 GB) is worse: about 29 GB read from RAM at best at 51 GB/s (dual-channel DDR4-3200), for a ceiling of about 1.7 tokens/s. MoEs such as Qwen3.6 35B-A3B are an exception, with 32 GB of RAM and llama.cpp.
Are two RTX 3060 12 GB cards as good as one RTX 3090 for local AI?+
In terms of capacity, yes: 24 GB of VRAM total, enough for a 32B model in Q4. In terms of speed, no: llama.cpp splits the model by default into layers executed sequentially, and each token rereads the weights at 360 GB/s. A 3090 offers 936 GB/s according to RunAIHome, for a 2.6-times-higher ceiling.
RTX 3060 12 GB or RTX 4060 8 GB for an LLM?+
The 3060 12 GB: 12 GB vs. 8 GB determines model size, and Gemma 4 12B (7.6 GB) leaves no headroom on 8 GB. The 4060 uses less power, with 115 W of board power vs. 170 W according to NVIDIA, but that isn't the deciding factor for an LLM.

Prices change quickly: every Monday and Thursday, our tracker records the lowest price for local AI graphics cards, along with the price per GB of VRAM.

Did this guide help you?

Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.

Prices in euros (€) are French market prices including VAT, as checked by BestLLMfor. US prices differ: the Amazon buttons show the current US price.