Beginner 11 minRTX 40

Which LLM on RTX 4060 (8 GB) ?

Direct response

A RTX 4060 (8 GB of GDDR6, 272 GB/s) runs 3- to 9-billion-parameter models in Q4 without difficulty: Qwen 3.5 4B (3.4 GB), Granite 4.2 8B (5.3 GB), or Qwen 3.5 9B (6.6 GB). It stalls as soon as a model and its context exceed 8 GB. Its bandwidth imposes a theoretical ceiling of about 50 tokens/s on an 8B in Q4. Beyond that, 12 or 16 GB changes how you use it, not the GPU's speed.

The RTX 4060 is the entry-level card in the Ada generation: 8 GB of VRAM, 115 W, and a low used price. It is enough to explore local LLMs, as long as you know its exact ceiling. This guide quantifies what fits in its 8 GB, what its bandwidth allows, how to configure Ollama to avoid overflowing it, and when it is better to move to a 12 or 16 GB card.

Choosing a machine? Our picks by budget →

By Mohamed Meguedmi·Update 2026-09-30·Tested on Windows, macOS, and Linux
Recommended hardware

Buying alternative for this guide: RTX 5060 Ti 16GB (ASUS Prime).

Compare all options by budget, from €800 to €3,500 →

On the go: which laptop for local AI →

Affiliate links — possible commission at no extra cost to you. As an Amazon Associate, BestLLMfor earns from qualifying purchases.

#The RTX 4060 for a local LLM: what it enables

On a RTX 4060, a local LLM works well as long as the model, its context, and the system fit within 8 GB of VRAM. In Q4, this covers models with 3 to 9 billion parameters: Qwen 3.5 4B weighs 3.4 GB, Granite 4.2 8B 5.3 GB, and Qwen 3.5 9B 6.6 GB according to the Ollama library. A 12B such as Gemma 4 12B (7.7 to 8 GB) already fills the card before the context is even counted. The 272 GB/s bandwidth determines speed: it sets a theoretical ceiling of about 51 tokens/s on a 5.3 GB model, and real-world performance remains below that. For a chat assistant, summarization, or a small RAG, it is sufficient. For code on a large repository or long-form reasoning, it falls short.

Architecture
Ada Lovelace, 3,072 CUDA cores, 4th-generation Tensor Cores, CUDA compute capability 8.9.
Memory
8 GB of GDDR6 on a 128-bit bus, for 272 GB/s of bandwidth (source: Wikipedia, GeForce 40-series table).
Power consumption
115 W of total graphics power; NVIDIA specifies a minimum 550 W power supply for the entire PC.
Interface
PCIe 4.0 limited to 8 lanes: a model that spills into RAM has to traverse a narrow link, which penalizes offloading even further.
Availability
Launched in 2023, the card is now found mostly on the used market; prices are tracked on the dedicated page.
!
Why bandwidth matters more than cores
During generation, each token forces the GPU to reread most of the model's weights. Speed therefore depends much more on the ratio between memory bandwidth and weight size than on the number of cores. In this respect, the 4060 is the most limited card in the consumer Ada lineup: 272 GB/s, compared with 672 GB/s for a 4070 Ti Super.

#Which models fit in 8 GB

The site's rule of thumb for Q4 is simple: about 0.6 GB per billion parameters, for weights alone. You must add the context cache (KV cache) and some headroom for the display and system. Ollama uses a context of 4,096 tokens by default, which remains light; at 16,000 tokens, the cache becomes a significant allocation. The table below uses the sizes published by Ollama and shows the remaining headroom on 8 GB, before context.

Common models on 8 GB of VRAM (sizes Ollama, weights only)
ModelSize OllamaHeadroom on 8 GBVerdict
Qwen 3.5 4B3.4 GB4.6 GBComfortable, long context possible
Granite 4.2 8B5.3 GB2.7 GBComfortable up to around 16,000 tokens with a quantized cache
Qwen 3.5 9B6.6 GB1.4 GBTight: short context, applications closed
Gemma 4 12B7.7 to 8.0 GB0 GB or lessDoesn't fit with a useful context
gpt-oss 20B14 GBNegativeOut of reach with VRAM alone
Mistral Small 24B14 GBNegativeOut of reach with VRAM alone

Two pitfalls come up often. First, a file size is not a memory requirement; the context is added on top. Second, a model that exceeds the limit doesn’t crash—it loads partly into system RAM and slows down significantly. The ollama ps command shows the split: a line at 100% GPU is healthy, while a shared CPU/GPU line indicates overflow.

#Install Ollama and run your first model

  1. 01
    Check the NVIDIA driver
    Ollama requires NVIDIA driver 550 or later for cards with compute capability 5.0 and above. A RTX 4060 (capability 8.9) is explicitly listed in Ollama documentation. Check the version with nvidia-smi.
  2. 02
    Install Ollama
    Follow the installation guide for your system. CUDA is included, so no separate toolkit installation is required.
  3. 03
    Run an adapted model
    Start with ollama run granite4.2:8b (5.3 GB) or ollama run qwen3.5:9b (6.6 GB). Then run ollama ps in a second terminal.
  4. 04
    Control placement
    The PROCESSOR column must show 100% GPU. Otherwise, reduce the context or switch to a smaller model.

#What speed to expect: the ceiling, not the promise

There is no in-house measurement here, and no throughput figure is presented as such. We can, however, calculate an upper bound: bandwidth divided by the size of the weights read for each token. A third-party report published on GitHub (LLaMA 3 8B in Q4_K_M under llama.cpp, May 2024 snapshot) gives observed-to-ceiling ratios of 68% on a RTX 4080 and 75% on a RTX 4070 Ti. A reasonable rule of thumb is 70% of the ceiling. These values remain estimates, valid for a short context.

Theoretical generation ceiling on RTX 4060 (272 GB/s)
Model (Q4)WeightsTheoretical ceilingEstimated at 70%
Qwen 3.5 4B3.4 GB80 tokens/sabout 56 tokens/s
Granite 4.2 8B5.3 GB51 tokens/sabout 36 tokens/s
Qwen 3.5 9B6.6 GB41 tokens/sapproximately 29 tokens/s

These estimates are consistent with using a chat assistant: reading is comfortable starting at 15 to 20 tokens/s. They do not account for context, which slows generation as it grows, or prompt-processing time, which affects compute rather than memory.

#Set Ollama so it stays within 8 GB

The settings that matter are the context and cache settings. Ollama quantizes the K/V cache to q8_0 to use about half the memory of f16, with very little loss of precision according to its documentation; this requires Flash Attention. On 8 GB, it is the most cost-effective lever: it lets you extend the context without changing models.

Ollama server settings (Linux, macOS)
export OLLAMA_FLASH_ATTENTION=1
export OLLAMA_KV_CACHE_TYPE=q8_0
export OLLAMA_CONTEXT_LENGTH=8192
ollama serve
Reasonable context
Increase the context to 8,192 or 16,384 tokens only if usage requires it: each context doubling doubles the K/V cache.
A single model loaded
On 8 GB, keep only one model in memory; an embeddings model for RAG should remain small.
Free up VRAM
The desktop, browser, and games reserve VRAM. Close them before loading and check with nvidia-smi.
Quantization
Stay with Q4_K_M: an 8B model’s Q8 (about 8.5 GB) exceeds the card’s capacity.

#What the 4060 does well, poorly, or not at all

The verdict depends more on the workload than on the GPU. A short chat or a summary of a few pages stays in the comfort zone with either a 4B or an 8B model. RAG requires a generation model, an embedding model, and a context long enough to hold the retrieved excerpts: on 8 GB, choose a 4B to 8B generation model and keep the embedding model small. A coding assistant is more demanding, because specialized models and their contexts of several thousand tokens quickly exceed the GPU's capacity.

Common uses on RTX 4060 (8 GB)
UsageRealistic?Recommended setting
Daily chat, short questionsYesGranite 4.2 8B or Qwen 3.5 4B, default context
Summarizing documents of 10 to 20 pagesYes, with an 8,192 to 16,384 contextK/V cache in q8_0, Flash Attention enabled
RAG over your filesYes, 4B to 8B modelLightweight embeddings, a single generation model loaded
Code assistant for a repositoryLimited4B to 8B model, short excerpts, no entire repository
20B model or larger, including gpt-oss 20BNoPlan for at least 16 GB of VRAM

The gpt-oss 20B case illustrates the boundary: Ollama indicates 14 GB for this model, and its specifications state that MXFP4 quantization allows it to run with as little as 16 GB of memory. An 8 GB card is not in the target range, regardless of its throughput.

#When to switch to another card

Three signals indicate that 8 GB is becoming a bottleneck. You want a 12B model or larger, including Gemma 4 12B. You work with long documents where the context exceeds 16 000 tokens. You use a coding assistant that loads a model of 14 to 24 GB. In these cases, adding speed solves nothing: memory capacity is what’s missing, and the page dedicated to 8 GB, 12 GB, and 16 GB helps you target the right tier.

#4060, 3060 12 GB, 5060: the right compromise

Entry-level cards: capacity and bandwidth
CardVRAMBandwidthWhat this changes
RTX 3060 12 GB12 GB360 GB/sMore VRAM: Gemma 4 12B fits, but with an older bus
RTX 40608 GB272 GB/sLow power consumption, 8B–9B ceiling
RTX 50608 GB448 GB/sSame capacity, 65% higher speed ceiling

The RTX 3060 12 GB provides 4 GB of additional memory for higher bandwidth (360 GB/s according to Wikipedia). It generally costs less used: see /prix-gpu-ia for current market conditions. The RTX 5060 still has 8 GB but reads memory 65% faster. If your workload fits within 8 GB, the 5060 is faster; if it exceeds 8 GB, the 3060 12 GB or 4060 Ti 16 GB can run models that the 4060 cannot load.

#2026 verdict

Keep your 4060 if
You use 3B to 9B models, a chat, or lightweight RAG. There is no reason to replace it as long as the memory is sufficient.
Don’t buy a 4060 for an LLM
If the LLM is the primary use case, a card with 12 GB or more is more useful, even if that means buying used.
Move up to 16 GB
If you are targeting 14B to 24B models, for which 8 GB will never be enough.

#Frequently asked questions

FAQ
Can RTX 4060 run an LLM locally?+
Yes, up to about 9 billion parameters in Q4. Granite 4.2 8B (5.3 GB) and Qwen 3.5 9B (6.6 GB) fit in VRAM, with a short context for the latter. Beyond 8 GB, Ollama spills into system RAM and generation slows sharply, making 12B and larger models impractical.
How many tokens per second on a RTX 4060?+
No reliable measurement is published here. The theoretical ceiling—bandwidth divided by weights—is about 51 tokens/s for a 5.3 GB model, or nearly 36 tokens/s at 70% of that ceiling. This is an estimate, not a measured result, and it decreases as the context gets longer.
RTX 4060 or RTX 3060 12 GB for an LLM?+
For LLMs, the 3060 12 GB is generally more attractive: 4 GB more memory and 360 GB/s of bandwidth versus 272 GB/s. It can load Gemma 4 12B, which the 4060 cannot fit. The 4060 retains the advantage in power consumption, at 115 W, and gaming.
What power supply do you need for a RTX 4060?+
NVIDIA indicates 115 W of total graphics power and a recommended minimum 550 W power supply for the entire PC. This figure includes headroom for the processor and power spikes. A quality power supply with this rating is sufficient, since the LLM puts less load on the card than some games do.
How can you tell whether your model fits in VRAM?+
Launch the model, then run ollama ps in a second terminal: the PROCESSOR column shows 100% GPU if everything is in VRAM, or a CPU/GPU split in case of overflow. Also monitor nvidia-smi during a long generation. If overflow occurs, reduce the context or choose a smaller model.
Can you improve 8 GB without changing the card?+
Three levers are involved: enable Flash Attention, then the K/V cache in q8_0, which cuts that cache by about half; limit the context to what you use; and close applications occupying VRAM. They let you extend the context, not load a larger model.
Did this guide help you?

Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.