Which LLM on RTX 4060 (8 GB) ?
A RTX 4060 (8 GB of GDDR6, 272 GB/s) runs 3- to 9-billion-parameter models in Q4 without difficulty: Qwen 3.5 4B (3.4 GB), Granite 4.2 8B (5.3 GB), or Qwen 3.5 9B (6.6 GB). It stalls as soon as a model and its context exceed 8 GB. Its bandwidth imposes a theoretical ceiling of about 50 tokens/s on an 8B in Q4. Beyond that, 12 or 16 GB changes how you use it, not the GPU's speed.
The RTX 4060 is the entry-level card in the Ada generation: 8 GB of VRAM, 115 W, and a low used price. It is enough to explore local LLMs, as long as you know its exact ceiling. This guide quantifies what fits in its 8 GB, what its bandwidth allows, how to configure Ollama to avoid overflowing it, and when it is better to move to a 12 or 16 GB card.
Choosing a machine? Our picks by budget →
Buying alternative for this guide: RTX 5060 Ti 16GB (ASUS Prime).
Compare all options by budget, from €800 to €3,500 →
On the go: which laptop for local AI →
Affiliate links — possible commission at no extra cost to you. As an Amazon Associate, BestLLMfor earns from qualifying purchases.
#The RTX 4060 for a local LLM: what it enables
On a RTX 4060, a local LLM works well as long as the model, its context, and the system fit within 8 GB of VRAM. In Q4, this covers models with 3 to 9 billion parameters: Qwen 3.5 4B weighs 3.4 GB, Granite 4.2 8B 5.3 GB, and Qwen 3.5 9B 6.6 GB according to the Ollama library. A 12B such as Gemma 4 12B (7.7 to 8 GB) already fills the card before the context is even counted. The 272 GB/s bandwidth determines speed: it sets a theoretical ceiling of about 51 tokens/s on a 5.3 GB model, and real-world performance remains below that. For a chat assistant, summarization, or a small RAG, it is sufficient. For code on a large repository or long-form reasoning, it falls short.
- Architecture
- Ada Lovelace, 3,072 CUDA cores, 4th-generation Tensor Cores, CUDA compute capability 8.9.
- Memory
- 8 GB of GDDR6 on a 128-bit bus, for 272 GB/s of bandwidth (source: Wikipedia, GeForce 40-series table).
- Power consumption
- 115 W of total graphics power; NVIDIA specifies a minimum 550 W power supply for the entire PC.
- Interface
- PCIe 4.0 limited to 8 lanes: a model that spills into RAM has to traverse a narrow link, which penalizes offloading even further.
- Availability
- Launched in 2023, the card is now found mostly on the used market; prices are tracked on the dedicated page.
#Which models fit in 8 GB
The site's rule of thumb for Q4 is simple: about 0.6 GB per billion parameters, for weights alone. You must add the context cache (KV cache) and some headroom for the display and system. Ollama uses a context of 4,096 tokens by default, which remains light; at 16,000 tokens, the cache becomes a significant allocation. The table below uses the sizes published by Ollama and shows the remaining headroom on 8 GB, before context.
| Model | Size Ollama | Headroom on 8 GB | Verdict |
|---|---|---|---|
| Qwen 3.5 4B | 3.4 GB | 4.6 GB | Comfortable, long context possible |
| Granite 4.2 8B | 5.3 GB | 2.7 GB | Comfortable up to around 16,000 tokens with a quantized cache |
| Qwen 3.5 9B | 6.6 GB | 1.4 GB | Tight: short context, applications closed |
| Gemma 4 12B | 7.7 to 8.0 GB | 0 GB or less | Doesn't fit with a useful context |
| gpt-oss 20B | 14 GB | Negative | Out of reach with VRAM alone |
| Mistral Small 24B | 14 GB | Negative | Out of reach with VRAM alone |
Two pitfalls come up often. First, a file size is not a memory requirement; the context is added on top. Second, a model that exceeds the limit doesn’t crash—it loads partly into system RAM and slows down significantly. The ollama ps command shows the split: a line at 100% GPU is healthy, while a shared CPU/GPU line indicates overflow.
#Install Ollama and run your first model
- 01Check the NVIDIA driverOllama requires NVIDIA driver 550 or later for cards with compute capability 5.0 and above. A RTX 4060 (capability 8.9) is explicitly listed in Ollama documentation. Check the version with nvidia-smi.
- 02Install OllamaFollow the installation guide for your system. CUDA is included, so no separate toolkit installation is required.
- 03Run an adapted modelStart with ollama run granite4.2:8b (5.3 GB) or ollama run qwen3.5:9b (6.6 GB). Then run ollama ps in a second terminal.
- 04Control placementThe PROCESSOR column must show 100% GPU. Otherwise, reduce the context or switch to a smaller model.
#What speed to expect: the ceiling, not the promise
There is no in-house measurement here, and no throughput figure is presented as such. We can, however, calculate an upper bound: bandwidth divided by the size of the weights read for each token. A third-party report published on GitHub (LLaMA 3 8B in Q4_K_M under llama.cpp, May 2024 snapshot) gives observed-to-ceiling ratios of 68% on a RTX 4080 and 75% on a RTX 4070 Ti. A reasonable rule of thumb is 70% of the ceiling. These values remain estimates, valid for a short context.
| Model (Q4) | Weights | Theoretical ceiling | Estimated at 70% |
|---|---|---|---|
| Qwen 3.5 4B | 3.4 GB | 80 tokens/s | about 56 tokens/s |
| Granite 4.2 8B | 5.3 GB | 51 tokens/s | about 36 tokens/s |
| Qwen 3.5 9B | 6.6 GB | 41 tokens/s | approximately 29 tokens/s |
These estimates are consistent with using a chat assistant: reading is comfortable starting at 15 to 20 tokens/s. They do not account for context, which slows generation as it grows, or prompt-processing time, which affects compute rather than memory.
#Set Ollama so it stays within 8 GB
The settings that matter are the context and cache settings. Ollama quantizes the K/V cache to q8_0 to use about half the memory of f16, with very little loss of precision according to its documentation; this requires Flash Attention. On 8 GB, it is the most cost-effective lever: it lets you extend the context without changing models.
- Reasonable context
- Increase the context to 8,192 or 16,384 tokens only if usage requires it: each context doubling doubles the K/V cache.
- A single model loaded
- On 8 GB, keep only one model in memory; an embeddings model for RAG should remain small.
- Free up VRAM
- The desktop, browser, and games reserve VRAM. Close them before loading and check with nvidia-smi.
- Quantization
- Stay with Q4_K_M: an 8B model’s Q8 (about 8.5 GB) exceeds the card’s capacity.
#What the 4060 does well, poorly, or not at all
The verdict depends more on the workload than on the GPU. A short chat or a summary of a few pages stays in the comfort zone with either a 4B or an 8B model. RAG requires a generation model, an embedding model, and a context long enough to hold the retrieved excerpts: on 8 GB, choose a 4B to 8B generation model and keep the embedding model small. A coding assistant is more demanding, because specialized models and their contexts of several thousand tokens quickly exceed the GPU's capacity.
| Usage | Realistic? | Recommended setting |
|---|---|---|
| Daily chat, short questions | Yes | Granite 4.2 8B or Qwen 3.5 4B, default context |
| Summarizing documents of 10 to 20 pages | Yes, with an 8,192 to 16,384 context | K/V cache in q8_0, Flash Attention enabled |
| RAG over your files | Yes, 4B to 8B model | Lightweight embeddings, a single generation model loaded |
| Code assistant for a repository | Limited | 4B to 8B model, short excerpts, no entire repository |
| 20B model or larger, including gpt-oss 20B | No | Plan for at least 16 GB of VRAM |
The gpt-oss 20B case illustrates the boundary: Ollama indicates 14 GB for this model, and its specifications state that MXFP4 quantization allows it to run with as little as 16 GB of memory. An 8 GB card is not in the target range, regardless of its throughput.
#When to switch to another card
Three signals indicate that 8 GB is becoming a bottleneck. You want a 12B model or larger, including Gemma 4 12B. You work with long documents where the context exceeds 16 000 tokens. You use a coding assistant that loads a model of 14 to 24 GB. In these cases, adding speed solves nothing: memory capacity is what’s missing, and the page dedicated to 8 GB, 12 GB, and 16 GB helps you target the right tier.
#4060, 3060 12 GB, 5060: the right compromise
| Card | VRAM | Bandwidth | What this changes |
|---|---|---|---|
| RTX 3060 12 GB | 12 GB | 360 GB/s | More VRAM: Gemma 4 12B fits, but with an older bus |
| RTX 4060 | 8 GB | 272 GB/s | Low power consumption, 8B–9B ceiling |
| RTX 5060 | 8 GB | 448 GB/s | Same capacity, 65% higher speed ceiling |
The RTX 3060 12 GB provides 4 GB of additional memory for higher bandwidth (360 GB/s according to Wikipedia). It generally costs less used: see /prix-gpu-ia for current market conditions. The RTX 5060 still has 8 GB but reads memory 65% faster. If your workload fits within 8 GB, the 5060 is faster; if it exceeds 8 GB, the 3060 12 GB or 4060 Ti 16 GB can run models that the 4060 cannot load.
#2026 verdict
- Keep your 4060 if
- You use 3B to 9B models, a chat, or lightweight RAG. There is no reason to replace it as long as the memory is sufficient.
- Don’t buy a 4060 for an LLM
- If the LLM is the primary use case, a card with 12 GB or more is more useful, even if that means buying used.
- Move up to 16 GB
- If you are targeting 14B to 24B models, for which 8 GB will never be enough.
- Source: official NVIDIA specifications (RTX 4060 and 4060 Ti)
- Source: official Ollama FAQ (context, Flash Attention, K/V cache)
- Source: model sizes Qwen 3.5 in the Ollama library
#Frequently asked questions
Can RTX 4060 run an LLM locally?+
How many tokens per second on a RTX 4060?+
RTX 4060 or RTX 3060 12 GB for an LLM?+
What power supply do you need for a RTX 4060?+
How can you tell whether your model fits in VRAM?+
Can you improve 8 GB without changing the card?+
Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.