Which LLM on RTX 3070 / 3070 Ti (8 GB) ?
A RTX 3070 or 3070 Ti offers 8 GB of VRAM: it keeps models up to 8 to 9 billion parameters in Q4 in VRAM (Granite 4.2 8B at 5.3 GB, Qwen 3.5 9B at 6.6 GB), but not 12-billion-parameter models or larger. The 3070 Ti, at 608 GB/s versus 448, generates faster with the same capacity. No published measurements: the throughput figures are estimates.
With 8 GB, the RTX 3070 and the 3070 Ti remain decent cards for an 8-billion-parameter model, but capacity is their limit. This guide shows which models actually fit, with their exact weight sizes, bandwidth-based throughput estimates, context limits, and the gap versus nearby cards. You will also know when to move up to 12 or 16 GB.
Choosing a machine? Our picks by budget →
Buying alternative for this guide: RTX 5060 Ti 16GB (ASUS Prime).
Compare all options by budget, from €800 to €3,500 →
On the go: which laptop for local AI →
Affiliate links — commission possible at no extra cost to you. As an Amazon Associate, BestLLMfor earns from qualifying purchases.
#RTX 3070 and 3070 Ti for a local LLM: what 8 GB enables
For a local LLM, a RTX 3070 or a 3070 Ti is valuable for its 8 GB of VRAM, and that capacity determines everything: models up to 8 to 9 billion parameters in Q4 fit, while models with 12 billion or more do not. According to the Ollama library, Granite 4.2 8B weighs 5.3 GB, Qwen3 8B weighs 5.2 GB, and Qwen 3.5 9B weighs 6.6 GB; Qwen3 14B (9.3 GB) and gpt-oss 20B (14 GB) exceed the card's capacity. The 3070 Ti, at 608 GB/s versus 448 for the 3070, generates faster but has the same capacity.
- Memory
- 8 GB on a 256-bit bus: GDDR6 at 448 GB/s for the 3070, GDDR6X at 608 GB/s for the 3070 Ti, according to the Wikipedia and NVIDIA table.
- Compute
- 5,888 CUDA cores for the 3070 and 6,144 for the 3070 Ti, according to Wikipedia.
- Power consumption
- 220 W of GPU power and a 650 W power supply required for the 3070; 290 W and 750 W for the 3070 Ti, according to NVIDIA. That’s more than the 550 W sometimes quoted.
#Models that fit in 8 GB
The weights are from the files in the Ollama library, accessed on September 29, 2026. The margin is what remains on 8 GB before the context, cache, and engine buffers; it must also cover display output if the card drives your monitor.
| Model | File Ollama | Gross margin | Verdict |
|---|---|---|---|
| Qwen 3.5 4B | 3.4 GB | 4.6 GB | Very comfortable; long context possible |
| Qwen3 8B | 5.2 GB | 2.8 GB | Comfortable |
| Granite 4.2 8B | 5.3 GB | 2.7 GB | Comfortable |
| Qwen 3.5 9B | 6.6 GB | 1.4 GB | Fits, context must be limited |
| Gemma 4 12B | 7.6 GB | 0.4 GB | No headroom: no useful context |
| Qwen3 14B | 9.3 GB | Short by 1.3 GB | Doesn't fit |
| gpt-oss 20B | 14 GB | 6 GB short | Doesn't fit |
The safest starting point is an 8-billion-parameter model: Granite 4.2 8B or Qwen3 8B leave nearly 3 GB for context. Qwen 3.5 9B is the practical upper limit: 1.4 GB of headroom is enough for a short conversation, but not a long document. A local RAG system adds an embedding model: bge-m3 weighs 1.2 GB in Ollama, for 6.5 GB total with Granite 4.2 8B, leaving 1.5 GB for everything else.
#Install Ollama on a RTX 3070
- 01Check the driverRun nvidia-smi in a terminal: the driver must be version 550 or later (551.61 on Windows), and total memory should show approximately 8 GB.
- 02Install OllamaInstall Ollama from its official website. The local API listens on http://localhost:11434.
- 03Run a modelRun ollama run granite4.2:8b: the 5.3 GB download happens once.
- 04Control the distributionIn a second terminal, run ollama ps: the Processor column should show 100% GPU. A CPU/GPU split means that the model or context spills into system RAM.
- 05Adjust the contextSet the context with OLLAMA_CONTEXT_LENGTH, then run ollama ps again. Set OLLAMA_FLASH_ATTENTION to 1 and OLLAMA_KV_CACHE_TYPE to q8_0 when starting the server to save VRAM.
#Throughput: what bandwidth can be expected to deliver
No authoritative benchmark site publishes a measurement for RTX 3070, so the figures below are estimates, not measurements. The method relies on the fact that generation is bandwidth-bound: the theoretical ceiling is approximately bandwidth divided by model size. For Granite 4.2 8B (5.3 GB), 448 ÷ 5.3 gives 85 tokens per second on the 3070, and 608 ÷ 5.3 gives 115 on the 3070 Ti. For Qwen 3.5 9B (6.6 GB), the ceilings are 68 and 92.
The llama.cpp community leaderboard lets us convert these limits into ballpark figures. On Llama 2 7B Q4_0, a RTX 3060 Ti 8 GB card with a 256-bit bus, comparable to the 3070, generates 92.23 tokens per second without Flash Attention and 96.50 with it. A RTX 3060 12 GB card at 360 GB/s generates 75.57 and 76.92: the ratio (1.22) tracks the bandwidth ratio (448 ÷ 360 = 1.24). A 3070 should therefore deliver rates close to those of the 3060 Ti, around 90 to 95 tokens per second on a 7B, and a 3070 Ti about 35% more, or 120 to 130. These are estimates.
| Model | File Ollama | 3070 ceiling (448 GB/s) | 3070 Ti limit (608 GB/s) |
|---|---|---|---|
| Qwen 3.5 4B | 3.4 GB | 132 | 179 |
| Granite 4.2 8B | 5.3 GB | 85 | 115 |
| Qwen 3.5 9B | 6.6 GB | 68 | 92 |
| Gemma 4 12B | 7.6 GB | 59 | 80 |
Real-world throughput is around 70 to 80% of those ceilings on the cards measured elsewhere: plan on 55 to 65 tokens per second for Granite 4.2 8B on a 3070. Whatever the exact figure, these speeds exceed human reading speed: the 3070's limitation is capacity, not speed.
#The limits of 8 GB: context, RAG, and large models
Ollama sets the default context based on video memory: its documentation specifies 4k tokens with 24 GiB of VRAM and recommends at least 64,000 tokens for agents, web research, and coding tools. On 8 GB, targeting 64,000 tokens with an 8-billion-parameter model is not realistic: the KV cache, which stores the attention keys and values for each token, consumes the available headroom. According to Ollama’s FAQ, q8_0 cache quantization reduces its memory use to approximately half that of f16, with a very small loss of precision: this is the first setting to enable.
- 12- to 24-billion-parameter models
- Gemma 4 12B (7.6 GB) leaves no room for context; Qwen3 14B (9.3 GB) and gpt-oss 20B (14 GB) do not fit. They spill into system RAM, and speed drops.
- Long context
- With Qwen 3.5 9B, 1.4 GB of headroom runs out quickly: keep the context to a few thousand tokens and monitor ollama ps.
- Multi-step RAG
- Granite 4.2 8B and bge-m3 use 6.5 GB: adding a reranker or a second model exceeds the limit.
- Fine-tuning
- According to Unsloth, the absolute minimum for 4-bit QLoRA is 3.5 GB for 3 billion parameters and 6 GB for 8 billion: an 8B model barely fits, with a batch size of 1 and a short context.
A simple way to stay under 8 GB takes three settings. First, choose a model whose file leaves at least 2 GB of headroom: Granite 4.2 8B or Qwen3 8B. Next, enable the q8_0 quantized cache and Flash Attention, which the official Flash Attention repository lists as compatible with Ampere GPUs such as the 3070. Finally, increase the context in stages, from 4,000 to 8,000 and then 16,000 tokens, restarting ollama ps at each step: as soon as any CPU usage appears, return to the previous level. This progression keeps you from discovering the limit in the middle of a conversation.
#3070 compared with other 8 to 16 GB cards
| Card | Memory | Bandwidth | What it changes |
|---|---|---|---|
| RTX 3070 | 8 GB GDDR6 | 448 GB/s | 8- to 9-billion-parameter models |
| RTX 3070 Ti | 8 GB GDDR6X | 608 GB/s | Same capacity, faster generation |
| RTX 3060 12 GB | 12 GB GDDR6 | 360 GB/s | Add 12- to 14-billion-parameter models |
| RTX 4060 Ti 16 GB | 16 GB | 288 GB/s | Adds gpt-oss 20B, but is slower |
This table highlights the trade-off. The 3060 12 GB has 20% lower bandwidth than the 3070, but 4 GB more: it can handle Qwen3 14B (9.3 GB), which the 3070 rejects. The 4060 Ti 16 GB adds the 20-billion-parameter class, with 288 GB/s of bandwidth that makes it slower on small models. For an LLM, capacity takes priority over speed as soon as the target model exceeds 7.5 GB.
#Verdict: keep, replace, or don't buy
| Situation | Decision | Quantified rationale |
|---|---|---|
| You already have a 3070 and are targeting an 8B | Keep it | Granite 4.2 8B: 5.3 GB, estimated at around 60 t/s |
| You want a 12B to 14B model | Upgrade to 12 or 16 GB | Qwen3 14B weighs 9.3 GB, while the 3070 has 8 GB |
| You want a 20B model or a long context | 16 GB or more card | gpt-oss 20B weighs 14 GB |
| You’re choosing between the 3070 and 3070 Ti | Choose the cheapest one | Same capacity, 608 versus 448 GB/s |
| You're looking for a card to get started | Compare with a 3060 12 GB | More memory for similar bandwidth |
The 3070 is an adequate card for an 8-billion-parameter model, but nothing more. It remains useful as an entry-level card: Ollama supports it, and an 8-billion-parameter model responds faster than you can read. It is not suitable if you plan to use coding agents or long documents, which require context that 8 GB cannot provide. If you already have one, it is enough to explore local LLMs. If you are choosing a card for this sole purpose, start with capacity: a larger model that fits is better than a smaller, faster model. For current prices, check our tracker, which records the lowest price for each card twice a week.
- Source: NVIDIA, specifications for RTX 3070 and 3070 Ti
- Source: llama.cpp community ranking on CUDA
- Source: Ollama documentation, context length
- Source: Ollama FAQ, Flash Attention, and KV cache
- Source: Unsloth, VRAM required for fine-tuning
#Frequently asked questions
Which LLM should you run on a RTX 3070?+
Can the RTX 3070 run a 14- or 24-billion-parameter model?+
RTX 3070 or RTX 3060 12 GB for an LLM?+
What’s the difference between the 3070 and 3070 Ti for an LLM?+
What power supply do you need for a RTX 3070 Ti?+
Can you fine-tune a model on a RTX 3070?+
Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.