Beginner 11 minRTX 30

Which LLM on RTX 3070 / 3070 Ti (8 GB) ?

Direct response

A RTX 3070 or 3070 Ti offers 8 GB of VRAM: it keeps models up to 8 to 9 billion parameters in Q4 in VRAM (Granite 4.2 8B at 5.3 GB, Qwen 3.5 9B at 6.6 GB), but not 12-billion-parameter models or larger. The 3070 Ti, at 608 GB/s versus 448, generates faster with the same capacity. No published measurements: the throughput figures are estimates.

With 8 GB, the RTX 3070 and the 3070 Ti remain decent cards for an 8-billion-parameter model, but capacity is their limit. This guide shows which models actually fit, with their exact weight sizes, bandwidth-based throughput estimates, context limits, and the gap versus nearby cards. You will also know when to move up to 12 or 16 GB.

Choosing a machine? Our picks by budget →

By Mohamed Meguedmi·Update 2026-09-29·Tested on Windows, macOS, and Linux
Recommended hardware

Buying alternative for this guide: RTX 5060 Ti 16GB (ASUS Prime).

Compare all options by budget, from €800 to €3,500 →

On the go: which laptop for local AI →

Affiliate links — commission possible at no extra cost to you. As an Amazon Associate, BestLLMfor earns from qualifying purchases.

#RTX 3070 and 3070 Ti for a local LLM: what 8 GB enables

For a local LLM, a RTX 3070 or a 3070 Ti is valuable for its 8 GB of VRAM, and that capacity determines everything: models up to 8 to 9 billion parameters in Q4 fit, while models with 12 billion or more do not. According to the Ollama library, Granite 4.2 8B weighs 5.3 GB, Qwen3 8B weighs 5.2 GB, and Qwen 3.5 9B weighs 6.6 GB; Qwen3 14B (9.3 GB) and gpt-oss 20B (14 GB) exceed the card's capacity. The 3070 Ti, at 608 GB/s versus 448 for the 3070, generates faster but has the same capacity.

Memory
8 GB on a 256-bit bus: GDDR6 at 448 GB/s for the 3070, GDDR6X at 608 GB/s for the 3070 Ti, according to the Wikipedia and NVIDIA table.
Compute
5,888 CUDA cores for the 3070 and 6,144 for the 3070 Ti, according to Wikipedia.
Power consumption
220 W of GPU power and a 650 W power supply required for the 3070; 290 W and 750 W for the 3070 Ti, according to NVIDIA. That’s more than the 550 W sometimes quoted.

#Models that fit in 8 GB

The weights are from the files in the Ollama library, accessed on September 29, 2026. The margin is what remains on 8 GB before the context, cache, and engine buffers; it must also cover display output if the card drives your monitor.

Models for RTX 3070 / 3070 Ti (files Ollama, Q4 unless noted)
ModelFile OllamaGross marginVerdict
Qwen 3.5 4B3.4 GB4.6 GBVery comfortable; long context possible
Qwen3 8B5.2 GB2.8 GBComfortable
Granite 4.2 8B5.3 GB2.7 GBComfortable
Qwen 3.5 9B6.6 GB1.4 GBFits, context must be limited
Gemma 4 12B7.6 GB0.4 GBNo headroom: no useful context
Qwen3 14B9.3 GBShort by 1.3 GBDoesn't fit
gpt-oss 20B14 GB6 GB shortDoesn't fit

The safest starting point is an 8-billion-parameter model: Granite 4.2 8B or Qwen3 8B leave nearly 3 GB for context. Qwen 3.5 9B is the practical upper limit: 1.4 GB of headroom is enough for a short conversation, but not a long document. A local RAG system adds an embedding model: bge-m3 weighs 1.2 GB in Ollama, for 6.5 GB total with Granite 4.2 8B, leaving 1.5 GB for everything else.

#Install Ollama on a RTX 3070

  1. 01
    Check the driver
    Run nvidia-smi in a terminal: the driver must be version 550 or later (551.61 on Windows), and total memory should show approximately 8 GB.
  2. 02
    Install Ollama
    Install Ollama from its official website. The local API listens on http://localhost:11434.
  3. 03
    Run a model
    Run ollama run granite4.2:8b: the 5.3 GB download happens once.
  4. 04
    Control the distribution
    In a second terminal, run ollama ps: the Processor column should show 100% GPU. A CPU/GPU split means that the model or context spills into system RAM.
  5. 05
    Adjust the context
    Set the context with OLLAMA_CONTEXT_LENGTH, then run ollama ps again. Set OLLAMA_FLASH_ATTENTION to 1 and OLLAMA_KV_CACHE_TYPE to q8_0 when starting the server to save VRAM.
Basic commands
ollama run granite4.2:8b
ollama ps

# Exemple de la documentation Ollama : contexte de 64 000 tokens
OLLAMA_CONTEXT_LENGTH=64000 ollama serve

#Throughput: what bandwidth can be expected to deliver

No authoritative benchmark site publishes a measurement for RTX 3070, so the figures below are estimates, not measurements. The method relies on the fact that generation is bandwidth-bound: the theoretical ceiling is approximately bandwidth divided by model size. For Granite 4.2 8B (5.3 GB), 448 ÷ 5.3 gives 85 tokens per second on the 3070, and 608 ÷ 5.3 gives 115 on the 3070 Ti. For Qwen 3.5 9B (6.6 GB), the ceilings are 68 and 92.

The llama.cpp community leaderboard lets us convert these limits into ballpark figures. On Llama 2 7B Q4_0, a RTX 3060 Ti 8 GB card with a 256-bit bus, comparable to the 3070, generates 92.23 tokens per second without Flash Attention and 96.50 with it. A RTX 3060 12 GB card at 360 GB/s generates 75.57 and 76.92: the ratio (1.22) tracks the bandwidth ratio (448 ÷ 360 = 1.24). A 3070 should therefore deliver rates close to those of the 3060 Ti, around 90 to 95 tokens per second on a 7B, and a 3070 Ti about 35% more, or 120 to 130. These are estimates.

Generation estimate, in tokens per second (theoretical ceiling, RTX 3070 / 3070 Ti)
ModelFile Ollama3070 ceiling (448 GB/s)3070 Ti limit (608 GB/s)
Qwen 3.5 4B3.4 GB132179
Granite 4.2 8B5.3 GB85115
Qwen 3.5 9B6.6 GB6892
Gemma 4 12B7.6 GB5980

Real-world throughput is around 70 to 80% of those ceilings on the cards measured elsewhere: plan on 55 to 65 tokens per second for Granite 4.2 8B on a 3070. Whatever the exact figure, these speeds exceed human reading speed: the 3070's limitation is capacity, not speed.

#The limits of 8 GB: context, RAG, and large models

Ollama sets the default context based on video memory: its documentation specifies 4k tokens with 24 GiB of VRAM and recommends at least 64,000 tokens for agents, web research, and coding tools. On 8 GB, targeting 64,000 tokens with an 8-billion-parameter model is not realistic: the KV cache, which stores the attention keys and values for each token, consumes the available headroom. According to Ollama’s FAQ, q8_0 cache quantization reduces its memory use to approximately half that of f16, with a very small loss of precision: this is the first setting to enable.

12- to 24-billion-parameter models
Gemma 4 12B (7.6 GB) leaves no room for context; Qwen3 14B (9.3 GB) and gpt-oss 20B (14 GB) do not fit. They spill into system RAM, and speed drops.
Long context
With Qwen 3.5 9B, 1.4 GB of headroom runs out quickly: keep the context to a few thousand tokens and monitor ollama ps.
Multi-step RAG
Granite 4.2 8B and bge-m3 use 6.5 GB: adding a reranker or a second model exceeds the limit.
Fine-tuning
According to Unsloth, the absolute minimum for 4-bit QLoRA is 3.5 GB for 3 billion parameters and 6 GB for 8 billion: an 8B model barely fits, with a batch size of 1 and a short context.

A simple way to stay under 8 GB takes three settings. First, choose a model whose file leaves at least 2 GB of headroom: Granite 4.2 8B or Qwen3 8B. Next, enable the q8_0 quantized cache and Flash Attention, which the official Flash Attention repository lists as compatible with Ampere GPUs such as the 3070. Finally, increase the context in stages, from 4,000 to 8,000 and then 16,000 tokens, restarting ollama ps at each step: as soon as any CPU usage appears, return to the previous level. This progression keeps you from discovering the limit in the middle of a conversation.

#3070 compared with other 8 to 16 GB cards

Neighboring cards for a local LLM (NVIDIA, Hardware Corner)
CardMemoryBandwidthWhat it changes
RTX 30708 GB GDDR6448 GB/s8- to 9-billion-parameter models
RTX 3070 Ti8 GB GDDR6X608 GB/sSame capacity, faster generation
RTX 3060 12 GB12 GB GDDR6360 GB/sAdd 12- to 14-billion-parameter models
RTX 4060 Ti 16 GB16 GB288 GB/sAdds gpt-oss 20B, but is slower

This table highlights the trade-off. The 3060 12 GB has 20% lower bandwidth than the 3070, but 4 GB more: it can handle Qwen3 14B (9.3 GB), which the 3070 rejects. The 4060 Ti 16 GB adds the 20-billion-parameter class, with 288 GB/s of bandwidth that makes it slower on small models. For an LLM, capacity takes priority over speed as soon as the target model exceeds 7.5 GB.

#Verdict: keep, replace, or don't buy

Which decision fits your situation
SituationDecisionQuantified rationale
You already have a 3070 and are targeting an 8BKeep itGranite 4.2 8B: 5.3 GB, estimated at around 60 t/s
You want a 12B to 14B modelUpgrade to 12 or 16 GBQwen3 14B weighs 9.3 GB, while the 3070 has 8 GB
You want a 20B model or a long context16 GB or more cardgpt-oss 20B weighs 14 GB
You’re choosing between the 3070 and 3070 TiChoose the cheapest oneSame capacity, 608 versus 448 GB/s
You're looking for a card to get startedCompare with a 3060 12 GBMore memory for similar bandwidth

The 3070 is an adequate card for an 8-billion-parameter model, but nothing more. It remains useful as an entry-level card: Ollama supports it, and an 8-billion-parameter model responds faster than you can read. It is not suitable if you plan to use coding agents or long documents, which require context that 8 GB cannot provide. If you already have one, it is enough to explore local LLMs. If you are choosing a card for this sole purpose, start with capacity: a larger model that fits is better than a smaller, faster model. For current prices, check our tracker, which records the lowest price for each card twice a week.

#Frequently asked questions

FAQ
Which LLM should you run on a RTX 3070?+
Start with Granite 4.2 8B (5.3 GB) or Qwen3 8B (5.2 GB): they leave nearly 3 GB for context. Qwen 3.5 9B (6.6 GB) is the upper limit, with 1.4 GB of headroom. Beyond 9 billion parameters in Q4, the model spills into RAM.
Can the RTX 3070 run a 14- or 24-billion-parameter model?+
No. Qwen3 14B weighs 9.3 GB in Ollama, which is 1.3 GB more than the card, while 24-billion-parameter models (15 GB) or gpt-oss 20B (14 GB) require nearly twice as much. They spill into system memory and performance drops. You need at least 12 GB for a 14B model.
RTX 3070 or RTX 3060 12 GB for an LLM?+
For an LLM, the 3060 12 GB is often the best choice: the extra 4 GB lets you run Qwen3 14B (9.3 GB). The 3070 has 448 GB/s of bandwidth versus 360, so it generates about one-quarter faster in theory, but with 8 GB it is limited to 8- to 9-billion-parameter models.
What’s the difference between the 3070 and 3070 Ti for an LLM?+
The capacity is the same at 8 GB, but the 3070 Ti has GDDR6X memory at 608 GB/s versus 448 GB/s for GDDR6, providing 36% more bandwidth. The bandwidth-limited generation theoretically gains by the same proportion. The Ti draws 290 W versus 220 W according to NVIDIA.
What power supply do you need for a RTX 3070 Ti?+
NVIDIA specifies a required system power supply of 750 W for the 3070 Ti (290 W graphics power) and 650 W for the 3070 (220 W). These are the manufacturer's figures for a complete PC: a 550 W power supply is below this recommendation, even though it may be sufficient for an efficient PC.
Can you fine-tune a model on a RTX 3070?+
Yes, with QLoRA on small models. According to Unsloth, the absolute minimum is 3.5 GB for a 3B and 6 GB for an 8B: on 8 GB, an 8B barely fits, with a batch size of 1 and a short context. A 14B requires 8.5 GB and does not fit.
Did this guide help you?

Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.

Prices in euros (€) are French market prices including VAT, as checked by BestLLMfor. US prices differ: the Amazon buttons show the current US price.