Intermediate 11 minRTX 50

Which LLM on RTX 5070 (12 GB) ?

Direct response

The RTX 5070 includes 12 GB of GDDR7 at 672 GB/s: it loads all models up to about 10 GB into VRAM, including Qwen 3.5 9B (6.6 GB) and Gemma 4 12B (7.7 to 8 GB), with real headroom for context. It cannot load Mistral Small 24B or gpt-oss 20B, which weigh 14 GB. Its theoretical ceiling reaches about 127 tokens/s on a 5.3 GB model, and 12 GB is the limit to keep in mind before buying.

The RTX 5070 is the fastest 12 GB card in the Blackwell generation, with 672 GB/s of bandwidth. Its challenge isn't speed but capacity: 12 GB is enough for 4B to 12B models, but not beyond. This guide quantifies what fits, what overflows, how to configure Ollama for context, and when the 16 GB 5070 Ti becomes the right buy.

Choosing a machine? Our picks by budget →

By Mohamed Meguedmi·Update 2026-09-30·Tested on Windows, macOS, and Linux
Recommended hardware

For this setup: RTX 5070 12GB (ASUS Prime OC).

Compare all options by budget, from €800 to €3,500 →

On the go: which laptop for local AI →

Affiliate links — commission possible at no extra cost to you. As an Amazon Associate, BestLLMfor earns from qualifying purchases.

#RTX 5070 for a local LLM: what 12 GB of GDDR7 enables

A RTX 5070 easily runs 4B to 12B models in Q4. Qwen 3.5 9B weighs 6.6 GB and Gemma 4 12B weighs 7.7 to 8 GB according to the Ollama library: both fit within 12 GB, leaving 4 to 5 GB of headroom for context, allowing 16,000 tokens or more with the quantized K/V cache. However, the card cannot load any 14 GB model: Mistral Small 24B and gpt-oss 20B exceed 12 GB. Its speed is excellent for interactive use: 672 GB/s on a 192-bit bus. The 5070 is therefore suitable for everyday use, but hits a clear ceiling as soon as you target higher-quality models.

Architecture
Blackwell, 6,144 CUDA cores and 5th-generation Tensor Cores according to NVIDIA.
Memory
12 GB of GDDR7 on a 192-bit bus, with 672 GB/s of bandwidth according to Wikipedia.
Power
250 W total graphics power; NVIDIA specifies a 650 W minimum power supply for the entire PC.
Launch price
549 dollars, released on March 5, 2025, according to Wikipedia.

#Which models fit in 12 GB

Models on RTX 5070 (sizes Ollama, weights only)
ModelSizeHeadroom on 12 GBVerdict
Granite 4.2 8B5.3 GB6.7 GBVery large, long context
Qwen 3.5 9B6.6 GB5.4 GBComfortable
Gemma 4 12B7.7 to 8.0 GB4 to 4.3 GBComfortable, with a possible 16,000-token context
Mistral Small 24B14 GBNegativeDoes not fit in VRAM
Qwen 3.5 27B17 GBNegativeDoes not fit in VRAM

The site's rule of thumb is about 0.6 GB per billion parameters in Q4, for weights alone. With 12 GB, that puts the practical limit around 14 to 16 billion parameters if the context stays short, and around 12 billion if you want to work with long documents. For an 8B to 9B model, the headroom is far greater than what is needed for a typical context.

#Install Ollama and check the card

  1. 01
    NVIDIA driver
    Ollama requires a NVIDIA driver version 550 or later. For an RTX 50, install the latest driver offered by NVIDIA and check with nvidia-smi.
  2. 02
    Install Ollama
    Download the installer for your system: CUDA is included.
  3. 03
    Run a model
    Try ollama run gemma4:12b to test the card’s limit, or ollama run qwen3.5:9b for a lighter model.
  4. 04
    Check
    Run ollama ps: the PROCESSOR column must display 100% GPU.

#What speed to expect: a ceiling, not a measurement

No throughput figure presented here is an in-house measurement. The generation ceiling is bandwidth divided by weight size. A public third-party measurement (LLaMA 3 8B in Q4_K_M under llama.cpp, May 2024) records 106 tokens/s on a RTX 4080 with a ceiling of 156 tokens/s, or about 68%. Using 70% as a rough benchmark, we get the following estimates for a short context.

Theoretical ceiling on RTX 5070 (672 GB/s)
Model (Q4)WeightsCapEstimated at 70%
Granite 4.2 8B5.3 GB127 tokens/sabout 89 tokens/s
Qwen 3.5 9B6.6 GB102 tokens/saround 71 tokens/s
Gemma 4 12B7.7 GB87 tokens/sabout 61 tokens/s

These ceilings are far above human reading speed: the 5070 will never be slow on these models. It is limited by capacity, not speed. Estimates drop with context, and prompt processing weighs more heavily on compute than on memory.

#Manage context on 12 GB

The 4 GB headroom of Gemma 4 12B disappears quickly with a long context. Ollama quantizes the K/V cache to q8_0 to use about half the memory of the default f16, with very little loss of precision according to its documentation; this requires Flash Attention. This is the setting that makes a 12B model usable on long documents.

Ollama server settings (Linux, macOS)
export OLLAMA_FLASH_ATTENTION=1
export OLLAMA_KV_CACHE_TYPE=q8_0
export OLLAMA_CONTEXT_LENGTH=16384
ollama serve

One tricky point if you serve multiple people: multiple parallel requests increase the reserved context by the same factor, according to the Ollama FAQ. A model that fits one user may overflow with three.

#What you can actually do with 12 GB

Twelve gigabytes changes the nature of the workloads compared with 8 GB. A 9B model leaves enough room for a 16,000-token context and a small embeddings model: this is a comfortable personal RAG setup. A 12B model provides higher-quality writing and summarization responses, at the cost of a more limited context. A coding assistant can make do with an 8B to 12B model and targeted excerpts, but not an entire repository loaded at once.

Uses on RTX 5070 (12 GB)
UsageRealistic?Recommended configuration
Daily chatYes, very comfortableQwen 3.5 9B or Gemma 4 12B
Summarizing long documentsYes, with a 16,000-token contextGemma 4 12B, q8_0 K/V cache
RAG over your filesYesQwen 3.5 9B and lightweight embeddings
Code assistantYes, with an 8B to 12B modelTargeted excerpts, 8,192-token context
14 GB and larger modelsNoPlan for 16 GB of VRAM

Choosing between 9B and 12B depends on the task. The 9B leaves more headroom and works well with long contexts; the 12B is more capable, but every gigabyte of context counts. Start with the smallest model that meets your needs, and move to the 12B only if quality requires it.

#When the 5070 is no longer enough

Three signals indicate that you need more memory. You want a 14B to 24B model, such as Mistral Small 24B or gpt-oss 20B. You need a context longer than 16,000 tokens on a 12B model. You load several models together: generation, embeddings, and reranking. In these cases, the 5070’s speed is not the issue, and the 5070 Ti, 5080, or a 24 GB card are the next steps up.

#What 12 GB is not enough to load

Models above the 12 GB threshold (sizes Ollama)
ModelSizeConsequence
Mistral Small 24B14 GBExceeds RAM by 2 GB
Qwen 3.5 27B17 GBClearly overflows
Qwen 3.5 35B24 GBOut of reach: 24 GB of VRAM minimum
70B models in Q4about 40 GBBeyond the reach of a 12 GB consumer GPU

A model that spills over does not crash: some of the weights are read from system RAM, and throughput drops sharply because the PCIe link is much slower than GDDR7. A 14 GB model on 12 GB is therefore usable for testing, but not for daily use. For these models, aim for 16 GB or more.

#Game and run an LLM on the same card

RTX 5070 is often used for both purposes. The loaded model occupies VRAM as long as it remains in memory, and a recent game also needs several gigabytes: on 12 GB, the two quickly exceed the limit together. The rule is simple: unload the model before launching a game, then reload it afterward. Ollama frees memory after a period of inactivity, and ollama ps shows what is still loaded. If you want an assistant available while you play, choose a small 4B model.

#A used 4070 Ti Super instead?

The question comes up often: a used RTX 4070 Ti Super with 16 GB versus a new RTX 5070 with 12 GB. For an LLM, capacity wins: 16 GB can load 14 GB models that the 5070 can't fit. The 5070 retains the advantages of a warranty, lower power consumption (250 W), and GDDR7 memory. If your models fit within 12 GB, the new 5070 is a reasonable choice; otherwise, capacity comes first. The dedicated guide covers the 4070 Ti Super in detail.

#5070 or 5070 Ti: 12 GB or 16 GB

RTX 5070 and RTX 5070 Ti: what changes for the LLM
CriterionRTX 5070RTX 5070 Ti
Memory12 GB GDDR7, 192-bit16 GB GDDR7, 256 bits
Bandwidth672 GB/s896 GB/s
CUDA cores6 1448 960
Graphics power250 W300 W
Minimum power supply650 W750 W
14 GB modelsNoYes, with 2 GB to spare

The 5070 Ti brings two things: 4 GB of additional memory, which enables Mistral Small 24B and gpt-oss 20B, and 33% more bandwidth. The recommended launch prices were 549 dollars for the 5070 and 749 dollars for the 5070 Ti, a 200-dollar difference. If your models fit within 12 GB, the 5070 is enough; if you are targeting 14 GB models, the difference is a good investment. Check the price tracker to compare prices when you buy.

#2026 verdict

Buy a 5070 if
Your models range from 4B to 12B and you want Blackwell-generation speed. It is highly capable for this use.
Choose the 5070 Ti if
Want to explore 14B to 24B models? The 16 GB makes the difference.
Watch out for the game and the LLM together
A recent game and a loaded model share the 12 GB: close one before launching the other.

In short, the 5070 is a fast card whose only drawback is stopping at 12 GB. If that limit does not bother you today, it will not bother you over the next few months either, as long as you stick to 4B to 12B models; if it does bother you, it is better to pay for the capacity now.

#Frequently asked questions

FAQ
Can the RTX 5070 12 GB run a 70B model?+
No. A 70B in Q4 weighs about 40 GB for the weights alone, more than three times the card's 12 GB. Even with offloading to system RAM, generation would be extremely slow. On 12 GB, stick to 4B to 12B models; for a 70B, plan for at least 48 GB of memory.
How many tokens per second on a RTX 5070?+
No reliable benchmark is published here. The theoretical ceiling—672 GB/s of bandwidth divided by the model size—is about 127 tokens/s for Granite 4.2 8B (5.3 GB), or nearly 89 tokens/s at 70% of that ceiling. This estimate decreases as the context grows.
Is 12 GB enough for professional use in 2026?+
For chat, summarization, RAG, or extraction with a 4B to 12B model, yes. For high-quality writing or a coding assistant working on a large repository, 14B to 24B models weighing 14 GB or more exceed the card's capacity: plan for 16 GB.
Does RTX 5070 support DLSS 4 and an LLM at the same time?+
Both work on the card, but they share its 12 GB of VRAM. A recent game uses a significant portion of the memory: load the model before the game, or close it while you play. The LLM and the game must not run together on 12 GB.
What power supply for a RTX 5070?+
NVIDIA specifies 250 W of total graphics power and a minimum 650 W power supply for the entire PC. A quality 650 W power supply is sufficient; 550 W is below the recommendation. Also check your card model's connectors.
How do you check whether the model fits in VRAM?+
After loading, run ollama ps: the PROCESSOR column should show 100% GPU. A CPU/GPU split means part of the model is in system RAM, which significantly slows generation. In that case, reduce the context, enable the K/V cache in q8_0, or choose a smaller model.
Did this guide help you?

Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.

Prices in euros (€) are French market prices including VAT, as checked by BestLLMfor. US prices differ: the Amazon buttons show the current US price.