Which LLM on RTX 4070 / 4070 Super / 4070 Ti (12 GB) ?
A RTX 4070, 4070 Super, or 4070 Ti offers 12 GB of VRAM and 504 GB/s of bandwidth: enough to run models of up to 14 billion parameters in Q4 (Qwen 3.5 9B, Gemma 4 12B, Qwen3 14B), but not models with 24 to 27 billion. The three cards generate at nearly the same speed because their memory is identical; only long-prompt processing differs.
The three cards in the RTX 4070 family share the same memory, but not the same compute power, and that difference matters less than you might think for an LLM. This guide lists the models that actually fit in 12 GB with their exact weight sizes, the context budget to plan for, third-party measured throughput, and what to do for models that are too large. You will also learn which variant to target used and when to move up to 16 GB.
Choosing a machine? Our picks by budget →
Buying alternative for this guide: GMKtec EVO-X2 64GB / 1TB (Ryzen AI Max+ 395).
A mini PC is a complete machine: check the available memory and engine compatibility. It does not replace macOS/MLX or CUDA.
Why this choice? Our complete guide on GMKtec EVO-X2 64GB / 1TB (Ryzen AI Max+ 395) →
Compare all options by budget, from €800 to €3,500 →
On the go: which laptop for local AI →
Affiliate links — possible commission at no extra cost to you. As an Amazon Associate, BestLLMfor earns from qualifying purchases.
#RTX 4070 and local LLMs: what 12 GB makes possible
For a local LLM, a RTX 4070 is valuable first for its 12 GB of memory and 504 GB/s of bandwidth: these two figures determine what fits and how fast it runs. Models with 7 to 14 billion parameters in Q4 fit entirely in VRAM: according to the Ollama library, Qwen 3.5 9B weighs 6.6 GB, Gemma 4 12B 7.6 GB, and Qwen3 14B 9.3 GB. Hardware Corner measures 71 tokens per second on Qwen3 8B and 42 on Qwen3 14B with a 4,000-token context on a RTX 4070. Models with 24 to 27 billion parameters, such as Mistral Small 24B (14 GB) or Qwen 3.8 27B (18 GB), do not fit: they spill into RAM and speed drops. The Super and Ti variants add no bandwidth.
- Outputs
- 4070 on January 5, 2023, 4070 on April 13, 2023, 4070 Super on January 17, 2024, according to Wikipedia.
- Memory
- 12 GB of GDDR6X, a 192-bit bus at 21 Gbit/s: 192 × 21 ÷ 8 = 504 GB/s, as measured by Puget Systems for all three cards.
- CUDA cores
- 5,888 for the 4070, 7,168 for the Super, 7,680 for the Ti, according to NVIDIA.
- Power consumption
- 200, 220, and 285 W of graphics power; required power supplies of 650, 650, and 700 W, respectively, according to NVIDIA.
#4070, Super, or Ti: same memory, different compute
| Card | CUDA cores | Bandwidth | FP16 (TFLOPS) | Graphics power | Required power supply |
|---|---|---|---|---|---|
| RTX 4070 | 5 888 | 504 GB/s | 29,15 | 200 W | 650 W |
| RTX 4070 Super | 7 168 | 504 GB/s | 35,48 | 220 W | 650 W |
| RTX 4070 Ti | 7 680 | 504 GB/s | 40,09 | 285 W | 700 W |
To produce a token, the GPU rereads all the model’s active weights: generation speed is limited by memory bandwidth, not core count. All three cards have the same bandwidth: 504.2 GB/s. Puget measured Phi-3-mini in 4-bit GGUF with llama.cpp: the 25% gap between the 4070 and 4070 Ti when reading the prompt becomes nearly identical generation scores. Hardware Corner, using other models, rates the Super at 114% and the Ti at 116% of the 4070. The real-world gap therefore ranges from almost zero to around 15%, while CUDA cores increase by 22% to 30%.
Compute is used to read the prompt, the phase that comes before the first word. In FP16, according to Puget, the Super offers 22% more TFLOPS than the 4070 and the Ti 38% more. This matters for RAG, a long document, or a coding agent that sends large prompts: the time before the response decreases. A short chat barely benefits.
#Models that fit in 12 GB
The weights are those of the Ollama library files reviewed on September 29, 2026. The headroom is what remains out of 12 GB before accounting for context, cache, and engine buffers.
| Model | File Ollama | Gross margin | What is it for |
|---|---|---|---|
| Granite 4.2 8B | 5.3 GB | 6.7 GB | RAG and tool calling, 128K context, Apache 2.0 |
| Qwen 3.5 9B (Q4_K_M) | 6.6 GB | 5.4 GB | Versatile, text and image, Apache 2.0 |
| Gemma 4 12B (Q4_K_M) | 7.6 GB | 4.4 GB | Text, image, audio; 7.2 GB QAT version |
| Qwen3 14B | 9.3 GB | 2.7 GB | The largest comfortable dense model |
| Qwen 3.5 9B (Q8_0) | 11 GB | 1 GB | Avoid: too little headroom for context and buffers |
Qwen 3.5 9B is the safest starting point. Gemma 4 12B, with 11.95 billion parameters according to Google, takes over if you also want audio and images. Granite 4.2 8B is suited to document search, where IBM highlights RAG and tool calling. Qwen3 14B is the upper limit: 2.7 GB of headroom is enough for a few thousand context tokens. The Q8_0 of Qwen 3.5 9B leaves only one GB: context, buffers, and memory used by Windows can make it overflow.
A local RAG adds an embedding model: bge-m3 weighs 1.2 GB in Ollama, for 6.5 GB total with Granite 4.2 8B and 8.8 GB with Gemma 4 12B.
#Context: the memory left after the model
Ollama adjusts the context to available video memory: its documentation specifies 4k tokens by default with less than 24 GiB of VRAM, and recommends at least 64,000 tokens for agents, coding tools, and web research. A RTX 4070 falls into the 4k range: an agent launched with the default settings loses its history long before the card is full. Increasing the context consumes VRAM in the form of a KV cache, which stores the attention keys and values for every token already read.
Its size is calculated as: 2 (keys and values) × full-attention layers × KV heads × head dimension × bytes per value × number of tokens. According to its Hugging Face model card, Qwen 3.5 9B has only 8 full-attention layers out of 32, with 4 KV heads of dimension 256. Gemma 4 12B combines 1,024-token sliding-window layers with global layers using unified keys and values. Their cache is much smaller than that of a conventional transformer.
| Model | 8,000 tokens | 32,000 tokens | 128,000 tokens |
|---|---|---|---|
| Qwen 3.5 9B | 0,25 | 1,0 | 4,0 |
| Gemma 4 12B | 0,4 | 0,6 à 0,8 | 1,3 à 2,3 |
| Classic transformer (32 layers, 8 KV heads with dimension 128, example) | 1,0 | 4,0 | 16,0 |
These are theoretical estimates, not measurements: they ignore compute buffers, the state of the DeltaNet layers, and the image encoder, and the Gemma range comes from uncertainty about key/value sharing. Qwen 3.5 9B (6.6 GB) with 32 000 tokens leaves about 4.4 GB; at 128 000 tokens, the cache alone reaches 4 GiB and the total comes close to 12 GB. Check the actual result with ollama ps.
The next lever is cache quantization: according to Ollama’s FAQ, q8_0 uses about half the memory of f16 with very little precision loss, and requires Flash Attention, which Ollama enables automatically when the GPU and model support it. A long context also slows generation: Hardware Corner goes from 71 to 52 to 38 tokens per second between 4k, 16k, and 32k with Qwen3 8B.
#Install Ollama and verify that the model fits
On Windows, Ollama requires an NVIDIA 551.61 driver or newer according to its documentation, and the 4070 family is listed among the supported GPUs. The procedure installs a model and verifies that it runs at 100% on the GPU.
- 01Check the driverOpen the NVIDIA application or run nvidia-smi in a terminal: the driver must be version 551.61 or later.
- 02Install OllamaDownload the Windows program from the Ollama website. The local API then listens on http://localhost:11434.
- 03Run a modelOpen a terminal and run ollama run qwen3.5:9b. The 6.6 GB download happens only once.
- 04Control memory allocationIn a second terminal, run ollama ps. The Processor column should show 100% GPU. A split such as 48% / 52% CPU/GPU means the model is spilling into system RAM.
- 05Configure the context and cacheIncrease the context in the Ollama app's settings or with OLLAMA_CONTEXT_LENGTH, then run ollama ps again. If you need more headroom, set OLLAMA_FLASH_ATTENTION to 1 and OLLAMA_KV_CACHE_TYPE to q8_0 when starting the server.
- Install Ollama on Windows 11: complete guide
- Troubleshoot Ollama: GPU not detected, slowdowns, memory errors
#Throughput measured by third parties and theoretical ceiling
QuelLLM does not publish an in-house measurement here: each throughput figure comes from a third-party source, with its own model and context. Hardware Corner measured a RTX 4070 in March 2026; XDA published a test on a 4070 Ti in June 2026.
| Source | Card | Model | Context | Generation | Prompt processing |
|---|---|---|---|---|---|
| Hardware Corner | RTX 4070 | Qwen3 8B Q4_K | 4k | 71,2 | 3 564 |
| Hardware Corner | RTX 4070 | Qwen3 8B Q4_K | 16k | 52,1 | 2 064 |
| Hardware Corner | RTX 4070 | Qwen3 8B Q4_K | 32k | 38,1 | 1 117 |
| Hardware Corner | RTX 4070 | Qwen3 14B Q4_K | 4k | 42,5 | 2 100 |
| Hardware Corner | RTX 4070 | Qwen3 14B Q4_K | 16k | 32,7 | 1 356 |
| XDA | RTX 4070 Ti | Qwen 3.5 9B (Ollama) | not specified | 65 à 67 | not specified |
A theoretical ceiling helps assess these figures: generation cannot exceed bandwidth divided by model weight. For Qwen3 8B (5.2 GB from Ollama), 504 ÷ 5.2 gives about 97 tokens per second; the 4k measurement, 71.2, reaches 73% of that. Qwen3 14B (9.3 GB) tops out at 54 and measures 42.5, or 78%. XDA reports 65 to 67 on Qwen 3.5 9B (ceiling 76), or 85 to 88%. For Gemma 4 12B (7.6 GB, ceiling 66), this 73 to 88% range gives 48 to 58 tokens per second: an estimate, not a measurement.
| Card | Bandwidth | Generation (t/s) |
|---|---|---|
| RTX 4060 Ti 8 GB | 288 GB/s | 64,03 |
| RTX 5070 12 GB | 672 GB/s | 128,21 |
| RTX 4070 Ti Super 16 GB | 672 GB/s | 132,85 |
| RTX 3080 10 GB | 760 GB/s | 139,95 |
These figures come from the llama.cpp community leaderboard, where no RTX 4070 appears. Generation follows bandwidth: two cards at 672 GB/s, one with 12 GB and the other with 16 GB, run at roughly the same speed; the 4060 Ti, at 288 GB/s, runs half as fast. By the rule of three, a 4070 at 504 GB/s would deliver approximately 99 tokens per second on this 3.56 GiB model: an estimate, not a measurement.
#Beyond 12 GB: Qwen 3.8, Mistral Small, gpt-oss
Qwen 3.8 exists in the Ollama library only with 27 billion parameters, at 18 GB in Q4_K_M: it exceeds the memory of a RTX 4070 by 6 GB and would be split between the card and system memory. Because it is dense, every token rereads all the weights. With 11 GB on the card at 504 GB/s and 7 GB in RAM at 60 GB/s (assuming dual-channel DDR5), one token takes 22 ms on the GPU and 117 ms in RAM: the ceiling falls to about 7 tokens per second, versus 76 for Qwen 3.5 9B.
| Model | File Ollama | Surplus | Nature | Possible approach |
|---|---|---|---|---|
| Mistral Small 24B | 14 GB | 2 GB | Dense | Not recommended: dense |
| gpt-oss 20B | 14 GB | 2 GB | MoE, 3.6 billion active | Offloaded experts (llama.cpp) |
| Gemma 4 26B | 19 GB | 7 GB | MoE (A4B) | Same here; measure it yourself |
| Qwen 3.8 27B | 18 GB | 6 GB | Dense | No: too slow |
Mixture-of-experts models are better suited to offloading. gpt-oss 20B has 21 billion parameters, 3.6 billion of them active per token, according to its model card. The llama.cpp --n-cpu-moe option keeps the experts from the first N layers in RAM, while attention and the cache remain on the card. The official llama.cpp guide for gpt-oss gives an example on an RTX 2060 with 8 GB: 16 expert layers on the CPU for a 32,000-token context. With 12 GB, lower N in steps until you hit an out-of-memory error, then raise it by one step.
We found no sourced measurements of this method on a RTX 4070: throughput depends on your RAM and N, so you must measure it yourself. One user reported about 10 tokens per second in August 2025 with gpt-oss 20B on a RTX 4070 in Ollama, with automatic allocation of 47% CPU and 53% GPU; the version and settings may have changed since then.
#Verdict: which 4070, and when to move to 16 GB
| Situation | Decision | Quantified rationale |
|---|---|---|
| You already have a 4070, Super, or Ti | Keep it for 7- to 14-billion-parameter models | Same bandwidth, generation gap from 0 to 16% |
| Chat, assistant, light coding | The base 4070 is sufficient | Bandwidth-bound generation |
| RAG, large documents, agents | The Super, or the Ti if unavailable | +22% and +38% FP16 TFLOPS |
| You want a 24B model entirely in VRAM (14 GB) | Move up to 16 GB (4070 Ti Super) | 16 GB and 672 GB/s |
| You want to go faster at 12 GB | One RTX 5070 | 672 GB/s versus 504, or 33% more |
The RTX 5070 fixes speed, not capacity: NVIDIA gives it 12 GB of GDDR7 on a 192-bit bus, or 672 GB/s at 28 Gbit/s versus 504 for the 4070. To fit a larger model, the 16 GB is what matters. The 5070 and 4070 Ti Super guides cover every case in detail.
Prices change quickly: every Monday and Thursday, our tracker records the lowest price for local AI graphics cards, along with the price per GB of VRAM.
- Source: NVIDIA, RTX 4070 family specifications
- Source: Puget Systems, consumer GPU LLM performance
- Source: Ollama documentation, context length
- Source: Ollama FAQ, Flash Attention, and KV cache
- Source: llama.cpp guide for running gpt-oss
#Frequently asked questions
Which model should you install on a RTX 4070?+
Does Qwen 3.8 run on a RTX 4070?+
RTX 4070, 4070 Super, or 4070 Ti: which one for a local LLM?+
Is 12 GB still enough for an LLM in 2026?+
What power supply do you need for a RTX 4070 Super?+
RTX 4070 or RTX 3080 10 GB for LLMs?+
Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.