Intermediate 14 minRTX 40

Which LLM on RTX 4070 / 4070 Super / 4070 Ti (12 GB) ?

Direct response

A RTX 4070, 4070 Super, or 4070 Ti offers 12 GB of VRAM and 504 GB/s of bandwidth: enough to run models of up to 14 billion parameters in Q4 (Qwen 3.5 9B, Gemma 4 12B, Qwen3 14B), but not models with 24 to 27 billion. The three cards generate at nearly the same speed because their memory is identical; only long-prompt processing differs.

The three cards in the RTX 4070 family share the same memory, but not the same compute power, and that difference matters less than you might think for an LLM. This guide lists the models that actually fit in 12 GB with their exact weight sizes, the context budget to plan for, third-party measured throughput, and what to do for models that are too large. You will also learn which variant to target used and when to move up to 16 GB.

Choosing a machine? Our picks by budget →

By Mohamed Meguedmi·Update 2026-09-29·Tested on Windows, macOS, and Linux
Recommended hardware

Buying alternative for this guide: GMKtec EVO-X2 64GB / 1TB (Ryzen AI Max+ 395).

A mini PC is a complete machine: check the available memory and engine compatibility. It does not replace macOS/MLX or CUDA.

Why this choice? Our complete guide on GMKtec EVO-X2 64GB / 1TB (Ryzen AI Max+ 395) →

Compare all options by budget, from €800 to €3,500 →

On the go: which laptop for local AI →

Affiliate links — possible commission at no extra cost to you. As an Amazon Associate, BestLLMfor earns from qualifying purchases.

#RTX 4070 and local LLMs: what 12 GB makes possible

For a local LLM, a RTX 4070 is valuable first for its 12 GB of memory and 504 GB/s of bandwidth: these two figures determine what fits and how fast it runs. Models with 7 to 14 billion parameters in Q4 fit entirely in VRAM: according to the Ollama library, Qwen 3.5 9B weighs 6.6 GB, Gemma 4 12B 7.6 GB, and Qwen3 14B 9.3 GB. Hardware Corner measures 71 tokens per second on Qwen3 8B and 42 on Qwen3 14B with a 4,000-token context on a RTX 4070. Models with 24 to 27 billion parameters, such as Mistral Small 24B (14 GB) or Qwen 3.8 27B (18 GB), do not fit: they spill into RAM and speed drops. The Super and Ti variants add no bandwidth.

Outputs
4070 on January 5, 2023, 4070 on April 13, 2023, 4070 Super on January 17, 2024, according to Wikipedia.
Memory
12 GB of GDDR6X, a 192-bit bus at 21 Gbit/s: 192 × 21 ÷ 8 = 504 GB/s, as measured by Puget Systems for all three cards.
CUDA cores
5,888 for the 4070, 7,168 for the Super, 7,680 for the Ti, according to NVIDIA.
Power consumption
200, 220, and 285 W of graphics power; required power supplies of 650, 650, and 700 W, respectively, according to NVIDIA.
!
Check the memory specification of a used 4070
Since August 2024, a RTX 4070 has also existed with GDDR6 at 20 Gbit/s instead of GDDR6X, according to Wikipedia. On the same 192-bit bus, that delivers 480 GB/s instead of 504, or 5% less. Ask for the exact part number.

#4070, Super, or Ti: same memory, different compute

The three RTX 4070 12 GB side by side (NVIDIA, Puget Systems)
CardCUDA coresBandwidthFP16 (TFLOPS)Graphics powerRequired power supply
RTX 40705 888504 GB/s29,15200 W650 W
RTX 4070 Super7 168504 GB/s35,48220 W650 W
RTX 4070 Ti7 680504 GB/s40,09285 W700 W

To produce a token, the GPU rereads all the model’s active weights: generation speed is limited by memory bandwidth, not core count. All three cards have the same bandwidth: 504.2 GB/s. Puget measured Phi-3-mini in 4-bit GGUF with llama.cpp: the 25% gap between the 4070 and 4070 Ti when reading the prompt becomes nearly identical generation scores. Hardware Corner, using other models, rates the Super at 114% and the Ti at 116% of the 4070. The real-world gap therefore ranges from almost zero to around 15%, while CUDA cores increase by 22% to 30%.

Compute is used to read the prompt, the phase that comes before the first word. In FP16, according to Puget, the Super offers 22% more TFLOPS than the 4070 and the Ti 38% more. This matters for RAG, a long document, or a coding agent that sends large prompts: the time before the response decreases. A short chat barely benefits.

→
The Ti is rarely worth it
It consumes 285 W versus 220 W for the Super, or 30% more, for 2 additional generation points according to Hardware Corner, and it is the oldest of the three. At a similar price, the Super is the best buy in the lineup.

#Models that fit in 12 GB

The weights are those of the Ollama library files reviewed on September 29, 2026. The headroom is what remains out of 12 GB before accounting for context, cache, and engine buffers.

Models testable on RTX 4070, Super, or Ti (files Ollama)
ModelFile OllamaGross marginWhat is it for
Granite 4.2 8B5.3 GB6.7 GBRAG and tool calling, 128K context, Apache 2.0
Qwen 3.5 9B (Q4_K_M)6.6 GB5.4 GBVersatile, text and image, Apache 2.0
Gemma 4 12B (Q4_K_M)7.6 GB4.4 GBText, image, audio; 7.2 GB QAT version
Qwen3 14B9.3 GB2.7 GBThe largest comfortable dense model
Qwen 3.5 9B (Q8_0)11 GB1 GBAvoid: too little headroom for context and buffers

Qwen 3.5 9B is the safest starting point. Gemma 4 12B, with 11.95 billion parameters according to Google, takes over if you also want audio and images. Granite 4.2 8B is suited to document search, where IBM highlights RAG and tool calling. Qwen3 14B is the upper limit: 2.7 GB of headroom is enough for a few thousand context tokens. The Q8_0 of Qwen 3.5 9B leaves only one GB: context, buffers, and memory used by Windows can make it overflow.

A local RAG adds an embedding model: bge-m3 weighs 1.2 GB in Ollama, for 6.5 GB total with Granite 4.2 8B and 8.8 GB with Gemma 4 12B.

#Context: the memory left after the model

Ollama adjusts the context to available video memory: its documentation specifies 4k tokens by default with less than 24 GiB of VRAM, and recommends at least 64,000 tokens for agents, coding tools, and web research. A RTX 4070 falls into the 4k range: an agent launched with the default settings loses its history long before the card is full. Increasing the context consumes VRAM in the form of a KV cache, which stores the attention keys and values for every token already read.

Its size is calculated as: 2 (keys and values) × full-attention layers × KV heads × head dimension × bytes per value × number of tokens. According to its Hugging Face model card, Qwen 3.5 9B has only 8 full-attention layers out of 32, with 4 KV heads of dimension 256. Gemma 4 12B combines 1,024-token sliding-window layers with global layers using unified keys and values. Their cache is much smaller than that of a conventional transformer.

Estimated KV cache in f16 (GiB), calculated from the published configuration
Model8,000 tokens32,000 tokens128,000 tokens
Qwen 3.5 9B0,251,04,0
Gemma 4 12B0,40,6 à 0,81,3 à 2,3
Classic transformer (32 layers, 8 KV heads with dimension 128, example)1,04,016,0

These are theoretical estimates, not measurements: they ignore compute buffers, the state of the DeltaNet layers, and the image encoder, and the Gemma range comes from uncertainty about key/value sharing. Qwen 3.5 9B (6.6 GB) with 32 000 tokens leaves about 4.4 GB; at 128 000 tokens, the cache alone reaches 4 GiB and the total comes close to 12 GB. Check the actual result with ollama ps.

The next lever is cache quantization: according to Ollama’s FAQ, q8_0 uses about half the memory of f16 with very little precision loss, and requires Flash Attention, which Ollama enables automatically when the GPU and model support it. A long context also slows generation: Hardware Corner goes from 71 to 52 to 38 tokens per second between 4k, 16k, and 32k with Qwen3 8B.

#Install Ollama and verify that the model fits

On Windows, Ollama requires an NVIDIA 551.61 driver or newer according to its documentation, and the 4070 family is listed among the supported GPUs. The procedure installs a model and verifies that it runs at 100% on the GPU.

  1. 01
    Check the driver
    Open the NVIDIA application or run nvidia-smi in a terminal: the driver must be version 551.61 or later.
  2. 02
    Install Ollama
    Download the Windows program from the Ollama website. The local API then listens on http://localhost:11434.
  3. 03
    Run a model
    Open a terminal and run ollama run qwen3.5:9b. The 6.6 GB download happens only once.
  4. 04
    Control memory allocation
    In a second terminal, run ollama ps. The Processor column should show 100% GPU. A split such as 48% / 52% CPU/GPU means the model is spilling into system RAM.
  5. 05
    Configure the context and cache
    Increase the context in the Ollama app's settings or with OLLAMA_CONTEXT_LENGTH, then run ollama ps again. If you need more headroom, set OLLAMA_FLASH_ATTENTION to 1 and OLLAMA_KV_CACHE_TYPE to q8_0 when starting the server.
Basic commands
ollama run qwen3.5:9b
ollama run gemma4:12b
ollama ps

# Exemple de la documentation Ollama pour fixer le contexte à 64 000 tokens
OLLAMA_CONTEXT_LENGTH=64000 ollama serve

#Throughput measured by third parties and theoretical ceiling

QuelLLM does not publish an in-house measurement here: each throughput figure comes from a third-party source, with its own model and context. Hardware Corner measured a RTX 4070 in March 2026; XDA published a test on a 4070 Ti in June 2026.

Published throughput on RTX 4070 and 4070 Ti (tokens per second)
SourceCardModelContextGenerationPrompt processing
Hardware CornerRTX 4070Qwen3 8B Q4_K4k71,23 564
Hardware CornerRTX 4070Qwen3 8B Q4_K16k52,12 064
Hardware CornerRTX 4070Qwen3 8B Q4_K32k38,11 117
Hardware CornerRTX 4070Qwen3 14B Q4_K4k42,52 100
Hardware CornerRTX 4070Qwen3 14B Q4_K16k32,71 356
XDARTX 4070 TiQwen 3.5 9B (Ollama)not specified65 à 67not specified

A theoretical ceiling helps assess these figures: generation cannot exceed bandwidth divided by model weight. For Qwen3 8B (5.2 GB from Ollama), 504 ÷ 5.2 gives about 97 tokens per second; the 4k measurement, 71.2, reaches 73% of that. Qwen3 14B (9.3 GB) tops out at 54 and measures 42.5, or 78%. XDA reports 65 to 67 on Qwen 3.5 9B (ceiling 76), or 85 to 88%. For Gemma 4 12B (7.6 GB, ceiling 66), this 73 to 88% range gives 48 to 58 tokens per second: an estimate, not a measurement.

Generation on llama.cpp based on bandwidth (Llama 2 7B Q4_0, tg128 test, Flash Attention enabled)
CardBandwidthGeneration (t/s)
RTX 4060 Ti 8 GB288 GB/s64,03
RTX 5070 12 GB672 GB/s128,21
RTX 4070 Ti Super 16 GB672 GB/s132,85
RTX 3080 10 GB760 GB/s139,95

These figures come from the llama.cpp community leaderboard, where no RTX 4070 appears. Generation follows bandwidth: two cards at 672 GB/s, one with 12 GB and the other with 16 GB, run at roughly the same speed; the 4060 Ti, at 288 GB/s, runs half as fast. By the rule of three, a 4070 at 504 GB/s would deliver approximately 99 tokens per second on this 3.56 GiB model: an estimate, not a measurement.

#Beyond 12 GB: Qwen 3.8, Mistral Small, gpt-oss

Qwen 3.8 exists in the Ollama library only with 27 billion parameters, at 18 GB in Q4_K_M: it exceeds the memory of a RTX 4070 by 6 GB and would be split between the card and system memory. Because it is dense, every token rereads all the weights. With 11 GB on the card at 504 GB/s and 7 GB in RAM at 60 GB/s (assuming dual-channel DDR5), one token takes 22 ms on the GPU and 117 ms in RAM: the ceiling falls to about 7 tokens per second, versus 76 for Qwen 3.5 9B.

Models exceeding 12 GB (files Ollama)
ModelFile OllamaSurplusNaturePossible approach
Mistral Small 24B14 GB2 GBDenseNot recommended: dense
gpt-oss 20B14 GB2 GBMoE, 3.6 billion activeOffloaded experts (llama.cpp)
Gemma 4 26B19 GB7 GBMoE (A4B)Same here; measure it yourself
Qwen 3.8 27B18 GB6 GBDenseNo: too slow

Mixture-of-experts models are better suited to offloading. gpt-oss 20B has 21 billion parameters, 3.6 billion of them active per token, according to its model card. The llama.cpp --n-cpu-moe option keeps the experts from the first N layers in RAM, while attention and the cache remain on the card. The official llama.cpp guide for gpt-oss gives an example on an RTX 2060 with 8 GB: 16 expert layers on the CPU for a 32,000-token context. With 12 GB, lower N in steps until you hit an out-of-memory error, then raise it by one step.

Example from the llama.cpp guide (8 GB card): adjust --n-cpu-moe
llama-server -hf ggml-org/gpt-oss-20b-GGUF --ctx-size 32768 --jinja -ub 2048 -b 2048 --n-cpu-moe 16

We found no sourced measurements of this method on a RTX 4070: throughput depends on your RAM and N, so you must measure it yourself. One user reported about 10 tokens per second in August 2025 with gpt-oss 20B on a RTX 4070 in Ollama, with automatic allocation of 47% CPU and 53% GPU; the version and settings may have changed since then.

i
QLoRA fine-tuning fits on 12 GB
According to Unsloth, the absolute minimum VRAM for 4-bit QLoRA is 6 GB for 8 billion parameters and 8.5 GB for 14 billion: a 14B remains feasible on a 4070, with a batch size of 1 and a short context.

#Verdict: which 4070, and when to move to 16 GB

Which decision fits your situation
SituationDecisionQuantified rationale
You already have a 4070, Super, or TiKeep it for 7- to 14-billion-parameter modelsSame bandwidth, generation gap from 0 to 16%
Chat, assistant, light codingThe base 4070 is sufficientBandwidth-bound generation
RAG, large documents, agentsThe Super, or the Ti if unavailable+22% and +38% FP16 TFLOPS
You want a 24B model entirely in VRAM (14 GB)Move up to 16 GB (4070 Ti Super)16 GB and 672 GB/s
You want to go faster at 12 GBOne RTX 5070672 GB/s versus 504, or 33% more

The RTX 5070 fixes speed, not capacity: NVIDIA gives it 12 GB of GDDR7 on a 192-bit bus, or 672 GB/s at 28 Gbit/s versus 504 for the 4070. To fit a larger model, the 16 GB is what matters. The 5070 and 4070 Ti Super guides cover every case in detail.

Prices change quickly: every Monday and Thursday, our tracker records the lowest price for local AI graphics cards, along with the price per GB of VRAM.

#Frequently asked questions

Which model should you install on a RTX 4070?+
Start with Qwen 3.5 9B: its 6.6 GB Ollama file leaves more than 5 GB of headroom for context, and it handles text and images. Gemma 4 12B, at 7.6 GB, adds audio. Qwen3 14B, at 9.3 GB, is the largest model that remains comfortable. Beyond 14 billion parameters in Q4, the model spills into system memory.
Does Qwen 3.8 run on a RTX 4070?+
Not comfortably. Qwen 3.8 is available only with 27 billion parameters, at 18 GB in Q4_K_M in Ollama, which is 6 GB more than the card has. Split between VRAM and RAM, it tops out at around 7 tokens per second in theory. On 12 GB, prefer Qwen 3.5 9B or Gemma 4 12B.
RTX 4070, 4070 Super, or 4070 Ti: which one for a local LLM?+
For text generation, they are nearly equivalent: the same 12 GB of memory, the same 504 GB/s bandwidth, and a difference of 0 to 16% depending on the cited measurements. The Super is worth its premium if you send long prompts, because it reads them faster. The Ti consumes 285 W for a minimal gain.
Is 12 GB still enough for an LLM in 2026?+
Yes for models with 7 to 14 billion parameters, which cover chat, light coding, and RAG. No for 24B to 27B models in Q4, which weigh 14 to 18 GB: a fully VRAM-resident 24B requires 16 GB. Models with offloaded experts remain an option.
What power supply do you need for a RTX 4070 Super?+
NVIDIA specifies a required power supply of 650 W for the 4070 and 4070 Super, and 700 W for the 4070 Ti, with 220 W of graphics power for the Super. These are the manufacturer’s figures for a complete PC: a 550 W power supply falls short of that recommendation, even though it may be enough for an energy-efficient PC.
RTX 4070 or RTX 3080 10 GB for LLMs?+
The 3080 10 GB has a 320-bit bus and 760 GB/s of bandwidth, compared with 504 for the 4070: it generates faster as long as the model fits within its 10 GB. The 4070 offers 2 GB more, which matters for Gemma 4 12B or Qwen3 14B. It all depends on the target model size.
Did this guide help you?

Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.