Which LLM on RTX 4070 Ti Super (16 GB) ?
The RTX 4070 Ti Super is a 16 GB GDDR6X card at 672 GB/s: it loads into VRAM Gemma 4 12B (7.7 to 8 GB), gpt-oss 20B, Mistral Small 24B (14 GB each), and any model up to about 14 GB. It is the only card in the 4070 family to exceed 12 GB, and its bandwidth is more than twice that of a RTX 4060 Ti 16 GB. Its theoretical ceiling is around 127 tokens/s on a 5.3 GB model.
When you look for a local LLM on a RTX 4070, you find four similar cards: 4070, 4070 Super, 4070 Ti, and 4070 Ti Super. Only one offers 16 GB. This guide explains what that capacity and 256-bit bus enable, what remains out of reach, and how to position the card against a RTX 4080 Super or a RTX 5070 Ti.
Choosing a machine? Our picks by budget →
Buying alternative for this guide: GMKtec EVO-X2 64GB / 1TB (Ryzen AI Max+ 395).
A mini PC is a complete machine: check the available memory and engine compatibility. It does not replace macOS/MLX or CUDA.
Why this choice? Our complete guide on GMKtec EVO-X2 64GB / 1TB (Ryzen AI Max+ 395) →
Compare all options by budget, from €800 to €3,500 →
On the go: which laptop for local AI →
Affiliate links — commission possible at no extra cost to you. As an Amazon Associate, BestLLMfor earns from qualifying purchases.
#RTX 4070 Ti Great for a local LLM: what 16 GB can do
A RTX 4070 Ti Super runs 4B to 24B models in Q4 with ease. It has 16 GB of GDDR6X on a 256-bit bus, or 672 GB/s, and 8,448 CUDA cores according to NVIDIA. This combination puts it above all 8 and 12 GB cards: Gemma 4 12B fits with plenty of room for context, gpt-oss 20B and Mistral Small 24B, each 14 GB, fit in VRAM with about 2 GB to spare, and a 5 to 8 GB model runs close to the limit of its bandwidth. What does not fit is the next tier: Qwen 3.5 27B weighs in at 17 GB and exceeds the card's capacity. The 4070 Ti Super is therefore not a card for 30B models, but one of the most balanced options for everything below that.
- Architecture
- Ada Lovelace, AD103 chip, the same as the RTX 4080 but with fewer active units.
- Memory
- 16 GB of GDDR6X, a 256-bit bus, and 672 GB/s of bandwidth according to Wikipedia.
- Compute
- 8,448 CUDA cores according to the NVIDIA product page, 4th-generation Tensor Cores.
- Power
- 285 W of total graphics power; NVIDIA specifies a 700 W minimum power supply for the entire PC.
#4070, 4070 Super, 4070 Ti, 4070 Ti Super: which one for an LLM
All four cards share the same product-line name but aren't equivalent for LLMs. Three of them are limited to 12 GB of VRAM; only the Ti Super moves up to 16 GB and a 256-bit bus. For local AI, this difference matters more than raw compute power: it determines which models the card can handle.
| Card | Memory | Bus | CUDA cores | Power |
|---|---|---|---|---|
| RTX 4070 Ti Super | 16 GB GDDR6X | 256-bit | 8 448 | 285 W |
| RTX 4070 Ti | 12 GB GDDR6X | 192 bits | 7 680 | 285 W |
| RTX 4070 Super | 12 GB GDDR6X | 192 bits | 7 168 | 220 W |
| RTX 4070 | 12 GB GDDR6 or GDDR6X | 192 bits | 5 888 | 200 W |
The RTX 4070 Ti lists 504 GB/s according to Wikipedia, compared with 672 GB/s for the Ti Super. If you were looking for a standard RTX 4070, the dedicated page covers the 12 GB variant. One thing they all have in common: 12B models fit, while those with 14 GB or more require 16 GB.
#Which models fit in 16 GB
| Model | Size | Headroom on 16 GB | Verdict |
|---|---|---|---|
| Granite 4.2 8B | 5.3 GB | 10.7 GB | Very large: long context |
| Gemma 4 12B | 7.7 to 8.0 GB | 8 GB | Comfortable |
| gpt-oss 20B | 14 GB | 2 GB | Worth watching, context |
| Mistral Small 24B | 14 GB | 2 GB | Worth watching, context |
| Devstral Small 2 24B | 15 GB | 1 GB | Barely fits, very short context |
| Qwen 3.5 27B | 17 GB | Negative | Doesn't fit |
The 14 and 15 GB models leave little room for the context cache: Ollama uses 4 096 tokens by default, which works, but 16 000 tokens require quantizing the K/V cache. Devstral Small 2, a coding model, is right at the limit: at 15 GB, almost nothing remains for context. For coding, Mistral Small 24B or a smaller model with a longer context is a better choice on 16 GB.
#What remains out of reach for 16 GB
Three model families exceed the card’s capacity. The 27B to 35B models in Q4: Qwen 3.5 27B weighs 17 GB and Qwen 3.5 35B weighs 24 GB according to Ollama, while Gemma 4 26B ranges from 16 to 19 GB. Then there are the 70B models, whose weights alone are around 40 GB according to the site’s reference point. Finally, long-context models pushed to their maximum: an 8 GB model advertised with 256,000 tokens of context will never fit with a full cache. For these cases, move to 24 GB of VRAM, a Mac with unified memory, or a dedicated machine: the hardware page compares these options.
#Install and verify VRAM placement
- 01DriverOllama requires NVIDIA driver 550 or later for recent GeForce cards. Check your version with nvidia-smi.
- 02InstallationInstall Ollama for your system, without a separate CUDA installation.
- 03First testLaunch ollama run mistral-small for a first model that uses the 16 GB, then ollama ps.
- 04ControlThe PROCESSOR column should show 100% GPU; otherwise, reduce the context.
The K/V cache in q8_0 uses about half the memory of f16, the default setting, and requires Flash Attention. This setting is what allows a 14 GB model to serve a context of 16 000 tokens on 16 GB.
#What speed to expect: ceilings and third-party measurements
No throughput is presented here as an in-house measurement. The generation ceiling is bandwidth divided by weight size. A public third-party benchmark (LLaMA 3 8B in Q4_K_M under llama.cpp, May 2024) records 106 tokens/s on a RTX 4080 with 716.8 GB/s of bandwidth, or about 68% of its ceiling, and 75% on a RTX 4070 Ti. The 4070 Ti Super’s bandwidth is close to that of the 4080, so its order of magnitude is too. The table applies 70% of the ceiling.
| Model | Weights | Cap | Estimated at 70% |
|---|---|---|---|
| Granite 4.2 8B | 5.3 GB | 127 tokens/s | about 89 tokens/s |
| Qwen 3.5 9B | 6.6 GB | 102 tokens/s | around 71 tokens/s |
| Gemma 4 12B | 7.7 GB | 87 tokens/s | about 61 tokens/s |
| Mistral Small 24B | 14 GB | 48 tokens/s | approximately 34 tokens/s |
gpt-oss 20B is a mixture-of-experts model: it reads only part of its weights per token, so its throughput exceeds that of a dense 14 GB model. No sourced figure is reproduced here. The main takeaway is that the card goes from fast on an 8B to adequate on a 24B.
#4070 Ti Super vs. the 4080 Super and 5070 Ti
| Card | VRAM | Bandwidth | Note |
|---|---|---|---|
| RTX 4070 Ti Super | 16 GB GDDR6X | 672 GB/s | Same AD103 chip as the 4080 |
| RTX 4080 Super | 16 GB GDDR6X | 736 GB/s | About 10% more bandwidth |
| RTX 5070 Ti | 16 GB GDDR7 | 896 GB/s | 33% more bandwidth |
All three cards offer 16 GB: capacity is identical, so the list of models is too. Only generation speed changes, in proportion to bandwidth. The 4080 Super gains about 10%; the 5070 Ti, about one-third. On the used market, the price difference matters more than these throughput differences: check the price tracker to compare. A used 4070 Ti Super at a good price remains an excellent choice; a new 5070 Ti is justified by its warranty and speed.
#Use cases that benefit from 16 GB
The 16 GB mainly changes what you can run at the same time. A local RAG combines a generation model, an embedding model, and a context window long enough to hold the retrieved excerpts: with Gemma 4 12B (7.7 to 8 GB) and a small embedding model, everything fits with several gigabytes to spare. A coding assistant needs only an 8B to 12B model and a 16,000-token context, with no constraints. A 14 GB model such as Mistral Small 24B takes up almost the entire card: it becomes an exclusive choice, not one model among several.
| Usage | Configuration | Remaining headroom |
|---|---|---|
| Personal RAG | Gemma 4 12B and lightweight embedding model | About 7 GB for the context |
| Code assistant | 8B model, 16,000-token context, q8_0 cache | Approximately 9 GB |
| Maximum-quality chat | Mistral Small 24B alone | About 2 GB, short context |
A potential pitfall for shared use: in Ollama, multiple parallel requests on the same model increase the size of the reserved context accordingly. Serving three colleagues with a 14 GB model can therefore overflow the card, while a single user works without difficulty. On 16 GB, limit parallelism or choose a lighter model for multi-user access.
#Buying a used 4070 Ti Super: what to check
The card was released in January 2024: most used examples are less than three years old, but nothing guarantees how they were used before. Prices change every week: check the price tracker rather than relying on a fixed figure in this guide. Before buying, verify that nvidia-smi shows 16 GB of memory, run a 14 GB model, monitor the temperature during a long generation, and listen to the fans. Ask for the invoice and the manufacturer's remaining warranty: they are often transferred with the card.
#2026 verdict
- Keep it if
- Your models fit in 16 GB. There is no reason to replace it before you need 24 GB.
- Buy it used if
- The price is significantly lower than that of a new 5070 Ti. Check the remaining warranty, temperature, and fan noise.
- Aim for 24 GB if
- You want Qwen 3.5 27B or more: no setting will fit 17 GB into 16 GB with context, and a 24 GB card saves you from buying again.
- Source: official NVIDIA specifications (RTX 4070 family)
- Source: Ollama's official FAQ (Flash Attention, K/V cache)
- Source: Mistral Small size in the Ollama library
#Frequently asked questions
Can the RTX 4070 Ti Super run Mistral Small 24B?+
What’s the difference between RTX 4070 Ti and the 4070 Ti Super for an LLM?+
RTX 4070 Ti Super or RTX 5070 Ti?+
The RTX 4070 Ti Super or the 4080 for an LLM?+
What power supply do you need for a RTX 4070 Ti Super?+
Can you run Qwen 3.5 27B on a RTX 4070 Ti Super?+
Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.