Intermediate 11 minRTX 40

Which LLM on RTX 4070 Ti Super (16 GB) ?

Direct response

The RTX 4070 Ti Super is a 16 GB GDDR6X card at 672 GB/s: it loads into VRAM Gemma 4 12B (7.7 to 8 GB), gpt-oss 20B, Mistral Small 24B (14 GB each), and any model up to about 14 GB. It is the only card in the 4070 family to exceed 12 GB, and its bandwidth is more than twice that of a RTX 4060 Ti 16 GB. Its theoretical ceiling is around 127 tokens/s on a 5.3 GB model.

When you look for a local LLM on a RTX 4070, you find four similar cards: 4070, 4070 Super, 4070 Ti, and 4070 Ti Super. Only one offers 16 GB. This guide explains what that capacity and 256-bit bus enable, what remains out of reach, and how to position the card against a RTX 4080 Super or a RTX 5070 Ti.

Choosing a machine? Our picks by budget →

By Mohamed Meguedmi·Update 2026-09-30·Tested on Windows, macOS, and Linux
Recommended hardware

Buying alternative for this guide: GMKtec EVO-X2 64GB / 1TB (Ryzen AI Max+ 395).

A mini PC is a complete machine: check the available memory and engine compatibility. It does not replace macOS/MLX or CUDA.

Why this choice? Our complete guide on GMKtec EVO-X2 64GB / 1TB (Ryzen AI Max+ 395) →

Compare all options by budget, from €800 to €3,500 →

On the go: which laptop for local AI →

Affiliate links — commission possible at no extra cost to you. As an Amazon Associate, BestLLMfor earns from qualifying purchases.

#RTX 4070 Ti Great for a local LLM: what 16 GB can do

A RTX 4070 Ti Super runs 4B to 24B models in Q4 with ease. It has 16 GB of GDDR6X on a 256-bit bus, or 672 GB/s, and 8,448 CUDA cores according to NVIDIA. This combination puts it above all 8 and 12 GB cards: Gemma 4 12B fits with plenty of room for context, gpt-oss 20B and Mistral Small 24B, each 14 GB, fit in VRAM with about 2 GB to spare, and a 5 to 8 GB model runs close to the limit of its bandwidth. What does not fit is the next tier: Qwen 3.5 27B weighs in at 17 GB and exceeds the card's capacity. The 4070 Ti Super is therefore not a card for 30B models, but one of the most balanced options for everything below that.

Architecture
Ada Lovelace, AD103 chip, the same as the RTX 4080 but with fewer active units.
Memory
16 GB of GDDR6X, a 256-bit bus, and 672 GB/s of bandwidth according to Wikipedia.
Compute
8,448 CUDA cores according to the NVIDIA product page, 4th-generation Tensor Cores.
Power
285 W of total graphics power; NVIDIA specifies a 700 W minimum power supply for the entire PC.

#4070, 4070 Super, 4070 Ti, 4070 Ti Super: which one for an LLM

All four cards share the same product-line name but aren't equivalent for LLMs. Three of them are limited to 12 GB of VRAM; only the Ti Super moves up to 16 GB and a 256-bit bus. For local AI, this difference matters more than raw compute power: it determines which models the card can handle.

RTX 4070 family: what matters for the LLM (source NVIDIA)
CardMemoryBusCUDA coresPower
RTX 4070 Ti Super16 GB GDDR6X256-bit8 448285 W
RTX 4070 Ti12 GB GDDR6X192 bits7 680285 W
RTX 4070 Super12 GB GDDR6X192 bits7 168220 W
RTX 407012 GB GDDR6 or GDDR6X192 bits5 888200 W

The RTX 4070 Ti lists 504 GB/s according to Wikipedia, compared with 672 GB/s for the Ti Super. If you were looking for a standard RTX 4070, the dedicated page covers the 12 GB variant. One thing they all have in common: 12B models fit, while those with 14 GB or more require 16 GB.

#Which models fit in 16 GB

Models on RTX 4070 Ti Super (sizes Ollama, weights only)
ModelSizeHeadroom on 16 GBVerdict
Granite 4.2 8B5.3 GB10.7 GBVery large: long context
Gemma 4 12B7.7 to 8.0 GB8 GBComfortable
gpt-oss 20B14 GB2 GBWorth watching, context
Mistral Small 24B14 GB2 GBWorth watching, context
Devstral Small 2 24B15 GB1 GBBarely fits, very short context
Qwen 3.5 27B17 GBNegativeDoesn't fit

The 14 and 15 GB models leave little room for the context cache: Ollama uses 4 096 tokens by default, which works, but 16 000 tokens require quantizing the K/V cache. Devstral Small 2, a coding model, is right at the limit: at 15 GB, almost nothing remains for context. For coding, Mistral Small 24B or a smaller model with a longer context is a better choice on 16 GB.

#What remains out of reach for 16 GB

Three model families exceed the card’s capacity. The 27B to 35B models in Q4: Qwen 3.5 27B weighs 17 GB and Qwen 3.5 35B weighs 24 GB according to Ollama, while Gemma 4 26B ranges from 16 to 19 GB. Then there are the 70B models, whose weights alone are around 40 GB according to the site’s reference point. Finally, long-context models pushed to their maximum: an 8 GB model advertised with 256,000 tokens of context will never fit with a full cache. For these cases, move to 24 GB of VRAM, a Mac with unified memory, or a dedicated machine: the hardware page compares these options.

#Install and verify VRAM placement

  1. 01
    Driver
    Ollama requires NVIDIA driver 550 or later for recent GeForce cards. Check your version with nvidia-smi.
  2. 02
    Installation
    Install Ollama for your system, without a separate CUDA installation.
  3. 03
    First test
    Launch ollama run mistral-small for a first model that uses the 16 GB, then ollama ps.
  4. 04
    Control
    The PROCESSOR column should show 100% GPU; otherwise, reduce the context.
Long context on 16 GB (Linux, macOS)
export OLLAMA_FLASH_ATTENTION=1
export OLLAMA_KV_CACHE_TYPE=q8_0
export OLLAMA_CONTEXT_LENGTH=16384
ollama serve

The K/V cache in q8_0 uses about half the memory of f16, the default setting, and requires Flash Attention. This setting is what allows a 14 GB model to serve a context of 16 000 tokens on 16 GB.

#What speed to expect: ceilings and third-party measurements

No throughput is presented here as an in-house measurement. The generation ceiling is bandwidth divided by weight size. A public third-party benchmark (LLaMA 3 8B in Q4_K_M under llama.cpp, May 2024) records 106 tokens/s on a RTX 4080 with 716.8 GB/s of bandwidth, or about 68% of its ceiling, and 75% on a RTX 4070 Ti. The 4070 Ti Super’s bandwidth is close to that of the 4080, so its order of magnitude is too. The table applies 70% of the ceiling.

Theoretical ceiling on RTX 4070 Ti Super (672 GB/s)
ModelWeightsCapEstimated at 70%
Granite 4.2 8B5.3 GB127 tokens/sabout 89 tokens/s
Qwen 3.5 9B6.6 GB102 tokens/saround 71 tokens/s
Gemma 4 12B7.7 GB87 tokens/sabout 61 tokens/s
Mistral Small 24B14 GB48 tokens/sapproximately 34 tokens/s

gpt-oss 20B is a mixture-of-experts model: it reads only part of its weights per token, so its throughput exceeds that of a dense 14 GB model. No sourced figure is reproduced here. The main takeaway is that the card goes from fast on an 8B to adequate on a 24B.

#4070 Ti Super vs. the 4080 Super and 5070 Ti

16 GB cards to compare
CardVRAMBandwidthNote
RTX 4070 Ti Super16 GB GDDR6X672 GB/sSame AD103 chip as the 4080
RTX 4080 Super16 GB GDDR6X736 GB/sAbout 10% more bandwidth
RTX 5070 Ti16 GB GDDR7896 GB/s33% more bandwidth

All three cards offer 16 GB: capacity is identical, so the list of models is too. Only generation speed changes, in proportion to bandwidth. The 4080 Super gains about 10%; the 5070 Ti, about one-third. On the used market, the price difference matters more than these throughput differences: check the price tracker to compare. A used 4070 Ti Super at a good price remains an excellent choice; a new 5070 Ti is justified by its warranty and speed.

#Use cases that benefit from 16 GB

The 16 GB mainly changes what you can run at the same time. A local RAG combines a generation model, an embedding model, and a context window long enough to hold the retrieved excerpts: with Gemma 4 12B (7.7 to 8 GB) and a small embedding model, everything fits with several gigabytes to spare. A coding assistant needs only an 8B to 12B model and a 16,000-token context, with no constraints. A 14 GB model such as Mistral Small 24B takes up almost the entire card: it becomes an exclusive choice, not one model among several.

Three typical configurations on 16 GB
UsageConfigurationRemaining headroom
Personal RAGGemma 4 12B and lightweight embedding modelAbout 7 GB for the context
Code assistant8B model, 16,000-token context, q8_0 cacheApproximately 9 GB
Maximum-quality chatMistral Small 24B aloneAbout 2 GB, short context

A potential pitfall for shared use: in Ollama, multiple parallel requests on the same model increase the size of the reserved context accordingly. Serving three colleagues with a 14 GB model can therefore overflow the card, while a single user works without difficulty. On 16 GB, limit parallelism or choose a lighter model for multi-user access.

#Buying a used 4070 Ti Super: what to check

The card was released in January 2024: most used examples are less than three years old, but nothing guarantees how they were used before. Prices change every week: check the price tracker rather than relying on a fixed figure in this guide. Before buying, verify that nvidia-smi shows 16 GB of memory, run a 14 GB model, monitor the temperature during a long generation, and listen to the fans. Ask for the invoice and the manufacturer's remaining warranty: they are often transferred with the card.

#2026 verdict

Keep it if
Your models fit in 16 GB. There is no reason to replace it before you need 24 GB.
Buy it used if
The price is significantly lower than that of a new 5070 Ti. Check the remaining warranty, temperature, and fan noise.
Aim for 24 GB if
You want Qwen 3.5 27B or more: no setting will fit 17 GB into 16 GB with context, and a 24 GB card saves you from buying again.

#Frequently asked questions

FAQ
Can the RTX 4070 Ti Super run Mistral Small 24B?+
Yes: the model weighs 14 GB according to Ollama; the card has 16 GB. That leaves about 2 GB for the context and system. The default 4,096-token context fits; for 16,000 tokens, enable Flash Attention and the K/V cache in q8_0 to reduce cache memory usage by about half.
What’s the difference between RTX 4070 Ti and the 4070 Ti Super for an LLM?+
The Ti Super has 16 GB of memory on a 256-bit bus, while the Ti has only 12 GB on a 192-bit bus. Bandwidth increases from 504 GB/s to 672 GB/s, according to Wikipedia. For an LLM, capacity matters most: the Ti Super can load 14 GB models that the Ti can't fit.
RTX 4070 Ti Super or RTX 5070 Ti?+
Both have 16 GB, so they support the same models. The 5070 Ti reads its memory at 896 GB/s versus 672 GB/s, or about one-third faster. It is faster and new with a warranty; the 4070 Ti Super mainly makes sense if its used price is significantly lower.
The RTX 4070 Ti Super or the 4080 for an LLM?+
Both have 16 GB of VRAM. The 4080 Super has 736 GB/s of bandwidth versus 672 GB/s, about 10% more, and the standard 4080 has 716.8 GB/s. For an LLM, the difference is modest: the used-price gap should decide, not the spec sheet.
What power supply do you need for a RTX 4070 Ti Super?+
NVIDIA specifies 285 W of total graphics power and a minimum 700 W power supply for the entire PC. A quality 750 W power supply provides comfortable headroom. Also check the connectors: the card requires two 8-pin PCIe cables or a PCIe Gen 5 cable rated for 300 W or more.
Can you run Qwen 3.5 27B on a RTX 4070 Ti Super?+
Not entirely in VRAM: the model weighs 17 GB according to Ollama, while the card has 16 GB. Part of it would be read from system RAM, which would significantly slow generation. For this model, aim for 24 GB of VRAM; with 16 GB, choose a model of 14 GB or less.
Did this guide help you?

Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.

Prices in euros (€) are French market prices including VAT, as checked by BestLLMfor. US prices differ: the Amazon buttons show the current US price.