Intermediate 11 minRTX 50

Which LLM on RTX 5070 Ti (16 GB) ?

The RTX 5070 Ti (February 2025) is probably the best 2025 card for mainstream local AI. With 16 GB of GDDR7 VRAM, 896 GB/s of bandwidth, and 8,960 CUDA cores, it runs Granite 4.2 8B at ~50 tokens/second and Mistral Small 24B Q4 comfortably, for ≈ €1,400 in late September 2026 — about €450 less than a 5080. This guide explains why it's the 2026 sweet spot and which models to install.

Choosing a machine? Our picks by budget →

By Mohamed Meguedmi·Update 2026-08-27·Tested on Windows, macOS, and Linux
Recommended hardware

Buying alternative for this guide: GMKtec EVO-X2 64GB / 1TB (Ryzen AI Max+ 395).

A mini PC is a complete machine: check the available memory and engine compatibility. It does not replace macOS/MLX or CUDA.

Why this choice? Our complete guide on GMKtec EVO-X2 64GB / 1TB (Ryzen AI Max+ 395) →

Compare all options by budget, from €800 to €3,500 →

On the go: which laptop for local AI →

Affiliate links — possible commission at no extra cost to you. As an Amazon Associate, BestLLMfor earns from qualifying purchases.

#The RTX 5070 Ti, the 16 GB sweet spot

Architecture
Blackwell GB203 (same die as the 5080, with more CUs disabled).
VRAM
16 GB GDDR7 at 28 Gbps, 256-bit bus. Bandwidth: 896 GB/s.
CUDA cores
8 960. 5th-gen Tensor Cores, native FP4 support.
TDP
300 W. A 750 W PSU is recommended.
MSRP
€884 at launch. ≈ €1,400 in late September 2026 (custom).
→
Why it’s the sweet spot
The same 16 GB of GDDR7 as a 5080, with 93% of its bandwidth, but ≈ €450 cheaper. For local AI, where VRAM and memory bandwidth dominate, the CU gap costs only 8-12% in speed—not a 30% price difference.

#1. Compatible LLM models

Recommended models — RTX 5070 Ti 16 GB
ModelQuantModel VRAMTokens/secUsage
Granite 4.2 8BQ4_K_M4.6 GB45-55Instant chat
Gemma 4 12BQ4_K_M7 GB25-31Daily chat/RAG
Qwen 3.5 9BQ4_K_M6 GB25-31VS Code coding assistant
gpt-oss 20BQ4_K_M13 GB50-61Chat + RAG + FR
Qwen 3.5 9BQ8_011 GB18-22Maximum quality / reasoning
Devstral Small 2 24BQ4_K_M14 GB14-17Chat quality
Mistral Small 24BQ4_K_M14 GB28-35Premium all-rounder
Qwen 3.6 35B-A3BQ3_K_M~15 GB22-28Advanced RAG, fast MoE (light offload)

#2. Installation (Ollama + CUDA)

  1. 01
    NVIDIA 570+ drivers
    Required for Blackwell. GeForce Experience on Windows, apt install nvidia-driver-570 on Ubuntu.
  2. 02
    Ollama
    Standard installer from ollama.com, CUDA included. ollama run granite4.2:8b for an immediate test.
  3. 03
    Verification
    nvidia-smi should show RTX 5070 Ti, 16376 MiB. ollama ps during a chat: PROCESSOR 100% GPU.
Complete Linux setup
sudo apt update
sudo apt install nvidia-driver-570
sudo reboot

curl -fsSL https://ollama.com/install.sh | sh
ollama run mistral-small

#3. Speed by quantization (orders of magnitude)

RTX 5070 Ti — tokens/sec (rough figures), recent Ollama, batch 1, Ryzen 9 7900X
ModelQ4_K_MQ5_K_MQ6_KQ8_0
Granite 4.2 8B50 t/s46 t/s42 t/s34 t/s
Qwen 3.5 9B30 t/s28 t/s26 t/s22 t/s
Gemma 4 12B28 t/s26 t/s24 t/s20 t/s
Mistral Small 24B32 t/s28 t/s——
Devstral 24B16 t/s14 t/s——

#4. Blackwell optimizations

Flash Attention 3
Enabled by default, +10–14% on 4k+ contexts. Visible in the llama.cpp logs.
KV cache Q8
OLLAMA_KV_CACHE_TYPE=q8_0. Enables 32k context on Granite 4.2 8B in ~5-6 GB instead of 8.
NVFP4 future-proof
Hardware present but underused in 2026. In 12–18 months, Ollama will support FP4 → +40% speed expected.
Undervolt
250 W instead of 300 W (via MSI Afterburner -80 mV): -1% LLM performance, -5 °C, quieter case.

#5. 5070 Ti vs. 4070 Ti Super vs. 5080

The three premium 16 GB cards
Criterion5070 Ti4070 Ti Super5080
VRAM16 GB GDDR716 GB GDDR6X16 GB GDDR7
Bandwidth896 GB/s672 GB/s960 GB/s
Granite 4.2 8B Q450 t/s42 t/s56 t/s
Gemma 4 12B Q428 t/s23 t/s32 t/s
2026 price≈ €1,400 at the end of September 2026used, price varies≈ €1,850 at the end of September 2026
LLM performance/€★★★★★★★★★★★★

#2026 buying verdict

Best mainstream LLM choice for 2026
For anyone with a GPU budget of ≈ €1,400 at the end of September 2026 who wants serious LLM performance without stepping up to a 5090. It crushes the 5080 on performance per euro.
Upgrade from 4070 Ti Super?
+20% LLM speed, +33% bandwidth. Not essential if it runs—but a defensible argument if you sell the 4070 Ti Super at a variable used price.
vs used RTX 4090
The used 4090 (variable used-market price) retains the advantage for 70B models (24 GB VRAM). If you don't need that, the 5070 Ti delivers nearly equivalent 14B LLM performance for considerably less.

#Frequently asked questions

RTX 5070 Ti or RTX 5080 for LLMs?+
Go with the 5070 Ti without hesitation if your budget is < €1,700. You lose ~10% speed on Granite 4.2 8B (50 vs 56 tok/s) while saving ≈ €450. Same VRAM (16 GB), same generation, same GDDR7. The 5080 makes sense for mixed 4K gaming + LLM use when you want the best.
Can the RTX 5070 Ti 16 GB run Mistral Small 24B?+
Yes, in Q4_K_M (14 GB) at 28–35 tokens/second. This versatile model makes the best use of its 16 GB. In Q5_K_M (17 GB), there is some offloading, but it remains usable at 20–25 tok/s.
What’s the difference between RTX 4070 Ti Super and RTX 5070 Ti?+
Same VRAM (16 GB), but GDDR7 instead of GDDR6X → +33% bandwidth. LLM result: +20–25% tokens/second. Similar launch price, but the 4070 Ti Super is cheaper used (~€750 vs. €950 new).
Can you run QLoRA on RTX 5070 Ti 16 GB?+
Yes, comfortably with 7B-8B (10-12 GB required with batch 4, context 2048). For 14B QLoRA, it is tight but workable with unsloth. For 24B QLoRA: impossible; you need 20 GB+.
Is a 650 W PSU enough for a RTX 5070 Ti?+
Yes with a CPU ≤ 150 W (Ryzen 7600/7700, Core i5). 300 W GPU TDP + 150 W CPU + 50 W system = 500 W nominal; 650 W provides headroom. If CPU ≥ 180 W (9950X, 14900K), target 750 W.
Does RTX 5070 Ti support FP4 / NVFP4?+
Yes, the 5th-gen Blackwell Tensor Cores natively handle FP4. In 2026, Ollama and llama.cpp still don't support it— that will come. With TensorRT-LLM 0.18+ or vLLM 0.7+, we already observe +30-40% vs Q4_K_M on models quantized in NVFP4.
I'm coming from a RTX 3080 10 GB — is the upgrade worth it?+
Yes, a huge generational leap. +60% VRAM (10 → 16 GB — opens up 14B–24B models), +3× the bandwidth (760 → 896 GB/s after GDDR7 inflation), +2× the raw speed. If you use an LLM every day, this is the most cost-effective upgrade of 2026.

Prices change quickly: every Monday and Thursday, our tracker records the lowest price for local AI graphics cards, along with the price per GB of VRAM.

Did this guide help you?

Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.