Which LLM on RTX 5070 Ti (16 GB) ?
The RTX 5070 Ti (February 2025) is probably the best 2025 card for mainstream local AI. With 16 GB of GDDR7 VRAM, 896 GB/s of bandwidth, and 8,960 CUDA cores, it runs Granite 4.2 8B at ~50 tokens/second and Mistral Small 24B Q4 comfortably, for ≈ €1,400 in late September 2026 — about €450 less than a 5080. This guide explains why it's the 2026 sweet spot and which models to install.
Choosing a machine? Our picks by budget →
Buying alternative for this guide: GMKtec EVO-X2 64GB / 1TB (Ryzen AI Max+ 395).
A mini PC is a complete machine: check the available memory and engine compatibility. It does not replace macOS/MLX or CUDA.
Why this choice? Our complete guide on GMKtec EVO-X2 64GB / 1TB (Ryzen AI Max+ 395) →
Compare all options by budget, from €800 to €3,500 →
On the go: which laptop for local AI →
Affiliate links — possible commission at no extra cost to you. As an Amazon Associate, BestLLMfor earns from qualifying purchases.
#The RTX 5070 Ti, the 16 GB sweet spot
- Architecture
- Blackwell GB203 (same die as the 5080, with more CUs disabled).
- VRAM
- 16 GB GDDR7 at 28 Gbps, 256-bit bus. Bandwidth: 896 GB/s.
- CUDA cores
- 8 960. 5th-gen Tensor Cores, native FP4 support.
- TDP
- 300 W. A 750 W PSU is recommended.
- MSRP
- €884 at launch. ≈ €1,400 in late September 2026 (custom).
#1. Compatible LLM models
| Model | Quant | Model VRAM | Tokens/sec | Usage |
|---|---|---|---|---|
| Granite 4.2 8B | Q4_K_M | 4.6 GB | 45-55 | Instant chat |
| Gemma 4 12B | Q4_K_M | 7 GB | 25-31 | Daily chat/RAG |
| Qwen 3.5 9B | Q4_K_M | 6 GB | 25-31 | VS Code coding assistant |
| gpt-oss 20B | Q4_K_M | 13 GB | 50-61 | Chat + RAG + FR |
| Qwen 3.5 9B | Q8_0 | 11 GB | 18-22 | Maximum quality / reasoning |
| Devstral Small 2 24B | Q4_K_M | 14 GB | 14-17 | Chat quality |
| Mistral Small 24B | Q4_K_M | 14 GB | 28-35 | Premium all-rounder |
| Qwen 3.6 35B-A3B | Q3_K_M | ~15 GB | 22-28 | Advanced RAG, fast MoE (light offload) |
#2. Installation (Ollama + CUDA)
- 01NVIDIA 570+ driversRequired for Blackwell. GeForce Experience on Windows, apt install nvidia-driver-570 on Ubuntu.
- 02OllamaStandard installer from ollama.com, CUDA included. ollama run granite4.2:8b for an immediate test.
- 03Verificationnvidia-smi should show RTX 5070 Ti, 16376 MiB. ollama ps during a chat: PROCESSOR 100% GPU.
#3. Speed by quantization (orders of magnitude)
| Model | Q4_K_M | Q5_K_M | Q6_K | Q8_0 |
|---|---|---|---|---|
| Granite 4.2 8B | 50 t/s | 46 t/s | 42 t/s | 34 t/s |
| Qwen 3.5 9B | 30 t/s | 28 t/s | 26 t/s | 22 t/s |
| Gemma 4 12B | 28 t/s | 26 t/s | 24 t/s | 20 t/s |
| Mistral Small 24B | 32 t/s | 28 t/s | — | — |
| Devstral 24B | 16 t/s | 14 t/s | — | — |
#4. Blackwell optimizations
- Flash Attention 3
- Enabled by default, +10–14% on 4k+ contexts. Visible in the llama.cpp logs.
- KV cache Q8
- OLLAMA_KV_CACHE_TYPE=q8_0. Enables 32k context on Granite 4.2 8B in ~5-6 GB instead of 8.
- NVFP4 future-proof
- Hardware present but underused in 2026. In 12–18 months, Ollama will support FP4 → +40% speed expected.
- Undervolt
- 250 W instead of 300 W (via MSI Afterburner -80 mV): -1% LLM performance, -5 °C, quieter case.
#5. 5070 Ti vs. 4070 Ti Super vs. 5080
| Criterion | 5070 Ti | 4070 Ti Super | 5080 |
|---|---|---|---|
| VRAM | 16 GB GDDR7 | 16 GB GDDR6X | 16 GB GDDR7 |
| Bandwidth | 896 GB/s | 672 GB/s | 960 GB/s |
| Granite 4.2 8B Q4 | 50 t/s | 42 t/s | 56 t/s |
| Gemma 4 12B Q4 | 28 t/s | 23 t/s | 32 t/s |
| 2026 price | ≈ €1,400 at the end of September 2026 | used, price varies | ≈ €1,850 at the end of September 2026 |
| LLM performance/€ | ★★★★★ | ★★★★ | ★★★ |
#2026 buying verdict
- Best mainstream LLM choice for 2026
- For anyone with a GPU budget of ≈ €1,400 at the end of September 2026 who wants serious LLM performance without stepping up to a 5090. It crushes the 5080 on performance per euro.
- Upgrade from 4070 Ti Super?
- +20% LLM speed, +33% bandwidth. Not essential if it runs—but a defensible argument if you sell the 4070 Ti Super at a variable used price.
- vs used RTX 4090
- The used 4090 (variable used-market price) retains the advantage for 70B models (24 GB VRAM). If you don't need that, the 5070 Ti delivers nearly equivalent 14B LLM performance for considerably less.
#Frequently asked questions
RTX 5070 Ti or RTX 5080 for LLMs?+
Can the RTX 5070 Ti 16 GB run Mistral Small 24B?+
What’s the difference between RTX 4070 Ti Super and RTX 5070 Ti?+
Can you run QLoRA on RTX 5070 Ti 16 GB?+
Is a 650 W PSU enough for a RTX 5070 Ti?+
Does RTX 5070 Ti support FP4 / NVFP4?+
I'm coming from a RTX 3080 10 GB — is the upgrade worth it?+
Prices change quickly: every Monday and Thursday, our tracker records the lowest price for local AI graphics cards, along with the price per GB of VRAM.
Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.