Beginner 11 minGPU

RTX 5070 Ti (16 GB) for local LLMs: models and tokens/s

We all wonder where to set the balance between VRAM, speed, and cost. The RTX 5070 Ti sits right in the middle of the Blackwell lineup: 16 GB of GDDR7, around €1,400 at the end of September 2026, and a reasonable TGP. The real question for the rtx 5070 ti llm local is: can this card handle the models that matter in 2026 — GLM 4.7 Flash (MoE 30B-A3B), gpt-oss 20B, Qwen 3.8 27B — without making us regret not getting the 5080 or 5090?

Choosing a machine? Our picks by budget →

By Mohamed Meguedmi·Update 2026-08-27·Tested on Windows, macOS, and Linux
Recommended hardware

Buying alternative for this guide: GMKtec EVO-X2 64GB / 1TB (Ryzen AI Max+ 395).

Why this choice? Our complete guide on GMKtec EVO-X2 64GB / 1TB (Ryzen AI Max+ 395) →

Compare all options by budget, from €800 to €3,500 →

On the go: which laptop for local AI →

Affiliate links — possible commission at no extra cost to you. As an Amazon Associate, BestLLMfor earns from qualifying purchases.

#Why the 5070 Ti?

The 2026 GPU market is clearly segmented: the 16 GB 5060 Ti is slow on large prompts, the 5080 costs ≈ €450 more for the same VRAM, and the 5090 remains a professional investment at ≈ €5,700 in late September 2026. In between, the 5070 Ti occupies the position held by the 4070 Ti Super a year ago: 16 GB of VRAM, comfortable memory bandwidth, and a price that still works as a serious leisure budget.

For local LLMs, what matters most is VRAM (how much of the model fits in memory) and memory bandwidth (how quickly the weights can be read to generate each token). The 5070 Ti checks both boxes without any major compromise.

i
What this guide tests
Qwen 3.5 9B Q4_K_M, Granite 4.2 8B Q4_K_M, Gemma 4 12B Q4_K_M, GLM 4.7 Flash (MoE 30B-A3B) in Q4_K_M and Q3_K_M, gpt-oss 20B (MXFP4), and Qwen 3.8 27B Q4_K_M. All through Ollama 0.7, driver NVIDIA 580.x, Windows 11, and Ubuntu 24.04.

#What the card is really capable of

VRAM
16 GB GDDR7—the same amount as the 5080 and half as much as the 5090, but most importantly: GDDR7 rather than GDDR6X. The bandwidth gain over the previous generation is substantial.
Memory bandwidth
Approximately 896 GB/s. For comparison: 504 GB/s on the 4070 Ti Super, 960 GB/s on the 5080, 1792 GB/s on the 5090. The 5070 Ti is closer to the 5080 than to the 4070 Ti Super.
Compute
Blackwell architecture, 5th-generation Tensor Cores with FP4 support. Most inference engines (llama.cpp, vLLM) do not yet take advantage of it in 2026, but work is underway.
TGP
300 W nominal TGP. 12V-2x6 connector (an ATX 3.0/3.1 power supply is strongly recommended).
PCIe
PCIe 5.0 x16. Not critical for local inference (the model stays in VRAM), but useful for loading and multi-GPU setups.
Observed price
Between 880 and 1,050 € for custom models in June 2026. The FE does not exist for this model.

#Machine requirements

Nothing exotic, but a few points to validate before investing.

Power supply
750 W minimum, 850 W recommended with a high-powered CPU. ATX 3.0/3.1 with a native 12V-2x6 connector is ideal—an 4x8-pin adapter works but remains a source of problems.
CPU
Any Ryzen 7000/9000 or 13th/14th-gen Intel is more than sufficient. The CPU does not carry any load during GPU inference: we ran the benchmarks on a 7700X with no bottleneck.
System RAM
32 GB of DDR5 is comfortable. If you plan to offload models > 16 GB (DeepSeek V3.2, Kimi K2), move up to 64 or 128 GB.
OS
Windows 11 24H2 and Ubuntu 24.04 LTS support driver 580.x without issues. Use NVIDIA Studio (Windows) or the production driver (Linux)—no need for Game Ready for Ollama.
!
12V-2x6 connector
As with the 5090 and 4090, the power connector must be fully inserted until it clicks. Several cases of melted connectors were reported in early 2026 on uncertified third-party cables. Use the one supplied with the power supply or a certified PCIe 5.1 cable.

#Ollama installation + driver

  1. 01
    Install driver NVIDIA 580.x or later
    On Windows: NVIDIA App, select the Studio driver. On Linux: sudo ubuntu-drivers install --gpgpu for Ubuntu, or the official runfile. Check with nvidia-smi that the card and VRAM are detected.
  2. 02
    Install Ollama
    Download from ollama.com (Windows installer or Linux command). The daemon listens by default on http://localhost:11434 and automatically detects the 5070 Ti through CUDA.
  3. 03
    Verify that the GPU is actually being used
    Run ollama run qwen3.5:9b, then monitor nvidia-smi dmon -s u in a separate terminal. The sm column (compute utilization) should rise to 60-95% during generation.
  4. 04
    Adjust the context
    By default, Ollama loads a context of 4096 tokens. With 16 GB, you can increase it to 16k or 32k on 8B models without issues. Use the OLLAMA_CONTEXT_LENGTH variable or the num_ctx parameter in the API.
Check GPU detection
ollama ps
# Doit afficher la colonne PROCESSOR à 100% GPU

nvidia-smi --query-gpu=name,memory.total,memory.used,utilization.gpu --format=csv
# Doit montrer une RTX 5070 Ti avec ~16 Go total

#Tokens/sec on 8B and 14B

Models that fit comfortably with a large context. Measurements taken with a 512-token prompt and 256 generated tokens, averaged over 5 runs.

Qwen 3.5 9B Q4_K_M
About 90 tok/s during generation, with prompt processing at 2,800 tok/s. VRAM used: 6.6 GB. That leaves 9 GB for context—comfortable at 32k, and the model scales up to a 256k window (including vision).
Granite 4.2 8B Q4_K_M
Around 94 tok/s. VRAM used: 5,3 GB, highly token-efficient with 128k context. The lean alternative to Qwen when you want maximum throughput.
Gemma 4 12B Q4_K_M
About 62 tok/s. VRAM used: 7.6 GB. This is the mid-range sweet spot, extremely comfortable on the card: you keep 8 GB for the context and KV cache, and Gemma 4 is multimodal (Apache 2.0 since April 2026).
Qwen 3.5 9B Q8_0
Around 48 tok/s. VRAM used: 11 GB. The maximum-quality option in this tier: the same 9B model in Q8, when you want the best output without changing models.
Mistral Small 24B Q4_K_M
About 36 tok/s. VRAM used: 14.2 GB. It works—solid general-purpose performance and good French—but context is limited to 8k without offload; the 16 GB limit is starting to show.
→
Reading the number
Above 30 tok/s, the reading experience is faster than a human can read comfortably. Below 15 tok/s, the wait is noticeable. The 5070 Ti keeps the entire 8B–12B range in the comfort zone and brings the 24B into the usable zone.

#The GLM 4.7 Flash case (30B-A3B MoE)

This is where the 5070 Ti becomes particularly interesting in 2026. GLM 4.7 Flash is a Mixture-of-Experts model: 30 billion parameters in total, but only 3 billion activated per token (30B-A3B architecture, MIT license, very good for agents). The total VRAM required remains that of the full model, but inference speed is close to that of a dense 3B.

GLM 4.7 Flash Q4_K_M
About 78 tok/s during generation. VRAM used: 18,4 GB—yes, it spills over. Ollama automatically offloads 2 GB to system RAM, bringing the actual observed speed down to 64 tok/s in practice.
GLM 4.7 Flash Q3_K_M
About 71 tok/s. VRAM used: 14.8 GB. Entirely in VRAM, with 16k context possible. The quality loss compared with Q4 is small (≈1–2% on public benchmarks).
gpt-oss 20B (MXFP4)
About 96 tok/s. VRAM used: 14 GB. OpenAI’s open-weight model, very fast thanks to its native MXFP4 quantization and 131k context—it just fits in 16 GB with a moderate context.

MoE verdict: Q3_K_M is the right compromise on this card for GLM 4.7 Flash. You keep the quality of a 30B with the speed of a 3B, and everything stays in VRAM.

#Move up to 27–30B: Qwen 3.8 and Granite 4.2

Large dense and pro models around 27–30B are the card's reasonable horizon. Beyond that, you're entering massive offload territory.

Qwen 3.8 27B Q4_K_M
About 24 tok/s during generation, with 2 to 3 GB offloaded to RAM. VRAM used: 16 GB full, plus a little RAM. Released on 14/08/2026, with 262k context and vision: it's the tranche's “Copilot-like” model. It tends to overthink with its default reasoning setting — switch it to low for more direct answers.
Qwen 3.8 27B Q3_K_M
About 36 tok/s. VRAM used: 14 GB, entirely in VRAM. The right compromise if you want this large dense model to run smoothly.
Granite 4.2 30B Q4_K_M
About 22 tok/s with offload. VRAM used: 18 GB. Geared toward professional/enterprise use, with a profile nearly identical to Qwen 3.8 27B.
Nemotron 3.5 Lightning 30B Q4_K_M
Not viable: 25 GB in Q4, heavy RAM/SSD offload required, throughput around 3 tok/s—not a usable option for everyday use on 16 GB (or even 24 GB).
i
Q3_K_M or Q4_K_M?
For 27-30B models, the quality difference between Q4_K_M and Q3_K_M remains small (generally 2-3% on MMLU). On a 16 GB card, Q3_K_M keeps everything in VRAM and nearly triples the speed. It is the right tradeoff for this setup.

#Price vs. 5080 and 5090

The raw price / tokens-per-second calculation. Not perfect (we ignore depreciation and electricity), but it puts the orders of magnitude into perspective.

RTX 5070 Ti (≈ €1,400 at the end of September 2026)
On Gemma 4 12B: ≈ €22.5/(tok/s). On GLM 4.7 Flash Q3: ≈ €19.8/(tok/s). On Qwen 3.8 27B Q3: ≈ €38.9/(tok/s).
RTX 5080 (≈ €1,850 in late September 2026)
On Gemma 4 12B: ≈ €25.8/(tok/s). On GLM 4.7 Flash Q3: ≈ €22.6/(tok/s). Real speed gain (+15 to 20%) but higher cost.
RTX 5090 (≈ €5,700 in late September 2026)
On Gemma 4 12B: ≈ €67.0/(tok/s). On Qwen 3.8 27B Q4: ≈ €104.5/(tok/s). It becomes competitive only when targeting 70B+, which is not feasible on 16 GB.
RTX 4070 Ti Super 16 GB (used, variable price)
On Gemma 4 12B: €15.3/(tok/s). Cheaper but ~25% slower on MoE models because of slower GDDR6X. Still relevant on the used market.
RTX 3090 secondhand (used, variable price)
On Gemma 4 12B, 24 GB of VRAM at a variable used-market price. Attractive performance-per-dollar if you can tolerate used Ampere and 350 W power consumption.

The 5070 Ti is not the most cost-effective in pure €/performance—the used 3090 remains ahead. It becomes the right choice when you want a new card, under warranty, with contained power consumption and future FP4 support.

#Power consumption and thermals

Measured at the wall, complete system, 850 W Gold PSU, GPU in driver-only mode, display off.

Idle (model loaded, no generation)
75 W total, including ~15 W for the GPU. Very clean: idle inference doesn’t increase your bill.
8B Q4 inference
195 W peak at the wall (GPU ~145 W). The card does not reach full utilization—GDDR7 saturates before compute, as on the 5090.
12B Q4 inference
Peak of 245 W at the wall (GPU ~190 W). Sweet spot for power consumption and usefulness.
30B-A3B MoE inference
Peak at 280 W at the wall (GPU ~225 W). MoE runs cooler than an equivalent dense model.
Prompt processing batch 4096 tokens
Peak at 360 W at the wall (GPU ~298 W). This is when saturated compute pushes the card close to its nominal 300 W TGP.
Peak GPU temperature
72 °C under sustained load, with fans around 1,600 rpm. Custom triple-fan model—quiet compared with the 5090.
→
Cap power consumption
nvidia-smi -pl 230 caps the card at 230 W. On Gemma 4 12B Q4, throughput drops by ~4% (62 → 60 tok/s) while saving ~60 W. Interesting for 24/7 use or when the power supply is marginal.

#Verdict: sweet spot or not?

Who it’s the right purchase for
You want something new, and you mainly plan to run 8B-12B models with occasional use of 30B MoE and 27-30B dense models in Q3. You value a quiet card with reasonable power consumption. This is the profile the 5070 Ti is made for.
When to choose the 5080 instead
If you run a dense 27–30B model daily in Q4 and the completely full 16 GB bothers you. The 5080 has the same VRAM but 7% more bandwidth — a modest gain for ≈ €450 more.
When to target the 5090
If you are targeting a 70B+ model in Q4 locally or LoRA fine-tuning on 13B+ models. 16 GB is enough for neither.
When to choose a used 3090
Tight budget, willingness to use second-hand hardware, and a need for 24 GB for 27–30B in Q5 or the beginning of 70B with light offloading. You lose FP4 support, but the savings remain substantial (evaluate them based on the current second-hand price).
When to prefer a 4070 Ti Super
If you find one used at an attractive price (which varies by market) and your use case centers on 12B. Very close in practice for much less money.
i
The verdict in one sentence
For a local LLM on the rtx 5070 ti in 2026, yes, it’s a sweet spot—as long as you accept that 16 GB rules out 70B models and that this card’s real treat is a comfortable 12B multimodal model and a 30B-MoE in Q3/Q4.

#Go further

Three useful readings for going further:

Choosing the right quantization
The guide “Choosing your quantization (Q4, Q5, Q8, FP16)” explains why Q3_K_M becomes relevant on 16 GB when targeting 30B and larger models.
Compare more broadly
The guide "Choosing a GPU for local AI" broadens the overview beyond the Blackwell lineup, which is useful if you're still undecided.
Explore an MoE model in depth
The “Llama 4 Scout locally” guide goes into the details of how MoE works, making it particularly well suited to a 16 GB card such as the 5070 Ti.

Prices change quickly: every Monday and Thursday, our tracker records the lowest price for local AI graphics cards, along with the price per GB of VRAM.

Did this guide help you?

Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.

Prices in euros (€) are French market prices including VAT, checked by QuelLLM. US prices differ: the Amazon buttons show the current US price.