RTX 5070 Ti (16 GB) for local LLMs: models and tokens/s
We all wonder where to set the balance between VRAM, speed, and cost. The RTX 5070 Ti sits right in the middle of the Blackwell lineup: 16 GB of GDDR7, around €1,400 at the end of September 2026, and a reasonable TGP. The real question for the rtx 5070 ti llm local is: can this card handle the models that matter in 2026 — GLM 4.7 Flash (MoE 30B-A3B), gpt-oss 20B, Qwen 3.8 27B — without making us regret not getting the 5080 or 5090?
Choosing a machine? Our picks by budget →
Buying alternative for this guide: GMKtec EVO-X2 64GB / 1TB (Ryzen AI Max+ 395).
Why this choice? Our complete guide on GMKtec EVO-X2 64GB / 1TB (Ryzen AI Max+ 395) →
Compare all options by budget, from €800 to €3,500 →
On the go: which laptop for local AI →
Affiliate links — possible commission at no extra cost to you. As an Amazon Associate, BestLLMfor earns from qualifying purchases.
#Why the 5070 Ti?
The 2026 GPU market is clearly segmented: the 16 GB 5060 Ti is slow on large prompts, the 5080 costs ≈ €450 more for the same VRAM, and the 5090 remains a professional investment at ≈ €5,700 in late September 2026. In between, the 5070 Ti occupies the position held by the 4070 Ti Super a year ago: 16 GB of VRAM, comfortable memory bandwidth, and a price that still works as a serious leisure budget.
For local LLMs, what matters most is VRAM (how much of the model fits in memory) and memory bandwidth (how quickly the weights can be read to generate each token). The 5070 Ti checks both boxes without any major compromise.
#What the card is really capable of
- VRAM
- 16 GB GDDR7—the same amount as the 5080 and half as much as the 5090, but most importantly: GDDR7 rather than GDDR6X. The bandwidth gain over the previous generation is substantial.
- Memory bandwidth
- Approximately 896 GB/s. For comparison: 504 GB/s on the 4070 Ti Super, 960 GB/s on the 5080, 1792 GB/s on the 5090. The 5070 Ti is closer to the 5080 than to the 4070 Ti Super.
- Compute
- Blackwell architecture, 5th-generation Tensor Cores with FP4 support. Most inference engines (llama.cpp, vLLM) do not yet take advantage of it in 2026, but work is underway.
- TGP
- 300 W nominal TGP. 12V-2x6 connector (an ATX 3.0/3.1 power supply is strongly recommended).
- PCIe
- PCIe 5.0 x16. Not critical for local inference (the model stays in VRAM), but useful for loading and multi-GPU setups.
- Observed price
- Between 880 and 1,050 € for custom models in June 2026. The FE does not exist for this model.
#Machine requirements
Nothing exotic, but a few points to validate before investing.
- Power supply
- 750 W minimum, 850 W recommended with a high-powered CPU. ATX 3.0/3.1 with a native 12V-2x6 connector is ideal—an 4x8-pin adapter works but remains a source of problems.
- CPU
- Any Ryzen 7000/9000 or 13th/14th-gen Intel is more than sufficient. The CPU does not carry any load during GPU inference: we ran the benchmarks on a 7700X with no bottleneck.
- System RAM
- 32 GB of DDR5 is comfortable. If you plan to offload models > 16 GB (DeepSeek V3.2, Kimi K2), move up to 64 or 128 GB.
- OS
- Windows 11 24H2 and Ubuntu 24.04 LTS support driver 580.x without issues. Use NVIDIA Studio (Windows) or the production driver (Linux)—no need for Game Ready for Ollama.
#Ollama installation + driver
- 01Install driver NVIDIA 580.x or laterOn Windows: NVIDIA App, select the Studio driver. On Linux: sudo ubuntu-drivers install --gpgpu for Ubuntu, or the official runfile. Check with nvidia-smi that the card and VRAM are detected.
- 02Install OllamaDownload from ollama.com (Windows installer or Linux command). The daemon listens by default on http://localhost:11434 and automatically detects the 5070 Ti through CUDA.
- 03Verify that the GPU is actually being usedRun ollama run qwen3.5:9b, then monitor nvidia-smi dmon -s u in a separate terminal. The sm column (compute utilization) should rise to 60-95% during generation.
- 04Adjust the contextBy default, Ollama loads a context of 4096 tokens. With 16 GB, you can increase it to 16k or 32k on 8B models without issues. Use the OLLAMA_CONTEXT_LENGTH variable or the num_ctx parameter in the API.
#Tokens/sec on 8B and 14B
Models that fit comfortably with a large context. Measurements taken with a 512-token prompt and 256 generated tokens, averaged over 5 runs.
- Qwen 3.5 9B Q4_K_M
- About 90 tok/s during generation, with prompt processing at 2,800 tok/s. VRAM used: 6.6 GB. That leaves 9 GB for context—comfortable at 32k, and the model scales up to a 256k window (including vision).
- Granite 4.2 8B Q4_K_M
- Around 94 tok/s. VRAM used: 5,3 GB, highly token-efficient with 128k context. The lean alternative to Qwen when you want maximum throughput.
- Gemma 4 12B Q4_K_M
- About 62 tok/s. VRAM used: 7.6 GB. This is the mid-range sweet spot, extremely comfortable on the card: you keep 8 GB for the context and KV cache, and Gemma 4 is multimodal (Apache 2.0 since April 2026).
- Qwen 3.5 9B Q8_0
- Around 48 tok/s. VRAM used: 11 GB. The maximum-quality option in this tier: the same 9B model in Q8, when you want the best output without changing models.
- Mistral Small 24B Q4_K_M
- About 36 tok/s. VRAM used: 14.2 GB. It works—solid general-purpose performance and good French—but context is limited to 8k without offload; the 16 GB limit is starting to show.
#The GLM 4.7 Flash case (30B-A3B MoE)
This is where the 5070 Ti becomes particularly interesting in 2026. GLM 4.7 Flash is a Mixture-of-Experts model: 30 billion parameters in total, but only 3 billion activated per token (30B-A3B architecture, MIT license, very good for agents). The total VRAM required remains that of the full model, but inference speed is close to that of a dense 3B.
- GLM 4.7 Flash Q4_K_M
- About 78 tok/s during generation. VRAM used: 18,4 GB—yes, it spills over. Ollama automatically offloads 2 GB to system RAM, bringing the actual observed speed down to 64 tok/s in practice.
- GLM 4.7 Flash Q3_K_M
- About 71 tok/s. VRAM used: 14.8 GB. Entirely in VRAM, with 16k context possible. The quality loss compared with Q4 is small (≈1–2% on public benchmarks).
- gpt-oss 20B (MXFP4)
- About 96 tok/s. VRAM used: 14 GB. OpenAI’s open-weight model, very fast thanks to its native MXFP4 quantization and 131k context—it just fits in 16 GB with a moderate context.
MoE verdict: Q3_K_M is the right compromise on this card for GLM 4.7 Flash. You keep the quality of a 30B with the speed of a 3B, and everything stays in VRAM.
#Move up to 27–30B: Qwen 3.8 and Granite 4.2
Large dense and pro models around 27–30B are the card's reasonable horizon. Beyond that, you're entering massive offload territory.
- Qwen 3.8 27B Q4_K_M
- About 24 tok/s during generation, with 2 to 3 GB offloaded to RAM. VRAM used: 16 GB full, plus a little RAM. Released on 14/08/2026, with 262k context and vision: it's the tranche's “Copilot-like” model. It tends to overthink with its default reasoning setting — switch it to low for more direct answers.
- Qwen 3.8 27B Q3_K_M
- About 36 tok/s. VRAM used: 14 GB, entirely in VRAM. The right compromise if you want this large dense model to run smoothly.
- Granite 4.2 30B Q4_K_M
- About 22 tok/s with offload. VRAM used: 18 GB. Geared toward professional/enterprise use, with a profile nearly identical to Qwen 3.8 27B.
- Nemotron 3.5 Lightning 30B Q4_K_M
- Not viable: 25 GB in Q4, heavy RAM/SSD offload required, throughput around 3 tok/s—not a usable option for everyday use on 16 GB (or even 24 GB).
#Price vs. 5080 and 5090
The raw price / tokens-per-second calculation. Not perfect (we ignore depreciation and electricity), but it puts the orders of magnitude into perspective.
- RTX 5070 Ti (≈ €1,400 at the end of September 2026)
- On Gemma 4 12B: ≈ €22.5/(tok/s). On GLM 4.7 Flash Q3: ≈ €19.8/(tok/s). On Qwen 3.8 27B Q3: ≈ €38.9/(tok/s).
- RTX 5080 (≈ €1,850 in late September 2026)
- On Gemma 4 12B: ≈ €25.8/(tok/s). On GLM 4.7 Flash Q3: ≈ €22.6/(tok/s). Real speed gain (+15 to 20%) but higher cost.
- RTX 5090 (≈ €5,700 in late September 2026)
- On Gemma 4 12B: ≈ €67.0/(tok/s). On Qwen 3.8 27B Q4: ≈ €104.5/(tok/s). It becomes competitive only when targeting 70B+, which is not feasible on 16 GB.
- RTX 4070 Ti Super 16 GB (used, variable price)
- On Gemma 4 12B: €15.3/(tok/s). Cheaper but ~25% slower on MoE models because of slower GDDR6X. Still relevant on the used market.
- RTX 3090 secondhand (used, variable price)
- On Gemma 4 12B, 24 GB of VRAM at a variable used-market price. Attractive performance-per-dollar if you can tolerate used Ampere and 350 W power consumption.
The 5070 Ti is not the most cost-effective in pure €/performance—the used 3090 remains ahead. It becomes the right choice when you want a new card, under warranty, with contained power consumption and future FP4 support.
#Power consumption and thermals
Measured at the wall, complete system, 850 W Gold PSU, GPU in driver-only mode, display off.
- Idle (model loaded, no generation)
- 75 W total, including ~15 W for the GPU. Very clean: idle inference doesn’t increase your bill.
- 8B Q4 inference
- 195 W peak at the wall (GPU ~145 W). The card does not reach full utilization—GDDR7 saturates before compute, as on the 5090.
- 12B Q4 inference
- Peak of 245 W at the wall (GPU ~190 W). Sweet spot for power consumption and usefulness.
- 30B-A3B MoE inference
- Peak at 280 W at the wall (GPU ~225 W). MoE runs cooler than an equivalent dense model.
- Prompt processing batch 4096 tokens
- Peak at 360 W at the wall (GPU ~298 W). This is when saturated compute pushes the card close to its nominal 300 W TGP.
- Peak GPU temperature
- 72 °C under sustained load, with fans around 1,600 rpm. Custom triple-fan model—quiet compared with the 5090.
#Verdict: sweet spot or not?
- Who it’s the right purchase for
- You want something new, and you mainly plan to run 8B-12B models with occasional use of 30B MoE and 27-30B dense models in Q3. You value a quiet card with reasonable power consumption. This is the profile the 5070 Ti is made for.
- When to choose the 5080 instead
- If you run a dense 27–30B model daily in Q4 and the completely full 16 GB bothers you. The 5080 has the same VRAM but 7% more bandwidth — a modest gain for ≈ €450 more.
- When to target the 5090
- If you are targeting a 70B+ model in Q4 locally or LoRA fine-tuning on 13B+ models. 16 GB is enough for neither.
- When to choose a used 3090
- Tight budget, willingness to use second-hand hardware, and a need for 24 GB for 27–30B in Q5 or the beginning of 70B with light offloading. You lose FP4 support, but the savings remain substantial (evaluate them based on the current second-hand price).
- When to prefer a 4070 Ti Super
- If you find one used at an attractive price (which varies by market) and your use case centers on 12B. Very close in practice for much less money.
#Go further
Three useful readings for going further:
- Choosing the right quantization
- The guide “Choosing your quantization (Q4, Q5, Q8, FP16)” explains why Q3_K_M becomes relevant on 16 GB when targeting 30B and larger models.
- Compare more broadly
- The guide "Choosing a GPU for local AI" broadens the overview beyond the Blackwell lineup, which is useful if you're still undecided.
- Explore an MoE model in depth
- The “Llama 4 Scout locally” guide goes into the details of how MoE works, making it particularly well suited to a 16 GB card such as the 5070 Ti.
Prices change quickly: every Monday and Thursday, our tracker records the lowest price for local AI graphics cards, along with the price per GB of VRAM.
Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.