RTX 5090 and local LLM: memory, models, benchmark
RTX 5090 and local LLMs: understand what 32 GB of VRAM enables, choose an identifiable model, and prepare a reproducible benchmark. Manufacturer specifications and estimates are distinguished from measurements.
Choosing a machine? Our picks by budget →
Good value for local AI: a GMKtec EVO-X2 64GB / 1TB (Ryzen AI Max+ 395).
A mini PC is a complete machine: check the available memory and engine compatibility. It does not replace macOS/MLX or CUDA.
Why this choice? Our complete guide on GMKtec EVO-X2 64GB / 1TB (Ryzen AI Max+ 395) →
Compare all options by budget, from €800 to €3,500 →
Small budget: RTX 5060 · Large models: RTX 5090 · Mac Studio.
On the go: which laptop for local AI →
Affiliate links — possible commission at no extra cost to you. As an Amazon Associate, BestLLMfor earns from qualifying purchases.
#What has been verified—and what hasn't
The NVIDIA specifications list 32 GB of GDDR7 for the RTX 5090, compared with 24 GB for the RTX 4090. This difference increases the memory budget; it does not prove throughput or response quality. Manufacturer source: https://www.nvidia.com/en-us/geforce/graphics-cards/50-series/rtx-5090/
#Weights, context cache, and runtime headroom
The quantized file, KV cache, and runtime buffers must all fit together. A 32 GB file will not fit entirely in 32 GB of VRAM with a useful context. Reducing the context, choosing a more compact quantization, or accepting CPU offload are distinct trade-offs.
Theoretical estimate for weights alone: total parameters × bits per weight / 8, before metadata and quantization overhead. This calculation is neither a measure of VRAM used nor a loading guarantee. The number of active parameters in an MoE never replaces the total number when estimating the weights that must be stored.
#Llama 4 Scout: the essential fix
The official identity is meta-llama/Llama-4-Scout-17B-16E-Instruct: 17 billion active parameters, 109 billion total, and 16 experts. The former designation “17B-A2B” and the claimed 17 GB footprint were incorrect. At four bits, 109 billion weights already represent approximately 54.5 decimal GB before overhead. This is not a fully GPU-resident model on a single RTX 5090 with 32 GB at this quantization.
Primary source: https://huggingface.co/meta-llama/Llama-4-Scout-17B-16E-Instruct . Offloading to RAM or multiple cards changes the protocol and throughput. We do not attribute any speed results to this unmeasured configuration.
#An identifiable first model
To get started, Qwen/Qwen3-8B and Qwen/Qwen3-14B are two identifiable models published by Qwen. Start with a Q4 quantization and a context of 4096 tokens, then check the actual placement. The 8B leaves more headroom; the 14B is an alternative to compare on your own tasks. This starting choice is not a quality ranking tested on RTX 5090.
Sources: https://huggingface.co/Qwen/Qwen3-8B and https://huggingface.co/Qwen/Qwen3-14B; Ollama distributions: https://ollama.com/library/qwen3:8b and https://ollama.com/library/qwen3:14b. The tag may change: record its digest and quantization with ollama show avant to compare results.
#RTX 5090 and RTX 4090: compare without a misleading benchmark
The RTX 4090 does not support NVLink. Two RTX 4090 cards do not automatically form a 48 GB pool usable as a single card. The runtime can distribute some workloads over PCIe, with communication overhead and model-specific constraints. Source NVIDIA: https://www.nvidia.com/en-us/geforce/graphics-cards/40-series/rtx-4090/
Comparing bandwidth alone does not provide a reliable tokens-per-second multiplier. Prompt processing, generation, quantization, context, kernels, and offloading must be identical. No 5090/4090/3090 performance advantage is claimed here without comparable measurements.
#Reproducible protocol to apply
Record the exact GPU and VRAM via nvidia-smi, CPU, RAM, OS, driver, runtime version, model identifier and hash, quantization, context, batch, and offload. Keep a single user and no other GPU workload during the test. Run a documented warm-up followed by five repetitions with the same inputs.
Keep prompt_eval_count / prompt_eval_duration and eval_count / eval_duration separate in the Ollama JSON responses. Durations are in nanoseconds. Publish every repetition, median, and spread, noting prompt-cache usage. Source: https://docs.ollama.com/api/generate. Generation throughput does not measure all perceived latency or quality.
#Power consumption and price: data not measured here
The manufacturer's advertised power is not wall power consumption. An energy comparison requires a wattmeter, the same system, the duration, and the outputs produced. The old price-per-token ratios and quantified power-limit recommendations have been removed for lack of traceable measurements. Compare current prices only after defining your memory requirements.
#Choose based on the memory actually required
A 32 GB card can be useful when your quantized model and its context exceed the budget of a 24 GB card while remaining within the 32 GB limit. For small models, first check the result on your current hardware. For Scout, total weight remains decisive: the active parameters do not make this model compatible with 32 GB.
#Continue to a free installation
The configurator provides a compatibility estimate. The model page and the Ollama guide then let you verify installation and CPU/GPU placement before making a purchase. Hardware results published on the Benchmarks page should be read together with their protocol and exact machine.
- Official Llama 4 Scout sheet
- Specifications RTX 4090 — NVIDIA
- Traceable measurements on our test machine
- Install Qwen3 8B
- Estimate hardware compatibility
Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.