Which LLM on RTX 5090 (32 GB) ?
A RTX 5090 provides 32 GB of GDDR7 VRAM and 1,792 GB/s of bandwidth: it keeps dense models with 27 to 32 billion parameters (Qwen 3.5 27B, Qwen3 32B) and 30- to 35-billion-parameter mixture-of-experts models in VRAM, with a long context. It generates about 60 tokens per second on a dense 27B. A 70B (43 GB in Q4) does not fit on a single card.
With 32 GB, the RTX 5090 can run models in VRAM that 24 GB cards cannot use with a long context. This guide provides the models that actually fit, with their exact weights, throughput measured by Hardware Corner, what FP4 changes, the context budget, and power limits. You'll also know when a 5090 is worth the premium over a 4090 or 3090.
Choosing a machine? Our picks by budget →
Buying alternative for this guide: GMKtec EVO-X2 64GB / 1TB (Ryzen AI Max+ 395).
A mini PC is a complete machine: check the available memory and engine compatibility. It does not replace macOS/MLX or CUDA.
Why this choice? Our complete guide on GMKtec EVO-X2 64GB / 1TB (Ryzen AI Max+ 395) →
Compare all options by budget, from €800 to €3,500 →
On the go: which laptop for local AI →
Affiliate links — possible commission at no extra cost to you. As an Amazon Associate, BestLLMfor earns from qualifying purchases.
#RTX 5090 for a local LLM: what 32 GB enables
For a local LLM, a RTX 5090 stands out with its 32 GB of GDDR7 and 1 792 GB/s of bandwidth: it can keep 27- to 35-billion-parameter models entirely in VRAM with a long context, while 24 GB cards leave them cramped. According to the Ollama library, Qwen 3.5 27B weighs 17 GB, Qwen 3.8 27B 18 GB, Qwen3 32B and Gemma 4 31B 20 GB, Qwen 3.5 35B and Nemotron 3 Nano 30B 24 GB. Hardware Corner measures approximately 59 to 61 tokens per second on dense 27- to 32-billion-parameter models with 4 000 tokens of context, and more than 160 on a 35-billion-parameter mixture-of-experts model. A 70B (43 GB for Llama 3.3 in Q4) still does not fit on a single card.
- Memory
- 32 GB of GDDR7 on a 512-bit bus, according to NVIDIA; 1,792 GB/s of bandwidth, according to Hardware Corner.
- Compute
- 21,760 CUDA cores and 5th-generation Tensor Cores, according to NVIDIA. Compute capability 12.0, supported by Ollama with driver 550 or later.
- Power consumption
- 575 W of graphics power and a required 1,000 W system power supply, according to NVIDIA, with a 600 W PCIe Gen 5 cable or included adapter.
#Models that fit in 32 GB
The weights are those of the files in the Ollama library, accessed on September 29, 2026. The margin is what remains on 32 GB before context, cache, and engine buffers.
| Model | File Ollama | Gross margin | Note |
|---|---|---|---|
| gpt-oss 20B | 14 GB | 18 GB | 128k context with no constraints |
| Qwen 3.5 27B | 17 GB | 15 GB | Dense, text, and image |
| Qwen 3.8 27B | 18 GB | 14 GB | Dense, recent generation |
| Qwen3 32B / Gemma 4 31B | 20 GB | 12 GB | Upper limit for comfortably usable dense models |
| Qwen 3.5 35B / Nemotron 3 Nano 30B | 24 GB | 8 GB | Do not fit in 24 GB with context |
| Llama 3.3 70B | 43 GB | Missing 11 GB | Doesn't fit |
| gpt-oss 120B | 65 GB | Short by 33 GB | Out of reach |
The 32 GB mainly changes the headroom: a dense 27B retains 14 to 15 GB for context, compared with 6 to 7 GB on a 24 GB card, and a model with 24 GB of weights can still be loaded with 8 GB of headroom. This is what makes the 128k contexts listed below possible.
#Measured throughput: what the 5090 brings compared with other cards
The figures below come from Hardware Corner, which tests with llama.cpp using contexts from 4k to 256k. They are not QuelLLM measurements.
| Model (Q4_K, or MXFP4 for Qwen 3.5 35B) | 4k context | 32k context | 128k context |
|---|---|---|---|
| Qwen3 14B | 123,8 | 82,4 | 37,2 |
| gpt-oss 20B | 298,2 | 215,1 | 112,0 |
| Qwen3 30B A3B | 226,1 | 143,1 | 76,8 |
| Qwen 3.5 27B | 58,8 | 53,8 | 44,1 |
| Gemma 4 31B | 61,1 | 55,4 | 43,4 |
| Qwen3 32B | 61,4 | 43,8 | not measured |
| Qwen 3.5 35B (MXFP4) | 165,2 | 143,2 | 118,2 |
Generation is bounded by bandwidth: the theoretical ceiling is approximately bandwidth divided by the model's weight. For Qwen 3.5 27B (17 GB), 1,792 ÷ 17 gives 105 tokens per second, and the measured rate, 58.8, reaches 56% of that. For Qwen3 32B (20 GB), the ceiling is 90 and the measured rate is 61.4, or 68%. A dense model with 27 to 32 billion parameters therefore runs at around 60 tokens per second on a 5090.
On Qwen 3.5 27B, the 5090 generates 1.76 times faster than the 3090 (58.8 versus 33.5 according to two Hardware Corner pages), while bandwidth is 1.91 times higher (1,792 versus 936 GB/s). On the llama.cpp leaderboard, the 5090 reaches approximately 290 to 300 tokens per second on Llama 2 7B Q4_0, versus approximately 160 for a 3090.
#5090 versus the 4090 and 3090: the measured gap
| Card | Memory | Relative generation |
|---|---|---|
| RTX 5090 | 32 GB | 100 % |
| RTX 4090 | 24 GB | 77 % |
| RTX 3090 Ti | 24 GB | 55 % |
| RTX 3090 | 24 GB | 51 % |
#Context and KV cache: sizing its 32 GB
Ollama sets the default context based on video memory: its documentation specifies 32k tokens between 24 and 48 GiB of VRAM and recommends at least 64,000 tokens for agents, web research, and coding tools. The 5090 falls into the 32k tier, more generous than the 4k on cards under 24 GiB. Increasing the context comes at the cost of KV cache, which stores the attention keys and values for every token read.
The main lever is cache quantization: according to the Ollama FAQ, q8_0 uses about half the memory of f16, with very little loss of accuracy, and requires Flash Attention, which Ollama enables automatically when the card and model support it. Calculation rule: memory left after the model minus 1 to 2 GB of buffers gives the cache budget. With Qwen 3.5 27B (17 GB), more than 13 GB remains for context, as confirmed by Hardware Corner's 128k measurement. Check the result with ollama ps.
#FP4, NVFP4, MXFP4: what Blackwell really changes
Blackwell introduces the NVFP4 format: according to NVIDIA, it’s a 4-bit floating-point format with a scale factor for every block of 16 values, compared with 32 for MXFP4, making it more granular. Hardware Corner reports that the 5090 accelerates NVFP4 in hardware. Q4_K_M, meanwhile, is a block quantization method from llama.cpp: the two aren’t mutually exclusive on paper; they’re different file formats, and a model exists in one or the other only if someone has converted it.
In practice, the Qwen 3.5 35B line in MXFP4 from the table above shows what a 4-bit floating-point format delivers on this card: 165 tokens per second at 4k for 24 GB of weights. Tool support for NVFP4 (llama.cpp, Ollama, vLLM) is evolving quickly: check the installed version and the model page before choosing a file, and keep Q4_K_M as a safe default.
#Install Ollama on a RTX 5090
- 01Check the driverRun nvidia-smi in a terminal: the driver must be version 550 or later (551.61 on Windows), and total memory should show approximately 32 GB.
- 02Install OllamaInstall Ollama from its official website. The local API listens on http://localhost:11434.
- 03Run a modelRun ollama run qwen3.5:27b: the 17 GB download happens once. For a larger mixture-of-experts model, try ollama run qwen3.5:35b (24 GB).
- 04Control the distributionIn a second terminal, run ollama ps: the Processor column should show 100% GPU.
- 05Adjust the contextSet the context with OLLAMA_CONTEXT_LENGTH, then run ollama ps again. If you are short on headroom, set OLLAMA_FLASH_ATTENTION to 1 and OLLAMA_KV_CACHE_TYPE to q8_0 at startup.
#Power, heat, and fine-tuning
NVIDIA specifies 575 W of graphics power, a 1,000 W system power supply, a maximum GPU temperature of 90 °C, and a 600 W PCIe Gen 5 cable: the 5090 is the most demanding card in the comparison. Check the cable and its connector before installing, and plan for a well-ventilated case. A power limit reduces consumption with little speed loss, since the LLM primarily stresses memory (see the 3090 measured at 250 W with a power limit in its guide).
For fine-tuning, Unsloth lists minimum VRAM requirements in 4-bit QLoRA: 22 GB for a 27-billion-parameter model, 26 GB for 32 billion, 30 GB for 40 billion, and 41 GB for 70 billion. The 5090 can therefore fine-tune up to 40B in QLoRA, with a batch size of 1 and a short context—not a 70B model.
#Verdict: 5090, 4090, or 3090 depending on the use case
| Situation | Decision | Quantified rationale |
|---|---|---|
| You’re targeting a dense 27B to 32B model with long context | 5090 | 32 GB: 12 to 15 GB of headroom, 44 t/s at 128k on Qwen 3.5 27B |
| You are targeting a 30- to 35-billion-parameter mixture-of-experts model | 5090 | 24 GB of weights: out of reach for a 24 GB card with context |
| You stay with 8B to 14B models | No 5090 | The 3090 or a 16 GB card gets the job done |
| You already have a 4090 | Keep it unless you need 32 GB | Generation at 77% of the 5090 |
| You want a 70B | Two cards or one unified-memory machine | 43 GB of weights in Q4 |
The 5090 replaces a 4090 first and foremost when capacity is insufficient: the speed gain is about 30% according to Hardware Corner, while 32 GB lets you load models and contexts beyond the reach of 24 GB. For current prices, check our tracker, which records the lowest price for each card twice a week.
- Source: NVIDIA, specifications for RTX 5090
- Source: Hardware Corner, LLM benchmarks on RTX 5090
- Source: NVIDIA, overview of the NVFP4 format
- Source: Ollama documentation, context length
- Source: Unsloth, VRAM required for fine-tuning
#Frequently asked questions
Which LLM should you install on a RTX 5090?+
Can a RTX 5090 run a 70B?+
Is the RTX 5090 much faster than a 4090 for LLMs?+
Does the RTX 5090 support FP4 and NVFP4?+
What power supply do you need for a RTX 5090?+
Can you fine-tune a model on a RTX 5090?+
Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.