Intermediate 11 minRTX 50

Which LLM on RTX 5090 (32 GB) ?

Direct response

A RTX 5090 provides 32 GB of GDDR7 VRAM and 1,792 GB/s of bandwidth: it keeps dense models with 27 to 32 billion parameters (Qwen 3.5 27B, Qwen3 32B) and 30- to 35-billion-parameter mixture-of-experts models in VRAM, with a long context. It generates about 60 tokens per second on a dense 27B. A 70B (43 GB in Q4) does not fit on a single card.

With 32 GB, the RTX 5090 can run models in VRAM that 24 GB cards cannot use with a long context. This guide provides the models that actually fit, with their exact weights, throughput measured by Hardware Corner, what FP4 changes, the context budget, and power limits. You'll also know when a 5090 is worth the premium over a 4090 or 3090.

Choosing a machine? Our picks by budget →

By Mohamed Meguedmi·Update 2026-09-29·Tested on Windows, macOS, and Linux
Recommended hardware

Buying alternative for this guide: GMKtec EVO-X2 64GB / 1TB (Ryzen AI Max+ 395).

A mini PC is a complete machine: check the available memory and engine compatibility. It does not replace macOS/MLX or CUDA.

Why this choice? Our complete guide on GMKtec EVO-X2 64GB / 1TB (Ryzen AI Max+ 395) →

Compare all options by budget, from €800 to €3,500 →

On the go: which laptop for local AI →

Affiliate links — possible commission at no extra cost to you. As an Amazon Associate, BestLLMfor earns from qualifying purchases.

#RTX 5090 for a local LLM: what 32 GB enables

For a local LLM, a RTX 5090 stands out with its 32 GB of GDDR7 and 1 792 GB/s of bandwidth: it can keep 27- to 35-billion-parameter models entirely in VRAM with a long context, while 24 GB cards leave them cramped. According to the Ollama library, Qwen 3.5 27B weighs 17 GB, Qwen 3.8 27B 18 GB, Qwen3 32B and Gemma 4 31B 20 GB, Qwen 3.5 35B and Nemotron 3 Nano 30B 24 GB. Hardware Corner measures approximately 59 to 61 tokens per second on dense 27- to 32-billion-parameter models with 4 000 tokens of context, and more than 160 on a 35-billion-parameter mixture-of-experts model. A 70B (43 GB for Llama 3.3 in Q4) still does not fit on a single card.

Memory
32 GB of GDDR7 on a 512-bit bus, according to NVIDIA; 1,792 GB/s of bandwidth, according to Hardware Corner.
Compute
21,760 CUDA cores and 5th-generation Tensor Cores, according to NVIDIA. Compute capability 12.0, supported by Ollama with driver 550 or later.
Power consumption
575 W of graphics power and a required 1,000 W system power supply, according to NVIDIA, with a 600 W PCIe Gen 5 cable or included adapter.

#Models that fit in 32 GB

The weights are those of the files in the Ollama library, accessed on September 29, 2026. The margin is what remains on 32 GB before context, cache, and engine buffers.

Models for RTX 5090 (files Ollama, Q4 unless noted)
ModelFile OllamaGross marginNote
gpt-oss 20B14 GB18 GB128k context with no constraints
Qwen 3.5 27B17 GB15 GBDense, text, and image
Qwen 3.8 27B18 GB14 GBDense, recent generation
Qwen3 32B / Gemma 4 31B20 GB12 GBUpper limit for comfortably usable dense models
Qwen 3.5 35B / Nemotron 3 Nano 30B24 GB8 GBDo not fit in 24 GB with context
Llama 3.3 70B43 GBMissing 11 GBDoesn't fit
gpt-oss 120B65 GBShort by 33 GBOut of reach

The 32 GB mainly changes the headroom: a dense 27B retains 14 to 15 GB for context, compared with 6 to 7 GB on a 24 GB card, and a model with 24 GB of weights can still be loaded with 8 GB of headroom. This is what makes the 128k contexts listed below possible.

#Measured throughput: what the 5090 brings compared with other cards

The figures below come from Hardware Corner, which tests with llama.cpp using contexts from 4k to 256k. They are not QuelLLM measurements.

Generation on RTX 5090, tokens per second (Hardware Corner)
Model (Q4_K, or MXFP4 for Qwen 3.5 35B)4k context32k context128k context
Qwen3 14B123,882,437,2
gpt-oss 20B298,2215,1112,0
Qwen3 30B A3B226,1143,176,8
Qwen 3.5 27B58,853,844,1
Gemma 4 31B61,155,443,4
Qwen3 32B61,443,8not measured
Qwen 3.5 35B (MXFP4)165,2143,2118,2

Generation is bounded by bandwidth: the theoretical ceiling is approximately bandwidth divided by the model's weight. For Qwen 3.5 27B (17 GB), 1,792 ÷ 17 gives 105 tokens per second, and the measured rate, 58.8, reaches 56% of that. For Qwen3 32B (20 GB), the ceiling is 90 and the measured rate is 61.4, or 68%. A dense model with 27 to 32 billion parameters therefore runs at around 60 tokens per second on a 5090.

On Qwen 3.5 27B, the 5090 generates 1.76 times faster than the 3090 (58.8 versus 33.5 according to two Hardware Corner pages), while bandwidth is 1.91 times higher (1,792 versus 936 GB/s). On the llama.cpp leaderboard, the 5090 reaches approximately 290 to 300 tokens per second on Llama 2 7B Q4_0, versus approximately 160 for a 3090.

#5090 versus the 4090 and 3090: the measured gap

Relative generation by card (Hardware Corner, 5090 = 100%)
CardMemoryRelative generation
RTX 509032 GB100 %
RTX 409024 GB77 %
RTX 3090 Ti24 GB55 %
RTX 309024 GB51 %
→
The real advantage: long context
The 5090 maintains its throughput as the context grows. Qwen 3.5 27B drops from 58.8 to 44.1 tokens per second between 4k and 128k (-25%), while Qwen 3.5 35B in MXFP4 still generates 97.3 tokens per second with 256,000 tokens of context. It is in this scenario, not in a short chat, that 32 GB justifies the card.

#Context and KV cache: sizing its 32 GB

Ollama sets the default context based on video memory: its documentation specifies 32k tokens between 24 and 48 GiB of VRAM and recommends at least 64,000 tokens for agents, web research, and coding tools. The 5090 falls into the 32k tier, more generous than the 4k on cards under 24 GiB. Increasing the context comes at the cost of KV cache, which stores the attention keys and values for every token read.

The main lever is cache quantization: according to the Ollama FAQ, q8_0 uses about half the memory of f16, with very little loss of accuracy, and requires Flash Attention, which Ollama enables automatically when the card and model support it. Calculation rule: memory left after the model minus 1 to 2 GB of buffers gives the cache budget. With Qwen 3.5 27B (17 GB), more than 13 GB remains for context, as confirmed by Hardware Corner's 128k measurement. Check the result with ollama ps.

#FP4, NVFP4, MXFP4: what Blackwell really changes

Blackwell introduces the NVFP4 format: according to NVIDIA, it’s a 4-bit floating-point format with a scale factor for every block of 16 values, compared with 32 for MXFP4, making it more granular. Hardware Corner reports that the 5090 accelerates NVFP4 in hardware. Q4_K_M, meanwhile, is a block quantization method from llama.cpp: the two aren’t mutually exclusive on paper; they’re different file formats, and a model exists in one or the other only if someone has converted it.

In practice, the Qwen 3.5 35B line in MXFP4 from the table above shows what a 4-bit floating-point format delivers on this card: 165 tokens per second at 4k for 24 GB of weights. Tool support for NVFP4 (llama.cpp, Ollama, vLLM) is evolving quickly: check the installed version and the model page before choosing a file, and keep Q4_K_M as a safe default.

#Install Ollama on a RTX 5090

  1. 01
    Check the driver
    Run nvidia-smi in a terminal: the driver must be version 550 or later (551.61 on Windows), and total memory should show approximately 32 GB.
  2. 02
    Install Ollama
    Install Ollama from its official website. The local API listens on http://localhost:11434.
  3. 03
    Run a model
    Run ollama run qwen3.5:27b: the 17 GB download happens once. For a larger mixture-of-experts model, try ollama run qwen3.5:35b (24 GB).
  4. 04
    Control the distribution
    In a second terminal, run ollama ps: the Processor column should show 100% GPU.
  5. 05
    Adjust the context
    Set the context with OLLAMA_CONTEXT_LENGTH, then run ollama ps again. If you are short on headroom, set OLLAMA_FLASH_ATTENTION to 1 and OLLAMA_KV_CACHE_TYPE to q8_0 at startup.
Basic commands
ollama run qwen3.5:27b
ollama run qwen3.5:35b
ollama ps

# Exemple de la documentation Ollama : contexte de 64 000 tokens
OLLAMA_CONTEXT_LENGTH=64000 ollama serve

#Power, heat, and fine-tuning

NVIDIA specifies 575 W of graphics power, a 1,000 W system power supply, a maximum GPU temperature of 90 °C, and a 600 W PCIe Gen 5 cable: the 5090 is the most demanding card in the comparison. Check the cable and its connector before installing, and plan for a well-ventilated case. A power limit reduces consumption with little speed loss, since the LLM primarily stresses memory (see the 3090 measured at 250 W with a power limit in its guide).

For fine-tuning, Unsloth lists minimum VRAM requirements in 4-bit QLoRA: 22 GB for a 27-billion-parameter model, 26 GB for 32 billion, 30 GB for 40 billion, and 41 GB for 70 billion. The 5090 can therefore fine-tune up to 40B in QLoRA, with a batch size of 1 and a short context—not a 70B model.

#Verdict: 5090, 4090, or 3090 depending on the use case

Which decision fits your situation
SituationDecisionQuantified rationale
You’re targeting a dense 27B to 32B model with long context509032 GB: 12 to 15 GB of headroom, 44 t/s at 128k on Qwen 3.5 27B
You are targeting a 30- to 35-billion-parameter mixture-of-experts model509024 GB of weights: out of reach for a 24 GB card with context
You stay with 8B to 14B modelsNo 5090The 3090 or a 16 GB card gets the job done
You already have a 4090Keep it unless you need 32 GBGeneration at 77% of the 5090
You want a 70BTwo cards or one unified-memory machine43 GB of weights in Q4

The 5090 replaces a 4090 first and foremost when capacity is insufficient: the speed gain is about 30% according to Hardware Corner, while 32 GB lets you load models and contexts beyond the reach of 24 GB. For current prices, check our tracker, which records the lowest price for each card twice a week.

#Frequently asked questions

FAQ
Which LLM should you install on a RTX 5090?+
Start with Qwen 3.5 27B: its 17 GB Ollama file leaves about 15 GB of headroom and generates around 59 tokens per second according to Hardware Corner. For a larger mixture-of-experts model, Qwen 3.5 35B (24 GB) exceeds 160 tokens per second. Qwen3 32B and Gemma 4 31B (20 GB) are the upper limit for dense models.
Can a RTX 5090 run a 70B?+
No, not entirely. Llama 3.3 70B takes up 43 GB in Ollama, which is 11 GB more than the card, and gpt-oss 120B takes up 65. When split across system RAM, throughput drops because every token rereads all the weights of a dense model. For a 70B in VRAM, you need two cards or a unified-memory machine.
Is the RTX 5090 much faster than a 4090 for LLMs?+
According to Hardware Corner, the 4090 produces 77% of the 5090’s throughput, a difference of about 30% in favor of the 5090. Its main advantage remains its 32 GB of memory, which lets you load larger models and contexts.
Does the RTX 5090 support FP4 and NVFP4?+
Yes: NVFP4 is a 4-bit format introduced with Blackwell, and Hardware Corner says the card accelerates it in hardware. Software support is evolving: verify that your version of Ollama, llama.cpp, or vLLM supports the model format. Q4_K_M remains the safe choice, with no need for FP4.
What power supply do you need for a RTX 5090?+
NVIDIA specifies a system power supply of 1,000 W for 575 W of graphics power, with a 600 W PCIe Gen 5 cable or the included adapter. Hardware Corner recommends at least 950 W Gold. Choose a quality power supply and check the cable before installation.
Can you fine-tune a model on a RTX 5090?+
Yes, with QLoRA. According to Unsloth, the minimum is 22 GB for a 27B and 30 GB for a 40B, which fits in 32 GB with a batch size of 1. A 70B requires 41 GB and does not fit. These are absolute minimums: the context and batch size increase them.
Did this guide help you?

Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.