Intermediate 13 minRTX 30

Which LLM on RTX 3090 / 3090 Ti (24 GB) ?

Direct response

A RTX 3090 offers 24 GB of VRAM and 936 GB/s of bandwidth—enough to run dense models with 27 to 32 billion parameters in Q4 (Qwen 3.5 27B, Qwen3 32B, Gemma 4 31B) and mixture-of-experts models such as gpt-oss 20B. The 3090 Ti adds only 6% to 8% more speed. A 70B model does not fit on a single card.

With 24 GB, the RTX 3090 and the 3090 Ti keep a dense 27- to 32-billion-parameter model in VRAM, which 12- and 16-GB cards cannot do. This guide covers the models that truly fit, with their exact weight sizes, throughput measured by Hardware Corner, the effect of context, the power limit that moderates power consumption, and the points to check before buying. You will also know when to choose a 4090 or a 5090.

Choosing a machine? Our picks by budget →

By Mohamed Meguedmi·Update 2026-09-29·Tested on Windows, macOS, and Linux
Recommended hardware

Buying alternative for this guide: GMKtec EVO-X2 64GB / 1TB (Ryzen AI Max+ 395).

A mini PC is a complete machine: check the available memory and engine compatibility. It does not replace macOS/MLX or CUDA.

Why this choice? Our complete guide on GMKtec EVO-X2 64GB / 1TB (Ryzen AI Max+ 395) →

Compare all options by budget, from €800 to €3,500 →

On the go: which laptop for local AI →

Affiliate links — commission possible at no extra cost to you. As an Amazon Associate, BestLLMfor earns from qualifying purchases.

#RTX 3090 for a local LLM: what 24 GB makes possible

For a local LLM, a RTX 3090 is primarily valuable for its 24 GB of GDDR6X and 936 GB/s of bandwidth: it keeps 27- to 32-billion-parameter models in Q4 entirely in VRAM, which 12 and 16 GB cards cannot accommodate. According to the Ollama library, Qwen 3.5 27B weighs 17 GB, Qwen 3.8 27B 18 GB, Qwen3 32B, and Gemma 4 31B 20 GB. Hardware Corner measures about 33 to 35 tokens per second on these dense models with 4 000 tokens of context, and more than 100 on expert models such as gpt-oss 20B. A 70B in Q4 (43 GB for Llama 3.3) remains beyond the reach of a single card. The 3090 Ti adds only 8% more bandwidth.

Memory
24 GB of GDDR6X on a 384-bit bus; 936 GB/s for the 3090 and 1,008 GB/s for the 3090 Ti, according to Wikipedia's table.
Power consumption
350 W of graphics power for the 3090, 450 W for the Ti; a 750 W system power supply is required for the 3090 and an 850 W supply for the Ti, according to NVIDIA.
Software
Ollama lists the RTX 3090 among compute capability 8.6 cards supported with driver 550 or later.
i
A specified bandwidth of 986 GB/s flows through the link
Some benchmark pages cite 986 GB/s for the 3090. The value derived from the 384-bit bus and 19.5 Gbit/s memory is 936 GB/s (384 × 19.5 ÷ 8), the figure shown in the Wikipedia table. This is the figure used by this guide.

#Models that fit in 24 GB

The weights below are those of the Ollama library files, checked on September 29, 2026. The margin is what remains on 24 GB before the context, cache, and engine buffers; also leave room for the system if the card drives your display.

Models for RTX 3090 / 3090 Ti (Ollama files, Q4 unless noted)
ModelFile OllamaGross marginWhat it brings
gpt-oss 20B14 GB10 GBFast reasoning, with room for a long context
Devstral Small 2 24B15 GB9 GBCode and agents
Qwen 3.5 27B17 GB7 GBVersatile dense model, text and image
Qwen 3.8 27B18 GB6 GBRecent generation, dense 27B
Qwen3-Coder 30B19 GB5 GBCode, mixture-of-experts model
Gemma 4 26B19 GB5 GBMixture-of-experts model, high throughput
Qwen3 32B / Gemma 4 31B20 GB4 GBUpper limit for dense models
Qwen 3.5 35B24 GB0 GBDoes not fit in Q4: context impossible
Llama 3.3 70B43 GBMissing 19 GBBeyond the capacity of a single card

Start with a dense 27-billion-parameter model: Qwen 3.5 27B leaves 7 GB for context. Above 20 GB, the headroom becomes too limited for a useful context. Mixture-of-experts models (Qwen3-Coder 30B, Gemma 4 26B, gpt-oss 20B) run much faster than dense models because each token activates only part of the weights.

#Context and KV cache: the 24 GB advantage

Ollama sets the default context based on video memory: its documentation specifies 4k tokens with less than 24 GiB of VRAM and 32k tokens between 24 and 48 GiB. A RTX 3090 sits on the boundary: depending on what the card reports to the engine, you get 32k or 4k, and an agent launched with the default settings may lose its history without warning. Check the actual value with ollama ps, and set it yourself for agents, for which the documentation recommends at least 64,000 tokens.

The KV cache costs memory: each context token stores attention keys and values. The main lever is cache quantization: according to the Ollama FAQ, q8_0 uses about half the memory of f16, with very little loss of precision. Context also costs speed: Hardware Corner goes from 70 to 52 to 39 tokens per second on Qwen3 14B at 4k, 16k, and 32k.

Prompt reading and generation on RTX 3090, tokens per second (Hardware Corner, third-party measurements)
Model (Q4_K, or MXFP4 for gpt-oss)4k generation16k generation32k generation4k prompt processing32k prompt processing
Qwen3 8B115,387,567,94 049,61 714,6
Qwen3 14B70,052,138,62 459,01 175,7
gpt-oss 20B147,5128,5112,64 400,32 547,2
Qwen3 30B A3B153,6113,887,22 988,61 336,8
Qwen 3.5 27B33,532,331,01 104,2848,2
Gemma 4 31B34,733,531,41 155,8723,7

Two behaviors stand out. Standard models such as Qwen3 14B lose nearly half their throughput between 4k and 32k. Qwen 3.5 27B, on the other hand, loses only 7% over the same range (33.5 to 31.0). The source does not explain why, but for long documents this behavior matters more than model size. These figures are from Hardware Corner, not measurements by QuelLLM.

#Install Ollama and verify that the model fits

The 3090 is an Ampere card with compute capability 8.6, supported by Ollama with a NVIDIA 550 or newer driver; on Windows, the documentation requires version 551.61 or later. Flash Attention 2, which reduces context memory usage, works on Ampere according to its official repository; version 3 is reserved for Hopper (H100) and is not available on a 3090.

  1. 01
    Check the driver
    Run nvidia-smi in a terminal: the driver version must be 550 or later (551.61 on Windows), and total memory should show approximately 24 GB.
  2. 02
    Install Ollama
    Install Ollama from its official website. The local API listens on http://localhost:11434.
  3. 03
    Running a 27-billion-parameter model
    Run ollama run qwen3.5:27b: the 17 GB download happens once. For a faster model, try ollama run gpt-oss:20b (14 GB).
  4. 04
    Control memory allocation
    In a second terminal, run ollama ps: the Processor column should display 100% GPU. A CPU/GPU split means the model or its context spills into system RAM.
  5. 05
    Adjust the context
    Set the context with OLLAMA_CONTEXT_LENGTH, then run ollama ps again. If there is not enough headroom, set OLLAMA_FLASH_ATTENTION to 1 and OLLAMA_KV_CACHE_TYPE to q8_0 when starting the server.
Basic commands
ollama run qwen3.5:27b
ollama run gpt-oss:20b
ollama ps

# Exemple de la documentation Ollama : contexte de 64 000 tokens
OLLAMA_CONTEXT_LENGTH=64000 ollama serve

#Throughput: what bandwidth really allows

Generation is bound by bandwidth: the GPU rereads all active weights for each token, so the theoretical ceiling is roughly bandwidth divided by model size. At 936 GB/s, Qwen3 14B (9.3 GB) tops out at about 100 tokens per second and Hardware Corner measures 70.0 at 4k, or 70%. Qwen 3.5 27B (17 GB) tops out at 55 and measures 33.5: 61%. Qwen3 32B (20 GB) tops out at 47 and measures 35.1: 75%. A dense 27B or 32B therefore runs at around 60 to 75% of its ceiling, or 33 to 35 tokens per second, which remains very comfortable to read.

The llama.cpp community leaderboard confirms the relationship between the two cards. On Llama 2 7B Q4_0 (3.56 GiB), the 3090 generates 158.16 tokens per second without Flash Attention and 161.89 with it, compared with 171.19 and 172.26 for the 3090 Ti: that’s +8% and +6%, consistent with the 8% increase in bandwidth. The two series come from different contributors, so the actual gap depends on the machine.

→
The prompt matters as much as the generation
Prompt processing depends on compute, not bandwidth. On Qwen 3.5 27B, Hardware Corner measures 1,104 tokens per second at 4k and 848 at 32k: a 32,000-token document is read in about forty seconds. For an agent that sends large contexts back every turn, this delay matters more than generation speed.

#Cooling, power limits, and buying used

The 3090 draws 350 W and the Ti 450 W, with maximum GPU temperatures of 93 and 92 °C respectively, according to NVIDIA. Under sustained LLM load, case noise and heat matter: plan for a well-ventilated case and a power supply that meets the 750 W (3090) or 850 W (Ti) requirements specified by NVIDIA. An LLM primarily stresses memory: the power limit costs little speed.

The llama.cpp leaderboard provides a rough benchmark: one contributor limits their 3090 to 250 W and records 137.72 tokens per second on Llama 2 7B Q4_0, compared with 158.16 in the stock series. That's about 13% less speed for 29% less power. The two measurements come from different machines: treat them as an indication, then measure on your own system.

Limit power (administrator rights required)
sudo nvidia-smi -pl 250

The nvidia-smi -pl option sets the maximum power draw in watts. On Windows, open the terminal as an administrator; the setting disappears after reboot unless you automate it. On a used card, also monitor memory temperature with a tool such as HWiNFO: if it climbs significantly higher than the core temperature, consider replacing the thermal pads.

#Buying a used 3090: what to check

The 3090 is an Ampere card found mostly used. For current prices, check our tracker, which records the lowest price for each card twice a week, along with the price per GB of VRAM.

Identity
Use nvidia-smi to verify that the card reports 24 GB and the correct model (3090 or 3090 Ti), not another model.
Stability
Run a 27-billion-parameter model in a loop for around thirty minutes: crashes, artifacts, or thermal throttling are warning signs.
Memory temperature
Compare it with the memory reading: a significant gap indicates tired pads.
Guarantee
Require a return option or seller warranty: replacing pads or thermal paste is a cost to plan for.
Power supply
350 W or 450 W peak: check the card's power supply and power connectors before buying.

#3090, 4090, 5090: what do you gain by paying more?

The RTX 4090 also offers 24 GB, with more compute: its advantage is more noticeable in prompt processing than generation, which remains limited by bandwidth. The RTX 5090 moves up to 32 GB, the category where a dense 32B model still has room for a long context. The guides for both cards detail every use case.

#vLLM, two cards, and fine-tuning on 3090s

Three use cases come up often with the 3090. vLLM works: according to its documentation, it requires compute capability 7.5 or higher, and the 3090 is 8.6. To choose between llama.cpp and vLLM, see the engine comparison. Two 3090s can be combined with an NVLink bridge for 30-series cards: the multi-GPU guide explains how the layers are split.

For fine-tuning, Unsloth lists minimum VRAM requirements in 4-bit QLoRA: 6 GB for an 8-billion-parameter model, 8.5 GB for 14 billion, 22 GB for 27 billion, and 26 GB for 32 billion. The 3090 can therefore fine-tune up to 27B in QLoRA, just barely, with a batch size of 1 and a short context, but not a 32B model. In 16-bit LoRA, 9 billion already requires 24 GB—the card’s entire capacity.

#Verdict: when the 3090 remains the right choice

Which decision fits your situation
SituationDecisionQuantified rationale
You want a dense 27B or 32B model locally3090 (or 4090)24 GB: 17 to 20 GB for weights and 4 to 7 GB of headroom
You are mainly targeting 7- to 14-billion-parameter modelsA 12 or 16 GB card is sufficientThe 3090 adds only speed: 936 GB/s
You want a 70B or very long contexts5090 or two cards43 GB of weights for a 70B Q4
Silence and power consumption come firstA recent 16 GB card750 W power supply required for the 3090
You're deciding between the 3090 and 3090 TiChoose the cheaper of the two+6 to +8% generation only

The 3090 remains the most common entry point for a local 27B: that's Hardware Corner's conclusion, calling it the best used-market entry point for serious use. The 4090 replaces it if your budget allows, and the 5090 if you want 32 GB.

#Frequently asked questions

FAQ
Which LLM should you run on a RTX 3090?+
Start with Qwen 3.5 27B: its 17 GB Ollama file leaves about 7 GB of headroom for context and generates around 33 tokens per second according to Hardware Corner. For more speed, gpt-oss 20B (14 GB) exceeds 100 tokens per second. Qwen3 32B and Gemma 4 31B (20 GB) are the upper limit.
Can a RTX 3090 run a 70B model?+
No, not entirely. Llama 3.3 70B weighs 43 GB in Ollama, which is 19 GB more than the card. Split across RAM, it becomes very slow because every token rereads all the weights. For a 70B model in VRAM, you need two 24 GB cards or one card with more than 43 GB.
RTX 3090 or 3090 Ti for a local LLM?+
They are nearly equivalent: the same 24 GB of memory, 936 versus 1,008 GB/s of bandwidth, and a 6 to 8% generation gain in the llama.cpp ranking. The Ti draws 450 W versus 350 W. Unless the price difference is minimal, the 3090 is the better choice.
What power supply does a RTX 3090 need?+
NVIDIA specifies a required system power supply of 750 W for the 3090 (350 W graphics power) and 850 W for the 3090 Ti (450 W). These values apply to a complete PC. A 250 W limit reduces the card's power consumption for about 13% less speed, according to a community benchmark.
Does the RTX 3090 support vLLM and Flash Attention?+
Yes to both. vLLM requires compute capability 7.5 or higher, and the 3090 is 8.6. Flash Attention 2 works on Ampere GPUs, including the 3090, according to its official repository. Flash Attention 3, however, is limited to Hopper: do not look for it on a consumer Ampere card.
Can you fine-tune a model on a RTX 3090?+
Yes, with QLoRA. According to Unsloth, the minimum is 22 GB for a 27B and 8,5 GB for a 14B, making the 27B just barely possible with a batch size of 1. A 32B requires 26 GB and does not fit. With 16-bit LoRA, the practical limit is around 9 billion.
Did this guide help you?

Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.

Prices in euros (€) are French market prices including VAT, as checked by BestLLMfor. US prices differ: the Amazon buttons show the current US price.