Which LLM on RTX 3090 / 3090 Ti (24 GB) ?
A RTX 3090 offers 24 GB of VRAM and 936 GB/s of bandwidth—enough to run dense models with 27 to 32 billion parameters in Q4 (Qwen 3.5 27B, Qwen3 32B, Gemma 4 31B) and mixture-of-experts models such as gpt-oss 20B. The 3090 Ti adds only 6% to 8% more speed. A 70B model does not fit on a single card.
With 24 GB, the RTX 3090 and the 3090 Ti keep a dense 27- to 32-billion-parameter model in VRAM, which 12- and 16-GB cards cannot do. This guide covers the models that truly fit, with their exact weight sizes, throughput measured by Hardware Corner, the effect of context, the power limit that moderates power consumption, and the points to check before buying. You will also know when to choose a 4090 or a 5090.
Choosing a machine? Our picks by budget →
Buying alternative for this guide: GMKtec EVO-X2 64GB / 1TB (Ryzen AI Max+ 395).
A mini PC is a complete machine: check the available memory and engine compatibility. It does not replace macOS/MLX or CUDA.
Why this choice? Our complete guide on GMKtec EVO-X2 64GB / 1TB (Ryzen AI Max+ 395) →
Compare all options by budget, from €800 to €3,500 →
On the go: which laptop for local AI →
Affiliate links — commission possible at no extra cost to you. As an Amazon Associate, BestLLMfor earns from qualifying purchases.
#RTX 3090 for a local LLM: what 24 GB makes possible
For a local LLM, a RTX 3090 is primarily valuable for its 24 GB of GDDR6X and 936 GB/s of bandwidth: it keeps 27- to 32-billion-parameter models in Q4 entirely in VRAM, which 12 and 16 GB cards cannot accommodate. According to the Ollama library, Qwen 3.5 27B weighs 17 GB, Qwen 3.8 27B 18 GB, Qwen3 32B, and Gemma 4 31B 20 GB. Hardware Corner measures about 33 to 35 tokens per second on these dense models with 4 000 tokens of context, and more than 100 on expert models such as gpt-oss 20B. A 70B in Q4 (43 GB for Llama 3.3) remains beyond the reach of a single card. The 3090 Ti adds only 8% more bandwidth.
- Memory
- 24 GB of GDDR6X on a 384-bit bus; 936 GB/s for the 3090 and 1,008 GB/s for the 3090 Ti, according to Wikipedia's table.
- Power consumption
- 350 W of graphics power for the 3090, 450 W for the Ti; a 750 W system power supply is required for the 3090 and an 850 W supply for the Ti, according to NVIDIA.
- Software
- Ollama lists the RTX 3090 among compute capability 8.6 cards supported with driver 550 or later.
#Models that fit in 24 GB
The weights below are those of the Ollama library files, checked on September 29, 2026. The margin is what remains on 24 GB before the context, cache, and engine buffers; also leave room for the system if the card drives your display.
| Model | File Ollama | Gross margin | What it brings |
|---|---|---|---|
| gpt-oss 20B | 14 GB | 10 GB | Fast reasoning, with room for a long context |
| Devstral Small 2 24B | 15 GB | 9 GB | Code and agents |
| Qwen 3.5 27B | 17 GB | 7 GB | Versatile dense model, text and image |
| Qwen 3.8 27B | 18 GB | 6 GB | Recent generation, dense 27B |
| Qwen3-Coder 30B | 19 GB | 5 GB | Code, mixture-of-experts model |
| Gemma 4 26B | 19 GB | 5 GB | Mixture-of-experts model, high throughput |
| Qwen3 32B / Gemma 4 31B | 20 GB | 4 GB | Upper limit for dense models |
| Qwen 3.5 35B | 24 GB | 0 GB | Does not fit in Q4: context impossible |
| Llama 3.3 70B | 43 GB | Missing 19 GB | Beyond the capacity of a single card |
Start with a dense 27-billion-parameter model: Qwen 3.5 27B leaves 7 GB for context. Above 20 GB, the headroom becomes too limited for a useful context. Mixture-of-experts models (Qwen3-Coder 30B, Gemma 4 26B, gpt-oss 20B) run much faster than dense models because each token activates only part of the weights.
#Context and KV cache: the 24 GB advantage
Ollama sets the default context based on video memory: its documentation specifies 4k tokens with less than 24 GiB of VRAM and 32k tokens between 24 and 48 GiB. A RTX 3090 sits on the boundary: depending on what the card reports to the engine, you get 32k or 4k, and an agent launched with the default settings may lose its history without warning. Check the actual value with ollama ps, and set it yourself for agents, for which the documentation recommends at least 64,000 tokens.
The KV cache costs memory: each context token stores attention keys and values. The main lever is cache quantization: according to the Ollama FAQ, q8_0 uses about half the memory of f16, with very little loss of precision. Context also costs speed: Hardware Corner goes from 70 to 52 to 39 tokens per second on Qwen3 14B at 4k, 16k, and 32k.
| Model (Q4_K, or MXFP4 for gpt-oss) | 4k generation | 16k generation | 32k generation | 4k prompt processing | 32k prompt processing |
|---|---|---|---|---|---|
| Qwen3 8B | 115,3 | 87,5 | 67,9 | 4 049,6 | 1 714,6 |
| Qwen3 14B | 70,0 | 52,1 | 38,6 | 2 459,0 | 1 175,7 |
| gpt-oss 20B | 147,5 | 128,5 | 112,6 | 4 400,3 | 2 547,2 |
| Qwen3 30B A3B | 153,6 | 113,8 | 87,2 | 2 988,6 | 1 336,8 |
| Qwen 3.5 27B | 33,5 | 32,3 | 31,0 | 1 104,2 | 848,2 |
| Gemma 4 31B | 34,7 | 33,5 | 31,4 | 1 155,8 | 723,7 |
Two behaviors stand out. Standard models such as Qwen3 14B lose nearly half their throughput between 4k and 32k. Qwen 3.5 27B, on the other hand, loses only 7% over the same range (33.5 to 31.0). The source does not explain why, but for long documents this behavior matters more than model size. These figures are from Hardware Corner, not measurements by QuelLLM.
#Install Ollama and verify that the model fits
The 3090 is an Ampere card with compute capability 8.6, supported by Ollama with a NVIDIA 550 or newer driver; on Windows, the documentation requires version 551.61 or later. Flash Attention 2, which reduces context memory usage, works on Ampere according to its official repository; version 3 is reserved for Hopper (H100) and is not available on a 3090.
- 01Check the driverRun nvidia-smi in a terminal: the driver version must be 550 or later (551.61 on Windows), and total memory should show approximately 24 GB.
- 02Install OllamaInstall Ollama from its official website. The local API listens on http://localhost:11434.
- 03Running a 27-billion-parameter modelRun ollama run qwen3.5:27b: the 17 GB download happens once. For a faster model, try ollama run gpt-oss:20b (14 GB).
- 04Control memory allocationIn a second terminal, run ollama ps: the Processor column should display 100% GPU. A CPU/GPU split means the model or its context spills into system RAM.
- 05Adjust the contextSet the context with OLLAMA_CONTEXT_LENGTH, then run ollama ps again. If there is not enough headroom, set OLLAMA_FLASH_ATTENTION to 1 and OLLAMA_KV_CACHE_TYPE to q8_0 when starting the server.
#Throughput: what bandwidth really allows
Generation is bound by bandwidth: the GPU rereads all active weights for each token, so the theoretical ceiling is roughly bandwidth divided by model size. At 936 GB/s, Qwen3 14B (9.3 GB) tops out at about 100 tokens per second and Hardware Corner measures 70.0 at 4k, or 70%. Qwen 3.5 27B (17 GB) tops out at 55 and measures 33.5: 61%. Qwen3 32B (20 GB) tops out at 47 and measures 35.1: 75%. A dense 27B or 32B therefore runs at around 60 to 75% of its ceiling, or 33 to 35 tokens per second, which remains very comfortable to read.
The llama.cpp community leaderboard confirms the relationship between the two cards. On Llama 2 7B Q4_0 (3.56 GiB), the 3090 generates 158.16 tokens per second without Flash Attention and 161.89 with it, compared with 171.19 and 172.26 for the 3090 Ti: that’s +8% and +6%, consistent with the 8% increase in bandwidth. The two series come from different contributors, so the actual gap depends on the machine.
#Cooling, power limits, and buying used
The 3090 draws 350 W and the Ti 450 W, with maximum GPU temperatures of 93 and 92 °C respectively, according to NVIDIA. Under sustained LLM load, case noise and heat matter: plan for a well-ventilated case and a power supply that meets the 750 W (3090) or 850 W (Ti) requirements specified by NVIDIA. An LLM primarily stresses memory: the power limit costs little speed.
The llama.cpp leaderboard provides a rough benchmark: one contributor limits their 3090 to 250 W and records 137.72 tokens per second on Llama 2 7B Q4_0, compared with 158.16 in the stock series. That's about 13% less speed for 29% less power. The two measurements come from different machines: treat them as an indication, then measure on your own system.
The nvidia-smi -pl option sets the maximum power draw in watts. On Windows, open the terminal as an administrator; the setting disappears after reboot unless you automate it. On a used card, also monitor memory temperature with a tool such as HWiNFO: if it climbs significantly higher than the core temperature, consider replacing the thermal pads.
#Buying a used 3090: what to check
The 3090 is an Ampere card found mostly used. For current prices, check our tracker, which records the lowest price for each card twice a week, along with the price per GB of VRAM.
- Identity
- Use nvidia-smi to verify that the card reports 24 GB and the correct model (3090 or 3090 Ti), not another model.
- Stability
- Run a 27-billion-parameter model in a loop for around thirty minutes: crashes, artifacts, or thermal throttling are warning signs.
- Memory temperature
- Compare it with the memory reading: a significant gap indicates tired pads.
- Guarantee
- Require a return option or seller warranty: replacing pads or thermal paste is a cost to plan for.
- Power supply
- 350 W or 450 W peak: check the card's power supply and power connectors before buying.
#3090, 4090, 5090: what do you gain by paying more?
The RTX 4090 also offers 24 GB, with more compute: its advantage is more noticeable in prompt processing than generation, which remains limited by bandwidth. The RTX 5090 moves up to 32 GB, the category where a dense 32B model still has room for a long context. The guides for both cards detail every use case.
#vLLM, two cards, and fine-tuning on 3090s
Three use cases come up often with the 3090. vLLM works: according to its documentation, it requires compute capability 7.5 or higher, and the 3090 is 8.6. To choose between llama.cpp and vLLM, see the engine comparison. Two 3090s can be combined with an NVLink bridge for 30-series cards: the multi-GPU guide explains how the layers are split.
For fine-tuning, Unsloth lists minimum VRAM requirements in 4-bit QLoRA: 6 GB for an 8-billion-parameter model, 8.5 GB for 14 billion, 22 GB for 27 billion, and 26 GB for 32 billion. The 3090 can therefore fine-tune up to 27B in QLoRA, just barely, with a batch size of 1 and a short context, but not a 32B model. In 16-bit LoRA, 9 billion already requires 24 GB—the card’s entire capacity.
- llama.cpp, vLLM, Exllama: which engine to choose
- Multi-GPU LLM with llama.cpp: tensor-split across 2 RTX 3090
#Verdict: when the 3090 remains the right choice
| Situation | Decision | Quantified rationale |
|---|---|---|
| You want a dense 27B or 32B model locally | 3090 (or 4090) | 24 GB: 17 to 20 GB for weights and 4 to 7 GB of headroom |
| You are mainly targeting 7- to 14-billion-parameter models | A 12 or 16 GB card is sufficient | The 3090 adds only speed: 936 GB/s |
| You want a 70B or very long contexts | 5090 or two cards | 43 GB of weights for a 70B Q4 |
| Silence and power consumption come first | A recent 16 GB card | 750 W power supply required for the 3090 |
| You're deciding between the 3090 and 3090 Ti | Choose the cheaper of the two | +6 to +8% generation only |
The 3090 remains the most common entry point for a local 27B: that's Hardware Corner's conclusion, calling it the best used-market entry point for serious use. The 4090 replaces it if your budget allows, and the 5090 if you want 32 GB.
- Source: NVIDIA, specifications for RTX 3090 and 3090 Ti
- Source: Hardware Corner, LLM benchmarks on RTX 3090
- Source: llama.cpp community ranking on CUDA
- Source: Ollama documentation, context length
- Source: Unsloth, VRAM required for fine-tuning
#Frequently asked questions
Which LLM should you run on a RTX 3090?+
Can a RTX 3090 run a 70B model?+
RTX 3090 or 3090 Ti for a local LLM?+
What power supply does a RTX 3090 need?+
Does the RTX 3090 support vLLM and Flash Attention?+
Can you fine-tune a model on a RTX 3090?+
Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.