Which LLM for 24 GB of VRAM ?
With 24 GB of VRAM, load a 27-billion-parameter dense model in Q4 such as Qwen 3.8 27B (16 GB of weights, 8 GB of headroom) or a 30- to 35-billion-parameter MoE (19 to 21 GB). A dense 70B in Q4 (approximately 40 GB) will not fit: it requires two cards or unified memory.
Twenty-four gigabytes of VRAM is the threshold for 27- to 35-billion-parameter models. This guide explains what fits, which cards actually have 24 GB, what speed ceiling their bandwidth allows, and how far fine-tuning can go.
Choosing a machine? Our picks by budget →
Good value for local AI: a GMKtec EVO-X2 64GB / 1TB (Ryzen AI Max+ 395).
A mini PC is a complete machine: check the available memory and engine compatibility. It does not replace macOS/MLX or CUDA.
Why this choice? Our complete guide on GMKtec EVO-X2 64GB / 1TB (Ryzen AI Max+ 395) →
Compare all options by budget, from €800 to €3,500 →
Small budget: RTX 5060 · Large models: RTX 5090 · Mac Studio.
On the go: which laptop for local AI →
Affiliate links — possible commission at no extra cost to you. As an Amazon Associate, BestLLMfor earns from qualifying purchases.
#24 GB of VRAM: the tier for 27- to 35-billion-parameter models
With 24 GB of VRAM, you can load a dense 27-billion-parameter model entirely into memory in Q4 (16 GB of weights according to the QuelLLM catalog), with 8 GB of headroom for context, or a 30- to 35-billion-parameter MoE model with 3 billion active parameters (19 to 21 GB), with a shorter context. A dense 70B model in Q4 weighs approximately 40 GB: it does not fit and requires two cards or unified memory. The 24 GB tier is where you move from 9- to 14-billion-parameter models to models that rival yesterday's large models.
| Model | Q4 weight | Margin on 24 GB | Note |
|---|---|---|---|
| gpt-oss 20B (MoE) | 13 GB | 11 GB | Very large, long context |
| Devstral Small 2 24B | 14 GB | 10 GB | Code agent |
| Qwen 3.8 27B | 16 GB | 8 GB | Dense, general-purpose, with a stated context of 262,144 tokens |
| Gemma 4 31B | 18 GB | 6 GB | Dense, Q5 at 22 GB is too tight |
| Qwen3-Coder 30B-A3B (MoE) | 19 GB | 5 GB | Code, 3 billion active parameters |
| Qwen 3.6 35B-A3B (MoE) | 21 GB | 3 GB | Short context, barely fits |
| Qwen 3.8 27B in Q8 | 29 GB | No | Overflows |
A correction to an earlier version of this guide: it sometimes listed 18 GB and sometimes 16 GB for Qwen 3.8 27B, and 23 GB for Qwen 3.6 35B-A3B. The catalog lists 16 GB and 21 GB. These figures are rounded Q4 weights; always add the KV cache and about one gigabyte of headroom.
#Which cards have 24 GB, and what the difference is
Three consumer cards dominate this tier. NVIDIA lists the RTX 3090 with 24 GB of GDDR6X, and the RTX 4090 with 24 GB of GDDR6X. AMD specifies 24 GB of GDDR6 and bandwidth of up to 960 GB/s for the RX 7900 XTX. The RTX 5090, with 32 GB of GDDR7 on a 512-bit bus according to NVIDIA, belongs to the next tier. Do not confuse it with the RX 7900 XT, which has 20 GB: llama.cpp's performance table lists it with 20 GB of GDDR6.
| Card | Software | Preferable if | Guide |
|---|---|---|---|
| RTX 3090 / 3090 Ti | CUDA | You want CUDA and a moderate entry price on the used market | LLM on RTX 3090 |
| RTX 4090 | CUDA | You want CUDA on a newer architecture | LLM on RTX 4090 |
| RX 7900 XTX | ROCm 7 or Vulkan | You are on Linux and accept ROCm | LLM on RX 7900 XTX |
The choice comes down to three criteria: the software ecosystem, the card’s age (condition, warranty), and the current price, which we do not lock in. The site’s price tracker shows the cost per gigabyte of VRAM, updated twice a week; that’s the right metric for choosing between a used 24 GB card and a new 16 GB card.
#Which model to choose on 24 GB depending on the use case
| Usage | Model | Why |
|---|---|---|
| Versatile general-purpose model | Qwen 3.8 27B Q4 | Dense, with 8 GB of headroom for context |
| Speed and reasoning | gpt-oss 20B | Light MoE, 11 GB of headroom |
| Development code and agents | Devstral Small 2 24B or Qwen3-Coder 30B-A3B | 14 or 19 GB, context depending on headroom |
| Largest MoE model | Qwen 3.6 35B-A3B Q4 | 21 GB, 3 billion active parameters, short context |
| Agents and tool calls | GLM 4.7 Flash Q4 | 19 GB, MIT license according to the catalog |
The choice between dense and MoE is the real trade-off at this tier. A 27-billion-parameter dense model reads all its weights for every token: it is slower than an MoE model, but leaves more VRAM for context. A 35-billion-parameter MoE reads only its active parameters, so it generates faster, but its total weights occupy almost the entire card. The guide to MoE models explains the mechanism.
#What about 70B? What 24 GB really allows
A dense 70B in Q4 weighs approximately 40 GB in weights according to the site’s benchmark: it exceeds 24 GB by at least 16 GB. This portion would have to be offloaded to system memory, so generation would run at the speed of RAM and the PCIe bus, not the graphics card. We do not provide a throughput figure for this case because we lack a verifiable source: assume it is too slow for comfortable interactive use. MoE models with 30 to 35 billion parameters are, at comparable quality, the practical answer to this limitation.
- A single 24 GB card
- 30–35B MoE in Q4, 27–31B dense in Q4 or Q5, with 3 to 8 GB for context.
- Two 24 GB cards (48 GB)
- 70B in Q4 (approximately 40 GB) with a little context, distributing the layers across the cards.
- A 32 GB card
- The RTX 5090: a 70B in Q4 still cannot run on its own, but dense models with 27 to 32 billion parameters become comfortable.
In llama.cpp, the default distribution mode splits the model by layer across the cards in a pipeline, with adjustable proportions for each card. The benefit is capacity, not speed per token, since the cards work one after another.
#Speed: the ceiling set by bandwidth
Generation reads the essential weights of a dense model at every token: the theoretical ceiling is bandwidth divided by weight size. AMD's product page lists 960 GB/s for the RX 7900 XTX; for NVIDIA cards, check the bandwidth on the exact model's product page. The table gives the ceiling in 100 GB/s increments for dense models; MoE models read fewer weights per token and run faster.
| Q4 model | Weights read | Per 100 GB/s | Example: 960 GB/s (7900 XTX) |
|---|---|---|---|
| Devstral Small 2 24B | 14 GB | ≈ 7 tokens/s | ≈ 69 tokens/s |
| Qwen 3.8 27B | 16 GB | ≈ 6 tokens/s | 60 tokens/s |
| Gemma 4 31B | 18 GB | ≈ 5.6 tokens/s | ≈ 53 tokens/s |
These values are ceilings, never measurements: actual speed is lower and varies with the driver, software, and context. We do not publish tokens per second by card because there is no comparable public benchmark for these models. To measure your own installation, run ollama run with the --verbose option.
#The context: the real limit on 24 GB
On 24 GB, the KV cache determines the usable context length. A Qwen 3.8 27B in Q4 leaves 8 GB, while a Qwen 3.6 35B-A3B leaves only 3 GB. By default, Ollama uses a context of 4,096 tokens according to its documentation, adjustable with OLLAMA_CONTEXT_LENGTH. To save space, quantize the cache: OLLAMA_KV_CACHE_TYPE=q8_0 uses about half the memory of the default f16, provided Flash Attention is enabled.
#A three-line margin calculation
To determine whether a model and its context will fit, add three terms: the model weights, the KV cache, and about one gigabyte of headroom. Let’s take two examples using the model weights. A Qwen 3.8 27B in Q5 weighs 19 GB: that leaves 24 minus 19 minus 1, or 4 GB for the KV cache. A Qwen 3.6 35B-A3B in Q4 weighs 21 GB: that leaves 24 minus 21 minus 1, or just 2 GB, meaning a short context unless you quantize the cache. The same 27B in Q4 (16 GB) frees up 7 GB, allowing a much longer context.
This calculation explains a common choice: prefer Q4 quantization over Q5 when context is the priority, and Q5 over Q4 when you work with short texts and model fidelity matters. The quantization selection guide details the quality differences, while the KV cache guide explains the space savings.
#Buying a used 24 GB: checks to perform
The most common 24 GB cards date from 2020 to 2022: many are sold used, and some have run continuously for a long time. Before buying, ask for a screenshot of a load test, check the memory temperature during the test, and listen for fan noise. Make sure the advertised capacity is actually 24 GB in the driver, that your power supply supports the card's power draw, and that there is enough room in the case. Prices for these cards change from week to week: the site's price tracker lets you compare the cost per gigabyte of VRAM between new and used cards.
#Fine-tuning on 24 GB: what works and what doesn’t
Unsloth's documentation publishes a table of minimum VRAM by model size and method, specifying that these are absolute minimums. In 4-bit QLoRA, a 7B model requires 5 GB, a 14B model 8.5 GB, and a 27B model 22 GB. In 16-bit LoRA, an 8B model requires 22 GB and a 14B model 33 GB. With 24 GB, QLoRA for a model of up to 14 billion parameters is therefore comfortable, QLoRA for a 27B model barely fits with no headroom, and 16-bit LoRA for a 14B model is impossible.
| Model | QLoRA (4-bit) | LoRA (16-bit) | On 24 GB |
|---|---|---|---|
| 7B | 5 GB | 19 GB | QLoRA and LoRA |
| 9B | 6.5 GB | 24 GB | QLoRA; LoRA without margin |
| 14B | 8.5 GB | 33 GB | QLoRA only |
| 27B | 22 GB | 64 GB | QLoRA at the limit |
| 32B | 26 GB | 76 GB | No |
| 70B | 41 GB | 164 GB | No |
- Which LLM for 32 GB of VRAM
- Which LLM for 16 GB of VRAM
- LLM on RTX 3090 / 3090 Ti (24 GB)
- LLM on RX 7900 XTX (24 GB)
- MoE explained: why a 30B-A3B runs like a small model
- Local AI graphics card prices
- Source: Unsloth VRAM requirements
- Source: Ollama FAQ (context, KV cache)
- Source: NVIDIA datasheet from RTX 4090
- Source: AMD specifications for the RX 7900 XTX
#Frequently asked questions
What's the best LLM for 24 GB of VRAM?+
With 24 GB, should you still target a 70B model?+
RTX 3090, RTX 4090, or RX 7900 XTX for 24 GB?+
How many tokens per second on 24 GB?+
Is 24 GB enough for fine-tuning?+
Can you use two 24 GB cards together?+
Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.