Intermediate 11 minBy VRAM

Which LLM for 24 GB of VRAM ?

Direct response

With 24 GB of VRAM, load a 27-billion-parameter dense model in Q4 such as Qwen 3.8 27B (16 GB of weights, 8 GB of headroom) or a 30- to 35-billion-parameter MoE (19 to 21 GB). A dense 70B in Q4 (approximately 40 GB) will not fit: it requires two cards or unified memory.

Twenty-four gigabytes of VRAM is the threshold for 27- to 35-billion-parameter models. This guide explains what fits, which cards actually have 24 GB, what speed ceiling their bandwidth allows, and how far fine-tuning can go.

Choosing a machine? Our picks by budget →

By Mohamed Meguedmi·Update 2026-09-29·Tested on Windows, macOS, and Linux
Recommended hardware

Good value for local AI: a GMKtec EVO-X2 64GB / 1TB (Ryzen AI Max+ 395).

A mini PC is a complete machine: check the available memory and engine compatibility. It does not replace macOS/MLX or CUDA.

Why this choice? Our complete guide on GMKtec EVO-X2 64GB / 1TB (Ryzen AI Max+ 395) →

Compare all options by budget, from €800 to €3,500 →

Small budget: RTX 5060 · Large models: RTX 5090 · Mac Studio.

On the go: which laptop for local AI →

Affiliate links — possible commission at no extra cost to you. As an Amazon Associate, BestLLMfor earns from qualifying purchases.

#24 GB of VRAM: the tier for 27- to 35-billion-parameter models

With 24 GB of VRAM, you can load a dense 27-billion-parameter model entirely into memory in Q4 (16 GB of weights according to the QuelLLM catalog), with 8 GB of headroom for context, or a 30- to 35-billion-parameter MoE model with 3 billion active parameters (19 to 21 GB), with a shorter context. A dense 70B model in Q4 weighs approximately 40 GB: it does not fit and requires two cards or unified memory. The 24 GB tier is where you move from 9- to 14-billion-parameter models to models that rival yesterday's large models.

What fits in 24 GB (QuelLLM catalog weights, Q4 unless noted)
ModelQ4 weightMargin on 24 GBNote
gpt-oss 20B (MoE)13 GB11 GBVery large, long context
Devstral Small 2 24B14 GB10 GBCode agent
Qwen 3.8 27B16 GB8 GBDense, general-purpose, with a stated context of 262,144 tokens
Gemma 4 31B18 GB6 GBDense, Q5 at 22 GB is too tight
Qwen3-Coder 30B-A3B (MoE)19 GB5 GBCode, 3 billion active parameters
Qwen 3.6 35B-A3B (MoE)21 GB3 GBShort context, barely fits
Qwen 3.8 27B in Q829 GBNoOverflows

A correction to an earlier version of this guide: it sometimes listed 18 GB and sometimes 16 GB for Qwen 3.8 27B, and 23 GB for Qwen 3.6 35B-A3B. The catalog lists 16 GB and 21 GB. These figures are rounded Q4 weights; always add the KV cache and about one gigabyte of headroom.

#Which cards have 24 GB, and what the difference is

Three consumer cards dominate this tier. NVIDIA lists the RTX 3090 with 24 GB of GDDR6X, and the RTX 4090 with 24 GB of GDDR6X. AMD specifies 24 GB of GDDR6 and bandwidth of up to 960 GB/s for the RX 7900 XTX. The RTX 5090, with 32 GB of GDDR7 on a 512-bit bus according to NVIDIA, belongs to the next tier. Do not confuse it with the RX 7900 XT, which has 20 GB: llama.cpp's performance table lists it with 20 GB of GDDR6.

Three ways to get 24 GB
CardSoftwarePreferable ifGuide
RTX 3090 / 3090 TiCUDAYou want CUDA and a moderate entry price on the used marketLLM on RTX 3090
RTX 4090CUDAYou want CUDA on a newer architectureLLM on RTX 4090
RX 7900 XTXROCm 7 or VulkanYou are on Linux and accept ROCmLLM on RX 7900 XTX

The choice comes down to three criteria: the software ecosystem, the card’s age (condition, warranty), and the current price, which we do not lock in. The site’s price tracker shows the cost per gigabyte of VRAM, updated twice a week; that’s the right metric for choosing between a used 24 GB card and a new 16 GB card.

#Which model to choose on 24 GB depending on the use case

A starting point for each use case
UsageModelWhy
Versatile general-purpose modelQwen 3.8 27B Q4Dense, with 8 GB of headroom for context
Speed and reasoninggpt-oss 20BLight MoE, 11 GB of headroom
Development code and agentsDevstral Small 2 24B or Qwen3-Coder 30B-A3B14 or 19 GB, context depending on headroom
Largest MoE modelQwen 3.6 35B-A3B Q421 GB, 3 billion active parameters, short context
Agents and tool callsGLM 4.7 Flash Q419 GB, MIT license according to the catalog

The choice between dense and MoE is the real trade-off at this tier. A 27-billion-parameter dense model reads all its weights for every token: it is slower than an MoE model, but leaves more VRAM for context. A 35-billion-parameter MoE reads only its active parameters, so it generates faster, but its total weights occupy almost the entire card. The guide to MoE models explains the mechanism.

#What about 70B? What 24 GB really allows

A dense 70B in Q4 weighs approximately 40 GB in weights according to the site’s benchmark: it exceeds 24 GB by at least 16 GB. This portion would have to be offloaded to system memory, so generation would run at the speed of RAM and the PCIe bus, not the graphics card. We do not provide a throughput figure for this case because we lack a verifiable source: assume it is too slow for comfortable interactive use. MoE models with 30 to 35 billion parameters are, at comparable quality, the practical answer to this limitation.

A single 24 GB card
30–35B MoE in Q4, 27–31B dense in Q4 or Q5, with 3 to 8 GB for context.
Two 24 GB cards (48 GB)
70B in Q4 (approximately 40 GB) with a little context, distributing the layers across the cards.
A 32 GB card
The RTX 5090: a 70B in Q4 still cannot run on its own, but dense models with 27 to 32 billion parameters become comfortable.

In llama.cpp, the default distribution mode splits the model by layer across the cards in a pipeline, with adjustable proportions for each card. The benefit is capacity, not speed per token, since the cards work one after another.

#Speed: the ceiling set by bandwidth

Generation reads the essential weights of a dense model at every token: the theoretical ceiling is bandwidth divided by weight size. AMD's product page lists 960 GB/s for the RX 7900 XTX; for NVIDIA cards, check the bandwidth on the exact model's product page. The table gives the ceiling in 100 GB/s increments for dense models; MoE models read fewer weights per token and run faster.

Theoretical ceiling for dense models per 100 GB/s slice
Q4 modelWeights readPer 100 GB/sExample: 960 GB/s (7900 XTX)
Devstral Small 2 24B14 GB≈ 7 tokens/s≈ 69 tokens/s
Qwen 3.8 27B16 GB≈ 6 tokens/s60 tokens/s
Gemma 4 31B18 GB≈ 5.6 tokens/s≈ 53 tokens/s

These values are ceilings, never measurements: actual speed is lower and varies with the driver, software, and context. We do not publish tokens per second by card because there is no comparable public benchmark for these models. To measure your own installation, run ollama run with the --verbose option.

#The context: the real limit on 24 GB

On 24 GB, the KV cache determines the usable context length. A Qwen 3.8 27B in Q4 leaves 8 GB, while a Qwen 3.6 35B-A3B leaves only 3 GB. By default, Ollama uses a context of 4,096 tokens according to its documentation, adjustable with OLLAMA_CONTEXT_LENGTH. To save space, quantize the cache: OLLAMA_KV_CACHE_TYPE=q8_0 uses about half the memory of the default f16, provided Flash Attention is enabled.

32k context with q8_0 KV cache
OLLAMA_FLASH_ATTENTION=1 OLLAMA_KV_CACHE_TYPE=q8_0 OLLAMA_CONTEXT_LENGTH=32768 ollama serve
ollama ps

#A three-line margin calculation

To determine whether a model and its context will fit, add three terms: the model weights, the KV cache, and about one gigabyte of headroom. Let’s take two examples using the model weights. A Qwen 3.8 27B in Q5 weighs 19 GB: that leaves 24 minus 19 minus 1, or 4 GB for the KV cache. A Qwen 3.6 35B-A3B in Q4 weighs 21 GB: that leaves 24 minus 21 minus 1, or just 2 GB, meaning a short context unless you quantize the cache. The same 27B in Q4 (16 GB) frees up 7 GB, allowing a much longer context.

This calculation explains a common choice: prefer Q4 quantization over Q5 when context is the priority, and Q5 over Q4 when you work with short texts and model fidelity matters. The quantization selection guide details the quality differences, while the KV cache guide explains the space savings.

#Buying a used 24 GB: checks to perform

The most common 24 GB cards date from 2020 to 2022: many are sold used, and some have run continuously for a long time. Before buying, ask for a screenshot of a load test, check the memory temperature during the test, and listen for fan noise. Make sure the advertised capacity is actually 24 GB in the driver, that your power supply supports the card's power draw, and that there is enough room in the case. Prices for these cards change from week to week: the site's price tracker lets you compare the cost per gigabyte of VRAM between new and used cards.

#Fine-tuning on 24 GB: what works and what doesn’t

Unsloth's documentation publishes a table of minimum VRAM by model size and method, specifying that these are absolute minimums. In 4-bit QLoRA, a 7B model requires 5 GB, a 14B model 8.5 GB, and a 27B model 22 GB. In 16-bit LoRA, an 8B model requires 22 GB and a 14B model 33 GB. With 24 GB, QLoRA for a model of up to 14 billion parameters is therefore comfortable, QLoRA for a 27B model barely fits with no headroom, and 16-bit LoRA for a 14B model is impossible.

Minimum fine-tuning VRAM according to Unsloth
ModelQLoRA (4-bit)LoRA (16-bit)On 24 GB
7B5 GB19 GBQLoRA and LoRA
9B6.5 GB24 GBQLoRA; LoRA without margin
14B8.5 GB33 GBQLoRA only
27B22 GB64 GBQLoRA at the limit
32B26 GB76 GBNo
70B41 GB164 GBNo
!
A corrected error: QLoRA 70B does not fit in 24 GB
An earlier version of this guide listed QLoRA for a “tangential” 70B at 22–23 GB. Unsloth indicates that QLoRA for a 70B requires at least 41 GB, versus 22 GB for a 27B. The 70B therefore requires at least two 24 GB cards, or one higher-capacity card.

#Frequently asked questions

Frequently asked questions
What's the best LLM for 24 GB of VRAM?+
For a general-purpose model, Qwen 3.8 27B in Q4: 16 GB of weights according to the QuelLLM catalog, leaving 8 GB for context. For speed, gpt-oss 20B (13 GB). For coding, Devstral Small 2 24B (14 GB) or Qwen3-Coder 30B-A3B (19 GB). For a large MoE, Qwen 3.6 35B-A3B (21 GB), with a short context.
With 24 GB, should you still target a 70B model?+
No. A dense 70B in Q4 weighs about 40 GB: it exceeds available RAM by at least 16 GB, and the speed becomes too low for interactive use. A 30- to 35-billion-parameter MoE fits in the card. For a 70B, plan on two 24 GB cards, or 48 GB.
RTX 3090, RTX 4090, or RX 7900 XTX for 24 GB?+
All three have 24 GB. The choice comes down to the ecosystem: CUDA for the two RTX cards, ROCm 7 or Vulkan for the RX 7900 XTX, and the current price, which we do not lock in. Check the site's price tracker, which calculates the cost per gigabyte of VRAM.
How many tokens per second on 24 GB?+
It depends on the model and the card. For a dense model, the theoretical ceiling is bandwidth divided by the weights: about 60 tokens/s for a 27B model at 16 GB on 960 GB/s, never reached exactly. We don't publish per-card measurements; ollama run with --verbose gives you yours.
Is 24 GB enough for fine-tuning?+
For QLoRA on a model of up to 14 billion parameters, yes, with room to spare: Unsloth indicates a minimum of 8.5 GB for a 14B. A 27B requires 22 GB, right at the limit. 16-bit LoRA for a 14B (33 GB) and QLoRA for a 70B (41 GB) do not fit in 24 GB.
Can you use two 24 GB cards together?+
Yes, with llama.cpp, whose default mode distributes the layers and KV cache across the GPUs. You get 48 GB, enough for a 70B model in Q4, but the per-token speed does not double because the cards work in a pipeline. Plan for the power supply and cooling required by two cards.
Did this guide help you?

Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.