The sizing math in 30 seconds
You can size any GGUF model with one multiplication. Take the parameter count — 27 — and multiply by the bytes-per-parameter of the quantization: FP16 is about 2.0 GB per billion, Q8_0 about 1.07, Q6_K about 0.82, Q5_K_M about 0.71, and Q4_K_M about 0.58. That gives you the weight footprint. Then add roughly 20% for the KV cache, activation buffers, and runtime overhead at an 8K context window. Double the context and the KV portion roughly doubles with it, so a 32K session costs several more gigabytes on top of the weights.
These are rules of thumb, not promises — actual file sizes vary a little by tokenizer and tensor layout. If you want the exact number for your context length, plug the figures into our VRAM calculator, and if the quant names above are unfamiliar, start with quantization explained.
Qwen 3.6 27B VRAM by quantization
Here is the full sizing table for the 27B model. "Weights" is the raw model footprint; "Total @8K" adds about 20% for the KV cache and runtime overhead at an 8,192-token context. Push the context to 32K and you should budget roughly 3–5 GB more on top of the 8K figure, since the KV cache grows with context length.
| Quant | Bytes/B | Weights (GB) | ~Total @8K (GB) | Fits 16 GB? |
|---|---|---|---|---|
Q4_K_M | 0.58 | 15.7 | 18.8 | Partial offload |
Q5_K_M | 0.71 | 19.2 | 23.0 | No |
Q6_K | 0.82 | 22.1 | 26.5 | No |
Q8_0 | 1.07 | 28.9 | 34.7 | No |
| FP16 | 2.00 | 54.0 | 64.8 | No |
The takeaway: 16 GB holds the Q4_K_M weights but not much else, 24 GB is the sweet spot for Q5/Q6, and full-precision FP16 is a data-center proposition at roughly 65 GB in use.
Which GPUs fit which quant
VRAM is the hard wall. The table below pairs common cards with the highest quant that runs cleanly at an 8K context. "Runs cleanly" means the weights plus KV cache stay on the GPU; where a quant is one notch too big, you can still run it by shortening the context or letting the runtime offload a few layers to system RAM, at a speed cost.
| GPU (VRAM) | Best quant @8K | Notes |
|---|---|---|
| RTX 5070 Ti / 4060 Ti 16GB (16 GB) | Q4_K_M | Weights fit; KV pushes past 16 GB, so expect a few offloaded layers or a shorter context |
| RTX 3090 / 4090 (24 GB) | Q5_K_M | Q4 is comfortable; Q6 needs a trimmed context |
| RTX 5090 (32 GB) | Q6_K | Q8 runs with a short context or light offload |
| RTX A6000 / dual 24 GB (48 GB) | Q8_0 | Full quality-preserving quant with headroom |
| A100 / H100 (80 GB) | FP16 | Only tier that holds full precision comfortably |
Quality loss per quant, from real runs
Numbers only tell half the story; here is what actually changed when we ran each quant on the RTX 5070 Ti. Q8_0 is effectively indistinguishable from FP16 — if you have the VRAM, there is no reason to prefer FP16 for inference. Q6_K is what we would call near-lossless: across chat, summarization, and everyday coding we could not reliably tell it apart from Q8 in blind spot-checks. Q5_K_M introduces the first perceptible softening, usually on long-chain reasoning and strict instruction-following, but it stays dependable for most work.
Q4_K_M is the honest trade-off. It is the quant most 16 GB owners will run, and it is genuinely usable — but the dip is real on multi-step code generation, where it drops a subtle edge case or misreads intent slightly more often than Q6. If your workload is code-heavy, budget for a 24 GB card so you can move up to Q5 or Q6. We keep updated head-to-head results on the leaderboard, and the trade-offs at tighter budgets in our 12 GB benchmark set.
Ollama and LM Studio tags
Both Ollama and LM Studio pull GGUF builds, and both expose the quantization through the model tag. The naming convention is stable even if a specific release is not: a tag like qwen3.6:27b-q4_K_M or a repository suffix such as -Q5_K_M tells you exactly which quant you are downloading. Because model publishers occasionally rename or re-roll builds, confirm the exact tag on the official library rather than trusting a number you saw in a blog post.
In Ollama the pattern is ollama run <model>:<tag>; in LM Studio you pick the quant from the download dropdown on the model card. For a curated list of what pulls cleanly, see our best Ollama models roundup. For exact, current tags check the Ollama library, the Qwen organization on Hugging Face, and the llama.cpp repository for the definitive bytes-per-weight of each GGUF quant type.
Frequently asked questions
How much VRAM does Qwen 3.6 27B need?
For the weights alone, plan on about 16 GB at Q4_K_M, 19 GB at Q5, 22 GB at Q6, 29 GB at Q8, and 54 GB at FP16. Add roughly 20% for the KV cache and overhead at an 8K context, so a Q4_K_M setup uses closer to 19 GB in practice.
Can I run Qwen 3.6 27B on a 16GB GPU?
Yes, but only at Q4_K_M, and it is tight. The weights fit in 16 GB, but the KV cache spills past it, so Ollama or llama.cpp will offload a few layers to system RAM or you will need to shorten the context window.
Which quantization is best for Qwen 3.6 27B?
Q6_K is the best quality-to-size balance — it is near-lossless in our testing while saving significant memory over Q8. If you are limited to 16 GB, Q4_K_M is the practical choice; with 24 GB or more, step up to Q5 or Q6 for cleaner reasoning and code.
How much RAM do I need to run Qwen 3.6 27B on CPU?
CPU inference needs roughly the same amount of system RAM as the GPU figures — about 16 to 19 GB for Q4_K_M — so 32 GB of RAM is a comfortable minimum. Expect much slower token generation than on a GPU, since 27B models are memory-bandwidth bound.
By Mohamed Meguedmi — independent comparator of locally-runnable LLMs, benchmarked on a real RTX 5070 Ti (data CC BY 4.0). See the local LLM leaderboard and the best Ollama models.