Family Qwen · 27B parameters

Qwen 3 27B

Dense 27B from Qwen, Apache 2.0, 32k context, ~16 GB VRAM in Q4: fits on a 16 GB card for chat, reasoning, and multilingual use.

🇨🇳 Alibaba·License Apache 2.0·Context 32k tokens·Output —·Fits within the 128 GB of the GIGABYTE AI TOP ATOM← Catalog

01What it can do

Strengths
  • Dense 27B: ~16 GB VRAM in Q4, fits on a 16 GB card
  • Permissive Apache 2.0 license (free commercial use)
  • Chat, reasoning, and multilingual support
  • Dense alternative to MoE Qwen of comparable size
Limitations to know
  • —No verified Ollama tag: installation via Hugging Face
  • —32k native context (default value, to be confirmed on the model card)
  • —Gated weights on Hugging Face (acceptance required)
Architecture
Dense transformer · 27B parameters · 32k context
Training
Dense 27B model from the Qwen 3 family (Alibaba), released under Apache 2.0. Corpus and method details not disclosed.
Ideal for
General-purpose assistantReasoningMultilingual

04Install

Install Ollama for your OS. Check the model and its quantization before downloading. Start with 4096 tokens of context, then check placement with ollama ps. A command below is not proof that a test was run on your machine.

$# HuggingFace : Qwen/Qwen3-27B
⚠
First download: between 2 and 40 GB depending on the selected quantization. Plan for sufficient disk space; a stable connection is recommended. Subsequent launches are instant.

02Required memory

Approximate GPU VRAM required to run this model, including 4k tokens of context overhead. For a longer context, add ~1 GB per 8k-token increment.

Q4_K_M
The lightest, ~5% loss
16 GB
Q5_K_M
Good quality/size compromise
19 GB
Q8_0
Nearly indistinguishable from FP16
29 GB
FP16
Full precision — server use
54 GB
Fallback CPU · If you don't have a GPU, allow 35 GB of RAM minimum to run this model at reduced speed.

What hardware do you need for Qwen 3 27B?

To run Qwen 3 27B locally with Q4 quantization, you need about 16 GB of VRAM. An option to compare: GMKtec EVO-X2 64GB / 1TB (Ryzen AI Max+ 395) — leave some headroom for the system and context; check engine compatibility with the GPU.

Current offer: GMKtec EVO-X2 64GB / 1TB (Ryzen AI Max+ 395)
AmazonSee price →

Why this choice? Our complete guide on GMKtec EVO-X2 64GB / 1TB (Ryzen AI Max+ 395) →

Affiliate links — commission possible at no extra cost to you; independent recommendations. As an Amazon Associate, BestLLMfor earns from qualifying purchases.

On the go: Qwen 3 27B also runs on a RTX laptop PC (16 GB of VRAM) →

This model in your private ChatGPT, without the cloud

Your private ChatGPT, free, on your own machine in an hour — LM Studio, Ollama, Open WebUI, your documents, no cloud.

  • Lifetime online access
  • PDF + files
  • Lifetime updates

03Expected speed

Tokens generated per second in Q4_K_M, 4k context. Beyond 20 t/s, reading is comfortable. Below 10 t/s, that's just for testing.

Entry-level
~14t/s
GTX 1650, RX 6600, MBA M2 8GB
Mid-range
~18t/s
RTX 4060, 4070, MBP M3 Pro
High-end
~22t/s
RTX 4090, M4 Max, Radeon 7900