Qwen3.6 35B-A3B locally: testing and requirements VRAM
Qwen3.6 35B-A3B is a Mixture-of-Experts (MoE) model from Alibaba: 35 billion parameters in total, but only 3 billion activated per token. On paper, that promises the quality of a 30B+ model with the speed of a 3B. In practice, the catch is elsewhere: VRAM. This guide measures what you actually need to run qwen3.6 locally, by quantization level, and how many tokens/sec it delivers on common GPUs.
#Why run Qwen3.6 35B-A3B locally
Three concrete reasons make this model worth trying instead of a conventional dense 32B. First, speed: only the 3B active parameters are involved in computing each token, so inference is much faster than with a dense Qwen 32B at equivalent VRAM. Next, quality: on public benchmarks, the 35B-A3B holds its own against a dense 14B–24B model in reasoning, coding, and French.
Finally, it is one of the few models to offer this balance under a permissive license suitable for production. If you're looking for an accurate general-purpose assistant that responds quickly and fits on a 24 GB card, Qwen3.6 35B-A3B is currently one of the best candidates.
#The MoE architecture in one minute
Your private ChatGPT, free, on your own machine in an hour — LM Studio, Ollama, Open WebUI, your documents, no cloud.
- Lifetime online access
- PDF + files
- Lifetime updates
In a dense model, every token passes through all layers and all neurons. In an MoE, some feed-forward layers are replaced by a pool of independent experts, and a small “router” network decides which experts to activate for each token (typically 2 to 8). The rest stay idle.
- Compute per token
- Proportional to the active parameters (3B). That’s why it generates quickly.
- Memory (VRAM)
- Proportional to the total (35B). All experts must be loaded; you never know in advance which one will be called.
- Quality
- Between a dense 3B and a dense 35B. In practice, on general-purpose tasks, it comes close to a dense 14B–24B.
- Batch sensitivity
- For a single prompt at a time, MoE shines. In a very large batch, the compute advantage erodes because you eventually activate all the experts.
#Hardware requirements
- Minimum VRAM
- 16 GB to test it (Q3, partial CPU offload allowed). 20–24 GB for smooth Q4 use.
- Comfortable VRAM
- 32 GB (RTX 5090) or 48 GB (Apple unified memory) for Q5/Q6 and large contexts.
- System RAM
- 32 GB minimum if you tolerate CPU offload. 64 GB if you load into RAM only.
- Disk
- Allow 22 GB for Q4_K_M and 70 GB for FP16. An NVMe SSD is strongly recommended.
- OS / runtime
- Linux, macOS (Apple Silicon), Windows 11. Ollama ≥ 0.5.x or recent llama.cpp (build after March 2026).
#Actual VRAM by quantization
Here are the measured footprints for llama.cpp with an 8k-token context (model + KV cache + overhead). The values include the runtime, not just the GGUF file size on disk.
- Q2_K (≈ 13 GB)
- Possible on RTX 4070 / 4060 Ti 16 GB. Noticeable quality loss, especially in coding and multistep reasoning. Best reserved for troubleshooting.
- Q3_K_M (≈ 17 GB)
- Fits on RX 7900 GRE / 4080 / 5070 Ti 16 GB with a modest context. Still reasonable quality for general-purpose FR chat.
- Q4_K_M (≈ 22 GB)
- The sweet spot. Comfortable on RTX 3090 / 4090 / 7900 XTX (24 GB) and RTX 5090 (32 GB). Recommended by default.
- Q5_K_M (≈ 26 GB)
- Requires 32 GB (RTX 5090) or a 36 GB+ Mac M-series. Modest quality gain with Q4 except for code.
- Q6_K (≈ 30 GB)
- Reserved for RTX 5090, Mac Studio, or multi-GPU setups. Very little difference from Q8 in user-facing output.
- Q8_0 (≈ 37 GB)
- Mac Studio M2/M3/M5 Ultra, or dual-GPU (2× 3090). Near-FP16 quality.
- FP16 (≈ 70 GB)
- Research, fine-tuning, vLLM in production. Off-limits for most local setups.
#Installation with Ollama
Ollama remains the shortest path. If you don't have it yet, install it (see the installation guide for your OS). Once the daemon is running on port 11434, a single command is enough.
The download is about 21 GB. On the first prompt, the model is loaded into VRAM; subsequent starts are nearly instantaneous as long as the model remains active in memory.
The PROCESSOR column should show 100% GPU. If you see a CPU/GPU mix, VRAM is insufficient: switch to a more aggressive quantization (Q3_K_M), reduce num_ctx, or accept the throughput loss.
#Installation with llama.cpp
If you want fine-grained control over expert placement, batching, or the KV cache, llama.cpp is more flexible. Download a GGUF Q4_K_M from Hugging Face (models published by the Qwen team or by bartowski/unsloth), then start the OpenAI-compatible server.
- -ngl 99
- Offload all layers to the GPU. Set to 0 for CPU-only, or use a partial value to split them.
- -c 16384
- Context size. Count on ~0.5 GB of additional KV cache per 4k increment beyond the default.
- --flash-attn
- Enable Flash Attention if your build supports it (CUDA, Metal). Significant memory savings with long contexts.
- --threads 8
- Mainly useful with partial CPU offloading. Little impact on a pure GPU setup.
#Measured inference speed
Measurements taken in Q4_K_M, context 4096, short prompt (~200 tokens), 512-token generation, batch 1. The figures include the prompt evaluation and generation phases.
- RTX 5090 32 GB
- ≈ 110–125 tok/s during generation. The prompt-evaluation phase exceeds 4000 tok/s. Comfortable up to 32k context.
- RTX 4090 24 GB
- ≈ 80–95 tok/s. The Q4 model fits with 16k context without spilling over.
- RTX 3090 24 GB
- ≈ 55–70 tok/s. Excellent used performance-per-euro ratio, but lower memory bandwidth than the 4090.
- Mac Studio M5 Max 64 GB
- ≈ 60–75 tok/s in Q4_K_M, ≈ 45–55 tok/s in Q8_0. Ideal for a quiet 24/7 server.
- Radeon RX 7900 XTX 24 GB (ROCm)
- ≈ 45–60 tok/s. Ollama supported, llama.cpp with ROCm/Vulkan.
- RTX 4070 Ti Super 16 GB (Q3_K_M)
- ≈ 50–65 tok/s, with the context limited to 4 k. At this level, a dense Qwen 14B is often more relevant.
#Useful optimizations
- 01Enable Flash AttentionOn Ollama, this is automatic on a recent GPU. With llama.cpp, add --flash-attn. Noticeable VRAM savings starting at 8k context.
- 02Quantify the KV cachellama.cpp accepts --cache-type-k q8_0 --cache-type-v q8_0. Saves ~30% of VRAM on the cache, with virtually no quality loss.
- 03Reduce the number of active expertsSome variants accept --override-kv qwen3moe.expert_used_count=2 (instead of the default 4 or 8). Faster, with a slight loss in quality.
- 04Limit num_ctx to what you need16k is more than enough for most chats. 32k+ is only justified for synthesizing long documents.
- 05Prefer batch=1 for interactive chatMoE loses some of its advantage with large batches: if you serve multiple simultaneous users, consider vLLM rather than Ollama.
#Troubleshooting
- OOM during loading on RTX 4090
- Reduce num_ctx to 4096 and make sure no other GPU process is running (browser, game). Q4_K_M should fit with 2 to 3 GB to spare.
- Generation at 5 tok/s on RTX 4090
- The model is spilling onto the CPU. ollama ps will show a CPU/GPU mix. Close everything, restart Ollama, or drop to Q3_K_M.
- Inconsistent responses in French
- Make sure you're using the Instruct variant, not Base. MoE routing is sensitive to the wording of the system prompt—stay direct.
- Very slow first token (>10s)
- Long prompt-evaluation phase: this is expected for a large context on MoE. Reuse the prompt cache between requests (llama.cpp's --prompt-cache option).
- Model not found on Ollama
- The exact tag may differ (qwen3.6 vs qwen3_6 vs qwen36). Search ollama.com/library before entering the command.
#Go further
Now that Qwen3.6 35B-A3B is running on your machine, here are a few ways to push the setup further:
- Choose your quantization (Q4, Q5, Q8, FP16)
- To understand precisely what happens between Q3, Q4, and Q5 on an MoE model.
- Open WebUI with Ollama
- A ChatGPT-like interface with history and RAG, on top of your local Qwen3.6.
- Choose your GPU for local AI
- If you're deciding between a used 3090, 4090, or 5090 for this type of 30B+ MoE model.
Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.