Intermediate 11 minQwen

Qwen3.6 35B-A3B locally: testing and requirements VRAM

Qwen3.6 35B-A3B is a Mixture-of-Experts (MoE) model from Alibaba: 35 billion parameters in total, but only 3 billion activated per token. On paper, that promises the quality of a 30B+ model with the speed of a 3B. In practice, the catch is elsewhere: VRAM. This guide measures what you actually need to run qwen3.6 locally, by quantization level, and how many tokens/sec it delivers on common GPUs.

By Mohamed Meguedmi·Update 2026-06-03·Tested on Windows, macOS, and Linux

#Why run Qwen3.6 35B-A3B locally

Three concrete reasons make this model worth trying instead of a conventional dense 32B. First, speed: only the 3B active parameters are involved in computing each token, so inference is much faster than with a dense Qwen 32B at equivalent VRAM. Next, quality: on public benchmarks, the 35B-A3B holds its own against a dense 14B–24B model in reasoning, coding, and French.

Finally, it is one of the few models to offer this balance under a permissive license suitable for production. If you're looking for an accurate general-purpose assistant that responds quickly and fits on a 24 GB card, Qwen3.6 35B-A3B is currently one of the best candidates.

i
The decoded name
35B = 35 billion total parameters. A3B = 3 billion "Active" parameters per token. The MoE router dynamically selects a few experts (out of ~64) on each pass, reducing compute cost but not memory cost.

#The MoE architecture in one minute

The Local AI Kit

Your private ChatGPT, free, on your own machine in an hour — LM Studio, Ollama, Open WebUI, your documents, no cloud.

  • Lifetime online access
  • PDF + files
  • Lifetime updates

In a dense model, every token passes through all layers and all neurons. In an MoE, some feed-forward layers are replaced by a pool of independent experts, and a small “router” network decides which experts to activate for each token (typically 2 to 8). The rest stay idle.

Compute per token
Proportional to the active parameters (3B). That’s why it generates quickly.
Memory (VRAM)
Proportional to the total (35B). All experts must be loaded; you never know in advance which one will be called.
Quality
Between a dense 3B and a dense 35B. In practice, on general-purpose tasks, it comes close to a dense 14B–24B.
Batch sensitivity
For a single prompt at a time, MoE shines. In a very large batch, the compute advantage erodes because you eventually activate all the experts.
→
Why it's interesting for self-hosting
On a RTX 4090 24 GB, a Qwen 32B dense model in Q4 often tops out at 25–35 tok/s. Qwen3.6 35B-A3B in the same VRAM runs at around 70–90 tok/s, with a comparable response level on most non-specialized tasks.

#Hardware requirements

Minimum VRAM
16 GB to test it (Q3, partial CPU offload allowed). 20–24 GB for smooth Q4 use.
Comfortable VRAM
32 GB (RTX 5090) or 48 GB (Apple unified memory) for Q5/Q6 and large contexts.
System RAM
32 GB minimum if you tolerate CPU offload. 64 GB if you load into RAM only.
Disk
Allow 22 GB for Q4_K_M and 70 GB for FP16. An NVMe SSD is strongly recommended.
OS / runtime
Linux, macOS (Apple Silicon), Windows 11. Ollama ≥ 0.5.x or recent llama.cpp (build after March 2026).
!
MoE ≠ small model
Do not be misled by “3B active”: you still have to load the entire model into VRAM. An RTX 4060 8 GB is not enough, even if inference “only computes 3B.” All experts must remain instantly accessible.

#Actual VRAM by quantization

Here are the measured footprints for llama.cpp with an 8k-token context (model + KV cache + overhead). The values include the runtime, not just the GGUF file size on disk.

Q2_K (≈ 13 GB)
Possible on RTX 4070 / 4060 Ti 16 GB. Noticeable quality loss, especially in coding and multistep reasoning. Best reserved for troubleshooting.
Q3_K_M (≈ 17 GB)
Fits on RX 7900 GRE / 4080 / 5070 Ti 16 GB with a modest context. Still reasonable quality for general-purpose FR chat.
Q4_K_M (≈ 22 GB)
The sweet spot. Comfortable on RTX 3090 / 4090 / 7900 XTX (24 GB) and RTX 5090 (32 GB). Recommended by default.
Q5_K_M (≈ 26 GB)
Requires 32 GB (RTX 5090) or a 36 GB+ Mac M-series. Modest quality gain with Q4 except for code.
Q6_K (≈ 30 GB)
Reserved for RTX 5090, Mac Studio, or multi-GPU setups. Very little difference from Q8 in user-facing output.
Q8_0 (≈ 37 GB)
Mac Studio M2/M3/M5 Ultra, or dual-GPU (2× 3090). Near-FP16 quality.
FP16 (≈ 70 GB)
Research, fine-tuning, vLLM in production. Off-limits for most local setups.
→
Which quantization to choose
For 95% of use cases, Q4_K_M is the right default: it preserves most of the quality, saves VRAM, and increases throughput. Move up to Q5/Q6 only if you have explicitly observed regressions on your prompts.

#Installation with Ollama

Ollama remains the shortest path. If you don't have it yet, install it (see the installation guide for your OS). Once the daemon is running on port 11434, a single command is enough.

Download and run
ollama run qwen3.6:35b-a3b-q4_K_M

The download is about 21 GB. On the first prompt, the model is loaded into VRAM; subsequent starts are nearly instantaneous as long as the model remains active in memory.

Check GPU placement
ollama ps

The PROCESSOR column should show 100% GPU. If you see a CPU/GPU mix, VRAM is insufficient: switch to a more aggressive quantization (Q3_K_M), reduce num_ctx, or accept the throughput loss.

i
Extend the context
By default, Ollama limits the context to 2048 tokens. For document summarization or long conversations, create a Modelfile that forces num_ctx to 8192 or 16384—each doubling costs a few hundred additional MB of VRAM.
Minimal Modelfile
# fichier : Modelfile.qwen36-long
FROM qwen3.6:35b-a3b-q4_K_M
PARAMETER num_ctx 16384
PARAMETER temperature 0.7

# build
# ollama create qwen36-long -f Modelfile.qwen36-long

#Installation with llama.cpp

If you want fine-grained control over expert placement, batching, or the KV cache, llama.cpp is more flexible. Download a GGUF Q4_K_M from Hugging Face (models published by the Qwen team or by bartowski/unsloth), then start the OpenAI-compatible server.

llama.cpp server
./llama-server \
  -m ./models/Qwen3.6-35B-A3B-Q4_K_M.gguf \
  --host 0.0.0.0 --port 8080 \
  -c 16384 \
  -ngl 99 \
  --flash-attn \
  --threads 8
-ngl 99
Offload all layers to the GPU. Set to 0 for CPU-only, or use a partial value to split them.
-c 16384
Context size. Count on ~0.5 GB of additional KV cache per 4k increment beyond the default.
--flash-attn
Enable Flash Attention if your build supports it (CUDA, Metal). Significant memory savings with long contexts.
--threads 8
Mainly useful with partial CPU offloading. Little impact on a pure GPU setup.
→
MoE and expert offloading
Recent llama.cpp builds accept --override-tensor to place certain experts in RAM. On a 4070 12 GB, this lets you fit Q4_K_M (at the cost of some throughput). Search the “expert offload” documentation if you want to dig deeper.

#Measured inference speed

Measurements taken in Q4_K_M, context 4096, short prompt (~200 tokens), 512-token generation, batch 1. The figures include the prompt evaluation and generation phases.

RTX 5090 32 GB
≈ 110–125 tok/s during generation. The prompt-evaluation phase exceeds 4000 tok/s. Comfortable up to 32k context.
RTX 4090 24 GB
≈ 80–95 tok/s. The Q4 model fits with 16k context without spilling over.
RTX 3090 24 GB
≈ 55–70 tok/s. Excellent used performance-per-euro ratio, but lower memory bandwidth than the 4090.
Mac Studio M5 Max 64 GB
≈ 60–75 tok/s in Q4_K_M, ≈ 45–55 tok/s in Q8_0. Ideal for a quiet 24/7 server.
Radeon RX 7900 XTX 24 GB (ROCm)
≈ 45–60 tok/s. Ollama supported, llama.cpp with ROCm/Vulkan.
RTX 4070 Ti Super 16 GB (Q3_K_M)
≈ 50–65 tok/s, with the context limited to 4 k. At this level, a dense Qwen 14B is often more relevant.
i
MoE and long context
Throughput remains high during generation, but prompt evaluation ramps up quickly when you inject 16k of context. That is normal: all experts are called upon at some point while processing the prompt. For RAG, batch your requests.

#Useful optimizations

  1. 01
    Enable Flash Attention
    On Ollama, this is automatic on a recent GPU. With llama.cpp, add --flash-attn. Noticeable VRAM savings starting at 8k context.
  2. 02
    Quantify the KV cache
    llama.cpp accepts --cache-type-k q8_0 --cache-type-v q8_0. Saves ~30% of VRAM on the cache, with virtually no quality loss.
  3. 03
    Reduce the number of active experts
    Some variants accept --override-kv qwen3moe.expert_used_count=2 (instead of the default 4 or 8). Faster, with a slight loss in quality.
  4. 04
    Limit num_ctx to what you need
    16k is more than enough for most chats. 32k+ is only justified for synthesizing long documents.
  5. 05
    Prefer batch=1 for interactive chat
    MoE loses some of its advantage with large batches: if you serve multiple simultaneous users, consider vLLM rather than Ollama.

#Troubleshooting

OOM during loading on RTX 4090
Reduce num_ctx to 4096 and make sure no other GPU process is running (browser, game). Q4_K_M should fit with 2 to 3 GB to spare.
Generation at 5 tok/s on RTX 4090
The model is spilling onto the CPU. ollama ps will show a CPU/GPU mix. Close everything, restart Ollama, or drop to Q3_K_M.
Inconsistent responses in French
Make sure you're using the Instruct variant, not Base. MoE routing is sensitive to the wording of the system prompt—stay direct.
Very slow first token (>10s)
Long prompt-evaluation phase: this is expected for a large context on MoE. Reuse the prompt cache between requests (llama.cpp's --prompt-cache option).
Model not found on Ollama
The exact tag may differ (qwen3.6 vs qwen3_6 vs qwen36). Search ollama.com/library before entering the command.

#Go further

Now that Qwen3.6 35B-A3B is running on your machine, here are a few ways to push the setup further:

Choose your quantization (Q4, Q5, Q8, FP16)
To understand precisely what happens between Q3, Q4, and Q5 on an MoE model.
Open WebUI with Ollama
A ChatGPT-like interface with history and RAG, on top of your local Qwen3.6.
Choose your GPU for local AI
If you're deciding between a used 3090, 4090, or 5090 for this type of 30B+ MoE model.
Did this guide help you?

Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.