Advanced 13 minQwen

Qwen3.7 Max locally: the top open-weights model self-hostable

Qwen3.7 Max is Alibaba's largest open-weight model to date: a frontier-class MoE designed to compete with top-end proprietary models. Running it locally on qwen3.7 max local is not a mainstream exercise—you need serious hardware and a bit of methodology. This guide shows which quantizations actually fit on 24 to 48 GB per card, how to split the model across multiple GPUs, the throughput you can expect, and the cases where it clearly beats cloud APIs.

By Mohamed Meguedmi·Update 2026-06-03·Tested on Windows, macOS, and Linux

#Why Qwen3.7 Max

On nearly all public benchmarks from late 2026, Qwen3.7 Max ranks first among open-weight models: multi-step reasoning, coding, math, long-instruction following in Chinese and English—and French that is no longer embarrassing. At this level, there’s no longer a technical reason to send your prompts to a cloud provider, except when you need massive throughput.

The model is released under the Tongyi Qianwen license (Apache-2.0 for most of the weights). The Instruct, Coder, and Thinking variants are published on Hugging Face and mirrored on ollama.com/library. For anyone with a well-equipped AI workstation, it is currently the most credible open-weight alternative to GPT-5 Mini or Claude Sonnet 4.

i
Who this guide is for
You have at least one RTX 4090, an M5 Ultra Mac Studio, or a dual-GPU 3090/4090 setup. If you have only one 12–16 GB card, stick with Qwen3.6 35B-A3B or Qwen3 32B—Qwen3.7 Max is not for you.

#The model at a glance

The Local AI Kit

Your private ChatGPT, free, on your own machine in an hour — LM Studio, Ollama, Open WebUI, your documents, no cloud.

  • Lifetime online access
  • PDF + files
  • Lifetime updates
Architecture
Mixture-of-Experts. Approximately 480 billion total parameters, ~46B active per token (8 experts routed out of 128).
Native context
256k tokens via YaRN, 1M tokens in experimental extension. The vast majority of use cases remain under 32k.
Tokenizer
Tiktoken-style multilingual BPE, compact ratio in French (≈ 1.4 tokens/word—comparable to GPT-4).
Variants
Qwen3.7-Max-Instruct (general chat), -Coder (programming), -Thinking (explicit reasoning like DeepSeek R1).
License
Apache-2.0 for the main weights. Commercial use is allowed, redistribution is permitted, and there is no strict anti-competition clause.
!
MoE = full weights in memory
Even if only 46B parameters are active per token, all 480B must remain accessible. The router switches experts for every token—it is impossible to “unload” inactive ones without incurring catastrophic I/O costs. Plan your total VRAM accordingly.

#Hardware requirements

Total VRAM (Q4_K_M)
Allow ≈ 270 GB for the model + KV cache. This is the target with multi-GPU or a 256 GB Mac Studio Ultra.
Minimum viable workstation
Mac Studio M2 Ultra 192 GB (Q3_K_M, 8k context) — the only consumer single-machine setup that works without dual GPUs.
Comfortable workstation
4× RTX 4090 24 GB (96 GB total, Q2_K + offload), or 2× RTX 6000 Ada 48 GB, or a Mac Studio M3 Ultra or M5 Ultra 256 GB for full Q4_K_M.
System RAM
128 GB minimum if you plan to offload experts. 256 GB recommended if you load into RAM alone (very slow but functional).
Disk
≈ 270 GB for Q4_K_M, ≈ 1 TB for FP16. A Gen4 NVMe SSD is mandatory—otherwise initial loading takes 20+ minutes.
Runtime
llama.cpp build after October 2026 (tensor-parallel MoE support), vLLM ≥ 0.7, or Ollama ≥ 0.6 (if the GGUF variant is released).

#Quantizations that fit on 24–48 GB

The practical question is: what fits without breaking everything? The figures below include the model, a KV cache for 8k context, and runtime overhead. They assume you add the VRAM of all your cards together.

IQ1_M (≈ 110 GB)
Fits on 2× 48 GB or a 128 GB Mac Studio. Degraded quality on code and long-form reasoning—useful for experimenting, not production.
IQ2_XS (≈ 140 GB)
Mac Studio 192 GB or 4× RTX 4090 96 GB (with ~40 GB of RAM offload). First honest tier for general-purpose chat in French.
Q3_K_M (≈ 200 GB)
Mac Studio M3 Ultra or M5 Ultra 256 GB, or 2× RTX 6000 Ada 96 GB + 128 GB RAM with offloading. A good quality/cost compromise.
Q4_K_M (≈ 270 GB)
The quality sweet spot. Mac Studio Ultra 256 GB handles it with a short context; otherwise, 4× A6000 48 GB (192 GB) + offload, or a DGX node.
Q5_K_M (≈ 330 GB)
Reserved for H100 80 GB ×4, MI300X 192 GB ×2, or a cluster. Marginal gain over Q4 in practice.
Q8_0 (≈ 510 GB)
Datacenter only. At this level, you might as well serve it with vLLM in tensor-parallel mode on 8× H100s.
→
The right default
If your goal is serious local use, target Q4_K_M and be prepared to invest in total VRAM. Quants below Q3 show visible regressions as soon as you move beyond trivial chat—typically in long code generation or following instructions with multiple constraints.

#Installation

Three paths depending on your hardware. Ollama for simplicity, llama.cpp for control, vLLM for serving multiple users.

#Path 1: Ollama (Mac Studio Ultra)

On a Mac Studio Ultra, Ollama manages unified memory automatically. The daemon listens on http://localhost:11434.

Terminal
# Vérifier qu'Ollama est à jour
ollama --version

# Télécharger la variante adaptée
ollama pull qwen3.7-max:q3_K_M

# Premier lancement (chargement long ~3 min)
ollama run qwen3.7-max:q3_K_M

While it loads, monitor ollama ps in a second terminal. The SIZE column should show the expected size (≈ 200 GB for Q3_K_M).

#Path 2: llama.cpp (multi-GPU NVIDIA)

Download the Q4_K_M GGUF from Hugging Face (search for: Qwen3.7-Max-Instruct-GGUF, models published by the Qwen team or by bartowski). The download is 270 GB—plan for the bandwidth.

llama.cpp tensor-parallel server
./llama-server \
  -m ./models/Qwen3.7-Max-Q4_K_M.gguf \
  --host 0.0.0.0 --port 8080 \
  -c 16384 \
  --tensor-split 24,24,24,24 \
  --main-gpu 0 \
  --flash-attn \
  --cache-type-k q8_0 --cache-type-v q8_0
--tensor-split 24,24,24,24
Distributes the layers evenly across 4 GPUs. Adapt to the GB available per card (e.g., 24,24 for 2 GPUs).
--cache-type-k q8_0
Quantizes the KV cache, frees about 30% of VRAM, with negligible quality loss.
--flash-attn
Essential beyond 8k context at this parameter count.
-c 16384
16k of context is a good compromise. Increasing it to 32k costs several additional GB on 480B.

#Path 3: vLLM (production, multi-user)

If you serve a team, vLLM with tensor parallelism makes better use of PCIe bandwidth and delivers significantly higher batch throughput.

Launching vLLM
vllm serve Qwen/Qwen3.7-Max-Instruct-AWQ \
  --tensor-parallel-size 4 \
  --max-model-len 32768 \
  --quantization awq \
  --gpu-memory-utilization 0.92 \
  --enable-expert-parallel
i
AWQ vs. GGUF
vLLM prefers AWQ or GPTQ quants, optimized for NVIDIA GPUs. GGUF remains king on CPU/Mac and in heterogeneous environments. For a single NVIDIA-only machine, 4-bit AWQ in vLLM outperforms GGUF Q4_K_M in llama.cpp on batch throughput (×2 to ×3 depending on the workload).

#Multi-GPU distribution

Three different strategies depending on your hardware and workload.

  1. 01
    Layer split (llama.cpp by default)
    Each GPU takes a contiguous block of layers. Simple, and it works with heterogeneous GPUs (e.g., 3090 + 4090). Bottleneck: only one GPU computes at a time while the others wait. Throughput is limited by the slowest card.
  2. 02
    Tensor parallel (vLLM, recent llama.cpp)
    Each layer is divided horizontally among all GPUs, which compute in parallel. Requires adequate PCIe bandwidth (Gen4 x16 or NVLink) and identical cards. Throughput 1.8× to 2.5× higher than layer split on 4 GPUs.
  3. 03
    Expert parallel (MoE-specific)
    Distributes experts across GPUs. For Qwen3.7 Max with 128 experts on 4 cards, each GPU handles 32 experts. Reduces VRAM per card, but each token activates experts on multiple GPUs—good option if you have PCIe Gen5 or NVLink.
→
PCIe really matters here
With an MoE of this size, routing tokens between experts generates constant inter-GPU traffic. A consumer motherboard with PCIe Gen4 x4 on the second slot will cut throughput by 30-50%. If you build a quad-GPU system, choose a ThreadRipper or Xeon chassis with full lanes.
Check multi-GPU placement
# Sous Linux, surveille la conso VRAM en direct
watch -n 1 nvidia-smi

# Avec llama.cpp, vérifie que tous les GPU travaillent
nvidia-smi dmon -s u -c 30

#Expected tokens/sec throughput

Measurements taken in Q4_K_M (unless indicated otherwise), with a 4k context, short prompt, 512-token generation, and batch 1. Throughput drops logarithmically as context length increases—the prompt-evaluation phase explodes beyond 16k tokens.

Mac Studio M3 Ultra or M5 Ultra 256 GB (Q3_K_M)
≈ 18–24 tok/s during generation. The prompt-evaluation phase remains adequate (~600 tok/s). Excellent for continuous solo use, quiet 24/7.
Mac Studio M2 Ultra 192 GB (IQ2_XS)
≈ 22–28 tok/s, degraded quality. Acceptable for testing, not for production.
2× RTX 6000 Ada 48 GB (Q3_K_M, layer split)
≈ 28–35 tok/s. Good pro workstation option.
4× RTX 4090 24 GB (Q4_K_M, tensor-parallel llama.cpp)
≈ 35–45 tok/s in chat, 80–120 tok/s in batch=4. The best performance/€ ratio in a DIY configuration.
4× RTX 4090 (vLLM AWQ tensor-parallel)
≈ 50–65 tok/s in solo chat, up to 300+ tok/s aggregated in large batches. Preferred for serving a team.
2× H100 80 GB (Q4_K_M, NVLink)
≈ 75–90 tok/s in chat. Pro market, but NVLink changes everything for this model profile.
i
What these figures do not show
On an MoE of this size, time to first token with a 32k context can reach 8 to 15 seconds. For RAG or long prompts, cache your prefixes (llama.cpp's --prompt-cache option or vLLM's prefix caching)—that's what makes the experience tolerable.

#Frontier open-weight vs. cloud

At comparable performance, what justifies running Qwen3.7 Max locally instead of paying for an API?

Real privacy
Prompts never leave your network. For lawyers, medical services, and HR teams, this is non-negotiable — and no cloud “zero retention” policy is a substitute for an air gap.
Zero marginal cost
Once the hardware investment has paid for itself, every token is free. For intensive workloads (synthetic dataset generation, analysis batches), the ROI surpasses the cloud's within a few months.
Deep customization
You can fine-tune, modify the default system prompt, connect custom tools, and intercept logprobs. No frontier API gives you that.
Vendor independence
No quota, no rate limit, no model disappearing at the next pricing refresh. The model stays on your disk.
Predictable latency
No server congestion, no “the model is overloaded” in the middle of a client demo.
!
Where the cloud still leads
For occasional interactive use at very low latency or for serving hundreds of requests per second, the cloud remains more economical. Qwen3.7 Max locally shines under sustained batch load, not at occasional peaks—it's the same logic as on-premises servers versus SaaS.

#Troubleshooting

OOM while loading on quad-4090
Make sure --tensor-split adds up to the VRAM actually available (not the displayed VRAM). Leave 1-2 GB of headroom per card for CUDA overhead.
Throughput below 10 tok/s
Most likely, a GPU is spilling over into system RAM. nvidia-smi will show 100% VRAM usage and a card capped at 30% utilization. Reduce the quantization or add VRAM.
Time-to-first-token > 30s
Prompt eval phase saturated. Enable --flash-attn, quantize the KV cache, and reduce num_ctx to the strict minimum. If possible, batch your requests.
Quality falls off a cliff vs. benchmarks
Make sure you are using the Instruct variant (not Base), that the chat template is applied (Ollama does this, llama.cpp does not without --chat-template), and that quantization has not dropped below Q3.
Crash after a few hours
Often thermal: 4 GPUs at 100% heat up the case. Check junction temperatures, adjust the fan curves, and undervolt by -50 mV to -100 mV. See the thermal guide.
Extreme slowness on the very first prompt
The model loads from the SSD — allow 2 to 5 minutes for 270 GB. Subsequent prompts are instantaneous as long as the model remains in memory.

#Go further

Qwen3.7 Max is a serious hardware commitment. Here are a few related guides to properly size your setup before buying:

Multi-GPU setup with llama.cpp
The practical guide to splitting a large model across 2 or 4 NVIDIA GPUs, including PCIe and NVLink pitfalls.
Which LLM on Mac Studio (M2/M3/M4 Ultra, 64–512 GB)?
Compare what unified memory Apple enables with multi-GPU NVIDIA for this model size.
Deploy vLLM in production
If Qwen3.7 Max needs to serve multiple simultaneous users, vLLM with tensor parallelism is the right component.
Did this guide help you?

Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.