Family Qwen · 125B parameters

Qwen3.8 Flash Next 125B-A6B

Qwen3.8 Flash Next: multimodal 125B/6B active MoE, 256k context, ~72 GB Q4 VRAM. Fast vision, coding, and chat for self-hosting.

🇨🇳 Qwen·License Other (open weights)·Context 250k tokens·Output 2026-08-24·Tested on the GIGABYTE AI TOP ATOM · our measurements← Catalog

01What it can do

Strengths
  • 125B/6B active MoE: high throughput for its size
  • Native 256k context
  • Multimodal (vision, code, chat)
  • Fits on an 80 GB GPU in Q4 (~72 GB)
Limitations to know
  • —One server-class option remains (~72 GB VRAM Q4)
  • —“other” license — check the terms of use
  • —Recent weights: the quantized ecosystem is still young
Architecture
Multimodal MoE · 125B total parameters / 6B active per token · 256k context
Training
Flash Next model from the Qwen3.8 family, with a multimodal Mixture-of-Experts architecture (vision + text). “Other” license (open weights).
Ideal for
Multimodal chat & code256k long contextSingle-GPU 80 GB server

04Install

Install Ollama for your OS. Check the model and its quantization before downloading. Start with 4096 tokens of context, then check placement with ollama ps. A command below is not proof that a test was run on your machine.

$ollama pull qwen3.8-flash-next
⚠
First download: between 2 and 40 GB depending on the selected quantization. Plan for sufficient disk space; a stable connection is recommended. Subsequent launches are instant.

02Required memory

Approximate GPU VRAM required to run this model, including 4k tokens of context overhead. For a longer context, add ~1 GB per 8k-token increment.

Q4_K_M
The lightest, ~5% loss
72 GB
Q5_K_M
Good quality/size compromise
89 GB
Q8_0
Nearly indistinguishable from FP16
134 GB
FP16
Full precision — server use
250 GB
Fallback CPU · If you don't have a GPU, allow 162 GB of RAM minimum to run this model at reduced speed.

What hardware do you need for Qwen3.8 Flash Next 125B-A6B?

To run Qwen3.8 Flash Next 125B-A6B locally with Q4 quantization, you need about 72 GB of VRAM. An option to compare: BOSGAME M5 128GB / 2TB (Ryzen AI Max+ 395) — leave some headroom for the system and context; check engine compatibility with the GPU.

Current offer: BOSGAME M5 128GB / 2TB (Ryzen AI Max+ 395)
AmazonSee price →

Why this choice? Our complete guide on BOSGAME M5 128GB / 2TB (Ryzen AI Max+ 395) →

Affiliate links — commission possible at no extra cost to you; independent recommendations. As an Amazon Associate, BestLLMfor earns from qualifying purchases.

03Expected speed

Tokens generated per second in Q4_K_M, 4k context. Beyond 20 t/s, reading is comfortable. Below 10 t/s, that's just for testing.

Entry-level
~32t/s
GTX 1650, RX 6600, MBA M2 8GB
Mid-range
~50t/s
RTX 4060, 4070, MBP M3 Pro
High-end
~75t/s
RTX 4090, M4 Max, Radeon 7900