Qwen 3.8 open weights: released August 14 2026
The open weights for Qwen 3.8 are out: Alibaba published Qwen3.8-27B on Hugging Face on August 14, 2026, under the Apache 2.0 license, followed shortly by an official Ollama tag. This guide traces the complete timeline—from the announcement at WAIC in late July to the launch of Qwen 3.8-Max in the API on August 3, followed by the weights release—and explains the persistent confusion between Max, proprietary and out of reach for local use, and the dense 27B model you can install at home. For step-by-step installation, see our dedicated guide to Qwen 3.8 27B locally.
#The one-minute overview
If you came looking for a Ollama tag for Qwen 3.8, it has existed since August 14, 2026: “ollama run qwen3.8:27b”. Before going further, understand that generation 3.8 covers two very different models—a giant proprietary model served through an API, and a dense model with 27 billion parameters that you can download.
- Qwen 3.8-Max
- Alibaba's flagship model, an MoE with approximately 2,400 billion parameters and a 1M context. Available through the Alibaba Cloud API since August 3, 2026. Proprietary, not downloadable — this is not a local model.
- Qwen3.8-27B (open weights)
- Dense multimodal version published on August 14, 2026, on Hugging Face under the Apache 2.0 license. 262,144 native context tokens, 18 GB in Q4_K_M, official Ollama tag qwen3.8:27b.
- What you can do today
- Install the 27B directly on a 24 GB card or a 32 GB Mac. Our installation guide details the command, quantizations, and settings.
#Verified timeline: from WAIC to the API
Your private ChatGPT, free, on your own machine in an hour — LM Studio, Ollama, Open WebUI, your documents, no cloud.
- Lifetime online access
- PDF + files
- Lifetime updates
Much of the confusion around Qwen 3.8 comes from conflating a conference demo with a product release. Here is the sequence, put back in order.
- 01Late July 2026 — WAIC (Shanghai)Alibaba introduces Qwen 3.8 at the World Artificial Intelligence Conference. This is where the first figures appear (1M context, massive MoE architecture) and where the intention to publish an open-weight variant is mentioned. At this stage, nothing is downloadable: these are announcements.
- 02August 3, 2026 — Launch of Qwen 3.8-Max via APIThe proprietary model becomes available through the Alibaba Cloud API (DashScope) and the Qwen app. This is a real release, but cloud-only. The reported benchmarks date from that point in time.
- 033 to 13 August 2026 — Awaiting weightsThe community is waiting for the Qwen3.8-27B open-weight model mentioned at WAIC. The usual repositories (huggingface.co/Qwen, ModelScope) have received no 3.8-generation weights during this period, fueling doubts about whether the announcement will hold up.
- 04August 14, 2026 — Qwen3.8-27B weights releasedAlibaba publishes the Qwen/Qwen3.8-27B repository on Hugging Face under the Apache 2.0 license: a dense 27-billion-parameter model with native vision-language capabilities and a 262 144-token context. The Ollama library references the official tag the same day, with twelve variants (Q4_K_M, Q8, BF16, and MLX builds for Apple Silicon). Thus, eleven days separate the Max API launch from the open-weights drop.
#What was actually released: Qwen 3.8-Max
Qwen 3.8-Max is the only model from this generation actually available today, and to be clear: it is not a model you can run at home. Its size rules it out for consumer self-hosting, and in any case its weights have not been published.
- Architecture
- A Mixture-of-Experts model with about 2,4 trillion parameters in total, with a fraction active per token. Impossible to load on consumer hardware, even high-end hardware.
- Context
- Advertised 1-million-token window, designed for long documents and entire codebases.
- Availability
- Alibaba Cloud API (DashScope) and Qwen application. Usage-based billing, no weight downloads.
- Positioning
- Direct competitor to frontier proprietary models. Compare it with cloud, not local.
#Open weights: what is promised vs. what is delivered
The distinction is now simple: what was announced at WAIC was delivered, with an eleven-day delay for the API model. Here is the factual summary.
- Promised
- A dense, open-weight Qwen3.8-27B under a permissive license, designed to fit on a 24 GB GPU when quantized.
- Delivered
- Exactly that, on August 14, 2026: safetensors weights on Hugging Face, Apache 2.0 license, GGUF and MLX builds distributed via Ollama, 18 GB in Q4_K_M—well within the 24 GB target.
- As a bonus
- Native vision (images and videos) and a 262,144-token context extendable to 1M via YaRN, which was not guaranteed by the initial announcements.
#The August 14 drop on Hugging Face
The official repository remains the best source for availability. Here's how to check for yourself, today and for the next generation—including against this article, which will eventually age too.
The official repository is Qwen/Qwen3.8-27B. Its model card lists 27 billion parameters, 64 layers, hybrid attention, native vision, and the Apache 2.0 license. On the Ollama side, twelve variants are published, from the 18 GB Q4_K_M to the 56 GB BF16 version.
#Should you move on from Qwen 3.6?
If you are already running on Qwen 3.6-27B, the question is no longer “what should I install in the meantime?” but “is the update worth the download?” Short answer: yes, with equivalent memory requirements—but keep the old model around while you revalidate your prompts.
- Qwen3.8-27B
- Same form factor as the 3.6-27B (24 GB of VRAM in Q4), Apache 2.0 license, native vision, 262k context. The reported scores improve mainly on agentic coding and computer use.
- Qwen 3.6-27B
- Still valid: stable, proven, excellent in French. There's no need to migrate if your use cases are conversational.
- Qwen3-Coder-Next
- The code specialist remains faster at autocompletion. Qwen 3.8 27B makes sense when an agent needs to read, plan, and modify multiple files.
Quick memory reference in Q4_K_M, for sizing without mistakes: a 14B model requires about 9 GB of VRAM, a 27-32B model about 19 GB, and a 70B model about 40 GB. A 27B in Q4 therefore targets a 24 GB card (RTX 4090, RTX 3090) or a Mac with 24 GB or more of unified memory.
- RTX 3060 12 GB
- Comfortable for 8–14B models. The 27B does not fit entirely — CPU offloading and slow performance.
- RTX 4090 / 3090 (24 GB)
- The ideal target for a 27–32B model in Q4_K_M, entirely on the GPU.
- Mac M4 Pro (24–48 GB unified)
- The 27B runs very well thanks to unified memory; aim for 32 GB or more for long context.
#Install: the shortest path
The complete installation—quantizations, the MLX variant on Mac, context configuration, thinking mode, and troubleshooting—is covered in a dedicated guide. Here is the summary in three commands.
- 011. Update OllamaThe hybrid layers of Qwen 3.8 require a recent engine version. A version earlier than August 2026 will return an architecture error.
- 022. Download the model18 GB in Q4_K_M, the default format. On Mac Apple Silicon, prefer the -mlx tag.
- 033. Check the distributionollama ps indique si le modèle tourne entièrement sur le GPU. Tant que ce n'est pas 100 % GPU, le débit reste médiocre.
For strictly coding-focused use, the specialized variant from the previous generation remains relevant and lighter.
#Key takeaways
The Qwen 3.8 episode is a textbook case of confusion between an announcement, an API release, and a weight release. Here are the verified, dated facts.
- The Ollama tag
- qwen3.8:27b, published on August 14, 2026, 18 GB in Q4_K_M. Twelve variants in total, including two MLX builds for Apple Silicon.
- The license
- Apache 2.0: commercial use, modification, and redistribution permitted, with no user threshold or restrictive use clause.
- Published sizes
- Only one dense size for now, the 27B. No smaller variant of the 3.8 generation has been published to date.
- Benchmarks
- The available scores come from the official model card. No independent reproduction has been published yet.
#Go further
To move from the timeline to hands-on practice, these guides build on this one:
- Qwen 3.8 27B locally
- The complete installation guide: Ollama command, VRAM by quantization, MLX variant on Mac, the 256k context pitfall, and thinking-mode settings.
- Install Ollama properly
- “Install Ollama: Windows, macOS, and Linux” details the RAM/GPU requirements and the first model to run if you are starting from scratch.
- Choose the right quantization
- “Choosing your quantization (Q4, Q5, Q8, FP16)” explains the quality/memory tradeoff—decisive for fitting a 27B model into 24 GB.
- GPU sizing
- “Choosing a GPU for local AI” keeps you from underestimating VRAM before buying for a future 27B model.
Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.