Beginner 8 minNews

Qwen 3.8 open weights: released August 14 2026

The open weights for Qwen 3.8 are out: Alibaba published Qwen3.8-27B on Hugging Face on August 14, 2026, under the Apache 2.0 license, followed shortly by an official Ollama tag. This guide traces the complete timeline—from the announcement at WAIC in late July to the launch of Qwen 3.8-Max in the API on August 3, followed by the weights release—and explains the persistent confusion between Max, proprietary and out of reach for local use, and the dense 27B model you can install at home. For step-by-step installation, see our dedicated guide to Qwen 3.8 27B locally.

By Mohamed Meguedmi·Update 2026-08-18·Tested on Windows, macOS, and Linux

#The one-minute overview

If you came looking for a Ollama tag for Qwen 3.8, it has existed since August 14, 2026: “ollama run qwen3.8:27b”. Before going further, understand that generation 3.8 covers two very different models—a giant proprietary model served through an API, and a dense model with 27 billion parameters that you can download.

Qwen 3.8-Max
Alibaba's flagship model, an MoE with approximately 2,400 billion parameters and a 1M context. Available through the Alibaba Cloud API since August 3, 2026. Proprietary, not downloadable — this is not a local model.
Qwen3.8-27B (open weights)
Dense multimodal version published on August 14, 2026, on Hugging Face under the Apache 2.0 license. 262,144 native context tokens, 18 GB in Q4_K_M, official Ollama tag qwen3.8:27b.
What you can do today
Install the 27B directly on a 24 GB card or a 32 GB Mac. Our installation guide details the command, quantizations, and settings.
i
Guide updated August 18, 2026
This overview was rewritten after the weights were actually released. The Qwen ecosystem is moving fast: the official huggingface.co/Qwen page and the Ollama library remain the only authoritative sources on model availability.

#Verified timeline: from WAIC to the API

The Local AI Kit

Your private ChatGPT, free, on your own machine in an hour — LM Studio, Ollama, Open WebUI, your documents, no cloud.

  • Lifetime online access
  • PDF + files
  • Lifetime updates

Much of the confusion around Qwen 3.8 comes from conflating a conference demo with a product release. Here is the sequence, put back in order.

  1. 01
    Late July 2026 — WAIC (Shanghai)
    Alibaba introduces Qwen 3.8 at the World Artificial Intelligence Conference. This is where the first figures appear (1M context, massive MoE architecture) and where the intention to publish an open-weight variant is mentioned. At this stage, nothing is downloadable: these are announcements.
  2. 02
    August 3, 2026 — Launch of Qwen 3.8-Max via API
    The proprietary model becomes available through the Alibaba Cloud API (DashScope) and the Qwen app. This is a real release, but cloud-only. The reported benchmarks date from that point in time.
  3. 03
    3 to 13 August 2026 — Awaiting weights
    The community is waiting for the Qwen3.8-27B open-weight model mentioned at WAIC. The usual repositories (huggingface.co/Qwen, ModelScope) have received no 3.8-generation weights during this period, fueling doubts about whether the announcement will hold up.
  4. 04
    August 14, 2026 — Qwen3.8-27B weights released
    Alibaba publishes the Qwen/Qwen3.8-27B repository on Hugging Face under the Apache 2.0 license: a dense 27-billion-parameter model with native vision-language capabilities and a 262 144-token context. The Ollama library references the official tag the same day, with twelve variants (Q4_K_M, Q8, BF16, and MLX builds for Apple Silicon). Thus, eleven days separate the Max API launch from the open-weights drop.
i
Why this gap is normal
At Alibaba and elsewhere, the proprietary “Max” model launches first as an API to monetize and control usage. Smaller, denser open-weight variants often arrive a few weeks later, once the flagship version has stabilized. A delay between the WAIC announcement and the weight release is nothing unusual.

#What was actually released: Qwen 3.8-Max

Qwen 3.8-Max is the only model from this generation actually available today, and to be clear: it is not a model you can run at home. Its size rules it out for consumer self-hosting, and in any case its weights have not been published.

Architecture
A Mixture-of-Experts model with about 2,4 trillion parameters in total, with a fraction active per token. Impossible to load on consumer hardware, even high-end hardware.
Context
Advertised 1-million-token window, designed for long documents and entire codebases.
Availability
Alibaba Cloud API (DashScope) and Qwen application. Usage-based billing, no weight downloads.
Positioning
Direct competitor to frontier proprietary models. Compare it with cloud, not local.
!
“Local” doesn’t apply to Max
If a tutorial claims to install Qwen 3.8-Max on a PC with Ollama, be wary: an MoE with several trillion parameters cannot fit on any consumer machine, and its weights are not distributed. What will be locally installable is the dense 27B variant—when it is released.

#Open weights: what is promised vs. what is delivered

The distinction is now simple: what was announced at WAIC was delivered, with an eleven-day delay for the API model. Here is the factual summary.

Promised
A dense, open-weight Qwen3.8-27B under a permissive license, designed to fit on a 24 GB GPU when quantized.
Delivered
Exactly that, on August 14, 2026: safetensors weights on Hugging Face, Apache 2.0 license, GGUF and MLX builds distributed via Ollama, 18 GB in Q4_K_M—well within the 24 GB target.
As a bonus
Native vision (images and videos) and a 262,144-token context extendable to 1M via YaRN, which was not guaranteed by the initial announcements.
→
The anti-hype reflex still applies
Before assuming that a model is available locally, use a simple test: can the weights be downloaded now, with a command, without a waiting list? Between 3 and 13 August 2026, Qwen 3.8 failed this test even though dozens of articles already presented it as installable. Since 14 August, it has passed.

#The August 14 drop on Hugging Face

The official repository remains the best source for availability. Here's how to check for yourself, today and for the next generation—including against this article, which will eventually age too.

Verify via the Hugging Face CLI
pip install -U "huggingface_hub[cli]"

# Lister les dépôts de l'organisation Qwen contenant "3.8"
huggingface-cli repo list Qwen --limit 200 | grep -i "3.8"

# Depuis le 14/08/2026, la commande renvoie bien Qwen/Qwen3.8-27B

The official repository is Qwen/Qwen3.8-27B. Its model card lists 27 billion parameters, 64 layers, hybrid attention, native vision, and the Apache 2.0 license. On the Ollama side, twelve variants are published, from the 18 GB Q4_K_M to the 56 GB BF16 version.

Check on the Ollama side
# Depuis le 14 août 2026, le tag officiel existe :
ollama pull qwen3.8:27b   # 18 Go (Q4_K_M)

# Variante Apple Silicon (Metal)
ollama pull qwen3.8:27b-mlx

#Should you move on from Qwen 3.6?

If you are already running on Qwen 3.6-27B, the question is no longer “what should I install in the meantime?” but “is the update worth the download?” Short answer: yes, with equivalent memory requirements—but keep the old model around while you revalidate your prompts.

Qwen3.8-27B
Same form factor as the 3.6-27B (24 GB of VRAM in Q4), Apache 2.0 license, native vision, 262k context. The reported scores improve mainly on agentic coding and computer use.
Qwen 3.6-27B
Still valid: stable, proven, excellent in French. There's no need to migrate if your use cases are conversational.
Qwen3-Coder-Next
The code specialist remains faster at autocompletion. Qwen 3.8 27B makes sense when an agent needs to read, plan, and modify multiple files.

Quick memory reference in Q4_K_M, for sizing without mistakes: a 14B model requires about 9 GB of VRAM, a 27-32B model about 19 GB, and a 70B model about 40 GB. A 27B in Q4 therefore targets a 24 GB card (RTX 4090, RTX 3090) or a Mac with 24 GB or more of unified memory.

RTX 3060 12 GB
Comfortable for 8–14B models. The 27B does not fit entirely — CPU offloading and slow performance.
RTX 4090 / 3090 (24 GB)
The ideal target for a 27–32B model in Q4_K_M, entirely on the GPU.
Mac M4 Pro (24–48 GB unified)
The 27B runs very well thanks to unified memory; aim for 32 GB or more for long context.

#Install: the shortest path

The complete installation—quantizations, the MLX variant on Mac, context configuration, thinking mode, and troubleshooting—is covered in a dedicated guide. Here is the summary in three commands.

  1. 01
    1. Update Ollama
    The hybrid layers of Qwen 3.8 require a recent engine version. A version earlier than August 2026 will return an architecture error.
  2. 02
    2. Download the model
    18 GB in Q4_K_M, the default format. On Mac Apple Silicon, prefer the -mlx tag.
  3. 03
    3. Check the distribution
    ollama ps indique si le modèle tourne entièrement sur le GPU. Tant que ce n'est pas 100 % GPU, le débit reste médiocre.
Terminal
# 1. Le daemon répond ?
curl http://localhost:11434/api/tags

# 2. Télécharger Qwen 3.8 27B (Q4_K_M, 18 Go)
ollama pull qwen3.8:27b

# 3. Premier chat
ollama run qwen3.8:27b "Explique en une phrase la différence entre un modèle MoE et un modèle dense."

# 4. Vérifier que tout est bien sur le GPU
ollama ps

For strictly coding-focused use, the specialized variant from the previous generation remains relevant and lighter.

Code variant
ollama pull qwen3-coder-next
ollama run qwen3-coder-next "Écris une fonction Python qui parse un fichier CSV et renvoie la somme d'une colonne."
→
The migration from 3.6 is trivial
You remain within the Qwen family: a “ollama pull qwen3.8:27b” and a model-name change in your configuration are enough. Your prompts, interface, and scripts remain unchanged — only the sampling settings deserve a review, since thinking mode is enabled by default in 3.8.

#Key takeaways

The Qwen 3.8 episode is a textbook case of confusion between an announcement, an API release, and a weight release. Here are the verified, dated facts.

The Ollama tag
qwen3.8:27b, published on August 14, 2026, 18 GB in Q4_K_M. Twelve variants in total, including two MLX builds for Apple Silicon.
The license
Apache 2.0: commercial use, modification, and redistribution permitted, with no user threshold or restrictive use clause.
Published sizes
Only one dense size for now, the 27B. No smaller variant of the 3.8 generation has been published to date.
Benchmarks
The available scores come from the official model card. No independent reproduction has been published yet.
i
In summary
Qwen 3.8-Max has been available since August 3, 2026 (API, proprietary, outside the scope of local deployment). The Qwen3.8-27B open weights were released on August 14, 2026, under the Apache 2.0 license and can be installed with one command. For the complete step-by-step guide — VRAM, quantizations, context, thinking mode — see the dedicated installation guide.

#Go further

To move from the timeline to hands-on practice, these guides build on this one:

Qwen 3.8 27B locally
The complete installation guide: Ollama command, VRAM by quantization, MLX variant on Mac, the 256k context pitfall, and thinking-mode settings.
Install Ollama properly
“Install Ollama: Windows, macOS, and Linux” details the RAM/GPU requirements and the first model to run if you are starting from scratch.
Choose the right quantization
“Choosing your quantization (Q4, Q5, Q8, FP16)” explains the quality/memory tradeoff—decisive for fitting a 27B model into 24 GB.
GPU sizing
“Choosing a GPU for local AI” keeps you from underestimating VRAM before buying for a future 27B model.
Did this guide help you?

Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.