Qwen 3.8 27B locally: Ollama installation and VRAM
The Qwen3.8-27B weights were released on August 14, 2026, on Hugging Face under the Apache 2.0 license, and the official Ollama tag followed the same day. This is the first model in the 3.8 generation that can genuinely be installed locally: dense, multimodal (image and video), with 262,144 tokens of native context. This guide provides the exact command, the actual VRAM required for each tag, the settings for “thinking” mode enabled by default, and the two pitfalls that ruin most first attempts—silent context truncation and spilling into RAM.
#The verdict in 30 seconds
Qwen 3.8 27B installs with one command: “ollama run qwen3.8:27b”. The default package (Q4_K_M) takes up 18 GB and requires a 24 GB graphics card—RTX 3090, 4090, 5090—or a Mac Apple Silicon with at least 32 GB of unified memory for comfortable use. Below that, the model still runs, but some layers spill over to the CPU and throughput collapses.
- What it is
- A dense model with 27 billion parameters, native vision-language support (images and videos), 262,144-token context, under the Apache 2.0 license—so it can be used commercially without restrictive clauses.
- The realistic prerequisite
- 24 GB of VRAM or 32 GB of unified memory for the Q4_K_M version. 30 GB of files and ~40 GB of VRAM for Q8, 56 GB for BF16.
- The command
- ollama pull qwen3.8:27b puis ollama run qwen3.8:27b. Sur Mac, préférez le tag -mlx, optimisé Metal.
- The #1 trap
- The advertised context is 256k, but Ollama applies a much smaller default window and truncates without warning. You need to set num_ctx manually.
- Pitfall No. 2
- “Thinking” mode is enabled by default and consumes many tokens before responding. Locally, lower reasoning_effort or disable it for simple tasks.
#What Qwen 3.8 27B really is
Your private ChatGPT, free, on your own machine in an hour — LM Studio, Ollama, Open WebUI, your documents, no cloud.
- Lifetime online access
- PDF + files
- Lifetime updates
Unlike the Mixture-of-Experts trend, Qwen3.8-27B is a dense model: all 27 billion parameters are activated for every token. That makes its memory usage predictable—no pleasant surprises in throughput, no unpleasant surprises in VRAM—and it also explains its lower local throughput than a comparably sized MoE.
- Architecture
- 64 layers organized into 16 blocks of “3 × (Gated DeltaNet → FFN) followed by 1 × (Gated Attention → FFN).” A hybrid attention design: most layers use memory-efficient linear attention, interspersed with true attention layers for precision.
- Terms
- Native vision-language: still images, documents, scientific diagrams, and videos. This is not an adapter bolted on afterward.
- Context
- 262,144 tokens natively, extendable to 1,000,000 via YaRN—but only on vLLM, SGLang, or TokenSpeed, not through Ollama.
- License
- Apache 2.0. Commercial use permitted, no user threshold as with Llama, and no use clause as with Gemma.
- Special feature
- The model is trained with Multi-Token Prediction (MTP)—hence the Ollama “mtp” tags, which speed up generation on engines that know how to use them.
#The hardware you need, tag by tag
The Ollama library publishes twelve variants. Download size is not the VRAM required: you also need to add the KV cache, which grows with context length, and the vision encoder. Allow 15 to 25% of headroom above the file size.
| Tag | Size | Comfortable memory | Who it's for |
|---|---|---|---|
| qwen3.8:27b (Q4_K_M, default) | 18 GB | 24 GB | The default choice: RTX 3090/4090/5090, 32 GB Mac |
| qwen3.8:27b-q8_0 | 30 GB | 40 GB and up | Near-maximum quality: RTX Pro 6000, dual-GPU 24 GB |
| qwen3.8:27b-bf16 | 56 GB | 80 GB | Reference weights, compute machines (H100, H200) |
| qwen3.8:27b-mlx | 18 GB | 32 GB unified memory | Apple Silicon: Metal build, the right default on Mac |
| qwen3.8:27b-mxfp8 | 32 GB | 48 GB unified | Mac Studio / M4 Max 64 GB, higher quality |
| qwen3.8:27b-mlx-bf16 | 56 GB | 96 GB unified | Mac Studio 128 GB, no quantization loss |
| qwen3.8:27b-mtp-q4_K_M | 18 GB | 24 GB | Same as the default, with explicit multi-token prediction |
For throughput, our estimates for a dense 27B in Q4 are around 14 tokens/second on a mid-range configuration and 22 tokens/second on recent high-end hardware. These are calculated estimates, not measurements: the “thinking” mode enabled by default can also double the perceived time before the first response line.
#Install Qwen 3.8 27B with Ollama
- 01Update OllamaThe hybrid layers and vision parser of Qwen 3.8 require a recent version of Ollama. A version older than August 2026 will return an architecture error or a misleading “model not found.” Reinstall from the official site, or rerun the installation script on Linux.
- 02Download the model18 GB are transferred: make sure there is room on the system drive, where Ollama stores its blobs. The pull can resume if the connection drops.
- 03Start a first exchangeThe first response is slower while the weights are loaded into memory. If startup takes more than a minute, that indicates an overflow onto the processor.
- 04Check where the model is runningThe ollama ps command displays the GPU/CPU split. Until you see 100% GPU, every token is expensive — drop down one quantization level or reduce the context.
#On Mac Apple Silicon: choose the MLX variant
Ollama publishes an MLX build compiled for Apple's Metal engine. At the same file size (18 GB), it makes better use of unified memory and delivers noticeably higher throughput than generic GGUF on M3, M4, and later chips.
- 16 GB Mac
- Insufficient. macOS only makes about 70% of unified memory available to the GPU: the model swaps and becomes unusable.
- Mac 24 GB
- It just fits with a short context and no other demanding application open. Acceptable for testing, frustrating in everyday use.
- 32 GB Mac
- The real entry point: the model fits, leaving room for the KV cache and the system.
- Mac 64 GB and up
- Comfortable, and lets you run in mxfp8 or work with long documents.
#The 256k context trap
This is the mistake almost all first-time users make. The model advertises 262,144 tokens, but Ollama applies a much shorter context window by default: beyond that limit, the beginning of your document is simply cut off, with no error message. The model confidently answers about text it has not seen in full.
As for the advertised million tokens: it relies on a YaRN extension, configured in the model's configuration file, and is supported only by vLLM, SGLang, or TokenSpeed. There is currently no way to access it from Ollama.
#Thinking mode and sampling parameters
Qwen 3.8 thinks before answering: it produces a reasoning block between « think » tags, followed by the final answer. This is enabled by default, and it explains most of the perceived latency. On engines that support it, the depth is controlled with reasoning_effort — xhigh by default, followed by medium and low.
| Mode | temperature | top_p | top_k | presence_penalty |
|---|---|---|---|---|
| Thinking (default) | 1.0 | 0.95 | 20 | 0.0 |
| Instruct (no reasoning) | 0.7 | 0.80 | 20 | 1.5 |
- Short local tasks
- Set reasoning_effort to low or medium: for rewriting or extraction, extended reasoning does not change quality and triples the wait time.
- Agentic tasks
- Keep xhigh. Alibaba notes that reducing the effort leads to insufficient analysis, and therefore retries—the gain per turn is canceled out by failures.
- Responses that get stuck in a loop
- Increase presence_penalty (between 0 and 2). Above 1.5, the model starts mixing languages.
- Output polluted by reasoning
- If your application displays reasoning tags, it is not separating reasoning_content from the final content. Disable reasoning for this use case instead of filtering the text afterward.
#Analyze an image or document
Vision is native: screenshots, whiteboard photos, technical diagrams, scanned pages. On the command line, just provide the file path in the prompt; through the API, images are passed as base64 in the field provided by Ollama.
#What the reported scores are worth
The claimed leap over Qwen 3.6–27B, the same-size model from the previous generation, is substantial — especially for agentic coding and computer use. Here are the figures published by Alibaba, keeping in mind that they have not been independently reproduced.
| Test | Qwen3.8-27B | Qwen3.6-27B |
|---|---|---|
| Terminal Bench 2.1 (agentic coding) | 73,0 | 63,4 |
| SWE-bench Pro | 61,7 | 53,5 |
| LiveCodeBench v6 | 90,3 | 83,9 |
| IFBench (instruction following) | 79,5 | 69,1 |
| GPQA Diamond (science) | 89,2 | 87,8 |
| OSWorld-Verified (computer use) | 84,3 | 63,9 |
| MathVision (visual math) | 90,0 | 85,1 |
| OmniDocBench 1.5 (documents) | 91,1 | 89,4 |
What to take away for local use: the gains are concentrated in long, tool-driven tasks — operating a terminal, fixing a repository, and chaining steps without going off track. In a simple conversation or summary, the difference from the previous generation will be far less noticeable than these figures suggest.
#Should you move on from Qwen 3.6, Gemma 4, or gpt-oss?
- You are on Qwen 3.6-27B
- Yes, the update is worth it: same memory footprint, same permissive license, clear gains in coding and agentic tasks. Keep the old tag while you validate your prompts; the outputs change in style.
- You're on Gemma 4 26B
- Qwen 3.8 retains the edge in coding and agentic capabilities; in terms of licensing, both are now Apache 2.0, with Gemma having become permissive in April 2026. Gemma 4 (MoE 26B-A4B, multimodal) often remains more natural in conversational French and is faster to start.
- You're running gpt-oss-20b
- Two philosophies: gpt-oss is lighter and faster, Qwen 3.8 sees images, handles a much longer context, and targets long-running tasks. Choose based on the available VRAM.
- You're mainly looking to code
- A specialized 30B code model in MoE remains faster for autocompletion. Qwen 3.8 27B makes sense when the agent needs to read, plan, and modify multiple files.
- You have less than 24 GB
- Don't force it. A 12-14B model in Q4 will give you a better experience than a 27B model that spills over onto the CPU.
#Troubleshooting
- “model not found” on pull
- Ollama is too old to know about this repository, or the tag is misspelled—it is “qwen3.8:27b,” with a dot, not a hyphen. Update Ollama and then run it again.
- Architecture error during loading
- Same cause: the hybrid layers of Qwen 3.8 are supported only by recent engine versions.
- Very slow generation (less than 5 tokens/s)
- The model spills into RAM. Check with ollama ps: if the split is not 100% GPU, reduce num_ctx, close other GPU applications, or switch to a smaller model.
- Responses that ignore the beginning of the document
- Context is silently truncated. Set num_ctx to match the actual size of your input, or split the document.
- The model “thinks” forever
- reasoning_effort is too high for the task. Lower it to medium or low, or disable reasoning for simple tasks.
- The image is ignored
- The file path must be accessible from the Ollama process, and the vision parser requires an up-to-date version. Test with a simple local image before blaming the model.
- Not enough disk space
- Between blobs and the cache, plan for twice the advertised size during installation. A pull interrupted by a full disk leaves fragments: ollama rm then pull again.
How much VRAM does Qwen 3.8 27B need?+
How do I install Qwen 3.8 27B locally?+
Is Qwen 3.8 27B free and suitable for business use?+
What is the difference between Qwen 3.8 27B and Qwen 3.8-Max?+
Can you run Qwen 3.8 27B with 16 GB of VRAM?+
Is Qwen 3.8 27B better than Qwen 3.6 27B?+
Can Qwen 3.8's reasoning mode be disabled?+
Is a one-million-token context available locally?+
#Go further
This guide assumes a working Ollama installation and a deliberate quantization choice. These pages provide additional coverage:
- Install Ollama
- The prerequisite if the engine is not yet set up on Windows, macOS, or Linux—and the update, required for Qwen 3.8.
- Choose your quantization
- To make informed choices between Q4, Q8, and BF16: what each step costs in memory and delivers in quality.
- Qwen 3.8: the open-weights timeline
- Where the confusion between the Max in the API and the open 27B comes from, and what was actually announced when.
- Which LLM for 24 GB of VRAM
- The comparison of models that fit on a 24 GB card, if you're still deciding between Qwen 3.8 27B and a lighter alternative.
Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.