Intermediate 14 minOllama

Qwen 3.8 27B locally: Ollama installation and VRAM

The Qwen3.8-27B weights were released on August 14, 2026, on Hugging Face under the Apache 2.0 license, and the official Ollama tag followed the same day. This is the first model in the 3.8 generation that can genuinely be installed locally: dense, multimodal (image and video), with 262,144 tokens of native context. This guide provides the exact command, the actual VRAM required for each tag, the settings for “thinking” mode enabled by default, and the two pitfalls that ruin most first attempts—silent context truncation and spilling into RAM.

By Mohamed Meguedmi·Update 2026-08-27·Tested on Windows, macOS, and Linux

#The verdict in 30 seconds

Qwen 3.8 27B installs with one command: “ollama run qwen3.8:27b”. The default package (Q4_K_M) takes up 18 GB and requires a 24 GB graphics card—RTX 3090, 4090, 5090—or a Mac Apple Silicon with at least 32 GB of unified memory for comfortable use. Below that, the model still runs, but some layers spill over to the CPU and throughput collapses.

What it is
A dense model with 27 billion parameters, native vision-language support (images and videos), 262,144-token context, under the Apache 2.0 license—so it can be used commercially without restrictive clauses.
The realistic prerequisite
24 GB of VRAM or 32 GB of unified memory for the Q4_K_M version. 30 GB of files and ~40 GB of VRAM for Q8, 56 GB for BF16.
The command
ollama pull qwen3.8:27b puis ollama run qwen3.8:27b. Sur Mac, préférez le tag -mlx, optimisé Metal.
The #1 trap
The advertised context is 256k, but Ollama applies a much smaller default window and truncates without warning. You need to set num_ctx manually.
Pitfall No. 2
“Thinking” mode is enabled by default and consumes many tokens before responding. Locally, lower reasoning_effort or disable it for simple tasks.
!
The performance figures come from Alibaba
As of the publication date of this guide, no independent benchmark has been reproduced for Qwen3.8-27B. All scores cited below come from the official model card. They are consistent with the previous generation, but should be treated as vendor figures.

#What Qwen 3.8 27B really is

The Local AI Kit

Your private ChatGPT, free, on your own machine in an hour — LM Studio, Ollama, Open WebUI, your documents, no cloud.

  • Lifetime online access
  • PDF + files
  • Lifetime updates

Unlike the Mixture-of-Experts trend, Qwen3.8-27B is a dense model: all 27 billion parameters are activated for every token. That makes its memory usage predictable—no pleasant surprises in throughput, no unpleasant surprises in VRAM—and it also explains its lower local throughput than a comparably sized MoE.

Architecture
64 layers organized into 16 blocks of “3 × (Gated DeltaNet → FFN) followed by 1 × (Gated Attention → FFN).” A hybrid attention design: most layers use memory-efficient linear attention, interspersed with true attention layers for precision.
Terms
Native vision-language: still images, documents, scientific diagrams, and videos. This is not an adapter bolted on afterward.
Context
262,144 tokens natively, extendable to 1,000,000 via YaRN—but only on vLLM, SGLang, or TokenSpeed, not through Ollama.
License
Apache 2.0. Commercial use permitted, no user threshold as with Llama, and no use clause as with Gemma.
Special feature
The model is trained with Multi-Token Prediction (MTP)—hence the Ollama “mtp” tags, which speed up generation on engines that know how to use them.
i
Do not confuse it with Qwen 3.8-Max
Qwen 3.8-Max, released on August 3, 2026, is an MoE with approximately 2,400 billion parameters, available through the API only: there is no way to run it on your own hardware. The 27B is the open-weight variant from the same generation. The complete timeline is covered in our guide to open weights, Qwen 3.8.

#The hardware you need, tag by tag

The Ollama library publishes twelve variants. Download size is not the VRAM required: you also need to add the KV cache, which grows with context length, and the vision encoder. Allow 15 to 25% of headroom above the file size.

Official variants of Qwen 3.8 27B on Ollama (data collected August 18, 2026)
TagSizeComfortable memoryWho it's for
qwen3.8:27b (Q4_K_M, default)18 GB24 GBThe default choice: RTX 3090/4090/5090, 32 GB Mac
qwen3.8:27b-q8_030 GB40 GB and upNear-maximum quality: RTX Pro 6000, dual-GPU 24 GB
qwen3.8:27b-bf1656 GB80 GBReference weights, compute machines (H100, H200)
qwen3.8:27b-mlx18 GB32 GB unified memoryApple Silicon: Metal build, the right default on Mac
qwen3.8:27b-mxfp832 GB48 GB unifiedMac Studio / M4 Max 64 GB, higher quality
qwen3.8:27b-mlx-bf1656 GB96 GB unifiedMac Studio 128 GB, no quantization loss
qwen3.8:27b-mtp-q4_K_M18 GB24 GBSame as the default, with explicit multi-token prediction
!
16 GB of VRAM: possible, but not comfortable
On a RTX 4060 Ti 16 GB or a 5070 Ti, Q4_K_M does not fit entirely in memory: Ollama places the remainder on the CPU, and generation often drops below 5 tokens per second. For 16 GB, a model with 12 to 14 billion parameters remains a much better use of the machine.

For throughput, our estimates for a dense 27B in Q4 are around 14 tokens/second on a mid-range configuration and 22 tokens/second on recent high-end hardware. These are calculated estimates, not measurements: the “thinking” mode enabled by default can also double the perceived time before the first response line.

#Install Qwen 3.8 27B with Ollama

  1. 01
    Update Ollama
    The hybrid layers and vision parser of Qwen 3.8 require a recent version of Ollama. A version older than August 2026 will return an architecture error or a misleading “model not found.” Reinstall from the official site, or rerun the installation script on Linux.
  2. 02
    Download the model
    18 GB are transferred: make sure there is room on the system drive, where Ollama stores its blobs. The pull can resume if the connection drops.
  3. 03
    Start a first exchange
    The first response is slower while the weights are loaded into memory. If startup takes more than a minute, that indicates an overflow onto the processor.
  4. 04
    Check where the model is running
    The ollama ps command displays the GPU/CPU split. Until you see 100% GPU, every token is expensive — drop down one quantization level or reduce the context.
Installation and first test
# 1. Récupérer le modèle (18 Go)
ollama pull qwen3.8:27b

# 2. Discuter avec
ollama run qwen3.8:27b

# 3. Vérifier la répartition GPU / CPU
ollama ps
i
qwen3.8 by itself = the 27B
The “qwen3.8:latest” tag currently points to the same blob as “qwen3.8:27b”: it is the only format published in the official library. Writing “ollama run qwen3.8” does exactly the same thing—but hard-coding the size in your scripts avoids unpleasant surprises when other variants arrive.

#On Mac Apple Silicon: choose the MLX variant

Ollama publishes an MLX build compiled for Apple's Metal engine. At the same file size (18 GB), it makes better use of unified memory and delivers noticeably higher throughput than generic GGUF on M3, M4, and later chips.

On Mac Apple Silicon
# Build MLX (Metal) — recommandé sur M1/M2/M3/M4
ollama run qwen3.8:27b-mlx

# Mac 64 Go et plus : qualité supérieure
ollama run qwen3.8:27b-mxfp8
16 GB Mac
Insufficient. macOS only makes about 70% of unified memory available to the GPU: the model swaps and becomes unusable.
Mac 24 GB
It just fits with a short context and no other demanding application open. Acceptable for testing, frustrating in everyday use.
32 GB Mac
The real entry point: the model fits, leaving room for the KV cache and the system.
Mac 64 GB and up
Comfortable, and lets you run in mxfp8 or work with long documents.

#The 256k context trap

This is the mistake almost all first-time users make. The model advertises 262,144 tokens, but Ollama applies a much shorter context window by default: beyond that limit, the beginning of your document is simply cut off, with no error message. The model confidently answers about text it has not seen in full.

Set the context window
# Dans une session interactive
/set parameter num_ctx 32768

# Ou via un Modelfile réutilisable
cat > Modelfile <<'EOF'
FROM qwen3.8:27b
PARAMETER num_ctx 32768
PARAMETER temperature 1.0
PARAMETER top_p 0.95
PARAMETER top_k 20
EOF
ollama create qwen38-long -f Modelfile
!
256k on paper, 16k to 64k in your setup
The KV cache grows roughly linearly with the context and is added to the 18 GB of weights. On a 24 GB card, a 32,000-token window is a sensible setting; 64,000 works at the cost of some slowness; 262,144 is out of reach. Enabling fast attention and quantizing the KV cache to q8_0 (OLLAMA_FLASH_ATTENTION and OLLAMA_KV_CACHE_TYPE variables) recovers several gigabytes.

As for the advertised million tokens: it relies on a YaRN extension, configured in the model's configuration file, and is supported only by vLLM, SGLang, or TokenSpeed. There is currently no way to access it from Ollama.

#Thinking mode and sampling parameters

Qwen 3.8 thinks before answering: it produces a reasoning block between « think » tags, followed by the final answer. This is enabled by default, and it explains most of the perceived latency. On engines that support it, the depth is controlled with reasoning_effort — xhigh by default, followed by medium and low.

Sampling parameters recommended by Alibaba
Modetemperaturetop_ptop_kpresence_penalty
Thinking (default)1.00.95200.0
Instruct (no reasoning)0.70.80201.5
Short local tasks
Set reasoning_effort to low or medium: for rewriting or extraction, extended reasoning does not change quality and triples the wait time.
Agentic tasks
Keep xhigh. Alibaba notes that reducing the effort leads to insufficient analysis, and therefore retries—the gain per turn is canceled out by failures.
Responses that get stuck in a loop
Increase presence_penalty (between 0 and 2). Above 1.5, the model starts mixing languages.
Output polluted by reasoning
If your application displays reasoning tags, it is not separating reasoning_content from the final content. Disable reasoning for this use case instead of filtering the text afterward.

#Analyze an image or document

Vision is native: screenshots, whiteboard photos, technical diagrams, scanned pages. On the command line, just provide the file path in the prompt; through the API, images are passed as base64 in the field provided by Ollama.

Command-line vision
ollama run qwen3.8:27b
>>> Décris ce schéma et liste les valeurs lisibles ./schema.png
!
Test vision before putting it into production
The multimodal path received a parser fix shortly after release. If your responses ignore the image or describe something else, update Ollama before investigating further—and validate it on about ten of your own documents before putting it into production.

#What the reported scores are worth

The claimed leap over Qwen 3.6–27B, the same-size model from the previous generation, is substantial — especially for agentic coding and computer use. Here are the figures published by Alibaba, keeping in mind that they have not been independently reproduced.

Qwen3.8-27B vs Qwen3.6-27B — scores reported by Alibaba (August 2026)
TestQwen3.8-27BQwen3.6-27B
Terminal Bench 2.1 (agentic coding)73,063,4
SWE-bench Pro61,753,5
LiveCodeBench v690,383,9
IFBench (instruction following)79,569,1
GPQA Diamond (science)89,287,8
OSWorld-Verified (computer use)84,363,9
MathVision (visual math)90,085,1
OmniDocBench 1.5 (documents)91,189,4

What to take away for local use: the gains are concentrated in long, tool-driven tasks — operating a terminal, fixing a repository, and chaining steps without going off track. In a simple conversation or summary, the difference from the previous generation will be far less noticeable than these figures suggest.

#Should you move on from Qwen 3.6, Gemma 4, or gpt-oss?

You are on Qwen 3.6-27B
Yes, the update is worth it: same memory footprint, same permissive license, clear gains in coding and agentic tasks. Keep the old tag while you validate your prompts; the outputs change in style.
You're on Gemma 4 26B
Qwen 3.8 retains the edge in coding and agentic capabilities; in terms of licensing, both are now Apache 2.0, with Gemma having become permissive in April 2026. Gemma 4 (MoE 26B-A4B, multimodal) often remains more natural in conversational French and is faster to start.
You're running gpt-oss-20b
Two philosophies: gpt-oss is lighter and faster, Qwen 3.8 sees images, handles a much longer context, and targets long-running tasks. Choose based on the available VRAM.
You're mainly looking to code
A specialized 30B code model in MoE remains faster for autocompletion. Qwen 3.8 27B makes sense when the agent needs to read, plan, and modify multiple files.
You have less than 24 GB
Don't force it. A 12-14B model in Q4 will give you a better experience than a 27B model that spills over onto the CPU.

#Troubleshooting

“model not found” on pull
Ollama is too old to know about this repository, or the tag is misspelled—it is “qwen3.8:27b,” with a dot, not a hyphen. Update Ollama and then run it again.
Architecture error during loading
Same cause: the hybrid layers of Qwen 3.8 are supported only by recent engine versions.
Very slow generation (less than 5 tokens/s)
The model spills into RAM. Check with ollama ps: if the split is not 100% GPU, reduce num_ctx, close other GPU applications, or switch to a smaller model.
Responses that ignore the beginning of the document
Context is silently truncated. Set num_ctx to match the actual size of your input, or split the document.
The model “thinks” forever
reasoning_effort is too high for the task. Lower it to medium or low, or disable reasoning for simple tasks.
The image is ignored
The file path must be accessible from the Ollama process, and the vision parser requires an up-to-date version. Test with a simple local image before blaming the model.
Not enough disk space
Between blobs and the cache, plan for twice the advertised size during installation. A pull interrupted by a full disk leaves fragments: ollama rm then pull again.
Frequently asked questions
How much VRAM does Qwen 3.8 27B need?+
24 GB of VRAM for the default Q4_K_M version, which takes 18 GB of files plus the KV cache. Allow 40 GB for the Q8 version and 56 GB for BF16 weights. On Mac, you need at least 32 GB of unified memory.
How do I install Qwen 3.8 27B locally?+
Install or update Ollama, then run “ollama pull qwen3.8:27b” and “ollama run qwen3.8:27b”. The download is 18 GB. On an Apple Apple Silicon Mac, use the qwen3.8:27b-mlx tag instead, optimized for Metal.
Is Qwen 3.8 27B free and suitable for business use?+
Yes. The weights are published under the Apache 2.0 license, which allows commercial use, modification, and redistribution, with no user threshold or restrictive-use clause. It is one of the most permissive licenses among open models.
What is the difference between Qwen 3.8 27B and Qwen 3.8-Max?+
Qwen 3.8-Max is a proprietary Mixture-of-Experts model with approximately 2,400 billion parameters, accessible only through an API. Qwen3.8-27B is the dense open-weight variant from the same generation: it is the only one of the two that you can run on your machine.
Can you run Qwen 3.8 27B with 16 GB of VRAM?+
Technically, yes, but part of the model spills onto the CPU and throughput generally falls below 5 tokens per second. With 16 GB, a model with 12 to 14 billion parameters provides a much better experience.
Is Qwen 3.8 27B better than Qwen 3.6 27B?+
Based on the figures published by Alibaba, yes—clearly for agentic coding (73.0 vs. 63.4 on Terminal Bench 2.1) and computer use (84.3 vs. 63.9 on OSWorld-Verified). These scores have not yet been independently reproduced, and the gap is less pronounced for conversational use cases.
Can Qwen 3.8's reasoning mode be disabled?+
Yes. Thinking mode is enabled by default but can be disabled per request, and its depth can be set with the reasoning_effort parameter (xhigh, medium, or low). In direct mode, the recommended settings are temperature 0.7 and top_p 0.80.
Is a one-million-token context available locally?+
Not via Ollama. The native context is 262 144 tokens, and extending it to one million requires a YaRN configuration supported only by vLLM, SGLang, and TokenSpeed. In practice, on a 24 GB card, a window of 16 000 to 64 000 tokens is the realistic setting.

#Go further

This guide assumes a working Ollama installation and a deliberate quantization choice. These pages provide additional coverage:

Install Ollama
The prerequisite if the engine is not yet set up on Windows, macOS, or Linux—and the update, required for Qwen 3.8.
Choose your quantization
To make informed choices between Q4, Q8, and BF16: what each step costs in memory and delivers in quality.
Qwen 3.8: the open-weights timeline
Where the confusion between the Max in the API and the open 27B comes from, and what was actually announced when.
Which LLM for 24 GB of VRAM
The comparison of models that fit on a 24 GB card, if you're still deciding between Qwen 3.8 27B and a lighter alternative.
Did this guide help you?

Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.