Intermediate 12 minOllama

Llama 4 Scout locally: installation and initial tests with Ollama

Llama 4 Scout is the smallest of the three models in Meta’s Llama 4 family. It is also the only one that fits locally on a high-end workstation—Maverick and Behemoth remain limited to the cloud or a dedicated server. This guide shows how to install Llama 4 Scout with Ollama, what VRAM it really requires (the MoE promise is often misunderstood), and how to use its two differentiating strengths: native multimodality and a 10M-token context.

By Mohamed Meguedmi·Update 2026-08-27·Tested on Windows, macOS, and Linux

#Why install Llama 4 Scout locally with Ollama?

Scout is the only model in the Llama 4 lineup designed to run on a single machine. Meta positioned it as the family’s “workstation model”: natively multimodal (text + images), with an advertised context window of 10 million tokens, and a Mixture of Experts (MoE) architecture that activates only a fraction of the parameters for each token.

In practical terms, compared with a dense model of equivalent size, Scout offers better latency (few active parameters per token), substantially more context memory, and built-in vision without a separate model. The price: a considerable disk footprint (all experts must remain loaded in VRAM for routing to work).

i
MoE in two words
An MoE model like Scout contains N “experts” (specialized subnetworks). For each token, a router chooses k of them (typically 1 or 2). Only those k experts are computed. The result: inference speed is that of a small model, but quality tends toward that of the large one. Total parameters remain in VRAM — only the computation is partial.

#Scout, Maverick, Behemoth: who does what

The Local AI Kit

Your private ChatGPT, free, on your own machine in an hour — LM Studio, Ollama, Open WebUI, your documents, no cloud.

  • Lifetime online access
  • PDF + files
  • Lifetime updates
Scout (17B active)
The workstation model. 16 experts, ~109B total parameters. Multimodal. 10M-token context. Target: RTX 4090, M3/M4 Max/Ultra, or a 48 GB+ multi-GPU setup.
Maverick (17B active)
The server model. 128 experts, ~400B total parameters. The same number of active experts as Scout but far more encoded knowledge. Target: 8x H100 server or cluster. Out of reach locally for most configurations.
Behemoth (288B active)
The frontier model. ~2T total parameters. Not runnable locally without a data center. Mainly useful as a “teacher” model for distilling the other two.
→
Which one should you choose?
If you're reading this guide and installing locally, it's Scout. Maverick requires at least a high-end multi-GPU server — at that level, the DeepSeek V4 Flash or Qwen3.7 Max guides are more relevant.

#Actual VRAM: what the spec sheet doesn’t tell you

The classic mistake with MoE models is to reason about active parameters. “17B active = it fits on 12 GB in Q4”: false. For routing to work, all experts must remain in memory. Scout in Q4_K_M requires much more than its 17B active parameters would suggest.

Q4_K_M (recommended)
≈ 65 GB of VRAM for the model alone. Add ~4 GB for a reasonable context (32k tokens). Total ≈ 70 GB.
Q5_K_M
≈ 78 GB. Marginal quality gain on most tasks compared with Q4_K_M.
Q8_0
≈ 115 GB. Reserved for reference benchmarks.
FP16
≈ 220 GB. Cloud or multi-server only.
!
Not for a 24 GB GPU
If you have a RTX 4090 or a single 3090, Scout won't run properly. You can force CPU offloading, but speed will collapse (< 2 tokens/sec). For a 24 GB GPU, consider Qwen3-30B-A3B or Mistral Small 24B instead. For a 32 GB RTX 5090, Qwen 3.6 35B-A3B is still more suitable than offloaded Scout.

Configurations that run Scout comfortably locally:

Mac Studio M3 Ultra / M5 Ultra 96 GB and above
The best consumer platform for Scout. Unified memory, no offloading, ~25–40 tokens/sec depending on the quantization.
2x RTX 4090 / 2x 5090
48 GB or 64 GB combined, enough for Q4_K_M with extended context. Rely on llama.cpp or Ollama with the OLLAMA_NUM_GPU variable.
RTX 6000 Ada workstation (48 GB) or A6000
Comfortably runs Scout Q4 with room to spare for context.

#Prerequisites

Ollama 0.6 or later
Support for Llama 4 was added starting with Ollama 0.6. Check with ollama --version and update if needed.
70 GB of free disk space
The Q4_K_M tag weighs about 65 GB. Allow plenty of room for cache overhead.
High memory bandwidth
On Mac: target ≥ 400 GB/s (M3 Max and above). On PC: DDR5 if CPU offload is planned.
A stable connection
The initial download can take 1 to 3 hours depending on your connection. ollama pull supporte resumption after an interruption.

#1. Install Llama 4 Scout with Ollama

If Ollama is not installed yet, first follow the guide for your system. Then Scout can be downloaded with a single command.

Terminal — model pull
ollama pull llama4:scout

This tag defaults to Q4_K_M quantization. If you want to force another variant:

Available variants
ollama pull llama4:scout-q5_K_M
ollama pull llama4:scout-q8_0
ollama pull llama4:scout-fp16

Once the download is complete, list the models to confirm the actual size:

Verification
ollama list
i
Be patient during the first load
Loading 65 GB into VRAM takes time: between 30 seconds and 2 minutes depending on the SSD. Subsequent loads are faster thanks to the OS cache.

#2. First test: text and reasoning

Start an interactive session to validate that everything works correctly before touching the rest.

Interactive session
ollama run llama4:scout

When the >>> prompt appears, check the language and reasoning quality with a simple test:

Test prompt
>>> Explique en 4 phrases ce qu'est un modèle MoE, puis donne un avantage et un inconvénient.

[Réponse attendue : explication structurée en français, distinction entre paramètres totaux et actifs, mention de la latence comme avantage et de la VRAM comme inconvénient.]

While the conversation is running, open a second terminal to monitor the load:

Monitor the load
ollama ps

The SIZE column should reflect the actual loaded weight. Under PROCESSOR, ideally 100% GPU. If you see a CPU/GPU mix, VRAM is insufficient and Ollama is offloading—速度 will suffer severely.

#3. Vision test in French

Scout is natively multimodal: image and text go through the same backbone, unlike an approach that attaches a separate visual encoder to a purely textual model. This produces better results on tasks that combine both modalities (contextual OCR, diagram reading, fine-grained description).

With Ollama's REST API, send the image as base64 in the images field:

Python vision test
import base64, requests, json

with open('facture.jpg', 'rb') as f:
    img_b64 = base64.b64encode(f.read()).decode()

response = requests.post(
    'http://localhost:11434/api/generate',
    json={
        'model': 'llama4:scout',
        'prompt': 'Décris cette image en français. Si c\'est un document, extrais les montants et la date.',
        'images': [img_b64],
        'stream': False,
    }
)
print(response.json()['response'])

For typical cases (interface screenshot, whiteboard photo, PDF scan), Scout produces a detailed description and remains consistent in French even when the image contains English text. It is a step above a small general-purpose vision model like Qwen 3.5 9B for structured details.

→
Image format
JPEG and PNG work. Avoid oversized images (> 4 MB): Ollama resizes them, but that is expensive. Resize them in advance to a maximum of 1024px on the longest side before sending.

#4. The 10M-token context: promise and reality

Meta announces a 10-million-token context window for Scout. In practice, two local limitations apply.

The KV cache explodes
At 1M tokens, the KV cache easily adds 30–50 GB to the VRAM used by the model. Beyond that, you need to enable KV-cache quantization (the num_ctx option plus experimental parameters) or accept CPU offloading for the cache.
Quality declines well before 10M
Independent benchmarks (RULER, NoLiMa) show a meaningful performance drop around 256k–512k tokens, despite long-context training. Beyond that, it is technically possible but unreliable.

To use an extended context without bringing everything down, configure Ollama through a Modelfile:

Modelfile for 128k context
FROM llama4:scout

PARAMETER num_ctx 131072
PARAMETER num_predict 4096
PARAMETER temperature 0.6
Variant creation
ollama create llama4-scout-128k -f Modelfile
ollama run llama4-scout-128k
!
128k is already a lot
A 128k-token context already represents about 250 pages of text. For 95% of local use cases (entire codebases, long PDFs, ticket archives), that is more than enough and much more stable than 1M+ configurations.

#Scout vs Qwen3-30B-A3B: the real comparison

Qwen3-30B-A3B is the other MoE that deserves serious consideration for local use in 2026: 30B total, 3B active, with a native 256k context. The comparison is worthwhile because the two aren't in the same hardware category.

Q4 VRAM footprint
Scout ≈ 65 GB. Qwen3-30B-A3B ≈ 18 GB. The difference is massive and completely changes the target machine profile.
Generation speed
With comparable active parameters (17B vs. 3B), Qwen3-30B-A3B is faster in tokens/sec—typically 2-3x faster on the same machine, when both fit.
Reasoning quality
On the public MMLU-Pro and GPQA benchmarks, Scout remains ahead of Qwen3-30B-A3B. The gap narrows considerably when Qwen3's thinking mode is enabled.
Multimodality
Scout is natively multimodal. Qwen3-30B-A3B is not—you need to run a separate vision model such as Qwen 3.5 9B in parallel if you want vision in the same stack.
Long context
Scout targets 10M (256k in practice, stably). Qwen3-30B-A3B offers native 256k, which is more predictable and less expensive in KV cache.
→
Practical verdict
If you have a Mac Studio Ultra, a 2x RTX 4090 setup, or an A6000, choose Scout for native vision and quality. If you're running a single RTX 4090, a 3090, or a 36–64 GB Mac M3 Max, choose Qwen3-30B-A3B—it makes sense and is significantly faster. Running Scout with CPU offload on this type of machine is a false good idea.

#Troubleshooting

"Error: model requires more system memory than is available"
Your combined VRAM + RAM is insufficient. Either switch to more aggressive quantization (rarely feasible in Q4, already tight), or change models. There is no magic on the Ollama side.
Speed < 5 tokens/sec on 24 GB GPU
You're using CPU offload. Check with ollama ps: the PROCESSOR column must show 100% GPU. If it doesn't, Scout isn't suitable for your machine.
Truncated responses
Increase num_predict (128 by default in some configurations). Use /set parameter num_predict 2048 in the session, or configure it through a Modelfile.
Crashes beyond 64k tokens
The KV cache saturates VRAM. Reduce num_ctx, or enable KV quantization with OLLAMA_FLASH_ATTENTION=1 + OLLAMA_KV_CACHE_TYPE=q8_0 in the environment.
Image rejected with "unsupported image format"
Convert to standard JPEG or PNG. WebP, HEIC, and AVIF are not all supported depending on the Ollama version.

#Go further

Once Scout is operational, several directions are worth exploring to take advantage of its specific capabilities.

Choosing the right quantization
Q4_K_M is the sweet spot for Scout, but knowing when to move up to Q5 or Q8 depends on the use case — especially for multimodal tasks, where quantization can degrade quality more than with text-only workloads.
Taking advantage of local vision
The dedicated vision guide covers prompt patterns that really work (contextual OCR, structured extraction, image comparison) with Scout and multimodal competitors such as Qwen 3.5 9B and Gemma 4 12B.
Modelfile for customization
Beyond num_ctx, a Modelfile lets you lock in a system prompt, a JSON output format, and a temperature suited to your use case—useful for avoiding reconfiguration every session.
Did this guide help you?

Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.