Llama 4 Scout locally: installation and initial tests with Ollama
Llama 4 Scout is the smallest of the three models in Meta’s Llama 4 family. It is also the only one that fits locally on a high-end workstation—Maverick and Behemoth remain limited to the cloud or a dedicated server. This guide shows how to install Llama 4 Scout with Ollama, what VRAM it really requires (the MoE promise is often misunderstood), and how to use its two differentiating strengths: native multimodality and a 10M-token context.
#Why install Llama 4 Scout locally with Ollama?
Scout is the only model in the Llama 4 lineup designed to run on a single machine. Meta positioned it as the family’s “workstation model”: natively multimodal (text + images), with an advertised context window of 10 million tokens, and a Mixture of Experts (MoE) architecture that activates only a fraction of the parameters for each token.
In practical terms, compared with a dense model of equivalent size, Scout offers better latency (few active parameters per token), substantially more context memory, and built-in vision without a separate model. The price: a considerable disk footprint (all experts must remain loaded in VRAM for routing to work).
#Scout, Maverick, Behemoth: who does what
Your private ChatGPT, free, on your own machine in an hour — LM Studio, Ollama, Open WebUI, your documents, no cloud.
- Lifetime online access
- PDF + files
- Lifetime updates
- Scout (17B active)
- The workstation model. 16 experts, ~109B total parameters. Multimodal. 10M-token context. Target: RTX 4090, M3/M4 Max/Ultra, or a 48 GB+ multi-GPU setup.
- Maverick (17B active)
- The server model. 128 experts, ~400B total parameters. The same number of active experts as Scout but far more encoded knowledge. Target: 8x H100 server or cluster. Out of reach locally for most configurations.
- Behemoth (288B active)
- The frontier model. ~2T total parameters. Not runnable locally without a data center. Mainly useful as a “teacher” model for distilling the other two.
#Actual VRAM: what the spec sheet doesn’t tell you
The classic mistake with MoE models is to reason about active parameters. “17B active = it fits on 12 GB in Q4”: false. For routing to work, all experts must remain in memory. Scout in Q4_K_M requires much more than its 17B active parameters would suggest.
- Q4_K_M (recommended)
- ≈ 65 GB of VRAM for the model alone. Add ~4 GB for a reasonable context (32k tokens). Total ≈ 70 GB.
- Q5_K_M
- ≈ 78 GB. Marginal quality gain on most tasks compared with Q4_K_M.
- Q8_0
- ≈ 115 GB. Reserved for reference benchmarks.
- FP16
- ≈ 220 GB. Cloud or multi-server only.
Configurations that run Scout comfortably locally:
- Mac Studio M3 Ultra / M5 Ultra 96 GB and above
- The best consumer platform for Scout. Unified memory, no offloading, ~25–40 tokens/sec depending on the quantization.
- 2x RTX 4090 / 2x 5090
- 48 GB or 64 GB combined, enough for Q4_K_M with extended context. Rely on llama.cpp or Ollama with the OLLAMA_NUM_GPU variable.
- RTX 6000 Ada workstation (48 GB) or A6000
- Comfortably runs Scout Q4 with room to spare for context.
#Prerequisites
- Ollama 0.6 or later
- Support for Llama 4 was added starting with Ollama 0.6. Check with ollama --version and update if needed.
- 70 GB of free disk space
- The Q4_K_M tag weighs about 65 GB. Allow plenty of room for cache overhead.
- High memory bandwidth
- On Mac: target ≥ 400 GB/s (M3 Max and above). On PC: DDR5 if CPU offload is planned.
- A stable connection
- The initial download can take 1 to 3 hours depending on your connection. ollama pull supporte resumption after an interruption.
#1. Install Llama 4 Scout with Ollama
If Ollama is not installed yet, first follow the guide for your system. Then Scout can be downloaded with a single command.
This tag defaults to Q4_K_M quantization. If you want to force another variant:
Once the download is complete, list the models to confirm the actual size:
#2. First test: text and reasoning
Start an interactive session to validate that everything works correctly before touching the rest.
When the >>> prompt appears, check the language and reasoning quality with a simple test:
While the conversation is running, open a second terminal to monitor the load:
The SIZE column should reflect the actual loaded weight. Under PROCESSOR, ideally 100% GPU. If you see a CPU/GPU mix, VRAM is insufficient and Ollama is offloading—速度 will suffer severely.
#3. Vision test in French
Scout is natively multimodal: image and text go through the same backbone, unlike an approach that attaches a separate visual encoder to a purely textual model. This produces better results on tasks that combine both modalities (contextual OCR, diagram reading, fine-grained description).
With Ollama's REST API, send the image as base64 in the images field:
For typical cases (interface screenshot, whiteboard photo, PDF scan), Scout produces a detailed description and remains consistent in French even when the image contains English text. It is a step above a small general-purpose vision model like Qwen 3.5 9B for structured details.
#4. The 10M-token context: promise and reality
Meta announces a 10-million-token context window for Scout. In practice, two local limitations apply.
- The KV cache explodes
- At 1M tokens, the KV cache easily adds 30–50 GB to the VRAM used by the model. Beyond that, you need to enable KV-cache quantization (the num_ctx option plus experimental parameters) or accept CPU offloading for the cache.
- Quality declines well before 10M
- Independent benchmarks (RULER, NoLiMa) show a meaningful performance drop around 256k–512k tokens, despite long-context training. Beyond that, it is technically possible but unreliable.
To use an extended context without bringing everything down, configure Ollama through a Modelfile:
#Scout vs Qwen3-30B-A3B: the real comparison
Qwen3-30B-A3B is the other MoE that deserves serious consideration for local use in 2026: 30B total, 3B active, with a native 256k context. The comparison is worthwhile because the two aren't in the same hardware category.
- Q4 VRAM footprint
- Scout ≈ 65 GB. Qwen3-30B-A3B ≈ 18 GB. The difference is massive and completely changes the target machine profile.
- Generation speed
- With comparable active parameters (17B vs. 3B), Qwen3-30B-A3B is faster in tokens/sec—typically 2-3x faster on the same machine, when both fit.
- Reasoning quality
- On the public MMLU-Pro and GPQA benchmarks, Scout remains ahead of Qwen3-30B-A3B. The gap narrows considerably when Qwen3's thinking mode is enabled.
- Multimodality
- Scout is natively multimodal. Qwen3-30B-A3B is not—you need to run a separate vision model such as Qwen 3.5 9B in parallel if you want vision in the same stack.
- Long context
- Scout targets 10M (256k in practice, stably). Qwen3-30B-A3B offers native 256k, which is more predictable and less expensive in KV cache.
#Troubleshooting
- "Error: model requires more system memory than is available"
- Your combined VRAM + RAM is insufficient. Either switch to more aggressive quantization (rarely feasible in Q4, already tight), or change models. There is no magic on the Ollama side.
- Speed < 5 tokens/sec on 24 GB GPU
- You're using CPU offload. Check with ollama ps: the PROCESSOR column must show 100% GPU. If it doesn't, Scout isn't suitable for your machine.
- Truncated responses
- Increase num_predict (128 by default in some configurations). Use /set parameter num_predict 2048 in the session, or configure it through a Modelfile.
- Crashes beyond 64k tokens
- The KV cache saturates VRAM. Reduce num_ctx, or enable KV quantization with OLLAMA_FLASH_ATTENTION=1 + OLLAMA_KV_CACHE_TYPE=q8_0 in the environment.
- Image rejected with "unsupported image format"
- Convert to standard JPEG or PNG. WebP, HEIC, and AVIF are not all supported depending on the Ollama version.
#Go further
Once Scout is operational, several directions are worth exploring to take advantage of its specific capabilities.
- Choosing the right quantization
- Q4_K_M is the sweet spot for Scout, but knowing when to move up to Q5 or Q8 depends on the use case — especially for multimodal tasks, where quantization can degrade quality more than with text-only workloads.
- Taking advantage of local vision
- The dedicated vision guide covers prompt patterns that really work (contextual OCR, structured extraction, image comparison) with Scout and multimodal competitors such as Qwen 3.5 9B and Gemma 4 12B.
- Modelfile for customization
- Beyond num_ctx, a Modelfile lets you lock in a system prompt, a JSON output format, and a temperature suited to your use case—useful for avoiding reconfiguration every session.
Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.