01What it can do
- European frontier model with open weights
- Versatile: chat, coding, reasoning, and vision
- MoE: only 49B active out of 1T parameters
- Official Ollama tag available
- —~580 GB VRAM in Q4: reserved for multi-GPU or datacenter servers
- —32k native context (default value, to be confirmed on the model card)
- —License not specified in the sources: verify before commercial use
04Install
Install Ollama for your OS. Check the model and its quantization before downloading. Start with 4096 tokens of context, then check placement with ollama ps. A command below is not proof that a test was run on your machine.
02Required memory
Approximate GPU VRAM required to run this model, including 4k tokens of context overhead. For a longer context, add ~1 GB per 8k-token increment.
What hardware do you need for Mistral Large 4?
To run Mistral Large 4 locally with Q4 quantization, you need about 580 GB of VRAM. An option to compare: BOSGAME M5 128GB / 2TB (Ryzen AI Max+ 395) — this model exceeds this mini-PC's GPU capacity: choose a smaller model or suitable infrastructure.
Why this choice? Our complete guide on BOSGAME M5 128GB / 2TB (Ryzen AI Max+ 395) →
Affiliate links — commission possible at no extra cost to you; independent recommendations. As an Amazon Associate, BestLLMfor earns from qualifying purchases.
03Expected speed
Tokens generated per second in Q4_K_M, 4k context. Beyond 20 t/s, reading is comfortable. Below 10 t/s, that's just for testing.