Intermediate 10 minOllama

Mistral Magistral: install the model from reasoning

Magistral is Mistral AI’s reasoning model: a model that “thinks out loud” before answering, like DeepSeek R1 or OpenAI o1, but with native French quality that no other open-weight reasoning model matches. This guide shows how to install Mistral Magistral locally through Ollama, configure it with the correct prompt template, and compare it honestly with distilled DeepSeek R1.

By Mohamed Meguedmi·Update 2026-08-27·Tested on Windows, macOS, and Linux

#Why Magistral rather than another reasoning model?

Reasoning models generate an explicit chain of thought before the final answer. On complex problems—math, logic, difficult code, and legal analysis—they score 10 to 30 points higher on benchmarks than a general-purpose model of the same size. The tradeoff: they are slower and more verbose.

Before Magistral, open-weight options were mostly Chinese: DeepSeek R1 and its distillations, Qwen3 in thinking mode, and GLM. They all reason perfectly in English and Chinese, but their chain of thought regularly switches to Chinese even when the prompt is in French. Magistral is trained specifically to keep its reasoning trace in the prompt's language.

Native multilingual reasoning
The chain of thought stays in French when you write in French. No wild switch to Chinese or English.
Apache 2.0 for the Small version
Commercial use allowed, fine-tuning unrestricted, redistribution permitted. No gray areas like those found with some restrictive “open” licenses.
Base Mistral Small 3.1
Good inherited French writing quality, evident in the drafting of legal or administrative responses.
40k-token context
Enough for a long problem statement, several code articles, or a contract of a few dozen pages.
i
Reasoning = additional tokens
Expect 500 to 3000 "thinking" tokens before each final answer. On a modest GPU, be patient: a complex question can take 30 to 60 seconds of continuous generation before the first useful line.

#Magistral Small vs Magistral Medium

The Local AI Kit

Your private ChatGPT, free, on your own machine in an hour — LM Studio, Ollama, Open WebUI, your documents, no cloud.

  • Lifetime online access
  • PDF + files
  • Lifetime updates

Mistral AI publishes Magistral in two versions, and the confusion is common. To install mistral magistral locally, only the Small version is usable—the Medium version is available only through the API.

Magistral Small (24B)
Open weights, Apache 2.0. This is the model you install locally. 24 billion parameters, derived from Mistral Small 3.1.
Magistral Medium
Proprietary, accessible only through the Mistral API or Le Chat. Higher performance but not self-hostable. Outside the scope of this guide.
→
A small naming tip
On Ollama Hub, the model is simply called `magistral`. There is no need to specify “small”: the Medium version is not distributed locally anyway.

#Hardware requirements

Magistral Small has 24B parameters. With Q4_K_M quantization (the Ollama standard), allow about 14 GB of VRAM to load the model, plus 1 to 3 GB for context depending on the length.

Comfortable minimum
RTX 4080 16 GB, RTX 4090 24 GB, RTX 3090 24 GB, or a Mac with Apple Silicon and 24 GB of unified memory.
Marginal but playable
RTX 4070 Ti Super 16 GB (Q4 fits exactly; short context required) or a 16 GB M-series Mac with partial offload.
Insufficient
Any GPU with 12 GB or less. The model spills into CPU RAM and drops to 2–4 tokens/sec, unusable for a model that must produce thousands before responding.
Ollama and drivers
Ollama ≥ 0.5.x, recent NVIDIA drivers (CUDA 12), or Metal on Mac. Check with `ollama --version`.
!
No 24 GB GPU? Reconsider
Magistral is not a model to try without a powerful GPU. On 12 GB or less, you will wait several minutes per response. Instead, consider a 7B-14B distillation of DeepSeek R1 — worse in French, but usable on modest hardware.

#1. Installation via Ollama

Magistral Small has been published on Ollama Hub since the official release of Mistral. Installation takes one command.

  1. 01
    Verify that Ollama is running
    On Windows/macOS, the application must be running (its icon appears in the taskbar). On Linux, `systemctl status ollama` confirms that the service is active.
  2. 02
    Download the model
    The default quantization served by Ollama is Q4_K_M. It's the right compromise for 16–24 GB of VRAM.
  3. 03
    First launch
    Ollama downloads ~14 GB on the first `pull`. The initial download can take 10 to 30 minutes depending on your connection.
  4. 04
    Load test
    Launch `ollama run magistral` and then ask a short question. If the model responds after 5-10 seconds of "warmup," everything is fine.
Terminal
ollama pull magistral
ollama run magistral

To verify that the model is actually using the GPU rather than CPU RAM, use a diagnostic command:

Runtime status
ollama ps

In the PROCESSOR column, you should see `100% GPU`. If you see `45% CPU / 55% GPU`, your VRAM is insufficient: Ollama has partially fallen back to RAM. The model will work, but at 1–3 tokens/sec.

#2. The official prompt template

Magistral uses a specific prompt format to correctly activate reasoning. Ollama already includes this template in the official Modelfile, but it is worth understanding so you do not accidentally mess it up with a misplaced system prompt.

The expected structure, simplified:

Raw format (internal to the model)
<s>[SYSTEM_PROMPT]Vous êtes un assistant qui raisonne soigneusement avant de répondre.[/SYSTEM_PROMPT][INST]Question utilisateur[/INST]
<think>
Chaîne de raisonnement interne...
</think>
Réponse finale</s>

In practice, you never write these tags by hand: Ollama inserts them automatically. What you need to know: the model expects you to ask it to reason. A short, clear system prompt significantly improves quality.

Configure the system prompt
ollama run magistral
>>> /set system "Tu es un assistant rigoureux. Avant chaque réponse, analyse le problème étape par étape en français. Réponds de manière concise après ton raisonnement."
→
No need to ask it to “think carefully”
Unlike a general-purpose model, Magistral reasons by default. There’s no need to write "take your time" or "think step by step": that’s already its normal mode. However, explicitly asking it to reason in French in the system prompt helps lock in the language.

#3. First practical tests in French

Here are three prompts for quickly evaluating Magistral on typical tasks. Note that the responses arrive in two stages: first the chain of thought (`<think>...</think>`), then the final answer.

Test 1 — Logic problem
>>> Trois personnes se partagent 17 chevaux selon les proportions 1/2, 1/3, 1/9. Comment faire sans couper aucun cheval ?

Magistral will run through the classic borrowed-camel reasoning. Expect 1500 to 2500 thinking tokens before the clear, numbered final answer.

Test 2 — Python code
>>> Écris une fonction Python qui trouve le plus long sous-tableau dont la somme vaut zéro. Complexité O(n).

A good opportunity to assess algorithmic quality. Magistral generally proposes the solution using a prefix-sum dictionary, explaining its data-structure choice during thinking.

Test 3 — FR legal writing
>>> Rédige une clause de non-concurrence pour un CDI de développeur en France, conforme au droit du travail français actuel.

For this type of request, Magistral's French language quality makes the difference. The generated clause is syntactically correct, mentions the validity conditions (limitations in time and geographic scope, and financial consideration), and uses French legal terminology—not an awkward translation from English.

#4. Performance: AIME, math, code

Mistral AI has published official figures for Magistral Small and Medium. The values below are the Small scores (the only version we self-host), rounded. They are only reference points—your results will vary depending on quantization and sampling.

AIME 2024 (Olympiad mathematics)
Magistral Small ~70%, versus ~50% for Mistral Small 3.1 non-reasoning. The gain from reasoning mode is massive for this type of problem.
LiveCodeBench (code)
A score equivalent to or slightly higher than Qwen3-Coder 30B-A3B on medium-level problems, below it on hard levels.
MMLU FR
Comparable to Mistral Small 3.1 base — reasoning adds nothing on general-knowledge multiple-choice questions, where the answer is either memorized or it isn't.
GPQA Diamond (scientific reasoning)
Net gain from thinking mode, around 50–55% versus 35–40% for the baseline without reasoning.
i
Quantization and quality loss
All official benchmarks are measured in FP16. In Q4_K_M (the default Ollama), expect a loss of 1 to 3 points on AIME. If you have 24 GB of VRAM, switching to Q5_K_M or Q6_K recovers that margin—at the cost of slightly lower throughput.

#5. Magistral Small vs DeepSeek distilled 14B R1

For French-speaking users, the useful comparison is with DeepSeek-R1-Distill-Qwen-14B, the direct open-weights competitor for anyone who wants local reasoning. Both have their strengths—here's what practical testing shows.

Memory footprint
DeepSeek R1 14B Q4 ≈ 9 GB, Magistral 24B Q4 ≈ 14 GB. Clear advantage DeepSeek if your GPU tops out at 12 GB.
Reasoning quality in FR
Magistral keeps French from beginning to end. DeepSeek Distilled R1 regularly switches to Chinese or English in its chain of thought, even with a strict French system prompt. Advantage: Magistral.
Final answer quality FR
Magistral produces higher-quality French (vocabulary, agreement, register). DeepSeek R1 14B makes the typical usage errors of a translated model. Magistral's advantage.
Pure math (AIME, AMC)
DeepSeek R1 distilled 14B is slightly better than Magistral Small on highly formal problems. Advantage: DeepSeek.
Code
A practical tie. Both generate correct code for medium-difficulty problems. For French-language code (comments, variable names), Magistral is cleaner.
Generation speed
DeepSeek 14B is ~1.7× faster at equivalent VRAM, simply due to its size. For an interactive session, that matters.
→
The practical verdict
If you work primarily in French — law, writing, user support, code commented in FR — Magistral Small is better. If you work on formal mathematics or have a 12 GB GPU, stick with DeepSeek distilled R1 14B.

#6. Use case: local legal analysis

The combination of French quality + reasoning + self-hosting makes Magistral an excellent candidate for sensitive legal tasks. Legal content never leaves your machine—a decisive argument when dealing with a law firm that cannot send its files to OpenAI.

Three concrete use cases that work well:

Contract clause analysis
Paste a clause and ask it to identify the risks for either party. Magistral explicitly explains its reasoning, making the analysis auditable—a human can verify where the model saw a problem.
Version comparison
Paste two versions of an amendment, ask for a list of the material differences (not cosmetic ones) and their impact. Reasoning mode helps separate noise from substance.
Internal memo drafting
Requesting a summary from a legal text or court ruling. French language quality is the differentiating factor here.
Legal prompt example
>>> /set system "Tu es un juriste qui analyse les contrats en droit français. Raisonne méthodiquement avant de conclure. Cite les principes juridiques mobilisés."
>>> Analyse cette clause de mobilité géographique : [coller le texte]
!
An LLM does not replace a lawyer
Magistral can hallucinate French law—citing an article that does not exist or misreferencing a ruling. Its value lies in structuring legal reasoning and speeding up a first draft, not in producing a conclusion ready to sign as is. Any passage intended for a client or judge must be reviewed by a qualified human.

#Troubleshooting

The model switches to English during thinking
Strengthen the system prompt with an explicit instruction: "Think and answer only in French." If that is not enough, add a short first message in FR to anchor the language before the actual question.
Responses truncated before the conclusion
The model used up its token budget during thinking. Increase the window with `/set parameter num_ctx 16384` and `/set parameter num_predict 4096`.
Generation at 1–2 tokens/sec
The model is running on the CPU or partially offloaded. `ollama ps` confirms it. Solutions: free up VRAM (close Chrome / Steam), switch to more aggressive quantization such as Q3_K_M, or accept that this GPU is not sufficient.
The <think> tags appear in the final response
Expected behavior in raw mode. An interface such as Open WebUI or LM Studio automatically folds them into a collapsible "reasoning" block. In the Ollama CLI, they remain visible—that is by design.
Model “hallucinates” on recent facts
Magistral Small is frozen at its training date. For up-to-date factual questions (recent case law, current events), you need a RAG pipeline, not just the bare model.

#Go further

Magistral is an excellent French model if you have the GPU. Here are a few ideas for getting the most out of your setup:

Compare with DeepSeek R1 on your machine
The DeepSeek R1 installation guide shows how to install the 14B distillation alongside Magistral to test them side by side with your own prompts.
Connect Magistral to your legal PDFs
The PrivateGPT v2 guide describes a 100% local RAG pipeline that perfectly complements Magistral for confidential document analysis.
Choose the most suitable quantization
If you're hesitating between Q4_K_M, Q5_K_M, or Q6_K, the “Choosing Your Quantization” guide lays out the trade-offs. On a reasoning model, quantization losses are more noticeable than on a general-purpose model.
Did this guide help you?

Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.