Beginner 8 minOllama

Phi-4 locally: install with Ollama, VRAM and Speed

Phi-4 is Microsoft's small model that made a lot of noise in late 2024: only 14 billion parameters, yet math and coding scores that challenge models three times its size. This guide shows how to install Phi-4 Ollama locally, what it is really like in practice, and where it falls short—especially in French.

By Mohamed Meguedmi·Update 2026-08-27·Tested on Windows, macOS, and Linux
i
In brief
Phi-4 is a Microsoft model with 14 billion parameters, strong at math and coding (GPQA, MATH, HumanEval), under the MIT license. · It uses approximately 9 GB of VRAM in Q4_K_M, accessible with as little as a RTX 3060 or RTX 4070 12 GB, and even in CPU-only mode (16 GB of RAM, approximately 5 tok/s). · Installation takes one command: ollama pull phi4. · Its weakness is French, with faulty agreement and switches to English—prefer Mistral Small or Qwen 3.5 9B for French writing.

#Why Phi-4?

Phi-4 is the logical successor to Microsoft Research's Phi family: small models trained primarily on highly curated synthetic data, designed to compete with much larger models on reasoning tasks. Version 4, released in late 2024, takes this approach to 14B parameters and achieves remarkable scores on GPQA, MATH, and HumanEval—often above Llama 3.3 70B on these specific benchmarks.

In practice, this is a model that fits in 9 GB of VRAM at Q4_K_M, making it accessible to an entry-level 12 GB GPU such as a RTX 3060 or a RTX 4070. It also performs honorably on pure CPU if you have 16 GB of RAM. This balance of size and quality makes it an excellent candidate for modest self-hosting.

i
In two words
Phi-4 = 14B, strong at math and coding, weak in general knowledge and French, MIT license. Ideal if you want a “small powerhouse” for technical tasks on a 12 GB GPU.

#Prerequisites

The Local AI Kit

Your private ChatGPT, free, on your own machine in an hour — LM Studio, Ollama, Open WebUI, your documents, no cloud.

  • Lifetime online access
  • PDF + files
  • Lifetime updates
Ollama installed
Windows, macOS, or Linux. If you have not done so yet, the Ollama installation guide takes 3 minutes.
VRAM for Q4_K_M quantization
≈ 9 GB. A RTX 3060 12 GB, RTX 4070 12 GB, or a Mac with 16 GB of unified memory will work without issues.
CPU RAM without a GPU
16 GB minimum for Q4. Expect 4 to 8 tokens/sec on a recent Ryzen 5 or Core i5—usable for testing, but not for smooth chat.
10 GB of disk space
The Q4_K_M model weighs about 9 GB. Allow more if you also want to test Phi-4-Mini in parallel.
→
CPU-only case
Phi-4 is one of the few 14B models that is genuinely usable without a GPU. On a Ryzen 7 5800X or an Intel i7-12700, you get around 5 tok/sec in Q4—slow but OK for one question at a time. For interactive use, you should still target a GPU.

#1. Install Ollama

If Ollama is already set up, skip directly to the next section. Otherwise, here is the one-command installation for Linux and macOS.

Linux / macOS
curl -fsSL https://ollama.com/install.sh | sh

On Windows, download the installer from the official website.

Official Windows link
https://ollama.com/download/windows

Once installed, the Ollama daemon listens on port 11434.

Verification
ollama --version
curl http://localhost:11434/api/tags

#2. Get Phi-4

Phi-4 is available in the official Ollama catalog. The default tag pulls the Q4_K_M quantization, which is the recommended compromise.

Terminal
ollama pull phi4

The download is about 9 GB. To explicitly target another quantization, specify the tag.

Variants
ollama pull phi4:14b-q4_K_M    # ≈ 9 Go  (défaut)
ollama pull phi4:14b-q5_K_M    # ≈ 10 Go
ollama pull phi4:14b-q8_0      # ≈ 15 Go
ollama pull phi4:14b-fp16      # ≈ 29 Go (sans quantization)
→
Which quantization to choose
For Phi-4, Q4_K_M remains the best compromise on consumer hardware. The difference from Q5_K_M is marginal in practice for this model, while Q8_0 requires 15 GB of VRAM for a quality gain that is difficult to perceive outside benchmarks.

#3. First tests: math, code, reasoning

Start an interactive session to verify that the model runs and test its strengths.

Session
ollama run phi4

You should see a `>>>` prompt. This is where Phi-4 shines: logic and technical exercises. A few revealing prompts:

Math test
>>> Une boutique vend des stylos à 0,80 € pièce. Si vous en achetez 12 ou plus, le prix tombe à 0,65 € pièce. À partir de combien de stylos est-il plus économique de commander une boîte de 12 ?

Phi-4 will work through the reasoning step by step and arrive at the result. For code, it handles standard Python and JavaScript without a hitch.

Test code
>>> Écris une fonction Python qui prend une liste d'entiers et renvoie le sous-tableau contigu de somme maximale (algorithme de Kadane). Avec docstring et 3 cas de test.
i
Expected speed
On RTX 4070 (12 GB) in Q4_K_M, expect 45–60 tok/sec. On RTX 3060 (12 GB), 25–35 tok/sec. On a 24 GB Mac M2, 20–30 tok/sec. On a Ryzen 5 5600X CPU, 4–6 tok/sec.

#4. French quality: the real weaknesses

This is where Phi-4 shows its limitations. The model was trained primarily on English (and a lot of synthetic data generated in English). French is present in the corpus but clearly secondary.

Grammar and agreement
Good on simple sentences, but faulty agreement, anglicisms, and awkward phrasing as soon as you request long-form writing.
FR technical vocabulary
A tendency to use English terms ("endpoint," "deployment," "thread") even when a common French equivalent exists.
Understanding
It understands a prompt in French very well. The problem is output.
Comments in FR
Code comments frequently switch to English during generation. State the requirement explicitly; it helps but doesn't solve everything.
French culture
Low. Superficial knowledge of French case law, history, or institutions. Anything but a model for drafting a legal memo in French.
!
The French-language verdict
For solid French writing, choose Mistral Small 24B (general-purpose and good at French) or, on a more modest machine, Qwen 3.5 9B — much more comfortable with French than Phi-4. Phi-4 remains excellent for tasks where the output is code or structured reasoning—and where the interaction language matters little.

A tip that helps a little: a strict system prompt forces the model to stay in French, at the cost of sometimes shorter responses.

French system prompt
>>> /set system "Tu réponds exclusivement en français, sans jamais basculer en anglais. Les commentaires de code, les noms de variables explicatifs et les titres restent en français."

#5. Phi-4-Mini: the 3.8B version

Microsoft also released Phi-4-Mini, a 3.8B variant aimed at highly constrained configurations: laptops without a dedicated GPU, Raspberry Pi 5, and edge devices. At 2.5 GB in Q4, it runs on any modern machine.

Installation
ollama pull phi4-mini
Required VRAM/RAM
≈ 2.5 GB in Q4_K_M. Runs on an 8 GB MacBook Air M1 or a PC with a 4 GB GPU.
Speed
Very fast. 80+ tok/sec on RTX 3060, 30+ tok/sec on a recent CPU.
Strengths
Follows short instructions well; fine for prompt assistance, simple rewriting, and classification.
Weaknesses
Complex reasoning and long code are beyond its capabilities. It is a supporting model, not a replacement for the 14B version.
→
Phi-4-Mini use cases
Excellent for connecting behind a local n8n workflow, classifying emails, or integrating an LLM into a Python script that needs to run quickly and with a small footprint. Not recommended for general-purpose chat.

#6. Phi-4 vs. Qwen 3.5 9B

In 2026, the most relevant direct competitor is Alibaba’s Qwen 3.5 9B. Slightly more compact (6.6 GB in Q4 versus 9 GB for Phi-4), it fits on the same 12 GB GPU and targets the same user category. What sets them apart:

Reasoning and math
Phi-4 has the edge on formal benchmarks (GPQA, MATH). In practical use on common problems, the gap is narrower than you might think.
Code
Qwen 3.5 9B is generally better at less common languages (Rust, Go, advanced SQL). Phi-4 remains competitive on Python and JS.
French
Qwen 3.5 is significantly better. Its training uses much more multilingual data, with a richer vocabulary and fewer switches to English. Qwen has a clear advantage on this criterion.
Context
Qwen 3.5 9B accepts up to 256k context tokens, while Phi-4 is capped at 16k. If you work with long documents, Qwen is much better suited.
Instruction following
Phi-4 follows structured instructions more precisely (JSON format, strict constraints). Qwen is a bit more "creative" and may take liberties.
→
How to choose
Work in English on reasoning, Python code, and structured output: Phi-4. Work in French, long context, and multilingual code: Qwen 3.5 9B. With a 12 GB GPU, you can install both and switch between them depending on the task.

#Troubleshooting

"Out of memory" while loading
Your VRAM is insufficient for the selected quantization. Switch to Q4_K_M (the lightest useful option), or let Ollama perform partial CPU offloading with `ollama run phi4 --num-gpu N`, where N is the number of layers to place on the GPU.
Generation at 1–2 tokens/sec on a machine equipped with a GPU
The model is probably running on the CPU. Check the "PROCESSOR" column with `ollama ps`. If it shows 100% CPU, free up VRAM (Chrome, games, other loaded models) and restart.
Responses that switch to English midway
Expected behavior on Phi-4 in French. Strengthening the system prompt helps somewhat. For recurring French-language use, switch models rather than fighting it.
Context truncated to 4096 tokens
Ollama loads a short context by default. Extend it explicitly: `/set parameter num_ctx 16384`. Beyond that, Phi-4 was trained only up to 16k—there is no point pushing it higher.
The model hallucinates recent facts
Phi-4 has knowledge frozen at its training cutoff (late 2024) and weaker general knowledge than larger models. For up-to-date factual information, you need a RAG pipeline, not the bare model.

#Go further

Phi-4 is a good entry point for local 14B LLMs. A few ways to go further with your setup:

Test a direct competitor
The Qwen 3 testing guide shows comparable benchmarks on consumer hardware, which is useful for comparing results on your own prompts.
Choosing the right quantization
The “Choosing your quantization (Q4, Q5, Q8, FP16)” guide details the quality/memory trade-offs that also apply to Phi-4.
Run Phi-4 without a GPU
If you’re running on pure CPU, the “Run an LLM locally without a GPU” guide complements what’s described here with specific optimizations.
Did this guide help you?

Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.