Phi-4 locally: install with Ollama, VRAM and Speed
Phi-4 is Microsoft's small model that made a lot of noise in late 2024: only 14 billion parameters, yet math and coding scores that challenge models three times its size. This guide shows how to install Phi-4 Ollama locally, what it is really like in practice, and where it falls short—especially in French.
#Why Phi-4?
Phi-4 is the logical successor to Microsoft Research's Phi family: small models trained primarily on highly curated synthetic data, designed to compete with much larger models on reasoning tasks. Version 4, released in late 2024, takes this approach to 14B parameters and achieves remarkable scores on GPQA, MATH, and HumanEval—often above Llama 3.3 70B on these specific benchmarks.
In practice, this is a model that fits in 9 GB of VRAM at Q4_K_M, making it accessible to an entry-level 12 GB GPU such as a RTX 3060 or a RTX 4070. It also performs honorably on pure CPU if you have 16 GB of RAM. This balance of size and quality makes it an excellent candidate for modest self-hosting.
#Prerequisites
Your private ChatGPT, free, on your own machine in an hour — LM Studio, Ollama, Open WebUI, your documents, no cloud.
- Lifetime online access
- PDF + files
- Lifetime updates
- Ollama installed
- Windows, macOS, or Linux. If you have not done so yet, the Ollama installation guide takes 3 minutes.
- VRAM for Q4_K_M quantization
- ≈ 9 GB. A RTX 3060 12 GB, RTX 4070 12 GB, or a Mac with 16 GB of unified memory will work without issues.
- CPU RAM without a GPU
- 16 GB minimum for Q4. Expect 4 to 8 tokens/sec on a recent Ryzen 5 or Core i5—usable for testing, but not for smooth chat.
- 10 GB of disk space
- The Q4_K_M model weighs about 9 GB. Allow more if you also want to test Phi-4-Mini in parallel.
#1. Install Ollama
If Ollama is already set up, skip directly to the next section. Otherwise, here is the one-command installation for Linux and macOS.
On Windows, download the installer from the official website.
Once installed, the Ollama daemon listens on port 11434.
#2. Get Phi-4
Phi-4 is available in the official Ollama catalog. The default tag pulls the Q4_K_M quantization, which is the recommended compromise.
The download is about 9 GB. To explicitly target another quantization, specify the tag.
#3. First tests: math, code, reasoning
Start an interactive session to verify that the model runs and test its strengths.
You should see a `>>>` prompt. This is where Phi-4 shines: logic and technical exercises. A few revealing prompts:
Phi-4 will work through the reasoning step by step and arrive at the result. For code, it handles standard Python and JavaScript without a hitch.
#4. French quality: the real weaknesses
This is where Phi-4 shows its limitations. The model was trained primarily on English (and a lot of synthetic data generated in English). French is present in the corpus but clearly secondary.
- Grammar and agreement
- Good on simple sentences, but faulty agreement, anglicisms, and awkward phrasing as soon as you request long-form writing.
- FR technical vocabulary
- A tendency to use English terms ("endpoint," "deployment," "thread") even when a common French equivalent exists.
- Understanding
- It understands a prompt in French very well. The problem is output.
- Comments in FR
- Code comments frequently switch to English during generation. State the requirement explicitly; it helps but doesn't solve everything.
- French culture
- Low. Superficial knowledge of French case law, history, or institutions. Anything but a model for drafting a legal memo in French.
A tip that helps a little: a strict system prompt forces the model to stay in French, at the cost of sometimes shorter responses.
#5. Phi-4-Mini: the 3.8B version
Microsoft also released Phi-4-Mini, a 3.8B variant aimed at highly constrained configurations: laptops without a dedicated GPU, Raspberry Pi 5, and edge devices. At 2.5 GB in Q4, it runs on any modern machine.
- Required VRAM/RAM
- ≈ 2.5 GB in Q4_K_M. Runs on an 8 GB MacBook Air M1 or a PC with a 4 GB GPU.
- Speed
- Very fast. 80+ tok/sec on RTX 3060, 30+ tok/sec on a recent CPU.
- Strengths
- Follows short instructions well; fine for prompt assistance, simple rewriting, and classification.
- Weaknesses
- Complex reasoning and long code are beyond its capabilities. It is a supporting model, not a replacement for the 14B version.
#6. Phi-4 vs. Qwen 3.5 9B
In 2026, the most relevant direct competitor is Alibaba’s Qwen 3.5 9B. Slightly more compact (6.6 GB in Q4 versus 9 GB for Phi-4), it fits on the same 12 GB GPU and targets the same user category. What sets them apart:
- Reasoning and math
- Phi-4 has the edge on formal benchmarks (GPQA, MATH). In practical use on common problems, the gap is narrower than you might think.
- Code
- Qwen 3.5 9B is generally better at less common languages (Rust, Go, advanced SQL). Phi-4 remains competitive on Python and JS.
- French
- Qwen 3.5 is significantly better. Its training uses much more multilingual data, with a richer vocabulary and fewer switches to English. Qwen has a clear advantage on this criterion.
- Context
- Qwen 3.5 9B accepts up to 256k context tokens, while Phi-4 is capped at 16k. If you work with long documents, Qwen is much better suited.
- Instruction following
- Phi-4 follows structured instructions more precisely (JSON format, strict constraints). Qwen is a bit more "creative" and may take liberties.
#Troubleshooting
- "Out of memory" while loading
- Your VRAM is insufficient for the selected quantization. Switch to Q4_K_M (the lightest useful option), or let Ollama perform partial CPU offloading with `ollama run phi4 --num-gpu N`, where N is the number of layers to place on the GPU.
- Generation at 1–2 tokens/sec on a machine equipped with a GPU
- The model is probably running on the CPU. Check the "PROCESSOR" column with `ollama ps`. If it shows 100% CPU, free up VRAM (Chrome, games, other loaded models) and restart.
- Responses that switch to English midway
- Expected behavior on Phi-4 in French. Strengthening the system prompt helps somewhat. For recurring French-language use, switch models rather than fighting it.
- Context truncated to 4096 tokens
- Ollama loads a short context by default. Extend it explicitly: `/set parameter num_ctx 16384`. Beyond that, Phi-4 was trained only up to 16k—there is no point pushing it higher.
- The model hallucinates recent facts
- Phi-4 has knowledge frozen at its training cutoff (late 2024) and weaker general knowledge than larger models. For up-to-date factual information, you need a RAG pipeline, not the bare model.
#Go further
Phi-4 is a good entry point for local 14B LLMs. A few ways to go further with your setup:
- Test a direct competitor
- The Qwen 3 testing guide shows comparable benchmarks on consumer hardware, which is useful for comparing results on your own prompts.
- Choosing the right quantization
- The “Choosing your quantization (Q4, Q5, Q8, FP16)” guide details the quality/memory trade-offs that also apply to Phi-4.
- Run Phi-4 without a GPU
- If you’re running on pure CPU, the “Run an LLM locally without a GPU” guide complements what’s described here with specific optimizations.
Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.