The short answer: is Phi-4 14B usable on CPU?
Yes, Phi-4 14B runs with no GPU, and on a modern quad-core it is fast enough to read along with the output as it streams. No, it will not feel like a cloud API. The single number that predicts your experience is not your CPU model but your memory bandwidth: token generation reads the entire set of active weights from RAM for every token, so speed is capped by how fast your RAM can feed the cores.
That gives a simple ceiling. Divide your usable memory bandwidth by the size of the model in RAM and you get the theoretical maximum tokens per second. A 4-bit Phi-4 is about 8-9 GB, so a dual-channel DDR4-3200 desktop (~51 GB/s) tops out near 6 tok/s and realistically delivers low single digits; dual-channel DDR5-5600 (~89 GB/s) roughly doubles that. Prompt processing (prefill) is compute-bound and much faster, so long prompts are cheap to ingest — the wait is in generation. If you want the underlying math for any model and quant, our quantization guide and the RAM/VRAM calculator cover it.
Tokens/s by quant and thread count
The table below is modeled from memory bandwidth, not a single benchmark run, because the honest answer depends far more on your RAM than on which 8-core chip you own. Treat these as expected ranges and measure your own box — llama-bench in llama.cpp gives a repeatable number in under a minute.
| Quant | RAM footprint (8K ctx) | DDR4-3200 gen (tok/s) | DDR5-5600 gen (tok/s) |
|---|---|---|---|
| Q4_K_M | ~9-10 GB | 2-5 | 5-9 |
| Q5_K_M | ~11-12 GB | 2-4 | 4-7 |
| Q8_0 | ~16-17 GB | 1-3 | 3-5 |
| FP16 | ~30-32 GB | <1-2 | 2-3 |
Thread count matters less than people expect. Generation saturates near your physical core count and then stops scaling, because you have run out of memory bandwidth, not cores. A few practical rules:
- Set threads to physical cores, not logical (hyper-threaded) cores. On an 8-core/16-thread chip, start at 8.
- Going past physical cores usually flat-lines or slightly regresses generation speed.
- Prefill benefits from more threads than generation, so tools that expose a separate batch-thread setting can process prompts faster.
- Close browsers and background apps: they compete for the same memory bandwidth you are trying to spend on the model.
How much RAM you actually need
Sizing uses the standard rules of thumb: Q4_K_M is about 0.58 GB per billion parameters, Q8 about 1.07, and FP16 about 2, plus roughly 20% for the KV cache and runtime overhead at an 8K context. For a ~14B model that works out as follows.
| Quant | Weights only | + KV/overhead (8K) | Comfortable RAM |
|---|---|---|---|
| Q4_K_M | ~8.1 GB | ~9.7 GB | 16 GB |
| Q5_K_M | ~9.7 GB | ~11.6 GB | 16 GB (tight) |
| Q8_0 | ~15 GB | ~18 GB | 32 GB |
| FP16 | ~28 GB | ~33 GB | 48 GB+ |
On 16 GB, Q4_K_M is the sweet spot: it leaves 5-6 GB for Windows and your editor. Q5_K_M works but you will want to close everything else, and larger contexts push the KV cache up fast. On 32 GB, Q8 becomes comfortable and is the quality-per-effort pick for CPU, while FP16 fits but is slow enough that it is rarely worth it over Q8. Remember the context window is not free — the KV cache grows with tokens, so a 32K context can add several gigabytes on top of the weights.
llama.cpp vs Ollama vs LM Studio on CPU
Here is the part that saves you a week of tinkering: all three run the same engine. Ollama and LM Studio both wrap llama.cpp's ggml backend, so on identical hardware, quant, and thread settings their generation speed lands within a few percent of each other. Choose on ergonomics, not on a myth that one is dramatically faster on CPU.
- llama.cpp — maximum control. You set thread count, batch size, and memory-mapping flags directly, and
llama-benchgives clean numbers. Best if you want to tune. Pull a GGUF from Hugging Face and check the model card for exact quant tags. - Ollama — simplest path.
ollama run phi-4pulls a sensible default quant and serves an API. Least to configure, slightly less visibility into settings. See what Ollama is. - LM Studio — GUI with a model browser, quant picker, and per-model thread and context sliders. Friendliest for non-terminal users.
If you are deciding between them, our comparisons of Ollama vs llama.cpp and LM Studio vs Ollama go deeper. One stable tip that survives version churn: whichever tool you pick, verify it is not silently offloading nothing to a GPU and confirm the thread count — those two settings explain most "why is it so slow" reports. Exact flags and model tags change between releases, so check each project's own docs.
When CPU-only Phi-4 makes sense (and when it doesn't)
CPU-only Phi-4 14B is genuinely good for offline, private, and non-interactive work: summarizing documents you cannot send to a cloud, drafting text, one-shot reasoning where you can wait a few seconds, or running on a laptop with no discrete GPU. Phi-4's strength is reasoning density for its size, so the quality you get from a 14B at Q4 on your own machine is a fair trade for the wait.
It is a poor fit for interactive coding assistants and long chats, where 3 tok/s feels glacial and long contexts balloon the KV cache. If that is your use case, either add a GPU or drop to a smaller model — a 7-8B or a 3-4B class model at Q4 will feel several times more responsive on the same CPU. See our picks for the best CPU-only models and the measured benchmarks to compare before you commit. The rule of thumb holds: on CPU, smaller-and-faster usually beats bigger-and-smarter for anything you interact with in real time. To read up on Phi-4 itself, its technical report and the model card are the authoritative sources.
Frequently asked questions
Can Phi-4 14B run on 16 GB of RAM with no GPU?
Yes, at Q4_K_M it uses roughly 9-10 GB including the KV cache at an 8K context, which leaves headroom for Windows and an editor on a 16 GB machine. Higher quants like Q8 need about 18 GB, so those want 32 GB. Keep the context modest, since a larger window grows the KV cache and can push you over 16 GB.
How many tokens per second does Phi-4 14B get on CPU only?
Expect roughly 2-5 tokens/s on a dual-channel DDR4-3200 system at Q4_K_M, and about 5-9 tokens/s on dual-channel DDR5-5600. Prompt processing is much faster than generation because it is compute-bound rather than bandwidth-bound. Your exact number depends far more on RAM speed than on your CPU model.
Does adding more CPU threads make Phi-4 faster?
Only up to a point. Generation is limited by memory bandwidth, so it typically saturates around your physical core count and then stops improving or slightly regresses. Set threads to physical cores rather than logical ones, and close background apps that are competing for the same memory bandwidth.
Is llama.cpp faster than Ollama or LM Studio for CPU inference?
Not meaningfully, because Ollama and LM Studio both wrap the same llama.cpp ggml backend. On identical hardware, quant, and thread settings they land within a few percent of each other. Pick based on how much control or GUI convenience you want, not on raw speed.
By Mohamed Meguedmi — independent comparator of locally-runnable LLMs, benchmarked on a real RTX 5070 Ti (data CC BY 4.0). See the local LLM leaderboard and the best Ollama models.