🇺🇸 Gemma 4 E4B
4B effective multimodal (text+image+audio). 140 languages. For laptops and edge devices.
ollama run gemma4:e4b
Ranking updated on 09/10/2026
No GPU? Modern 1–7B LLMs run reasonably well on the CPU thanks to llama.cpp and Q4 quantization. You need 8–16 GB of RAM and must accept speeds of 5–15 tokens/sec. Here are the best options.
4B effective multimodal (text+image+audio). 140 languages. For laptops and edge devices.
ollama run gemma4:e4b
Dense 3B Apache 2.0, 12 languages including FR, 131k ctx, GQA 40Q/8KV. Tool calling and code FIM. Released April 29, 2026.
ollama run granite4.1:3b
Gemma 4 E2B: 2B active (5.1B total), ~3 GB VRAM Q4 (weights Ollama 4.3 GB in QAT, 7.2 GB by default). Text-and-image multimodal, 128k context, Apache 2.0.
ollama run gemma4:e2b
Granite 4.1 (3B Apache 2.0): generic Ollama tag from the IBM Granite 4.1 family, 128k ctx, tool calling, and code. Released May 2026.
ollama run granite4.1
SFT “thinking” fine-tune of OLMo 3 7B: step-by-step reasoning, 16k context, ~4.2 GB VRAM in Q4. 100% open, Apache 2.0 license.
# HuggingFace : zimplex/olmo3-7b-think-sft-eosfix-16k-3ep-euc
GLM 5.3 (Zhipu): dense 7B specialized in code and reasoning, 128k context, ~4.1 GB VRAM in Q4. Lightweight, runs on a 6–8 GB GPU, MIT license.
ollama pull glm-5.3
3B VLM specialized in enterprise document extraction. OCR, tables, forms.
# HuggingFace : ibm-granite/granite-4.0-3b-vision
| Rank | Model | Params | Q4 VRAM | Context | License | On integrated GPU / none |
|---|---|---|---|---|---|---|
| #1 | Gemma 4 E4B | 4B | 10 GB | 128 000 | Apache 2.0 | ✗ |
| #2 | Granite 4.1 3B Instruct | 3B | 2 GB | 131 072 | Apache 2.0 | ✗ |
| #3 | Gemma 4 E2B | 2B | 3 GB | 128 000 | Apache 2.0 | ✗ |
| #4 | Granite 4.1 | 3B | 1.7 GB | 128 000 | Apache 2.0 | ✗ |
| #5 | OLMo 3 7B Think (SFT) | 7B | 4.2 GB | 16 000 | Apache 2.0 | ✗ |
| #6 | GLM 5.3 7B | 7B | 4.1 GB | 128 000 | MIT | ✗ |
| #7 | Granite 4.0 3B Vision | 3B | 2.2 GB | 16 384 | Apache 2.0 | ✗ |
Your private, free ChatGPT on your machine in 1 hour — LM Studio, Ollama, Open WebUI, your documents, no cloud.
Models ≤ 8B only—beyond that, CPU throughput becomes too low. Big bonus for ≤ 3B (very fast on modern CPUs) and permissive licenses.
Criteria considered:
The scoring is fully transparent: see our methodology for details on VRAM/tokens/sec calculations.
Can you really run an LLM without a GPU?
Yes — thanks to llama.cpp + Q4_K_M quantization, a 7B model runs at 5–10 tokens/sec on a Ryzen 7 CPU or Apple M1. For interactive use, target a 1–3B model (30–50 tokens/sec).
How much RAM do you need?
For a 7B in Q4: 8 GB of RAM minimum, 16 GB recommended (the model uses ~5 GB, with the rest for the OS + context). For a 3B: 6–8 GB is enough. For a 1B: 4 GB.
Which CPU for LLMs?
The more cores and AVX2/AVX-512, the better. Recent Ryzen 7/9, recent Intel Core i7/i9, Apple M1/M2/M3—all excellent. Favor fast memory (DDR5 > DDR4): RAM bandwidth is often more limiting than core count.
Can an iGPU (Intel / AMD) help?
Marginally—Vulkan on an iGPU provides a 20–40% gain over CPU alone. Not transformative, but worth taking. On Mac Apple Silicon, the integrated GPU can be used through Metal and makes a big difference (this is the 'unified memory' mode).