Local LLM without a GPU (CPU): models by RAM capacity 8/16/32 GB
Running a local LLM without a GPU, using only the CPU, is entirely possible in 2026. With a recent Ryzen 7 or i7 and 16 GB of RAM, you can chat with a 3B model in near-real time, or let a 7B model run for less time-sensitive responses. This guide gives you the models to choose based on your RAM, real-world figures measured on common machines, and the settings that genuinely make a difference.
Choosing a machine? Our picks by budget →
#Why run a local LLM without a GPU?
All local AI guides assume you have a RTX 4070 in your desktop. The reality is that a large majority of laptops, professional desktop PCs, and mini PCs run on integrated graphics or have no dedicated card at all. The good news: a pure-CPU LLM works.
Three typical reasons to target CPU-only: a laptop without a GPU NVIDIA (most Dell and Lenovo systems, and older Intel Macs), a professional desktop with an Intel or AMD iGPU, or a headless Linux server you don’t want to equip. In all three cases, the question isn’t “does it work?” but “which model remains usable?”
#What you can expect in practice
Your private ChatGPT, free, on your own machine in an hour — LM Studio, Ollama, Open WebUI, your documents, no cloud.
- Lifetime online access
- PDF + files
- Lifetime updates
Before downloading anything, set your expectations. A local LLM running on a CPU without a GPU is not as responsive as a ChatGPT or a model on RTX 4090. Here are realistic ballpark figures for 2026:
- 1B–2B Q4 model
- 20 to 40 tokens/sec on a recent CPU. Very fluid, almost like an online chat. Perfect for simple tasks: rewriting, short summaries, classification.
- 3B Q4 model
- 10 to 20 tokens/sec. Still comfortable: you read as the model writes. The CPU sweet spot for everyday use.
- 7B–8B Q4 model
- 4 to 10 tokens/sec. Usable, but you wait. Good for asynchronous tasks (analyzing text, generating a draft).
- 13B–14B Q4 model
- 2 to 5 tokens/sec. Barely tolerable. Reserve it for batch tasks, not interactive chat.
- Beyond 14B
- Possible but painful. It’s better to rent a cloud GPU for an hour than run a 32B model on the CPU.
#Recommended models by available RAM
On CPU, all system RAM can be used by the model, but you need to leave headroom for the OS and your applications. Reserve 4 to 6 GB for the system. The rest determines the accessible model size.
#8 GB of RAM: 2B to 4B models
- Qwen 3.5 2B Q4_K_M
- ≈1.9 GB. Ultra-fast, multimodal, with up to 256k context and an Apache 2.0 license. Good for paraphrasing, translation, and classification.
- Granite 4.2 3B Q4_K_M
- ≈2.2 GB. Very lightweight and token-efficient, from IBM, Apache 2.0 licensed. Speaks good French.
- Qwen 3.5 4B Q4_K_M
- ≈3.4 GB. The new small default model. Solid at coding and French.
- Gemma 4 E2B Q4 (QAT)
- ≈4.3 GB. Compact and multimodal, made by Google, and licensed under Apache 2.0 since 2026.
#16 GB of RAM: comfortable for 8B–12B
- Granite 4.2 8B Q4_K_M
- ≈5.3 GB. Highly token-efficient, 128k context, Apache 2.0 license. Fast to load.
- Qwen 3.5 9B Q4_K_M
- ≈6.6 GB. THE 8 GB choice for 2026: 256k context, vision, excellent reasoning, and French.
- Gemma 4 12B Q4_K_M
- ≈7.6 GB. Multimodal and efficient, with an Apache 2.0 license. A step up if your RAM allows it.
- Qwen 2.5 Coder 7B base Q4_K_M
- ≈4.7 GB. The exception that still holds true: the 2026 benchmark for local inline code autocompletion (FIM).
#32 GB of RAM: you can target a 24B model
- Mistral Small 24B Q4_K_M
- ≈14 GB. General-purpose, very good in French, but 3 to 5 tokens/sec on CPU.
- gpt-oss 20B Q4 (MXFP4)
- ≈14 GB. OpenAI open-weight model, very fast thanks to the MXFP4 format, 131k context.
- Qwen 3.8 27B Q4_K_M (borderline)
- ≈18 GB. 262k context, vision, Apache 2.0 license. Possible but slow, ~2 tokens/sec. More useful in batch—consider switching its reasoning to low to prevent it from overthinking.
#1. Install Ollama (automatic CPU mode)
Ollama is the simplest tool to get started. It automatically detects the absence of a GPU and switches to CPU without any special configuration. The daemon listens on http://localhost:11434 by default.
On Windows and macOS, download the installer from ollama.com. No special setting is needed for CPU mode—Ollama makes the right choice automatically.
The ollama ps command should display the daemon status. If a conversation is running, the PROCESSOR column will show 100% CPU—that is exactly what we want here.
#2. Three models to compare in CPU-only mode
With 16 GB of RAM, the question is not “which model” but “which of the three leading small models of 2026.” Download all three and form your own opinion in an hour.
- Qwen 3.5 4B
- The best all-rounder at this size. Very strong in French, good at coding, 256k context, and follows instructions well. The slowest of the three (4B has its price).
- Granite 4.2 3B
- Very lean and token-efficient, from IBM. Good instruction following, Apache 2.0 license. The right balance of speed and quality.
- Gemma 4 E2B
- The fastest of the three. Multimodal, with astonishing quality for its size. Ideal if you want near-real-time performance on a modest CPU. Worse at coding than the other two.
The --verbose option displays statistics at the bottom of each response: prompt eval rate, eval rate (tokens/sec during generation), total duration. This is your reference metric on this machine.
#3. Tokens/sec benchmarks: ballpark figures
Here are rough figures for representative machines, without a GPU, with Ollama (which uses llama.cpp under the hood) in Q4_K_M quantization. Your results will vary by ±20% depending on context, memory, and DDR frequency.
#Intel Core i5-12400 + DDR4-3200 16 GB
- Gemma 4 E2B Q4
- ≈ 26 tokens/sec during generation
- Granite 4.2 3B Q4_K_M
- ≈ 20 tokens/sec
- Qwen 3.5 4B Q4_K_M
- ≈ 15 tokens/sec
- Granite 4.2 8B Q4_K_M
- ≈ 8 tokens/sec
- Qwen 3.5 9B Q4_K_M
- ≈ 6 tokens/sec
#Intel Core i7-13700K + DDR5-5600 32 GB
- Gemma 4 E2B Q4
- ≈ 40 tokens/sec
- Granite 4.2 3B Q4_K_M
- ≈ 30 tokens/sec
- Qwen 3.5 4B Q4_K_M
- ≈ 23 tokens/sec
- Qwen 3.5 9B Q4_K_M
- ≈ 12 tokens/sec
- Mistral Small 24B Q4_K_M
- ≈ 4 tokens/sec
#AMD Ryzen 7 7700X + 32 GB DDR5-6000
- Gemma 4 E2B Q4
- ≈ 44 tokens/sec
- Granite 4.2 3B Q4_K_M
- ≈ 32 tokens/sec
- Qwen 3.5 4B Q4_K_M
- ≈ 24 tokens/sec
- Granite 4.2 8B Q4_K_M
- ≈ 13 tokens/sec
- Qwen 3.5 9B Q4_K_M
- ≈ 11 tokens/sec
- Mistral Small 24B Q4_K_M
- ≈ 5 tokens/sec
#4. Compile llama.cpp with AVX-512 (advanced)
Ollama includes generic precompiled llama.cpp binaries. By compiling llama.cpp yourself with your CPU's instruction sets (AVX2, AVX-512, AMX), you can gain 10 to 30% tokens/sec on some processors. For Intel Core 11th gen+ (Ice Lake / Rocket Lake / Sapphire Rapids) processors that support AVX-512.
If the command returns lines (avx512f, avx512dq, etc.), your CPU supports AVX-512. Otherwise, stick with standard Ollama; you have nothing to gain.
The -DGGML_NATIVE=ON option lets the compiler automatically detect your CPU's instruction sets and enables everything available. It is the simplest and most reliable method.
The -t option sets the number of threads (use the number of physical cores, not logical cores). llama-bench returns a table with pp512 (prompt eval) and tg128 (generation) in tokens/sec—your new baseline for comparison.
#Tips for getting more tokens/sec on CPU
- 01Set the number of threadsBy default, Ollama uses all logical cores. On some CPUs with hyper-threading, limiting it to physical cores (OLLAMA_NUM_THREADS=8 for an 8-core CPU) increases speed by 5 to 15%.
- 02Keep the model loadedLoading the model takes several seconds. OLLAMA_KEEP_ALIVE=30m keeps the model in RAM for 30 minutes after the last request. Without a GPU, this is even more valuable because reloading is slow.
- 03Reduce the context if possiblenum_ctx 2048 instead of 8192 saves RAM and noticeably improves speed. Keep a large context only for use cases that truly need it (RAG, long documents).
- 04Close Chrome and SlackA 7B LLM on the CPU saturates memory bandwidth. Anything else that also uses RAM (a browser with 50 tabs, Slack, Teams) takes cycles away from it. On 16 GB, that can make the difference between 5 and 8 tokens/sec.
- 05Choose the fastest compatible DDRIf you upgrade: DDR4-3200 → DDR4-3600 = +10%. DDR4 → DDR5-5600 = +30 to 50%. The CPU matters much less than memory for LLM inference.
#When the CPU is no longer enough
Let's be honest: without a GPU, some use cases remain out of reach. If you recognize your situation in the list below, it is time to consider even a modest GPU (RTX 3060 12 GB used for €250 can be life-changing) or rent cloud capacity by the hour.
- Interactive chat with a 13B+
- 2 to 5 tokens/sec is too slow for conversational AI. A 12 GB GPU fixes that instantly.
- RAG with a large context (16k+)
- Prompt evaluation time explodes on CPU. A RTX 3060 processes an 8k prompt in 1 sec, while an i7 takes 30 sec.
- Real-time code completion
- Inline autocompletion extensions (Tabby and similar tools, in FIM mode with Qwen 2.5 Coder 7B as the base) need responses in under 200 ms. On CPU, you will not get below 1 to 2 sec. A GPU is mandatory.
- High-volume generation
- Processing 1000 documents takes days on a CPU and hours on a GPU. For a one-off batch, RunPod or Vast.ai at €0.30/h gets the job done overnight.
#Go further
You have a CPU-only local LLM that responds. A few natural next steps, depending on your next question:
- Choose the right quantization
- Q4_K_M is the sensible default, but Q5_K_M or Q3 have their place depending on your RAM. The quantization guide details the trade-offs.
- Wrap it all in an interface
- The terminal is fine for testing. Open WebUI or LM Studio give you a local ChatGPT-like interface in minutes.
- When to add a GPU
- If you take the plunge, the GPU selection guide compares RTX 3060 vs 4060 vs 4070, along with the associated LLM benchmarks.
Recommended hardware: Radeon RX 9070 XT 16GB (ASUS Prime OC) — to move from CPU-only to a dedicated GPU. All AI hardware →
Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.