Beginner 10 minConfiguration

Local LLM without a GPU (CPU): models by RAM capacity 8/16/32 GB

Running a local LLM without a GPU, using only the CPU, is entirely possible in 2026. With a recent Ryzen 7 or i7 and 16 GB of RAM, you can chat with a 3B model in near-real time, or let a 7B model run for less time-sensitive responses. This guide gives you the models to choose based on your RAM, real-world figures measured on common machines, and the settings that genuinely make a difference.

Choosing a machine? Our picks by budget →

By Mohamed Meguedmi·Update 2026-08-27·Tested on Windows, macOS, and Linux

#Why run a local LLM without a GPU?

All local AI guides assume you have a RTX 4070 in your desktop. The reality is that a large majority of laptops, professional desktop PCs, and mini PCs run on integrated graphics or have no dedicated card at all. The good news: a pure-CPU LLM works.

Three typical reasons to target CPU-only: a laptop without a GPU NVIDIA (most Dell and Lenovo systems, and older Intel Macs), a professional desktop with an Intel or AMD iGPU, or a headless Linux server you don’t want to equip. In all three cases, the question isn’t “does it work?” but “which model remains usable?”

i
The real limiting factor: RAM, not the CPU
On pure CPU, memory bandwidth caps tokens/sec, not raw processor power. A recent Ryzen 7 and i5 often deliver very similar results on the same model.

#What you can expect in practice

The Local AI Kit

Your private ChatGPT, free, on your own machine in an hour — LM Studio, Ollama, Open WebUI, your documents, no cloud.

  • Lifetime online access
  • PDF + files
  • Lifetime updates

Before downloading anything, set your expectations. A local LLM running on a CPU without a GPU is not as responsive as a ChatGPT or a model on RTX 4090. Here are realistic ballpark figures for 2026:

1B–2B Q4 model
20 to 40 tokens/sec on a recent CPU. Very fluid, almost like an online chat. Perfect for simple tasks: rewriting, short summaries, classification.
3B Q4 model
10 to 20 tokens/sec. Still comfortable: you read as the model writes. The CPU sweet spot for everyday use.
7B–8B Q4 model
4 to 10 tokens/sec. Usable, but you wait. Good for asynchronous tasks (analyzing text, generating a draft).
13B–14B Q4 model
2 to 5 tokens/sec. Barely tolerable. Reserve it for batch tasks, not interactive chat.
Beyond 14B
Possible but painful. It’s better to rent a cloud GPU for an hour than run a 32B model on the CPU.
→
Rule of thumb
Below 5 tokens/sec, interactive chat becomes frustrating. Above 15 tokens/sec, you read as fast as the model writes. Aim for the second range.

#Recommended models by available RAM

On CPU, all system RAM can be used by the model, but you need to leave headroom for the OS and your applications. Reserve 4 to 6 GB for the system. The rest determines the accessible model size.

#8 GB of RAM: 2B to 4B models

Qwen 3.5 2B Q4_K_M
≈1.9 GB. Ultra-fast, multimodal, with up to 256k context and an Apache 2.0 license. Good for paraphrasing, translation, and classification.
Granite 4.2 3B Q4_K_M
≈2.2 GB. Very lightweight and token-efficient, from IBM, Apache 2.0 licensed. Speaks good French.
Qwen 3.5 4B Q4_K_M
≈3.4 GB. The new small default model. Solid at coding and French.
Gemma 4 E2B Q4 (QAT)
≈4.3 GB. Compact and multimodal, made by Google, and licensed under Apache 2.0 since 2026.

#16 GB of RAM: comfortable for 8B–12B

Granite 4.2 8B Q4_K_M
≈5.3 GB. Highly token-efficient, 128k context, Apache 2.0 license. Fast to load.
Qwen 3.5 9B Q4_K_M
≈6.6 GB. THE 8 GB choice for 2026: 256k context, vision, excellent reasoning, and French.
Gemma 4 12B Q4_K_M
≈7.6 GB. Multimodal and efficient, with an Apache 2.0 license. A step up if your RAM allows it.
Qwen 2.5 Coder 7B base Q4_K_M
≈4.7 GB. The exception that still holds true: the 2026 benchmark for local inline code autocompletion (FIM).

#32 GB of RAM: you can target a 24B model

Mistral Small 24B Q4_K_M
≈14 GB. General-purpose, very good in French, but 3 to 5 tokens/sec on CPU.
gpt-oss 20B Q4 (MXFP4)
≈14 GB. OpenAI open-weight model, very fast thanks to the MXFP4 format, 131k context.
Qwen 3.8 27B Q4_K_M (borderline)
≈18 GB. 262k context, vision, Apache 2.0 license. Possible but slow, ~2 tokens/sec. More useful in batch—consider switching its reasoning to low to prevent it from overthinking.
!
More aggressive quantization?
Q3_K_M saves ~20% of RAM compared with Q4_K_M, but quality visibly declines on models <7B. For a 1B–3B model, stay with at least Q4_K_M. For a 13B model on 16 GB, Q3 may save the day.

#1. Install Ollama (automatic CPU mode)

Ollama is the simplest tool to get started. It automatically detects the absence of a GPU and switches to CPU without any special configuration. The daemon listens on http://localhost:11434 by default.

Linux — official install script
curl -fsSL https://ollama.com/install.sh | sh

On Windows and macOS, download the installer from ollama.com. No special setting is needed for CPU mode—Ollama makes the right choice automatically.

Verify that Ollama is running
ollama --version
ollama ps

The ollama ps command should display the daemon status. If a conversation is running, the PROCESSOR column will show 100% CPU—that is exactly what we want here.

#2. Three models to compare in CPU-only mode

With 16 GB of RAM, the question is not “which model” but “which of the three leading small models of 2026.” Download all three and form your own opinion in an hour.

Download the three challengers
ollama pull qwen3.5:4b
ollama pull granite4.2:3b
ollama pull gemma4:e2b-it-qat
Qwen 3.5 4B
The best all-rounder at this size. Very strong in French, good at coding, 256k context, and follows instructions well. The slowest of the three (4B has its price).
Granite 4.2 3B
Very lean and token-efficient, from IBM. Good instruction following, Apache 2.0 license. The right balance of speed and quality.
Gemma 4 E2B
The fastest of the three. Multimodal, with astonishing quality for its size. Ideal if you want near-real-time performance on a modest CPU. Worse at coding than the other two.
Run a conversational benchmark
ollama run gemma4:e2b-it-qat --verbose
>>> Explique en 3 phrases la différence entre une LLC et une SAS.

The --verbose option displays statistics at the bottom of each response: prompt eval rate, eval rate (tokens/sec during generation), total duration. This is your reference metric on this machine.

→
Compare scientifically
Ask all three models exactly the same question, in the same order, cold (first launch). Compare response quality, displayed eval rate, and total time. The “best” one depends on your use case, not on an absolute ranking.

#3. Tokens/sec benchmarks: ballpark figures

Here are rough figures for representative machines, without a GPU, with Ollama (which uses llama.cpp under the hood) in Q4_K_M quantization. Your results will vary by ±20% depending on context, memory, and DDR frequency.

#Intel Core i5-12400 + DDR4-3200 16 GB

Gemma 4 E2B Q4
≈ 26 tokens/sec during generation
Granite 4.2 3B Q4_K_M
≈ 20 tokens/sec
Qwen 3.5 4B Q4_K_M
≈ 15 tokens/sec
Granite 4.2 8B Q4_K_M
≈ 8 tokens/sec
Qwen 3.5 9B Q4_K_M
≈ 6 tokens/sec

#Intel Core i7-13700K + DDR5-5600 32 GB

Gemma 4 E2B Q4
≈ 40 tokens/sec
Granite 4.2 3B Q4_K_M
≈ 30 tokens/sec
Qwen 3.5 4B Q4_K_M
≈ 23 tokens/sec
Qwen 3.5 9B Q4_K_M
≈ 12 tokens/sec
Mistral Small 24B Q4_K_M
≈ 4 tokens/sec

#AMD Ryzen 7 7700X + 32 GB DDR5-6000

Gemma 4 E2B Q4
≈ 44 tokens/sec
Granite 4.2 3B Q4_K_M
≈ 32 tokens/sec
Qwen 3.5 4B Q4_K_M
≈ 24 tokens/sec
Granite 4.2 8B Q4_K_M
≈ 13 tokens/sec
Qwen 3.5 9B Q4_K_M
≈ 11 tokens/sec
Mistral Small 24B Q4_K_M
≈ 5 tokens/sec
i
How to read these figures
Jump 50% between DDR4-3200 and DDR5-6000 on the same model. On CPUs, your RAM matters more than your processor. Fast DDR5 is often worth it before a CPU upgrade.

#4. Compile llama.cpp with AVX-512 (advanced)

Ollama includes generic precompiled llama.cpp binaries. By compiling llama.cpp yourself with your CPU's instruction sets (AVX2, AVX-512, AMX), you can gain 10 to 30% tokens/sec on some processors. For Intel Core 11th gen+ (Ice Lake / Rocket Lake / Sapphire Rapids) processors that support AVX-512.

Check AVX-512 support (Linux)
grep -o 'avx512[a-z_]*' /proc/cpuinfo | sort -u

If the command returns lines (avx512f, avx512dq, etc.), your CPU supports AVX-512. Otherwise, stick with standard Ollama; you have nothing to gain.

Clone and compile llama.cpp with AVX-512
git clone https://github.com/ggerganov/llama.cpp
cd llama.cpp
cmake -B build \
  -DGGML_NATIVE=ON \
  -DGGML_AVX512=ON \
  -DGGML_AVX512_VBMI=ON \
  -DGGML_AVX512_VNNI=ON
cmake --build build --config Release -j

The -DGGML_NATIVE=ON option lets the compiler automatically detect your CPU's instruction sets and enables everything available. It is the simplest and most reliable method.

Test with a GGUF model
./build/bin/llama-bench -m qwen3.5-9b-Q4_K_M.gguf -t 8

The -t option sets the number of threads (use the number of physical cores, not logical cores). llama-bench returns a table with pp512 (prompt eval) and tg128 (generation) in tokens/sec—your new baseline for comparison.

!
AVX-512 on mainstream Intel, beware
Intel Core 12th- and 13th-generation processors (Alder Lake, Raptor Lake) have AVX-512 disabled by default in the BIOS since a 2022 microcode update. On these CPUs, AVX-512 compilation won’t provide any benefit—stick with AVX2.

#Tips for getting more tokens/sec on CPU

  1. 01
    Set the number of threads
    By default, Ollama uses all logical cores. On some CPUs with hyper-threading, limiting it to physical cores (OLLAMA_NUM_THREADS=8 for an 8-core CPU) increases speed by 5 to 15%.
  2. 02
    Keep the model loaded
    Loading the model takes several seconds. OLLAMA_KEEP_ALIVE=30m keeps the model in RAM for 30 minutes after the last request. Without a GPU, this is even more valuable because reloading is slow.
  3. 03
    Reduce the context if possible
    num_ctx 2048 instead of 8192 saves RAM and noticeably improves speed. Keep a large context only for use cases that truly need it (RAG, long documents).
  4. 04
    Close Chrome and Slack
    A 7B LLM on the CPU saturates memory bandwidth. Anything else that also uses RAM (a browser with 50 tabs, Slack, Teams) takes cycles away from it. On 16 GB, that can make the difference between 5 and 8 tokens/sec.
  5. 05
    Choose the fastest compatible DDR
    If you upgrade: DDR4-3200 → DDR4-3600 = +10%. DDR4 → DDR5-5600 = +30 to 50%. The CPU matters much less than memory for LLM inference.

#When the CPU is no longer enough

Let's be honest: without a GPU, some use cases remain out of reach. If you recognize your situation in the list below, it is time to consider even a modest GPU (RTX 3060 12 GB used for €250 can be life-changing) or rent cloud capacity by the hour.

Interactive chat with a 13B+
2 to 5 tokens/sec is too slow for conversational AI. A 12 GB GPU fixes that instantly.
RAG with a large context (16k+)
Prompt evaluation time explodes on CPU. A RTX 3060 processes an 8k prompt in 1 sec, while an i7 takes 30 sec.
Real-time code completion
Inline autocompletion extensions (Tabby and similar tools, in FIM mode with Qwen 2.5 Coder 7B as the base) need responses in under 200 ms. On CPU, you will not get below 1 to 2 sec. A GPU is mandatory.
High-volume generation
Processing 1000 documents takes days on a CPU and hours on a GPU. For a one-off batch, RunPod or Vast.ai at €0.30/h gets the job done overnight.

#Go further

You have a CPU-only local LLM that responds. A few natural next steps, depending on your next question:

Choose the right quantization
Q4_K_M is the sensible default, but Q5_K_M or Q3 have their place depending on your RAM. The quantization guide details the trade-offs.
Wrap it all in an interface
The terminal is fine for testing. Open WebUI or LM Studio give you a local ChatGPT-like interface in minutes.
When to add a GPU
If you take the plunge, the GPU selection guide compares RTX 3060 vs 4060 vs 4070, along with the associated LLM benchmarks.

Recommended hardware: Radeon RX 9070 XT 16GB (ASUS Prime OC) — to move from CPU-only to a dedicated GPU. All AI hardware →

Did this guide help you?

Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.