BestLLMfor Your hardware. Your LLM. Your call.
The Local Copilot Kit APIOpen data Find my LLM
Editorial ranking · 2026

Best Ollama models

Last updated 2026-08-28 · Page updated 2026-08-31

Top 9 open-source picks for running locally with Ollama, ranked by benchmark performance and real-world fit. Updated monthly.

Verdict (August 2026): Qwen 3.5 27B is our current top pick for running locally with Ollama (27B params · Apache 2.0). The full ranking and per-model reasoning follow.

This ranking is built from real specs — parameter count, VRAM at 4-bit quantization, license terms, and available benchmark scores — pulled from BestLLMfor's tracked catalog of 239 models, not vendor marketing. Data reflects the catalog as of 2026-08-28; see our full methodology for how models are scored and re-ranked.

#1

Qwen 3.5 27B

27B · Alibaba · Apache 2.0

Alibaba's dense 27B Qwen 3.5 with a 262K context window and calibrated thinking mode. One of the best quality-to-size trade-offs in the open 25B-30B class.

VRAM Q4: 16 GB · Context: 255k
Read full fiche →
#2

Qwen 2.5 Coder 32B

32B · Alibaba · Apache 2.0

Alibaba's Qwen 2.5 Coder 32B — the strongest open-weight code model we've benchmarked, trading punches with Claude 3.5 Sonnet on HumanEval.

VRAM Q4: 19 GB · Context: 128k
Read full fiche →
#3

Gemma 4 31B

31B · Google · Gemma

Google's dense 31B multimodal model with native text, image, and audio support across 140+ languages. Ranked #3 on Chatbot Arena's open leaderboard with a 256K context window.

VRAM Q4: 18 GB · Context: 250k
Read full fiche →
#4

gpt-oss 20B

21B · OpenAI · Apache 2.0

OpenAI's compact open-weight MoE with 3.6B active out of 21B total parameters. Matches o3-mini on a laptop-class GPU under Apache 2.0.

VRAM Q4: 13 GB · Context: 125k
Read full fiche →
#5

Mistral Small 3.2 24B

24B · Mistral AI · Apache 2.0

Mistral AI's June 2025 refresh of Small 3.1: a 24B Apache 2.0 dense model with vision input, sharper function calling, and roughly half the rate of runaway generations seen in 3.1.

VRAM Q4: 14 GB · Context: 125k
Read full fiche →
#6

DeepSeek R1 Distill 32B

32B · DeepSeek · MIT

The 32B DeepSeek R1 distill — the best accessible open-weight reasoner we've tested. Explicit chain-of-thought, MIT-licensed, runs on a single 24GB GPU.

VRAM Q4: 19 GB · Context: 32k
Read full fiche →
#7

Llama 3.3 70B Instruct

70B · Meta · Llama 3.3 Community

Meta's Llama 3.3 70B — same quality tier as Llama 3.1 405B at one-sixth the size, thanks to improved post-training. Weights are gated on Hugging Face.

VRAM Q4: 40 GB · Context: 125k
Read full fiche →
#8

Qwen 3 8B

8B · Alibaba · Apache 2.0

Alibaba's 8B dense model with a toggleable thinking mode and broad multilingual coverage. Punches well above its weight for an 8B and runs comfortably on a single consumer GPU.

VRAM Q4: 5 GB · Context: 128k
Read full fiche →
#9

Gemma 4 E4B

4B · Google · Gemma

Google's 4B-effective multimodal Gemma variant tuned for laptops and edge devices, handling text, image, and audio across 140 languages with a 128K context.

VRAM Q4: 10 GB · Context: 125k
Read full fiche →

Which hardware should you buy to run Qwen 3.5 27B?

To run Qwen 3.5 27B locally at Q4, you need ~16 GB of VRAM. The best value for this today is a RTX 5070 Ti 16GB (GIGABYTE Gaming OC) (16 GB VRAM).

As an Amazon Associate, BestLLMfor earns from qualifying purchases, at no extra cost to you. It does not influence our independent rankings.

Frequently asked questions

What is the best local LLM for running locally with Ollama?

Qwen 3.5 27B tops this ranking — a 27B model, licensed under Apache 2.0, needing about 16 GB of VRAM at Q4 quantization. See the full list below for the runner-ups and how they compare.

How much VRAM do I need to run Qwen 3.5 27B?

At Q4 quantization, Qwen 3.5 27B needs about 16 GB of VRAM and fits comfortably on a single 24 GB GPU.

Which of these models fit an 8 GB GPU?

At Q4 quantization, Qwen 3 8B fit within 8 GB of VRAM.

Are the models on this running locally with Ollama list free for commercial use?

Licenses across this list include Apache 2.0, Gemma, Llama 3.3 Community, MIT. Check the specific license of each model on its catalog page before deploying commercially, as terms vary by author.

What context window do these models support?

Context windows on this list range from 32k to 255k tokens, depending on the model.