🇨🇳 Qwen 3.6 27B
Dense multimodal 27B released April 22, 2026. 262k ctx (1M YaRN). SWE-bench Verified 77.2%.
ollama run qwen3.6:27b
Ranking updated on 09/10/2026
Ranking of the most reliable LLMs for building autonomous local agents: function-calling accuracy, multi-turn robustness, JSON schema comprehension, and enough context to maintain an action plan.
Dense multimodal 27B released April 22, 2026. 262k ctx (1M YaRN). SWE-bench Verified 77.2%.
ollama run qwen3.6:27b
Qwen 3.8 27B: dense multimodal (text + vision), 262k context, ~16 GB Q4 VRAM (18 GB of Ollama weights). Apache 2.0, agentic coding and vision.
ollama run qwen3.8:27b
GLM-4.7-Flash (MoE 31B, ~3B active): the best code/VRAM ratio in the 30B class. MIT, 128k ctx, very fast on 3090/4090.
ollama run glm-4.7-flash
MoE with 30B/3B active: thinking mode + instruct. Gold medalist at IMO 2025 and IOI 2025. Fast inference thanks to the 3B active parameters, with 30B-level reasoning capabilities. Released April 2026.
ollama run nemotron-cascade-2
Cohere 30.5B model focused on agentic coding and reasoning. 488k context, Apache 2.0, ~18 GB VRAM in Q4.
ollama run north-mini-code-1.0
MoE with 35B/3B active parameters for agentic coding. 73.4% SWE-Bench. Release: April 16, 2026.
ollama run qwen3.6:35b-a3b
MoE 30B/3B active hybrid thinking. MMLU 81.4, AIME24 80.4. 100+ languages.
ollama run qwen3:30b-a3b
| Rank | Model | Params | Q4 VRAM | Context | License |
|---|---|---|---|---|---|
| #1 | Qwen 3.6 27B | 27B | 16 GB | 262 144 | Apache 2.0 |
| #2 | Qwen 3.8 27B | 27B | 16 GB | 262 144 | Apache 2.0 |
| #3 | GLM 4.7 Flash | 31B | 19 GB | 128 000 | MIT |
| #4 | Nemotron Cascade 2 30B-A3B | 30B | 17 GB | 128 000 | NVIDIA Open Model License |
| #5 | North Mini Code 1.0 | 30.5B | 18 GB | 488 000 | Apache 2.0 |
| #6 | Qwen 3.6 35B-A3B | 35B | 21 GB | 262 000 | Apache 2.0 |
| #7 | Qwen 3 30B-A3B | 30B | 19 GB | 131 072 | Apache 2.0 |
Agents that act on your machine: agentic Cline, MCP, n8n + Ollama, local automations.
We favor models ≥ 7B with context ≥ 32k (to maintain an action history) and a chat or reasoning tag. The score favors recent models and medium-to-large sizes (where function calling becomes reliable).
Criteria considered:
The scoring is fully transparent: see our methodology for details on VRAM/tokens/sec calculations.
Which LLM for a local autonomous agent?
Qwen 3.6 27B is our #1. Mistral Small 3.1 is particularly well known for its reliability in tool use and low latency—critical for an agent that makes 20+ calls per task.
Does Ollama support function calling?
Yes since 0.3.0 via the parameter tools. LM Studio and vLLM as well. Make sure the model was trained for (the recent Mistral, Qwen, Llama models were).
Which framework should you use to build a local agent?
LangGraph, CrewAI, AutoGen, Smolagents — all of them can point to a local Ollama through the OpenAI-compatible endpoint (http://localhost:11434/v1).
What’s the difference between reasoning and agents?
A reasoning model (DeepSeek R1, QwQ) works through a chain of thought before responding. An agent model (Mistral Small, Qwen) chooses and calls external tools. The two can be combined: reasoning for planning, agent for execution.