Home › Catalog › Best local LLM for agents / tool use in 2026

Best local LLM for agents / tool use in 2026

◆ Local Agents — Agents that act on your machine, without the cloud · $24 · or all kits $49 →

Ranking updated on 09/10/2026

Ranking of the most reliable LLMs for building autonomous local agents: function-calling accuracy, multi-turn robustness, JSON schema comprehension, and enough context to maintain an action plan.

Ranking

1

🇨🇳 Qwen 3.6 27B

Alibaba · 27B parameters · Apache 2.0 · 262,144-token context

Dense multimodal 27B released April 22, 2026. 262k ctx (1M YaRN). SWE-bench Verified 77.2%.

Why this ranking Reliable function calling, 262,144 context tokens to preserve the action history. Reasoning capability as a bonus.
ollama run qwen3.6:27b
Q4 VRAM
16 GB
29 GB in Q8
2

🇨🇳 Qwen 3.8 27B

Alibaba · 27B parameters · Apache 2.0 · 262,144-token context

Qwen 3.8 27B: dense multimodal (text + vision), 262k context, ~16 GB Q4 VRAM (18 GB of Ollama weights). Apache 2.0, agentic coding and vision.

Why this ranking Reliable function calling, 262,144 context tokens to preserve the action history. Reasoning capability as a bonus.
ollama run qwen3.8:27b
Q4 VRAM
16 GB
29 GB in Q8
3

🇨🇳 GLM 4.7 Flash

Zhipu AI · 31B parameters · MIT · 128,000 tokens ctx

GLM-4.7-Flash (MoE 31B, ~3B active): the best code/VRAM ratio in the 30B class. MIT, 128k ctx, very fast on 3090/4090.

Why this ranking Reliable function calling, 128,000-token context to retain the action history. Reasoning capability as a bonus.
ollama run glm-4.7-flash
Q4 VRAM
19 GB
35 GB in Q8
4

🇺🇸 Nemotron Cascade 2 30B-A3B

NVIDIA · 30B parameters · NVIDIA Open Model License · 128,000 tokens ctx

MoE with 30B/3B active: thinking mode + instruct. Gold medalist at IMO 2025 and IOI 2025. Fast inference thanks to the 3B active parameters, with 30B-level reasoning capabilities. Released April 2026.

Why this ranking Reliable function calling, 128,000-token context to retain the action history. Reasoning capability as a bonus.
ollama run nemotron-cascade-2
Q4 VRAM
17 GB
32 GB in Q8
5

🇺🇸 North Mini Code 1.0

Cohere · 30.5B parameters · Apache 2.0 · 488,000 tokens ctx

Cohere 30.5B model focused on agentic coding and reasoning. 488k context, Apache 2.0, ~18 GB VRAM in Q4.

Why this ranking Reliable function calling, 488,000-token context to preserve the action history. Bonus reasoning capability.
ollama run north-mini-code-1.0
Q4 VRAM
18 GB
33 GB in Q8
6

🇨🇳 Qwen 3.6 35B-A3B

Alibaba · 35B parameters · Apache 2.0 · 262,000-token context

MoE with 35B/3B active parameters for agentic coding. 73.4% SWE-Bench. Release: April 16, 2026.

Why this ranking Reliable function calling, 262,000-token context to retain the action history. Reasoning capability as a bonus.
ollama run qwen3.6:35b-a3b
Q4 VRAM
21 GB
38 GB in Q8
7

🇨🇳 Qwen 3 30B-A3B

Alibaba · 30B parameters · Apache 2.0 · 131,072 tokens ctx

MoE 30B/3B active hybrid thinking. MMLU 81.4, AIME24 80.4. 100+ languages.

Why this ranking Reliable function calling, 131,072 tokens of context to retain the action history. Reasoning capability as a bonus.
ollama run qwen3:30b-a3b
Q4 VRAM
19 GB
35 GB in Q8

Comparison table

Rank Model Params Q4 VRAM Context License
#1 Qwen 3.6 27B 27B 16 GB 262 144 Apache 2.0
#2 Qwen 3.8 27B 27B 16 GB 262 144 Apache 2.0
#3 GLM 4.7 Flash 31B 19 GB 128 000 MIT
#4 Nemotron Cascade 2 30B-A3B 30B 17 GB 128 000 NVIDIA Open Model License
#5 North Mini Code 1.0 30.5B 18 GB 488 000 Apache 2.0
#6 Qwen 3.6 35B-A3B 35B 21 GB 262 000 Apache 2.0
#7 Qwen 3 30B-A3B 30B 19 GB 131 072 Apache 2.0
The Local Agents Kit

Agents that act on your machine: agentic Cline, MCP, n8n + Ollama, local automations.

  • Lifetime online access
  • PDF + files
  • Lifetime updates

Ranking methodology

We favor models ≥ 7B with context ≥ 32k (to maintain an action history) and a chat or reasoning tag. The score favors recent models and medium-to-large sizes (where function calling becomes reliable).

Criteria considered:

  • Precise function calling
  • Context ≥ 32k tokens
  • Multi-turn reliability
  • JSON format respected

The scoring is fully transparent: see our methodology for details on VRAM/tokens/sec calculations.

Frequently asked questions

Which LLM for a local autonomous agent?

Qwen 3.6 27B is our #1. Mistral Small 3.1 is particularly well known for its reliability in tool use and low latency—critical for an agent that makes 20+ calls per task.

Does Ollama support function calling?

Yes since 0.3.0 via the parameter tools. LM Studio and vLLM as well. Make sure the model was trained for (the recent Mistral, Qwen, Llama models were).

Which framework should you use to build a local agent?

LangGraph, CrewAI, AutoGen, Smolagents — all of them can point to a local Ollama through the OpenAI-compatible endpoint (http://localhost:11434/v1).

What’s the difference between reasoning and agents?

A reasoning model (DeepSeek R1, QwQ) works through a chain of thought before responding. An agent model (Mistral Small, Qwen) chooses and calls external tools. The two can be combined: reasoning for planning, agent for execution.

Go further

QuelLLM Kits The reference guide by use case
All kits for life — $49