BestLLMfor Your hardware. Your LLM. Your call.
◆ The kits◆ Kits APIOpen data Find my LLM
Guide · 2026-09-26

AutoGen on a Local Model

◆ Local Agents — Agents that act on your machine, without the cloud · $24 · or all kits $49 →

AutoGen builds systems out of agents that talk to each other. Pointed at a local model it works — up to a point that depends almost entirely on model size and on how much rope you give the conversation.

By Mohamed Meguedmi·Last updated 2026-09-26·10 min read·Tested on Windows, macOS, Linux

Key takeaways

  • AutoGen is a framework for multi-agent conversations: agents exchange messages, and the work emerges from the exchange rather than from a fixed pipeline.
  • It reaches local models through OpenAI-compatible endpoints, so Ollama, vLLM or llama.cpp all work with a configuration change.
  • Conversations need hard termination conditions. Two polite agents will thank each other indefinitely, and every turn costs a full generation on your GPU.
  • Code execution is the feature and the risk. An agent that writes and runs code needs a container, not your shell.
  • Below roughly 14B parameters the pattern degrades; at 27B–32B it becomes genuinely useful. Small models produce fluent conversation and unreliable structure.

The conversation model

The Local Agents Kit

Agents that act on your machine: agentic Cline, MCP, n8n + Ollama, local automations.

  • Lifetime online access
  • PDF + files
  • 30-day refund

Where a pipeline framework asks you to define steps, AutoGen asks you to define participants. An assistant agent proposes; a user-proxy agent represents you, optionally executing code and returning results; a group chat lets several specialists speak under a manager that decides who goes next. The behaviour emerges from those exchanges.

That is genuinely powerful for open-ended tasks — debugging, iterating on an analysis — and it is exactly what makes cost unpredictable. A fixed pipeline runs N model calls. A conversation runs as many as it takes, and "as it takes" is decided by the same model you are unsure about.

Pointing it at a local model

  1. Serve the model with an OpenAI-compatible endpoint. Ollama exposes one; for several agents talking at once, vLLM handles concurrency far better, which is the whole point of that comparison.
  2. Configure the client with the base URL, the model name and a placeholder API key. Local servers ignore the key; the client library often insists one exists.
  3. Declare the model's capabilities honestly. Frameworks ask whether the model supports function calling, vision or JSON mode. Claiming a capability the model lacks produces confusing downstream errors rather than a clean failure.
  4. Raise the context window. Conversation history grows every turn. A default local context silently truncates the beginning, and agents then repeat work they have already done.

Set a maximum number of turns before your first run. Not after. The characteristic first experience with a local multi-agent conversation is discovering, twenty minutes later, that two agents have been agreeing with each other since the third message while your GPU sat at full load.

Code execution: contain it

The pattern that makes AutoGen compelling — an agent writes code, a proxy runs it, the error comes back, the agent fixes it — is also the one that runs model-authored code on your machine. Run it in a container with no credentials and no network unless it needs one, never against your home directory, and treat any instruction that arrives from a fetched document as hostile. Why that last point matters is in prompt injection.

Model size, again

Model classWhat happens in a conversation
3B–8BFluent messages, unreliable tool calls, poor termination discipline. Loops.
12B–14BTwo-agent exchanges with clear roles work; group chats wander.
24B–32BThe practical floor for group chats and code execution loops.
70B+Closest to cloud behaviour, slow enough that long conversations become batch jobs.

Prefer models trained for tool use and structured output; the agents and tool-use ranking is a reasonable shortlist, and the VRAM calculator tells you which of them fits.

AutoGen or CrewAI

You wantReach for
Defined roles, ordered hand-offs, predictable costCrewAI
Open-ended conversation, iteration, code-and-fix loopsAutoGen
Explicit graph with state you controlA graph framework — see LangChain vs LangGraph
One agent editing a repositoryA dedicated coding agent

Verdict

AutoGen suits problems where the path is not known in advance and iteration is the point. Locally, that flexibility meets a fixed budget — your GPU — so the discipline is all in the limits: hard turn caps, explicit termination phrases, a contained executor, and a model large enough to follow structure. Give it a 27B-class model and firm boundaries and it earns its place; give it an 8B model and an open-ended conversation and you will be watching a loop.

Frequently asked questions

Does AutoGen work with Ollama?

Yes, through Ollama's OpenAI-compatible endpoint: base URL, model name and a placeholder key. For several agents running concurrently, a server built for concurrency behaves much better.

Why do my agents keep talking forever?

No effective termination condition, and often a model too small to recognise one. Set a maximum number of turns, define an explicit termination string, and give the manager role a larger model.

Is it safe to let an agent run code?

Only inside a container with no credentials and restricted network access. The executor runs code written by a model that may have read attacker-controlled text.

What is the smallest usable model?

Around 12B–14B for simple two-agent exchanges, and 24B–32B for group chats or code loops. Below that, conversations look plausible and fail on structure.

AutoGen or CrewAI?

CrewAI for defined roles and predictable hand-offs; AutoGen for open-ended conversation and iterative problem solving. Both are equally sensitive to model quality when run locally.

Does it need a GPU?

The framework does not; the model does. Multi-agent work multiplies generations, so weak hardware turns a two-minute task into an hour.

Recommended hardware

A current option for local AI: GMKtec EVO-X2 64GB / 1TB (Ryzen AI Max+ 395). Match memory to your model and software. A mini PC is a complete PC alternative; Mac/MLX and CUDA instructions require compatible hardware.

Amazon Check GMKtec EVO-X2 64GB / 1TB (Ryzen AI Max+ 395) price →

As an Amazon Associate, BestLLMfor earns from qualifying purchases, at no extra cost to you. It does not influence our independent rankings.

Did this guide help?

Found an error or have feedback? Let us know — it helps everyone who reads this guide.