AutoGen locally: what works and what casse
AutoGen works locally through an OpenAI-compatible endpoint (Ollama, vLLM), but Microsoft placed the project in maintenance mode on October 2, 2025: no more new features, and Microsoft recommends migrating to Microsoft Agent Framework. AG2, the Apache-2.0 fork from the original creators, is now the actively developed version to consider for a new project.
AutoGen builds systems from agents that talk to one another. Connected to a local model, it works—up to a point that depends almost entirely on the model size and how much rope you give the conversation. Locally, each turn is a generation on your own graphics card: the discipline is not in the prompt but in the limits. Another thing has changed since late 2025, and few tutorials mention it yet: the name “AutoGen” no longer refers to a single active project, but to three distinct paths that you need to distinguish before deciding where to invest your time.
#The conversation model
Where a chain framework asks you to define steps, AutoGen asks you to define participants. An assistant agent proposes; a proxy agent represents you, may execute code, and returns the result; a group discussion brings in several specialists under a manager who decides who speaks. Behavior emerges from these exchanges.
It is genuinely powerful for open-ended tasks—debugging, iterating on an analysis—and that is exactly what makes the cost unpredictable. A fixed pipeline makes N calls to the model. A conversation makes as many as needed, and “as many as needed” is decided by the model you are rightly questioning.
This design has a direct consequence for role selection: the more distinct agents there are in the conversation, the more turns are possible before a final answer emerges, and the greater the risk of drift. Two agents responding to each other in a loop without ever converging on a conclusion are the most common symptom locally, precisely because no limit was set before the first attempt. The number of agents is therefore not just an architectural choice; it directly multiplies the compute time consumed on your own graphics card.
#AutoGen in maintenance mode: what changed
Agents that act on your machine: agentic Cline, MCP, n8n + Ollama, local automations.
- Lifetime online access
- PDF + files
- Lifetime updates
This is the point most tutorials still online leave out. The official microsoft/autogen repository has stated it plainly since October 2, 2025: the project is in maintenance mode, will receive no new features, and will be community-managed. Microsoft recommends that new projects start directly with Microsoft Agent Framework (MAF), its successor, which merges AutoGen’s components with those of Semantic Kernel in a single development kit that reached general availability in early April 2026.
AutoGen’s original creators, Chi Wang and Qingyun Wu, left Microsoft at the end of 2024 and continued development under the name AG2, an Apache 2.0-licensed fork hosted by the ag2ai organization. AG2 presents itself as the project’s active continuation rather than a break: the historical codebase remains available as AG2 Classic for anyone who doesn’t want to change their code, while AG2 1.0 evolves the architecture toward a more explicit protocol between agents. For a project starting today, AG2 is the most coherent choice if you want to stay close to the original AutoGen API; Microsoft Agent Framework is the one to consider if you’re already in the Microsoft ecosystem and its integration with Semantic Kernel tools directly concerns you.
#Connect it to a local model
- 01Serve the modelOn an OpenAI-compatible endpoint. For multiple agents speaking at the same time, a server designed for concurrency performs much better than one designed for a single workstation.
- 02Configure the clientBase address (for example, http://localhost:11434/v1 for Ollama), the model name exactly as it was retrieved locally, and a dummy key. Local servers ignore the key; the client library often requires it to exist.
- 03State capabilities honestlyFrameworks ask whether the model supports function calling or JSON mode. Claiming an unsupported capability produces incomprehensible errors later instead of a clean failure.
- 04Increase the context windowThe history grows with every turn. A default context silently truncates the beginning, so agents redo work that has already been completed.
#Code execution: contain it
The pattern that makes AutoGen appealing—an agent writes code, an executor runs it, the error comes back, and the agent fixes it—is also what runs model-written code on your machine. Run it in a container, without credentials, without network access unless the task requires it, and never against your personal directory. Any content retrieved from outside—a ticket, a web page, or a documentation file—must be treated as potentially hostile.
Observability matters as much as containment. In a conversation with three or four agents, you diagnose an incident by rereading the exact prompt received by each agent on every turn, not a summary afterward. Without a complete trace, you can never know whether an agent misinterpreted an instruction or simply received truncated context because of insufficient memory. Logging every exchange, including tool calls and their exact arguments, costs little and avoids guessing what happened once the graphics card has gone quiet again.
#The model size, again
| Model class | Behavior |
|---|---|
| 3 to 8 billion | Fluid messages, unreliable tool calls, no stop discipline. Loop. |
| 12 to 14 billion | Two-agent exchanges with clearly defined roles work; group discussions wander. |
| 24 to 32 billion | The practical baseline for group discussions and code-execution loops. |
| 70 billion and more | The behavior closest to the cloud, slow enough for long conversations to become batch processing. |
Prefer models trained for tool calling and structured output. This is the same ceiling faced by all local agent frameworks: prose is easy, strict formats are not. In a group discussion, the manager role generally deserves the most capable model you have: it decides who speaks next, and a poor routing decision costs an entire turn for every mistake, repeated as many times as the conversation keeps going without converging.
- CrewAI: ordered roles and deliverables
- Architecture of a local agent
- Create a local agent in Python with LangChain
- The QuelLLM toolkit for setting up local agents
- Official repository: maintenance mode announcement
- AG2: the actively developed fork
- Ollama's official OpenAI-compatible API
#AutoGen, AG2, or CrewAI
| You want | Take |
|---|---|
| Defined roles, orderly handoffs, predictable cost | CrewAI |
| An open conversation, iteration, and fix-and-retry loops | AG2 (the active fork of what was AutoGen) |
| An explicit graph with a state you control | A graph-oriented framework |
| A single agent that modifies a repository | A dedicated coding agent |
| A project already built on the Microsoft ecosystem | Microsoft Agent Framework |
For code already written with AutoGen’s legacy API, migrating to AG2 is generally minor: the original autogen.* namespace remains available in a dedicated branch (AG2 Classic) while you transition at your own pace. Starting from scratch today is a different decision: you might as well choose AG2 directly or evaluate Microsoft Agent Framework instead of learning a codebase that its own publisher describes as community-maintained.
#What a conversation really costs
On a paid API, a three-agent conversation that stretches over twenty turns shows up on the bill. Locally, it shows up in the case fan and on the wall clock. Each turn is a full generation: if a turn takes an average of five seconds on your card, twenty turns without a stop condition represent at least one minute and forty seconds of continuous computation, often much more once the context has grown and each generation rereads a longer history than the previous one. Nothing alerts you while this happens—no counter, no threshold, just a GPU that stays hot.
That is why the turn limit is not just another configuration detail: it is the only mechanism that turns a potentially unlimited cost into a bounded, predictable cost, exactly as a token budget would on a commercial API. Setting this limit before the first test, rather than discovering it afterward, is the difference between a five-minute test and a graphics card monopolized all evening by a conversation that had gone nowhere since the tenth message. On a shared machine with other workloads—a workstation, not a dedicated server—this monopolization also has a real opportunity cost: you cannot launch anything else GPU-intensive while an unlimited multi-agent conversation runs in the background.
#FAQ
Is AutoGen still being developed by Microsoft?+
What is AG2, and should you use it instead?+
Does AutoGen work with Ollama?+
Why do my agents talk endlessly?+
Is it safe to let an agent execute code?+
What is the smallest usable model?+
AutoGen (AG2) or CrewAI?+
Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.