Advanced 11 minAgents

AutoGen locally: what works and what casse

Direct response

AutoGen works locally through an OpenAI-compatible endpoint (Ollama, vLLM), but Microsoft placed the project in maintenance mode on October 2, 2025: no more new features, and Microsoft recommends migrating to Microsoft Agent Framework. AG2, the Apache-2.0 fork from the original creators, is now the actively developed version to consider for a new project.

AutoGen builds systems from agents that talk to one another. Connected to a local model, it works—up to a point that depends almost entirely on the model size and how much rope you give the conversation. Locally, each turn is a generation on your own graphics card: the discipline is not in the prompt but in the limits. Another thing has changed since late 2025, and few tutorials mention it yet: the name “AutoGen” no longer refers to a single active project, but to three distinct paths that you need to distinguish before deciding where to invest your time.

By Mohamed Meguedmi·Update 2026-09-28·Tested on Windows, macOS, and Linux

#The conversation model

Where a chain framework asks you to define steps, AutoGen asks you to define participants. An assistant agent proposes; a proxy agent represents you, may execute code, and returns the result; a group discussion brings in several specialists under a manager who decides who speaks. Behavior emerges from these exchanges.

It is genuinely powerful for open-ended tasks—debugging, iterating on an analysis—and that is exactly what makes the cost unpredictable. A fixed pipeline makes N calls to the model. A conversation makes as many as needed, and “as many as needed” is decided by the model you are rightly questioning.

This design has a direct consequence for role selection: the more distinct agents there are in the conversation, the more turns are possible before a final answer emerges, and the greater the risk of drift. Two agents responding to each other in a loop without ever converging on a conclusion are the most common symptom locally, precisely because no limit was set before the first attempt. The number of agents is therefore not just an architectural choice; it directly multiplies the compute time consumed on your own graphics card.

#AutoGen in maintenance mode: what changed

The Local Agents Kit

Agents that act on your machine: agentic Cline, MCP, n8n + Ollama, local automations.

  • Lifetime online access
  • PDF + files
  • Lifetime updates

This is the point most tutorials still online leave out. The official microsoft/autogen repository has stated it plainly since October 2, 2025: the project is in maintenance mode, will receive no new features, and will be community-managed. Microsoft recommends that new projects start directly with Microsoft Agent Framework (MAF), its successor, which merges AutoGen’s components with those of Semantic Kernel in a single development kit that reached general availability in early April 2026.

AutoGen’s original creators, Chi Wang and Qingyun Wu, left Microsoft at the end of 2024 and continued development under the name AG2, an Apache 2.0-licensed fork hosted by the ag2ai organization. AG2 presents itself as the project’s active continuation rather than a break: the historical codebase remains available as AG2 Classic for anyone who doesn’t want to change their code, while AG2 1.0 evolves the architecture toward a more explicit protocol between agents. For a project starting today, AG2 is the most coherent choice if you want to stay close to the original AutoGen API; Microsoft Agent Framework is the one to consider if you’re already in the Microsoft ecosystem and its integration with Semantic Kernel tools directly concerns you.

!
What this concretely changes for you
The legacy AutoGen code continues to work: it receives security fixes, not new features. Nothing that follows in this guide (configuration Ollama, code-execution isolation, model sizes) changes with AG2, whose API remains very similar. But starting today with “AutoGen” without knowing that the name now covers three distinct projects—the frozen Microsoft repository, the active AG2 fork, and Microsoft Agent Framework—is the kind of tool-selection mistake that becomes costly six months later, when it's time to evolve the code.

#Connect it to a local model

  1. 01
    Serve the model
    On an OpenAI-compatible endpoint. For multiple agents speaking at the same time, a server designed for concurrency performs much better than one designed for a single workstation.
  2. 02
    Configure the client
    Base address (for example, http://localhost:11434/v1 for Ollama), the model name exactly as it was retrieved locally, and a dummy key. Local servers ignore the key; the client library often requires it to exist.
  3. 03
    State capabilities honestly
    Frameworks ask whether the model supports function calling or JSON mode. Claiming an unsupported capability produces incomprehensible errors later instead of a clean failure.
  4. 04
    Increase the context window
    The history grows with every turn. A default context silently truncates the beginning, so agents redo work that has already been completed.
!
Set a maximum number of turns before the first attempt
Not afterward. The typical opening experience of a local multi-agent conversation is discovering, twenty minutes later, that two agents have been congratulating each other since the third message while the graphics card runs at full capacity.

#Code execution: contain it

The pattern that makes AutoGen appealing—an agent writes code, an executor runs it, the error comes back, and the agent fixes it—is also what runs model-written code on your machine. Run it in a container, without credentials, without network access unless the task requires it, and never against your personal directory. Any content retrieved from outside—a ticket, a web page, or a documentation file—must be treated as potentially hostile.

Observability matters as much as containment. In a conversation with three or four agents, you diagnose an incident by rereading the exact prompt received by each agent on every turn, not a summary afterward. Without a complete trace, you can never know whether an agent misinterpreted an instruction or simply received truncated context because of insufficient memory. Logging every exchange, including tool calls and their exact arguments, costs little and avoids guessing what happened once the graphics card has gone quiet again.

#The model size, again

What really happens in a conversation
Model classBehavior
3 to 8 billionFluid messages, unreliable tool calls, no stop discipline. Loop.
12 to 14 billionTwo-agent exchanges with clearly defined roles work; group discussions wander.
24 to 32 billionThe practical baseline for group discussions and code-execution loops.
70 billion and moreThe behavior closest to the cloud, slow enough for long conversations to become batch processing.

Prefer models trained for tool calling and structured output. This is the same ceiling faced by all local agent frameworks: prose is easy, strict formats are not. In a group discussion, the manager role generally deserves the most capable model you have: it decides who speaks next, and a poor routing decision costs an entire turn for every mistake, repeated as many times as the conversation keeps going without converging.

#AutoGen, AG2, or CrewAI

Two different mental models
You wantTake
Defined roles, orderly handoffs, predictable costCrewAI
An open conversation, iteration, and fix-and-retry loopsAG2 (the active fork of what was AutoGen)
An explicit graph with a state you controlA graph-oriented framework
A single agent that modifies a repositoryA dedicated coding agent
A project already built on the Microsoft ecosystemMicrosoft Agent Framework

For code already written with AutoGen’s legacy API, migrating to AG2 is generally minor: the original autogen.* namespace remains available in a dedicated branch (AG2 Classic) while you transition at your own pace. Starting from scratch today is a different decision: you might as well choose AG2 directly or evaluate Microsoft Agent Framework instead of learning a codebase that its own publisher describes as community-maintained.

#What a conversation really costs

On a paid API, a three-agent conversation that stretches over twenty turns shows up on the bill. Locally, it shows up in the case fan and on the wall clock. Each turn is a full generation: if a turn takes an average of five seconds on your card, twenty turns without a stop condition represent at least one minute and forty seconds of continuous computation, often much more once the context has grown and each generation rereads a longer history than the previous one. Nothing alerts you while this happens—no counter, no threshold, just a GPU that stays hot.

That is why the turn limit is not just another configuration detail: it is the only mechanism that turns a potentially unlimited cost into a bounded, predictable cost, exactly as a token budget would on a commercial API. Setting this limit before the first test, rather than discovering it afterward, is the difference between a five-minute test and a graphics card monopolized all evening by a conversation that had gone nowhere since the tenth message. On a shared machine with other workloads—a workstation, not a dedicated server—this monopolization also has a real opportunity cost: you cannot launch anything else GPU-intensive while an unlimited multi-agent conversation runs in the background.

#FAQ

Is AutoGen still being developed by Microsoft?+
No, not actively anymore. Microsoft placed the microsoft/autogen repository in maintenance mode on October 2, 2025: security fixes only, no new features, and community management for everything else. Microsoft explicitly recommends that new projects use Microsoft Agent Framework, its official successor, which reached general availability in early April 2026, with a dedicated migration guide.
What is AG2, and should you use it instead?+
AG2 is the Apache 2.0-licensed fork created by Chi Wang and Qingyun Wu, the original authors of AutoGen, after they left Microsoft in late 2024. It is now the actively developed version closest to the historical AutoGen API: a reasonable choice for a new project that wants to stay in that spirit without depending on the Microsoft ecosystem.
Does AutoGen work with Ollama?+
Yes, through Ollama's OpenAI-compatible endpoint, exposed at http://localhost:11434/v1: base URL, the model name exactly as retrieved locally, and a dummy key because Ollama does not verify it. For multiple simultaneous agents, a server designed for concurrency performs significantly better than Ollama alone on a typical workstation.
Why do my agents talk endlessly?+
No effective stopping condition was set, and often the model is too small to recognize one even when it exists in the prompt. Set a maximum number of turns before the first attempt, define an explicit, tested stopping phrase, and give a larger model the manager role that decides who speaks next.
Is it safe to let an agent execute code?+
Only inside a container, without credentials and with network access restricted to what is strictly necessary. The executor runs code written by a model that may have read third-party-controlled text during the conversation, making it a classic entry point for indirect prompt injection on a local installation.
What is the smallest usable model?+
Around 12 to 14 billion parameters for simple exchanges between two agents with clearly defined roles, and 24 to 32 billion for group discussions or code-execution loops. Below this threshold, messages remain apparently smooth, but stop discipline and tool calls become unreliable.
AutoGen (AG2) or CrewAI?+
CrewAI is better suited to predefined roles and predictable handoffs from one agent to another; AG2 is better for open-ended conversation and iterative problem-solving where the path isn’t known at the outset. Both remain equally sensitive to model quality once run locally.
Did this guide help you?

Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.