Advanced 14 minAgents

CrewAI + Ollama: orchestrating multiple AI agents in local

A single AI agent answers a question; a team of agents solves a problem. CrewAI orchestrates several specialized LLMs — a researcher, a writer, and a reviewer — that pass work between them like colleagues. This guide builds a CrewAI crew that runs entirely locally on Ollama: no data reaches a cloud API, and there is no per-token cost. You'll see how to connect CrewAI to the local endpoint, define roles and tasks that actually work, equip agents with tools, and, above all, which local models can handle the load of multiple agents without collapsing.

By Thomas P.·Update 2026-08-27·Tested on Windows, macOS, and Linux

#Why orchestrate a CrewAI crew locally

The multi-agent pattern starts with a simple observation: splitting a complex task among several specialized agents produces better results than a single giant prompt. Each agent has a clear role, a specific objective, and sees only its part of the work. CrewAI is the Python framework that formalizes this split—roles, tasks, sequential or hierarchical collaboration—without the overhead of wiring up a state graph yourself.

Running this crew on Ollama instead of GPT-4 or Claude changes three things. First, privacy: a multi-agent pipeline multiplies calls to the model, and therefore multiplies potential leaks to a third party; locally, nothing leaves the machine. Next, cost: a chatty crew can burn hundreds of thousands of tokens per run, which quickly becomes expensive on a token-priced API—in local operation, the marginal cost is zero. Finally, control: you choose the model, quantization, context, and can iterate without quotas or rate limits.

i
Who this guide is for
Advanced level: we assume you're comfortable with Python, that Ollama is already running (daemon on http://localhost:11434), and that you've launched at least one model. If you need the basics of orchestrating a single agent, the LangChain + Ollama guide lays them out before tackling multi-agent orchestration.

#The 4 building blocks of CrewAI

The Local Copilot Kit

This guide gets you to the model. The kit gets you to the coding copilot in your editor.

  • Lifetime online access
  • PDF + files
  • Lifetime updates

Before writing a line, you need the vocabulary. CrewAI is built around four objects that fit together. Understanding them prevents 90% of design errors in a crew.

Agent
An LLM with an identity: a role (“Market Analyst”), a goal, and a backstory that shape its behavior. Each agent can use its own Ollama model.
Task
A unit of work assigned to an agent: a description, an expected result (expected_output), and often, tools. The task—not the agent—carries the concrete instruction.
Tool
An external capability that an agent can call: web search, file reading, SQL queries, or computation. Without tools, an agent can only reason about what it already knows.
Crew
The crew: the list of agents, the list of tasks, and the process that determines the execution order (sequential or hierarchical). This is the object you launch with kickoff().

The process deserves one clarification. In sequential mode, tasks run in the declared order, and the output of one feeds the next—perfect for a research → writing → proofreading pipeline. In hierarchical mode, a “manager” agent (a dedicated LLM) delegates to and coordinates the others. Hierarchical mode is more powerful but much more demanding for a local model, because the manager must reason about who does what: always start with sequential mode.

#Prerequisites and model selection

Ollama working
Daemon running, endpoint at http://localhost:11434. Verify with “ollama list” that at least one tool-use-capable model is available.
Python 3.10+ and a venv
CrewAI deploys cleanly in an isolated virtual environment. Avoid installing it in the system Python.
VRAM—and patience
Multi-agent systems chain calls together: each task means one or more round trips to the model. Plan on at least an 8B model, ideally in the 12B to 24B range, so the reasoning holds up.
A tool-capable model
To equip agents with tools, you need a model that supports function calling: Qwen 3.5, Qwen 3.8, Mistral Small, or GLM 4.7 Flash. A model without tool use can only reason in text.
!
Multi-agent setups aren't magic locally
A crew multiplies the number of tokens and reasoning turns. A 3B model that answers an isolated question correctly can fall apart in a three-agent chain: it forgets the context, makes up outputs, and loops. Local multi-agent setups require larger models than simple chat. We’ll return to this in the models section.

#Install CrewAI and connect it to Ollama

CrewAI installs via pip. The crewai-tools package also provides a library of ready-to-use tools (search, files, scraping).

Terminal — installation
# Créer et activer un environnement virtuel
python -m venv .venv
source .venv/bin/activate   # Windows : .venv\Scripts\activate

# Installer CrewAI et sa bibliothèque d'outils
pip install crewai crewai-tools

The key point about the local connection: CrewAI relies on LiteLLM to talk to models. To target Ollama, prefix the model name with “ollama/” and point the base URL to the local daemon. Make sure you have pulled the model beforehand.

Terminal — downloading the models
# Un modèle solide et tool-capable pour les agents
ollama pull qwen3.5:9b

# Un modèle plus léger pour les tâches simples (optionnel)
ollama pull qwen3.5:4b
Python — configure the local LLM
from crewai import LLM

# Un LLM CrewAI branché sur Ollama en local
local_llm = LLM(
    model="ollama/qwen3.5:9b",
    base_url="http://localhost:11434",
    temperature=0.3,   # bas = plus déterministe, utile pour une crew
)
→
Lower agent temperatures
In multi-agent setups, a high temperature makes agents drift: they improvise instead of following their task. A value between 0.2 and 0.4 makes the crew more reliable and reproducible. Reserve high temperatures for purely creative tasks isolated to a single agent.

#First crew: define roles and tasks

Let’s assemble a classic, immediately useful crew: a researcher who gathers information, followed by a writer who formats it. This is the framework we reuse for monitoring, document summarization, or content generation. We first define the agents, with clear roles.

Python — defining agents
from crewai import Agent

chercheur = Agent(
    role="Analyste de recherche",
    goal="Rassembler des faits précis et vérifiables sur {sujet}",
    backstory=(
        "Tu es un analyste rigoureux. Tu distingues les faits "
        "des opinions et tu ne retiens que l'information solide."
    ),
    llm=local_llm,
    verbose=True,
)

redacteur = Agent(
    role="Rédacteur technique francophone",
    goal="Produire une synthèse claire et structurée en français",
    backstory=(
        "Tu écris pour un public technique. Tu vas droit au but, "
        "sans jargon inutile ni remplissage marketing."
    ),
    llm=local_llm,
    verbose=True,
)

The role, goal, and backstory aren't decorative: they form each agent's system prompt. A vague role (“Assistant”) produces a vague agent. Be specific and establish a stance—that's what keeps a small model from going off track. Note the {sujet} placeholder: CrewAI injects it at launch from the inputs.

Next comes the core of the work: the tasks. Each task points to an agent, describes what it must produce, and, most importantly, specifies an expected_output. This field is CrewAI’s most underestimated quality lever: the more concrete it is, the more constrained the output.

Python — define the tasks
from crewai import Task

tache_recherche = Task(
    description=(
        "Recherche les points clés sur {sujet}. Identifie 5 à 7 "
        "faits marquants, chiffrés quand c'est possible."
    ),
    expected_output="Une liste à puces de 5 à 7 faits sourcés et concis.",
    agent=chercheur,
)

tache_redaction = Task(
    description=(
        "À partir des faits rassemblés, rédige une synthèse "
        "structurée de 300 mots sur {sujet}, en français."
    ),
    expected_output="Un texte de ~300 mots avec un titre et 3 sections.",
    agent=redacteur,
    context=[tache_recherche],   # nourrie par la sortie de la recherche
)

The context parameter explicitly links the tasks: the writing task receives the research output. In sequential mode, the chain is already implicit, but declaring context makes the dependency clear and makes information transfer more reliable. Finally, we assemble the crew and launch it.

Python — assemble and launch
from crewai import Crew, Process

crew = Crew(
    agents=[chercheur, redacteur],
    tasks=[tache_recherche, tache_redaction],
    process=Process.sequential,
    verbose=True,
)

resultat = crew.kickoff(inputs={"sujet": "les LLM open-weight en 2026"})
print(resultat)
  1. 01
    The retriever runs
    Its local LLM receives its role plus the description of the research task and produces the expected list of facts.
  2. 02
    The output passes through
    CrewAI passes the search result to the writing task as context, as specified by the context field.
  3. 03
    The writer executes
    Its agent receives the facts and writes the 300-word summary, constrained by its expected_output.
  4. 04
    kickoff() returns the final result
    The output of the last task is returned. verbose=True displays all intermediate reasoning in the terminal.

#Give agents tools

An agent without tools only reasons about what the model already has in memory—quickly limiting and prone to hallucinations. Tools give it concrete capabilities: read a file, search the web, query a database. crewai-tools provides a ready-to-use set, and you can write your own.

Python — one built-in tool and one custom tool
from crewai_tools import FileReadTool
from crewai.tools import tool

# Outil intégré : lire un fichier local
lecteur = FileReadTool()

# Outil maison : décoré avec @tool
@tool("Compteur de mots")
def compter_mots(texte: str) -> int:
    """Compte le nombre de mots d'un texte donné."""
    return len(texte.split())

# On équipe l'agent des deux outils
chercheur = Agent(
    role="Analyste de recherche",
    goal="Analyser des documents locaux sur {sujet}",
    backstory="Tu analyses des sources internes avec rigueur.",
    tools=[lecteur, compter_mots],
    llm=local_llm,
    verbose=True,
)

This is where the model choice becomes decisive. To use a tool, the agent must generate a structured function call that CrewAI intercepts and executes. A model that does not handle tool use well will ignore the tool or produce invalid JSON. Qwen 3.5, Qwen 3.8, Mistral Small, and GLM 4.7 Flash handle this mechanism well; many small general-purpose models do not.

!
One tool per need, not ten per agent
Each added tool expands the agent's system prompt and complicates its decisions. Locally, an agent buried under ten tools makes poor choices or even loops. Give each agent only the tools its task requires—it's more reliable and faster.

#Which local models can handle multi-agent workloads?

This is the real question addressed by this guide. Multi-agent workloads are far more demanding than chat: each agent must follow a role, respect an output format, and often call tools—all while chaining tasks without losing track. A model that is too small loses the thread. Here are the realistic tiers, in Q4_K_M, with the associated VRAM.

8B (≈5–7 GB) — floor
Granite 4.2 8B (5.3 GB), Qwen 3.5 9B (6.6 GB). They can handle a simple sequential crew of 2 agents with basic tools. RTX 3060 12 GB, RTX 4070. Below this size, multi-agent operation becomes unreliable.
16 GB (≈14 GB) — recommended
gpt-oss 20B or Mistral Small 24B (≈14 GB, the latter very good in French), or Qwen 3.5 9B in Q8 (11 GB). The right compromise: solid reasoning, reliable tool use, follows roles without drifting. RTX 4070 12 GB (tight) to RTX 4080 16 GB. This is the sweet spot for most crews.
24 GB (≈18–19 GB) — comfortable
Qwen 3.8 27B (18 GB, 262k ctx) or the Qwen3-Coder 30B-A3B MoE (19 GB). Handles longer crews, multiple tools, and a lightweight hierarchical mode. RTX 4090 24 GB, or a Mac M4 Pro with unified memory. Set the reasoning level of Qwen 3.8 to “low,” otherwise it overthinks in a crew.
MoE 35B+ (≈23-32 GB) — close to cloud performance
Qwen 3.6 35B-A3B (23 GB) or Qwen3-Coder 30B-A3B in Q8 (32 GB). Coordination quality approaches that of cloud APIs, now achievable from 32 GB thanks to MoE architectures. Mac Studio with substantial unified memory or multiple GPUs. Intended for ambitious crews.

Architecture tip: there is no requirement for all agents to share the same model. Assign simple tasks (rewriting, counting, extraction) to a small, fast model (Qwen 3.5 4B, Granite 4.2 8B), and reserve a Qwen 3.8 27B or a 35B MoE for agents that reason or orchestrate. Simply instantiate two LLM objects and assign them by agent.

Python — different models per agent
gros = LLM(model="ollama/qwen3.8:27b", base_url="http://localhost:11434")
petit = LLM(model="ollama/granite4.2:8b", base_url="http://localhost:11434")

stratege = Agent(role="Stratège", goal="...", llm=gros)
assistant = Agent(role="Assistant", goal="...", llm=petit)
→
Watch out for loading and unloading
Running multiple distinct models side by side forces Ollama to juggle VRAM: if they do not fit together, it unloads one to load the other each time you switch, which slows the team down. On a single card, one well-chosen model is often faster than two fighting for memory.

#Costs and limitations compared with a cloud API crew

The strongest argument for local is cost. A crew is chatty by nature: each agent rereads the context, reasons, calls tools, and expands the context as tasks progress, increasing the token count. A single somewhat ambitious run can consume hundreds of thousands of tokens. With a token-billed API, a pipeline running in a loop during development quickly becomes painful; locally, every iteration is free after you buy the hardware.

Cost—local advantage
Zero cost per token. You can iterate, restart, and debug without a counter running. Multi-agent workloads, which are heavy consumers, are where local execution pays off the GPU fastest.
Privacy — the local advantage
None of the calls—and there are many—leave the machine. Decisive for proprietary code, customer data, or anything covered by an NDA or GDPR.
Coordination quality — cloud advantage
GPT-4 and Claude handle hierarchical mode, long chains, and complex tool use with a level of reliability that a local 14B model cannot match. The gap widens as the crew becomes more complex.
Speed — hardware-dependent
The cloud is often faster to respond than a consumer GPU on large models. A crew of 4 agents on a local 32B model can take several minutes per run.

The honest take: local excels at well-scoped sequential crews, where each agent has a clear role and a narrowly defined task. It shows its limits with ambitious hierarchical orchestration, where a 14B model struggles to act as the manager delegating work. The best strategy is often hybrid—prototype and run locally, reserving the cloud for stages where coordination exceeds what your model can handle. A proxy such as LiteLLM lets you route between the two.

#Troubleshooting

The agent ignores its tools
The model does not support function calling. Switch to Qwen 3.5, Qwen 3.8, Mistral Small, or GLM 4.7 Flash, and verify that the agent has the tools=[...] list.
“Connection refused” / litellm error
The Ollama daemon is not running, or base_url is incorrect. Check “ollama ps” and verify that the URL is http://localhost:11434.
The agent loops or does not stop
The model is too small for the task, or there are too many tools. Move up in size (Qwen 3.5 9B at minimum), reduce the number of tools, lower the temperature, and set max_iter on the agent.
Malformed outputs / expected_output ignored
Role too vague or expected_output unclear. Make them highly concrete, and prefer a more capable model (Qwen 3.8 27B, Mistral Small 24B) that follows formatting instructions more reliably.
Very slow crew
The model spills into RAM/CPU because of insufficient VRAM (“ollama ps” shows it), or multiple models unload one another. Drop down one size tier or consolidate on a single model.
The hierarchical mode goes off the rails
The LLM manager isn't up to the task. Go back to Process.sequential, or assign a Qwen 3.8 27B or a 35B MoE to the manager role.

#Go further

A local CrewAI crew relies on building blocks already covered on the site. These guides build on this one:

Build a local AI agent in Python with LangChain and Ollama
The foundations of a single agent—tools and the reasoning loop—before moving to multi-agent systems.
Function calling and structured JSON outputs with Ollama
To understand the tool-use mechanism that your agents' tools depend on.
LiteLLM: a unified local and cloud proxy
To route a crew between Ollama locally and a cloud API based on the task, as part of a hybrid strategy.
Did this guide help you?

Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.