CrewAI + Ollama: orchestrating multiple AI agents in local
A single AI agent answers a question; a team of agents solves a problem. CrewAI orchestrates several specialized LLMs — a researcher, a writer, and a reviewer — that pass work between them like colleagues. This guide builds a CrewAI crew that runs entirely locally on Ollama: no data reaches a cloud API, and there is no per-token cost. You'll see how to connect CrewAI to the local endpoint, define roles and tasks that actually work, equip agents with tools, and, above all, which local models can handle the load of multiple agents without collapsing.
#Why orchestrate a CrewAI crew locally
The multi-agent pattern starts with a simple observation: splitting a complex task among several specialized agents produces better results than a single giant prompt. Each agent has a clear role, a specific objective, and sees only its part of the work. CrewAI is the Python framework that formalizes this split—roles, tasks, sequential or hierarchical collaboration—without the overhead of wiring up a state graph yourself.
Running this crew on Ollama instead of GPT-4 or Claude changes three things. First, privacy: a multi-agent pipeline multiplies calls to the model, and therefore multiplies potential leaks to a third party; locally, nothing leaves the machine. Next, cost: a chatty crew can burn hundreds of thousands of tokens per run, which quickly becomes expensive on a token-priced API—in local operation, the marginal cost is zero. Finally, control: you choose the model, quantization, context, and can iterate without quotas or rate limits.
#The 4 building blocks of CrewAI
This guide gets you to the model. The kit gets you to the coding copilot in your editor.
- Lifetime online access
- PDF + files
- Lifetime updates
Before writing a line, you need the vocabulary. CrewAI is built around four objects that fit together. Understanding them prevents 90% of design errors in a crew.
- Agent
- An LLM with an identity: a role (“Market Analyst”), a goal, and a backstory that shape its behavior. Each agent can use its own Ollama model.
- Task
- A unit of work assigned to an agent: a description, an expected result (expected_output), and often, tools. The task—not the agent—carries the concrete instruction.
- Tool
- An external capability that an agent can call: web search, file reading, SQL queries, or computation. Without tools, an agent can only reason about what it already knows.
- Crew
- The crew: the list of agents, the list of tasks, and the process that determines the execution order (sequential or hierarchical). This is the object you launch with kickoff().
The process deserves one clarification. In sequential mode, tasks run in the declared order, and the output of one feeds the next—perfect for a research → writing → proofreading pipeline. In hierarchical mode, a “manager” agent (a dedicated LLM) delegates to and coordinates the others. Hierarchical mode is more powerful but much more demanding for a local model, because the manager must reason about who does what: always start with sequential mode.
#Prerequisites and model selection
- Ollama working
- Daemon running, endpoint at http://localhost:11434. Verify with “ollama list” that at least one tool-use-capable model is available.
- Python 3.10+ and a venv
- CrewAI deploys cleanly in an isolated virtual environment. Avoid installing it in the system Python.
- VRAM—and patience
- Multi-agent systems chain calls together: each task means one or more round trips to the model. Plan on at least an 8B model, ideally in the 12B to 24B range, so the reasoning holds up.
- A tool-capable model
- To equip agents with tools, you need a model that supports function calling: Qwen 3.5, Qwen 3.8, Mistral Small, or GLM 4.7 Flash. A model without tool use can only reason in text.
#Install CrewAI and connect it to Ollama
CrewAI installs via pip. The crewai-tools package also provides a library of ready-to-use tools (search, files, scraping).
The key point about the local connection: CrewAI relies on LiteLLM to talk to models. To target Ollama, prefix the model name with “ollama/” and point the base URL to the local daemon. Make sure you have pulled the model beforehand.
#First crew: define roles and tasks
Let’s assemble a classic, immediately useful crew: a researcher who gathers information, followed by a writer who formats it. This is the framework we reuse for monitoring, document summarization, or content generation. We first define the agents, with clear roles.
The role, goal, and backstory aren't decorative: they form each agent's system prompt. A vague role (“Assistant”) produces a vague agent. Be specific and establish a stance—that's what keeps a small model from going off track. Note the {sujet} placeholder: CrewAI injects it at launch from the inputs.
Next comes the core of the work: the tasks. Each task points to an agent, describes what it must produce, and, most importantly, specifies an expected_output. This field is CrewAI’s most underestimated quality lever: the more concrete it is, the more constrained the output.
The context parameter explicitly links the tasks: the writing task receives the research output. In sequential mode, the chain is already implicit, but declaring context makes the dependency clear and makes information transfer more reliable. Finally, we assemble the crew and launch it.
- 01The retriever runsIts local LLM receives its role plus the description of the research task and produces the expected list of facts.
- 02The output passes throughCrewAI passes the search result to the writing task as context, as specified by the context field.
- 03The writer executesIts agent receives the facts and writes the 300-word summary, constrained by its expected_output.
- 04kickoff() returns the final resultThe output of the last task is returned. verbose=True displays all intermediate reasoning in the terminal.
#Give agents tools
An agent without tools only reasons about what the model already has in memory—quickly limiting and prone to hallucinations. Tools give it concrete capabilities: read a file, search the web, query a database. crewai-tools provides a ready-to-use set, and you can write your own.
This is where the model choice becomes decisive. To use a tool, the agent must generate a structured function call that CrewAI intercepts and executes. A model that does not handle tool use well will ignore the tool or produce invalid JSON. Qwen 3.5, Qwen 3.8, Mistral Small, and GLM 4.7 Flash handle this mechanism well; many small general-purpose models do not.
#Which local models can handle multi-agent workloads?
This is the real question addressed by this guide. Multi-agent workloads are far more demanding than chat: each agent must follow a role, respect an output format, and often call tools—all while chaining tasks without losing track. A model that is too small loses the thread. Here are the realistic tiers, in Q4_K_M, with the associated VRAM.
- 8B (≈5–7 GB) — floor
- Granite 4.2 8B (5.3 GB), Qwen 3.5 9B (6.6 GB). They can handle a simple sequential crew of 2 agents with basic tools. RTX 3060 12 GB, RTX 4070. Below this size, multi-agent operation becomes unreliable.
- 16 GB (≈14 GB) — recommended
- gpt-oss 20B or Mistral Small 24B (≈14 GB, the latter very good in French), or Qwen 3.5 9B in Q8 (11 GB). The right compromise: solid reasoning, reliable tool use, follows roles without drifting. RTX 4070 12 GB (tight) to RTX 4080 16 GB. This is the sweet spot for most crews.
- 24 GB (≈18–19 GB) — comfortable
- Qwen 3.8 27B (18 GB, 262k ctx) or the Qwen3-Coder 30B-A3B MoE (19 GB). Handles longer crews, multiple tools, and a lightweight hierarchical mode. RTX 4090 24 GB, or a Mac M4 Pro with unified memory. Set the reasoning level of Qwen 3.8 to “low,” otherwise it overthinks in a crew.
- MoE 35B+ (≈23-32 GB) — close to cloud performance
- Qwen 3.6 35B-A3B (23 GB) or Qwen3-Coder 30B-A3B in Q8 (32 GB). Coordination quality approaches that of cloud APIs, now achievable from 32 GB thanks to MoE architectures. Mac Studio with substantial unified memory or multiple GPUs. Intended for ambitious crews.
Architecture tip: there is no requirement for all agents to share the same model. Assign simple tasks (rewriting, counting, extraction) to a small, fast model (Qwen 3.5 4B, Granite 4.2 8B), and reserve a Qwen 3.8 27B or a 35B MoE for agents that reason or orchestrate. Simply instantiate two LLM objects and assign them by agent.
#Costs and limitations compared with a cloud API crew
The strongest argument for local is cost. A crew is chatty by nature: each agent rereads the context, reasons, calls tools, and expands the context as tasks progress, increasing the token count. A single somewhat ambitious run can consume hundreds of thousands of tokens. With a token-billed API, a pipeline running in a loop during development quickly becomes painful; locally, every iteration is free after you buy the hardware.
- Cost—local advantage
- Zero cost per token. You can iterate, restart, and debug without a counter running. Multi-agent workloads, which are heavy consumers, are where local execution pays off the GPU fastest.
- Privacy — the local advantage
- None of the calls—and there are many—leave the machine. Decisive for proprietary code, customer data, or anything covered by an NDA or GDPR.
- Coordination quality — cloud advantage
- GPT-4 and Claude handle hierarchical mode, long chains, and complex tool use with a level of reliability that a local 14B model cannot match. The gap widens as the crew becomes more complex.
- Speed — hardware-dependent
- The cloud is often faster to respond than a consumer GPU on large models. A crew of 4 agents on a local 32B model can take several minutes per run.
The honest take: local excels at well-scoped sequential crews, where each agent has a clear role and a narrowly defined task. It shows its limits with ambitious hierarchical orchestration, where a 14B model struggles to act as the manager delegating work. The best strategy is often hybrid—prototype and run locally, reserving the cloud for stages where coordination exceeds what your model can handle. A proxy such as LiteLLM lets you route between the two.
#Troubleshooting
- The agent ignores its tools
- The model does not support function calling. Switch to Qwen 3.5, Qwen 3.8, Mistral Small, or GLM 4.7 Flash, and verify that the agent has the tools=[...] list.
- “Connection refused” / litellm error
- The Ollama daemon is not running, or base_url is incorrect. Check “ollama ps” and verify that the URL is http://localhost:11434.
- The agent loops or does not stop
- The model is too small for the task, or there are too many tools. Move up in size (Qwen 3.5 9B at minimum), reduce the number of tools, lower the temperature, and set max_iter on the agent.
- Malformed outputs / expected_output ignored
- Role too vague or expected_output unclear. Make them highly concrete, and prefer a more capable model (Qwen 3.8 27B, Mistral Small 24B) that follows formatting instructions more reliably.
- Very slow crew
- The model spills into RAM/CPU because of insufficient VRAM (“ollama ps” shows it), or multiple models unload one another. Drop down one size tier or consolidate on a single model.
- The hierarchical mode goes off the rails
- The LLM manager isn't up to the task. Go back to Process.sequential, or assign a Qwen 3.8 27B or a 35B MoE to the manager role.
#Go further
A local CrewAI crew relies on building blocks already covered on the site. These guides build on this one:
- Build a local AI agent in Python with LangChain and Ollama
- The foundations of a single agent—tools and the reasoning loop—before moving to multi-agent systems.
- Function calling and structured JSON outputs with Ollama
- To understand the tool-use mechanism that your agents' tools depend on.
- LiteLLM: a unified local and cloud proxy
- To route a crew between Ollama locally and a cloud API based on the task, as part of a hybrid strategy.
Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.