Running CrewAI on a Local LLM
CrewAI orchestrates several agents with roles and tasks. Pointing it at Ollama instead of a cloud API works — but only above a certain model size. Setup, the failure modes, and how to pick the model.
Key takeaways
- CrewAI is a Python framework for multi-agent workflows: you declare agents with a role, a goal and tools, give them tasks, and a process runs them in order or under a manager agent.
- It reaches local models through a provider layer, so pointing it at Ollama is a configuration change, not a rewrite.
- Model size is the whole game. Multi-agent frameworks lean on reliable instruction-following and structured tool calls; small models produce plausible text and malformed calls, which turns into loops.
- Realistic floor on consumer hardware: a 14B-class model for simple crews, 27B–32B for anything with tools. Below 8B, expect to debug the framework rather than use it.
- The cost model flips: no per-token bill, but a crew of four agents doing three passes each is dozens of full-context generations on your own GPU. Wall-clock time replaces money as the constraint.
What CrewAI is, in one pass
Agents that act on your machine: agentic Cline, MCP, n8n + Ollama, local automations.
- Lifetime online access
- PDF + files
- 30-day refund
CrewAI models a workflow the way an org chart does. An agent has a role ("research analyst"), a goal, a backstory that conditions its behaviour, and optionally tools it may call. A task is a unit of work with a description and an expected output, assigned to an agent. A crew is the set of agents and tasks plus a process: sequential, where tasks run in order and each sees the previous output, or hierarchical, where a manager agent delegates and reviews.
The appeal is that this decomposition often beats one long prompt: a writer agent that receives a researcher's notes produces better output than a single model asked to research and write in one shot. The cost is that every hand-off is another generation, and every tool call is another chance to fail.
Pointing a crew at a local model
Three pieces have to line up: a running local server, a model that is good at following instructions, and CrewAI configured to use that endpoint rather than a cloud provider.
- Serve the model. Ollama is the usual choice — see what Ollama is and the service guide for exposing it properly. For several agents hammering the endpoint in parallel, vLLM serves concurrent requests far better; that difference is the point of vLLM vs Ollama.
- Configure the provider. CrewAI routes model calls through a provider abstraction that understands Ollama and OpenAI-compatible endpoints. In practice you set the model identifier with its provider prefix and the base URL of your server, in the environment or when constructing the LLM object.
- Raise the context window. This is the step people skip. Agent prompts carry role, backstory, task description, tool schemas and prior outputs; the default context of a local model is often smaller than that. A truncated prompt is the single most common cause of a crew that behaves erratically. Budget the memory for it — what a long context costs in VRAM has the arithmetic.
Which model actually works
| Model class | Typical VRAM at 4-bit | What to expect in a crew |
|---|---|---|
| 3B–8B | 2–6 GB | Fluent prose, unreliable tool calls and format adherence. Fine for a single-agent summariser, frustrating as a crew. |
| 12B–14B | 7–10 GB | The practical floor. Sequential crews without tools work; tool use is hit and miss. |
| 24B–32B | 14–20 GB | Where local multi-agent becomes usable. Structured output and delegation hold up across several turns. |
| 70B and up | 40 GB+ | Closest to cloud behaviour, at a speed that makes long crews a batch job rather than an interaction. |
Prefer models trained for tool use and structured output; a coding-oriented model is often a better crew member than a general chat model of the same size, because it has seen more strict formats. The agents and tool-use ranking is the shortlist to start from, and the VRAM calculator tells you which of them your card holds.
The four failure modes, and what causes them
- The crew loops. An agent repeats the same step because it cannot produce the exact format the framework expects to mark the task done. Cause: model too small, or the context silently truncated. Fix: bigger model, longer context, fewer tools per agent.
- Tool calls fail to parse. The model writes something that reads like a function call but is not valid JSON. Fix: a model with genuine tool-calling training; constrain output where your server supports it; give the agent one tool rather than six.
- Agents agree with each other endlessly. Hierarchical processes with a weak manager model degenerate into mutual approval. Fix: sequential process, or a larger model as manager.
- It is slow beyond usefulness. Four agents × three tasks × a long prompt each is a lot of tokens. Fix: shorten backstories, cut the crew to the agents that change the output, and serve with something built for concurrency.
CrewAI or something else
| You want | Reach for |
|---|---|
| Role-based teams of agents in Python | CrewAI |
| Conversational agents that negotiate turn by turn | AutoGen |
| Explicit graphs with state and cycles you control | LangGraph |
| A visual builder instead of code | A no-code agent platform |
| One agent that edits your repository | A dedicated coding agent — see local coding assistants |
Verdict
CrewAI runs locally without drama; what does not run locally is the assumption that any model will do. Give it a 27B-class model with real tool-calling ability, a context window sized for agent prompts, and a server that handles concurrent requests, and a local crew does useful work with no API bill. Try the same crew on an 8B model and you will conclude the framework is broken when the model is.
Frequently asked questions
Can CrewAI run fully offline?
Yes. With a local server such as Ollama and no tools that reach the internet, nothing leaves the machine. Add a web-search tool and that stops being true for those calls.
What is the smallest model that works with CrewAI?
In practice a 12B–14B model at 4-bit is the floor for simple sequential crews, and 24B–32B for crews that use tools. Smaller models generate convincing text but fail the strict formats the framework depends on.
Does CrewAI need LangChain?
No. CrewAI is a standalone framework with its own agent and task model; it does not require you to build on LangChain, though it can interoperate with common tool ecosystems.
Why does my local crew keep repeating the same step?
Almost always the model failing to emit the completion format the framework expects, or a prompt that exceeded the context window and was truncated. Raise the context limit first, then try a larger model, then reduce the number of tools per agent.
Is CrewAI better than AutoGen for local models?
Neither is strictly better; they model different things. CrewAI suits pipelines with defined roles and hand-offs, AutoGen suits conversational multi-agent patterns. Both are equally sensitive to model quality when run locally.
How much VRAM do I need for a local crew?
Enough for the model plus a generous context: roughly 16 GB for a 14B-class model with a long context, and 24 GB or more for a 27B–32B model, which is where local multi-agent starts working reliably.
A current option for local AI: GMKtec EVO-X2 64GB / 1TB (Ryzen AI Max+ 395). Match memory to your model and software. A mini PC is a complete PC alternative; Mac/MLX and CUDA instructions require compatible hardware.
Amazon Check GMKtec EVO-X2 64GB / 1TB (Ryzen AI Max+ 395) price →As an Amazon Associate, BestLLMfor earns from qualifying purchases, at no extra cost to you. It does not influence our independent rankings.
Found an error or have feedback? Let us know — it helps everyone who reads this guide.