Advanced 25 minAgents

Building a local AI agent: architecture and recommended tools

A local AI agent is a self-hosted LLM given tools, memory, and a decision loop to perform tasks without constant supervision. Unlike a simple chatbot, it plans, acts, observes the result, and tries again. This guide describes the architecture of such an agent and the recommended frameworks—CrewAI and AutoGen on Ollama—so nothing leaves your machine.

By Mohamed Meguedmi·Update 2026-08-31·Tested on Windows, macOS, and Linux

#Why build a local AI agent

A local AI agent addresses three needs that the cloud handles poorly. First, confidentiality: when an agent reads your emails, queries your database, or browses your files, every call sent to an external API is a potential leak. Locally, the context never leaves the machine. Second, cost: an agent chains dozens of model calls for a single task, and the bill for a usage-based API quickly explodes. Finally, autonomy: no rate limits, no network outages, and no overnight pricing-policy changes.

The trade-off is real: a local 14B or 32B model does not reason as finely as the best proprietary models. Agent design—well-defined tools, strict prompts, guardrails—therefore matters more than it does in the cloud. That is precisely what this guide to local AI agents covers.

i
Agent ≠ chatbot
A chatbot responds. An agent decides what to do, performs an action (tool call), reads the result, then repeats until it reaches the goal. This “reason → act → observe” loop is the heart of the subject.

#Anatomy of an agent

The Local Copilot Kit

This guide gets you to the model. The kit gets you to the coding copilot in your editor.

  • Lifetime online access
  • PDF + files
  • Lifetime updates

Regardless of the framework, a local AI agent always relies on the same building blocks. Understanding them lets you choose your tools knowledgeably instead of blindly following a tutorial.

The model (reasoning)
The LLM that decides the next action. It must handle tool calling reliably: Qwen 3.x, Granite 4.x, or Mistral Small are good local candidates.
Tools (actions)
Python functions the agent can call: read a file, query an API, run a search, write to a database. Each tool is described by a JSON schema that the model reads.
Memory (state)
Short term (the history of the current conversation) and long term (a vector store that persists between sessions). That is RAG’s role, detailed below.
The orchestrator (loop)
The code that chains reasoning, tool calls, and observations until the stopping condition is met. This is what a framework such as CrewAI or AutoGen provides.
The planner (strategy)
The logic that breaks a complex goal into ordered subtasks. It can be explicit (a dedicated planning agent) or implicit (the model reasons step by step).

A minimal agent uses one model, two or three tools, and a loop. An advanced system adds persistent memory, multi-step planning, and several specialized agents that collaborate. Start simple: most tasks don’t justify a team of ten agents.

#Requirements and recommended stack

The reference stack for a local AI agent has three layers: Ollama to serve the model, a Python orchestration framework, and a model capable of tool calling. Ollama listens on http://localhost:11434 by default and exposes an OpenAI-compatible endpoint, simplifying integration with most frameworks.

GPU / VRAM
An agent reasons better with a 14B+ model than with a 7B model. Plan on ~9 GB of VRAM for a 14B in Q4_K_M, and ~19 GB for a 32B. A RTX 4070 12 GB runs a 14B comfortably; a RTX 4090 24 GB or an M4 Pro Mac targets the 32B.
Model
Choose a model known to be reliable for tool calling. Function-calling quality matters more than raw size: a rigorous 14B model beats a 32B model that makes up arguments.
Python 3.10+
CrewAI and AutoGen are Python libraries. Work in a dedicated virtual environment to avoid dependency conflicts.
Quantization
Q4_K_M is the right default compromise. Move up to Q5_K_M or Q8_0 if the model makes reasoning errors and the VRAM allows it.
Terminal — prepare the stack
# 1. Vérifier qu'Ollama tourne
curl http://localhost:11434/api/tags

# 2. Récupérer un modèle capable de tool calling
ollama pull qwen3:14b

# 3. Environnement Python isolé
python -m venv .venv && source .venv/bin/activate
pip install crewai crewai-tools
→
Test tool calling first
Before connecting a framework, verify that your model correctly calls a simple function through the Ollama API. An agent built on a model that hallucinates tool calls will be unmanageable—better to find out right away.

#CrewAI or AutoGen: which should you choose?

Both frameworks orchestrate agents, but with different philosophies. The right choice depends on your task, not on an absolute ranking.

CrewAI
Team-oriented: define agents with a role, objective, and tools, then assign them ordered tasks. A declarative, readable approach that’s ideal for business pipelines (research → draft → review). Built-in RAG memory.
AutoGen
Conversation-oriented: agents talk to one another until they converge. More flexible for open-ended problems and collaborative reasoning, but requires more tuning to stay bounded locally. v0.4 connects to Ollama through its OpenAI-compatible client.
When to keep it simple
For a single agent with a few tools, a lightweight framework (or LangChain) is enough. CrewAI and AutoGen really shine when there are multiple roles or nontrivial orchestration.
Python — CrewAI connected to Ollama
from crewai import Agent, Task, Crew, LLM

# CrewAI passe par LiteLLM : préfixe 'ollama/' + base_url local
llm = LLM(
    model="ollama/qwen3:14b",
    base_url="http://localhost:11434",
)

chercheur = Agent(
    role="Analyste documentaire",
    goal="Extraire les faits clés des documents fournis",
    backstory="Expert méthodique, ne répond que sur la base des sources.",
    llm=llm,
    verbose=True,
)

tache = Task(
    description="Résume les 3 points essentiels du rapport fourni.",
    expected_output="Une liste à puces de 3 points, sourcés.",
    agent=chercheur,
)

equipe = Crew(agents=[chercheur], tasks=[tache])
print(equipe.kickoff())
Python — AutoGen 0.4 toward the Ollama endpoint
from autogen_ext.models.openai import OpenAIChatCompletionClient
from autogen_agentchat.agents import AssistantAgent

# Endpoint OpenAI-compatible d'Ollama : /v1
client = OpenAIChatCompletionClient(
    model="qwen3:14b",
    base_url="http://localhost:11434/v1",
    api_key="ollama",  # ignore par Ollama, mais requis par le client
    model_info={
        "function_calling": True,
        "json_output": True,
        "vision": False,
        "family": "unknown",
    },
)

agent = AssistantAgent(name="assistant", model_client=client)

#Build the agent step by step

Here's the process for going from a raw model to a local AI agent that performs a real task. The order matters: each step validates the previous one.

  1. 01
    Define the objective and stopping condition
    Write one sentence describing what the agent should produce and how you know it is finished. A vague goal (“help me”) produces an agent that goes in circles; a bounded goal (“sort these 20 invoices by supplier in a CSV”) produces a controllable agent.
  2. 02
    Break it down into atomic tools
    Every external action becomes a Python function with an explicit name, typed arguments, and a clear docstring — this is the description the model reads. Prefer several small, precise tools over one large catch-all tool.
  3. 03
    Write the system prompt
    Define the role, available tools, and rules (never make things up, always cite the source, stop when in doubt). Locally, a strict prompt offsets the model's lower reasoning finesse.
  4. 04
    Wire up the orchestration loop
    Let the framework manage the reason → act → observe loop, but set an iteration limit (for example, 10) to prevent a stuck agent from consuming resources indefinitely. This is an essential safeguard for autonomous operation.
  5. 05
    Add memory
    Connect a vector store for long-term memory if the agent needs to remember between sessions (see the next section). Without that need, the conversation history is enough.
  6. 06
    Test on real-world cases and iterate
    Run the agent on varied inputs, read the execution traces (verbose), and refine the prompts and tool descriptions. 80% of the work on a local agent happens here, not in the initial code.

#Long-term memory with RAG

An agent without memory starts from scratch in every session. Long-term memory relies on the same principle as RAG (Retrieval-Augmented Generation): information is stored as vectors in a database, and the most relevant items are retrieved when needed and fed back into the context.

Local embeddings
Generate the vectors with an embedding model served by Ollama (for example, nomic-embed-text or mxbai-embed-large): the data being memorized never leaves the machine, consistent with the privacy goal.
Vector store
ChromaDB is the default choice locally: lightweight, persistent on disk, and natively integrated with CrewAI. For larger volumes, self-hosted Qdrant takes over.
CrewAI integrated memory
CrewAI provides ready-to-use memory (short-term, long-term, and entity memory) that you can configure to use an embedder Ollama, without building the RAG pipeline by hand.
Python — CrewAI memory with Ollama embeddings
from crewai import Crew

equipe = Crew(
    agents=[chercheur],
    tasks=[tache],
    memory=True,  # active la memoire long terme (ChromaDB sous le capot)
    embedder={
        "provider": "ollama",
        "config": {"model": "nomic-embed-text"},
    },
)
→
Don't memorize everything
Memory that grows without filtering eventually brings noise into every request and degrades responses. Explicitly decide what is worth retaining (durable facts, user preferences) and leave the rest in ephemeral session memory.

#Task scheduling

Planning is the agent's ability to break a complex goal into ordered steps before acting. Without it, a local model tends to rush into the first action that comes to mind and get lost. Two approaches coexist.

Implicit planning (ReAct)
The model reasons out loud step by step, chooses an action, observes, then replans. Easy to set up, but fragile on smaller models that lose track after a few turns.
Explicit planning
A dedicated planning agent (or an initial task) produces a list of steps, which execution agents then process. More robust locally: “thinking” and “doing” are separated, reducing the load on each call.
Hierarchical decomposition
For long-running tasks, CrewAI enables a hierarchical process in which a “manager” agent delegates and supervises. Powerful, but reserve it for cases that warrant it—the coordination costs tokens.

A practical rule for local deployments: the smaller the model, the more explicit the planning must be and the narrower the scope of each step. A 14B handling a well-defined micro-task is more reliable than a 32B given a vague objective.

#Data security

Running everything locally eliminates leaks to third-party APIs, but an autonomous agent introduces its own risks: it executes actions, sometimes destructive ones, based on generated text. Privacy doesn’t eliminate the need for safeguards.

Principle of least privilege
Give the agent only the tools it strictly needs. An agent that does not need to write to disk must not have a write tool — this is the first barrier against damage.
Human validation of sensitive actions
For anything irreversible (deleting, sending, paying, modifying a database), add a manual confirmation. Full autonomy is justified only for safe, reversible actions.
Code execution isolation
If the agent runs code, do it in a container or sandboxed environment, never directly on the host machine. A malicious prompt embedded in a document can hijack an agent (prompt injection).
Action logging
Trace every tool call and every decision. When behavior is unexpected, the trace is your only way to understand what the agent actually did.
!
Prompt injection remains the primary threat
An agent that reads untrusted content (emails, web pages, received files) can be manipulated by instructions hidden in that content. Never mix trusted data and external data in the same context without treating the latter as hostile.

#Tips and troubleshooting

The agent loops endlessly
Check the stopping condition and iteration limit. Often the goal is too vague, or the agent doesn’t recognize that it has finished: make the expected result explicit in the prompt.
Malformed tool calls
The model invents arguments or forgets fields. Simplify the tool schemas, move up one quantization level (Q4 → Q5), or switch to a more reliable model for tool calling.
Slow responses in multi-agent setups
Each agent is a full model call. Locally, reduce the number of agents, shorten the system prompts, and verify that the model fits entirely in VRAM (otherwise CPU offloading crushes throughput).
Memory retrieves nothing relevant
Wrong embedding model or chunks that are too large or too small. Verify that the Ollama embedder is running, and adjust the size of the stored chunks.
Connection refused on :11434
Ollama is not running or is listening on another interface. Confirm with “curl http://localhost:11434/api/tags” and check the base_url passed to the framework.

#Go further

This guide lays out the architecture; these tutorials cover the concrete implementation of each component:

Multi-agent systems with CrewAI
“CrewAI + Ollama: orchestrating multiple local AI agents” provides a detailed look at setting up a team of specialized agents, with roles and tasks included.
An end-to-end Python agent
“Build a local AI agent in Python with LangChain and Ollama” walks you step by step through building an agent that can call tools and read files.
The memory building block (RAG)
“Local RAG with ChromaDB and Ollama: Python tutorial” covers the complete embeddings → search → answer pipeline, the core of long-term memory.
Did this guide help you?

Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.