Intermediate 20 minPython

Building a local AI agent in Python with LangChain and Ollama

A local AI agent in Python with LangChain and Ollama is more than a chatbot: it is a program that decides on its own when to call a function, read a file, or chain multiple steps to respond. This guide builds a working agent step by step in about twenty minutes, using an Qwen 3.5 9B model that runs entirely on your machine. Zero API keys, zero data sent to a third party.

By Mohamed Meguedmi·Update 2026-08-27·Tested on Windows, macOS, and Linux

#Why run a local AI agent in Python?

An agent, in the LangChain sense, is a simple loop: the LLM receives a question and its list of tools, chooses whether to call one, reads the result, and repeats until it can answer. The entire “decision-making” mechanism relies on the model’s ability to emit a structured tool call.

Doing this locally with Ollama changes two concrete things: your data never leaves the machine, and every call costs zero euros. That’s the difference between prototyping with OpenAI and ending the week with a €50 bill, and iterating without counting the cost.

Privacy
The files the agent reads (contracts, proprietary code, medical notes) do not leave the workstation. No DPA to sign, no transfer outside the EU.
Zero marginal cost
Once the model is downloaded, you can iterate hundreds of times a day without your bill climbing.
Reproducibility
You pin the model's exact version (qwen3.5:9b, granite4.2:8b, etc.). No silent drift like with gpt-4o-2024-11-20, which becomes something else a month later.
Predictable latency
No network round trip. On a decent GPU, the first token arrives in under a second.
i
It's not magic either
A local 9B model is still weaker than GPT-5 or Claude 4.7 on highly complex tasks. For 80% of useful agents (reading a file, calling an internal API, performing a calculation, sorting an email), it’s more than sufficient. For everything else, it’s an excellent learning ground before paying for tokens.

#Prerequisites

The Local Copilot Kit

This guide gets you to the model. The kit gets you to the coding copilot in your editor.

  • Lifetime online access
  • PDF + files
  • Lifetime updates
Python 3.10+
LangChain is no longer tested on 3.9. Check with python --version.
Ollama installed and started
It needs to listen on http://localhost:11434. See the Ollama installation guides (Windows, macOS, Linux) if you haven't already.
A model that can call tools
Not all LLMs can do this. Qwen 3.5, Granite 4.2, Gemma 4, Devstral, and GLM 4.7 Flash natively support tool calling. Avoid models that are now outdated (Llama 2/3, Qwen 2.5, Mistral 7B).
Hardware
Qwen 3.5 9B Q4 weighs about 6.6 GB in VRAM. An 8 GB GPU (RTX 3060, 4060) is sufficient; a 12 GB one (4070) provides headroom. On Mac, plan on 16 GB of unified memory for comfortable operation.
→
Model choice is critical
With a model that cannot properly call tools, your agent will hallucinate arguments or respond in free-form text instead of producing a tool call. If you are just starting out, stick with qwen3.5:9b—it is the quality/VRAM sweet spot in 2026.

#1. Initialize the Python project

A virtual environment, three packages, and that's it. We avoid installing LangChain in system Python — it changes quickly and creates clutter.

Create and activate the venv
mkdir agent-local && cd agent-local
python -m venv .venv

# macOS / Linux
source .venv/bin/activate

# Windows PowerShell
# .venv\Scripts\Activate.ps1
Install the dependencies
pip install --upgrade pip
pip install langchain langchain-ollama langgraph
langchain
The core: prompt abstractions, tools, messages.
langchain-ollama
The official integration Ollama. Maintained by the LangChain team since 2024.
langgraph
For the agent loop. This is the recommended engine today, more stable than the older AgentExecutors.
i
Why langgraph instead of AgentExecutor?
Older LangChain tutorials use AgentExecutor + create_react_agent (from langchain.agents). This API is in maintenance mode. The official docs now point to langgraph.prebuilt.create_react_agent—that's what we'll use here. Simpler, better typed, free streaming.

#2. Connect to Ollama from Python

Before setting up an agent, verify that you're actually talking to the model. Download the model if you haven't already, then test the simplest possible call.

Download Qwen 3.5 9B
ollama pull qwen3.5:9b

The download is about 6.6 GB in Q4_K_M (the default quantization in Ollama). Once it's in place, create the first script:

test_ollama.py
from langchain_ollama import ChatOllama

llm = ChatOllama(
    model="qwen3.5:9b",
    temperature=0,
    # base_url="http://localhost:11434",  # par défaut, à changer si Ollama est ailleurs
)

reponse = llm.invoke("En une phrase : qu'est-ce qu'un agent IA ?")
print(reponse.content)
Run the test
python test_ollama.py

If you see a coherent sentence, the Python ↔ Ollama connection is working. If you get a ConnectionError, check that Ollama is running properly (ollama ps should list an active service).

→
temperature=0 for agents
We want deterministic behavior when the model chooses a tool. A high temperature makes tool calls vary from one run to the next—it’s a nightmare to debug. For creative responses, raise it to 0.7 later.

#3. Define the agent's tools

A LangChain tool is simply a Python function decorated with @tool. The docstring becomes the description the LLM sees—it uses it to decide when to call the tool. Be precise: a vague docstring produces random calls.

We will create two representative tools: an arithmetic-expression evaluator and a file reader.

tools.py
from pathlib import Path
from langchain_core.tools import tool


@tool
def calculer(expression: str) -> str:
    """Évalue une expression arithmétique simple.

    Args:
        expression: une expression contenant uniquement des chiffres,
                    des espaces et les opérateurs + - * / ( ).

    Returns:
        Le résultat numérique sous forme de chaîne, ou un message d'erreur.
    """
    autorise = set("0123456789+-*/(). ")
    if not all(c in autorise for c in expression):
        return "Erreur : caractère non autorisé. Seuls 0-9 et + - * / ( ) sont permis."
    try:
        resultat = eval(expression, {"__builtins__": {}}, {})
        return str(resultat)
    except Exception as e:
        return f"Erreur de calcul : {e}"


@tool
def lire_fichier(chemin: str) -> str:
    """Lit le contenu d'un fichier texte du répertoire courant.

    Args:
        chemin: chemin relatif ou absolu vers un fichier texte (.txt, .md, .py, etc.).

    Returns:
        Le contenu du fichier, ou un message d'erreur si introuvable.
    """
    p = Path(chemin)
    if not p.exists():
        return f"Fichier introuvable : {chemin}"
    if not p.is_file():
        return f"Ce n'est pas un fichier : {chemin}"
    try:
        return p.read_text(encoding="utf-8")
    except UnicodeDecodeError:
        return "Fichier binaire ou encodage non UTF-8."
    except Exception as e:
        return f"Erreur de lecture : {e}"
!
eval() is dangerous in production
Using eval(), even with __builtins__ emptied, is not a real sandbox. For an agent running on your machine that you control, it is acceptable. For anything exposed to third-party users, use ast.parse with an operator whitelist, or the simpleeval library.

Three rules for tools the model uses correctly:

Explicit name
calculate rather than process, read_file rather than get. The LLM chooses based first on the name.
Detailed docstring
Describe what the tool does, what it expects, and what it returns. LangChain reads the Python type annotations and exposes them to the model.
Return a string
Always. If the function returns a dict or an object, LangChain serializes it, but you lose readability on the model side.

#4. Assemble the agent

We have an LLM and we have tools. langgraph's create_react_agent connects the two and manages the loop: as long as the model wants to call tools, we continue; when it responds with text, we stop.

agent.py
from langchain_ollama import ChatOllama
from langgraph.prebuilt import create_react_agent
from tools import calculer, lire_fichier

llm = ChatOllama(model="qwen3.5:9b", temperature=0)

SYSTEM_PROMPT = (
    "Tu es un assistant en français. Tu disposes d'outils pour calculer "
    "et lire des fichiers. Utilise-les dès que c'est pertinent, sans jamais "
    "inventer un résultat. Réponds toujours en français."
)

agent = create_react_agent(
    model=llm,
    tools=[calculer, lire_fichier],
    prompt=SYSTEM_PROMPT,
)

if __name__ == "__main__":
    question = (
        "Combien fait 1234 * 5678 ? "
        "Ensuite, lis le fichier notes.txt et résume-le en deux phrases."
    )
    reponse = agent.invoke({"messages": [("user", question)]})

    # Le dernier message est la réponse finale du modèle
    print(reponse["messages"][-1].content)

Create a small notes.txt file next to it for testing:

Test file
echo "Réunion projet Hermes : on garde Ollama comme runtime principal, on évalue vLLM pour la prod, RAG sur ChromaDB. Décision : POC en 2 semaines." > notes.txt

#5. Run and observe the loop

Launch the agent
python agent.py

You should see a response containing both the calculation result (7,006,652) and a summary of the file. But it is more instructive to see what happens during execution. Add this verbose mode to follow the loop step by step:

Verbose streaming mode
for evenement in agent.stream(
    {"messages": [("user", question)]},
    stream_mode="values",
):
    dernier = evenement["messages"][-1]
    dernier.pretty_print()
    print("---")

You'll observe the typical agent sequence: the model produces a call to calculate, receives the result, produces a call to read_file, receives the content, then generates the final answer. Three iterations for a single user question.

i
If the model does not call tools
Two common causes: (1) tool calling is not enabled for the model on the Ollama side—pull ollama pull qwen3.5:9b again to get the latest version. (2) The system prompt is too vague. Explicitly say "use the tools for calculations" instead of hoping it figures that out.

#Tips and troubleshooting

Context too short
By default, Ollama truncates to 2048 tokens. If your agent chains together several tools, it quickly overflows. Set num_ctx=8192 in ChatOllama(model="...", num_ctx=8192).
Model that hallucinates tools
If the agent invents function names, lower the temperature to 0 and rewrite the system prompt to explicitly list the available tools.
Infinite loop
Set a limit: create_react_agent(..., recursion_limit=10). Beyond that, the agent stops cleanly.
Too much latency
On a CPU, a 9B model runs at 5–10 tok/s. Switch to qwen3.5:4b (3.4 GB VRAM, 30+ tok/s on a modest GPU) if the quality remains acceptable for your use case.
Error: "context length exceeded"
The summary of a long file exceeds num_ctx. Add an intermediate tool that chunks the file, or increase num_ctx to 32768 if your VRAM can handle it.
→
Trace your agents with LangSmith
For serious debugging, LangSmith traces every call, every token, and every tool. It's free in development. Set LANGSMITH_TRACING=true and LANGSMITH_API_KEY in your environment, and you get a complete timeline. No data is sent if you don't define the key.

#Go further

You have an agent that calculates, reads, and reasons locally. Three natural directions to explore further:

Give it access to your documents
Pair the agent with a vector database so it can answer questions about an internal corpus — that’s exactly what the introduction to local RAG guide is about.
Live in the CLI for coding
Aider is a development agent that edits your files directly from the terminal. You can connect it to the same Ollama and use Qwen3-Coder 30B or Devstral for assisted editing.
Adjust the model quantization
If you find Qwen 3.5 9B Q4 too slow or not good enough, the quantization guide explains when to move to Q5_K_M or drop to a smaller model.
Did this guide help you?

Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.