Building a local AI agent in Python with LangChain and Ollama
A local AI agent in Python with LangChain and Ollama is more than a chatbot: it is a program that decides on its own when to call a function, read a file, or chain multiple steps to respond. This guide builds a working agent step by step in about twenty minutes, using an Qwen 3.5 9B model that runs entirely on your machine. Zero API keys, zero data sent to a third party.
#Why run a local AI agent in Python?
An agent, in the LangChain sense, is a simple loop: the LLM receives a question and its list of tools, chooses whether to call one, reads the result, and repeats until it can answer. The entire “decision-making” mechanism relies on the model’s ability to emit a structured tool call.
Doing this locally with Ollama changes two concrete things: your data never leaves the machine, and every call costs zero euros. That’s the difference between prototyping with OpenAI and ending the week with a €50 bill, and iterating without counting the cost.
- Privacy
- The files the agent reads (contracts, proprietary code, medical notes) do not leave the workstation. No DPA to sign, no transfer outside the EU.
- Zero marginal cost
- Once the model is downloaded, you can iterate hundreds of times a day without your bill climbing.
- Reproducibility
- You pin the model's exact version (qwen3.5:9b, granite4.2:8b, etc.). No silent drift like with gpt-4o-2024-11-20, which becomes something else a month later.
- Predictable latency
- No network round trip. On a decent GPU, the first token arrives in under a second.
#Prerequisites
This guide gets you to the model. The kit gets you to the coding copilot in your editor.
- Lifetime online access
- PDF + files
- Lifetime updates
- Python 3.10+
- LangChain is no longer tested on 3.9. Check with python --version.
- Ollama installed and started
- It needs to listen on http://localhost:11434. See the Ollama installation guides (Windows, macOS, Linux) if you haven't already.
- A model that can call tools
- Not all LLMs can do this. Qwen 3.5, Granite 4.2, Gemma 4, Devstral, and GLM 4.7 Flash natively support tool calling. Avoid models that are now outdated (Llama 2/3, Qwen 2.5, Mistral 7B).
- Hardware
- Qwen 3.5 9B Q4 weighs about 6.6 GB in VRAM. An 8 GB GPU (RTX 3060, 4060) is sufficient; a 12 GB one (4070) provides headroom. On Mac, plan on 16 GB of unified memory for comfortable operation.
#1. Initialize the Python project
A virtual environment, three packages, and that's it. We avoid installing LangChain in system Python — it changes quickly and creates clutter.
- langchain
- The core: prompt abstractions, tools, messages.
- langchain-ollama
- The official integration Ollama. Maintained by the LangChain team since 2024.
- langgraph
- For the agent loop. This is the recommended engine today, more stable than the older AgentExecutors.
#2. Connect to Ollama from Python
Before setting up an agent, verify that you're actually talking to the model. Download the model if you haven't already, then test the simplest possible call.
The download is about 6.6 GB in Q4_K_M (the default quantization in Ollama). Once it's in place, create the first script:
If you see a coherent sentence, the Python ↔ Ollama connection is working. If you get a ConnectionError, check that Ollama is running properly (ollama ps should list an active service).
#3. Define the agent's tools
A LangChain tool is simply a Python function decorated with @tool. The docstring becomes the description the LLM sees—it uses it to decide when to call the tool. Be precise: a vague docstring produces random calls.
We will create two representative tools: an arithmetic-expression evaluator and a file reader.
Three rules for tools the model uses correctly:
- Explicit name
- calculate rather than process, read_file rather than get. The LLM chooses based first on the name.
- Detailed docstring
- Describe what the tool does, what it expects, and what it returns. LangChain reads the Python type annotations and exposes them to the model.
- Return a string
- Always. If the function returns a dict or an object, LangChain serializes it, but you lose readability on the model side.
#4. Assemble the agent
We have an LLM and we have tools. langgraph's create_react_agent connects the two and manages the loop: as long as the model wants to call tools, we continue; when it responds with text, we stop.
Create a small notes.txt file next to it for testing:
#5. Run and observe the loop
You should see a response containing both the calculation result (7,006,652) and a summary of the file. But it is more instructive to see what happens during execution. Add this verbose mode to follow the loop step by step:
You'll observe the typical agent sequence: the model produces a call to calculate, receives the result, produces a call to read_file, receives the content, then generates the final answer. Three iterations for a single user question.
#Tips and troubleshooting
- Context too short
- By default, Ollama truncates to 2048 tokens. If your agent chains together several tools, it quickly overflows. Set num_ctx=8192 in ChatOllama(model="...", num_ctx=8192).
- Model that hallucinates tools
- If the agent invents function names, lower the temperature to 0 and rewrite the system prompt to explicitly list the available tools.
- Infinite loop
- Set a limit: create_react_agent(..., recursion_limit=10). Beyond that, the agent stops cleanly.
- Too much latency
- On a CPU, a 9B model runs at 5–10 tok/s. Switch to qwen3.5:4b (3.4 GB VRAM, 30+ tok/s on a modest GPU) if the quality remains acceptable for your use case.
- Error: "context length exceeded"
- The summary of a long file exceeds num_ctx. Add an intermediate tool that chunks the file, or increase num_ctx to 32768 if your VRAM can handle it.
#Go further
You have an agent that calculates, reads, and reasons locally. Three natural directions to explore further:
- Give it access to your documents
- Pair the agent with a vector database so it can answer questions about an internal corpus — that’s exactly what the introduction to local RAG guide is about.
- Live in the CLI for coding
- Aider is a development agent that edits your files directly from the terminal. You can connect it to the same Ollama and use Qwen3-Coder 30B or Devstral for assisted editing.
- Adjust the model quantization
- If you find Qwen 3.5 9B Q4 too slow or not good enough, the quantization guide explains when to move to Q5_K_M or drop to a smaller model.
Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.