Mastering Tool Calling with Ollama and Python
Tool calling turns a model that merely generates text into an agent capable of triggering real code: calling a weather API, querying a database, or running a calculation. This guide shows how to master tool calling Ollama in Python from end to end—tool JSON format, execution loop, tool-call streaming (0.17 series), and JSON Schema-constrained structured outputs applied directly during decoding. Everything runs locally on http://localhost:11434, with no API key or data leakage.
#Why tool calling?
A standalone LLM knows nothing about the real world after training: it knows neither today's weather, nor an account balance, nor the contents of your database. Tool calling fills that gap. You describe a list of available functions to the model; it decides which ones to call and with what arguments, your code executes them, then returns the result to the model so it can write an informed response.
The crucial point to understand: the model never executes anything itself. It only produces a structured request — “call get_meteo with city='Lyon'.” Your Python program executes the function and retains full control. This separation is what makes tool calling safe and predictable.
- Fresh data
- The model queries a real-time API instead of guessing from its training memories.
- Concrete actions
- Create a ticket, send an email, write to a database — the LLM orchestrates, your code acts.
- Reliability
- Calculations and exact lookups are delegated to deterministic code, not hallucinated by the model.
- 100% local
- With Ollama, the entire pipeline stays on your machine: no API key, no outbound request, and no per-token bill.
#How the Ollama tool call works
This guide gets you to the model. The kit gets you to the coding copilot in your editor.
- Lifetime online access
- PDF + files
- Lifetime updates
The complete cycle takes five steps. Visualizing it clearly avoids the most common misunderstanding: thinking a single call is enough. You need at least two — one to obtain the tool request, and one to obtain the final response.
- 01You send the question and the toolsThe chat request contains the user's message and the list of available tools (the tools parameter).
- 02The model returns a tool requestInstead of responding with text, it returns one or more tool_calls with the function name and arguments.
- 03Your code executes the functionYou retrieve name and arguments, call the corresponding real Python function, and get a result.
- 04You return the resultThe result is added to the history as a message with the «tool» role, then you run chat again.
- 05The model writes the final answerBuoyed by the result, it now produces a natural-language response for the user.
#Prerequisites
Three building blocks: the running Ollama daemon, a model that genuinely supports tools, and the official Python library. Pay attention to the second point—all models cannot perform tool calling. Target recent families designed for it.
- Ollama up to date
- Version 0.17 or later to take advantage of tool-call streaming. The daemon listens on http://localhost:11434.
- A tool-compatible model
- Qwen 3.5, Granite 4.2, Mistral Small, Devstral, gpt-oss. Models marked “tools” on ollama.com/library.
- Enough VRAM
- A small model (Qwen 3.5 4B ≈ 3.4 GB) is enough for testing; a Granite 4.2 8B (≈ 5.3 GB) or a Qwen 3.5 9B (≈ 6.6 GB) follows multi-tool instructions more reliably. RTX 3060 12 GB as an entry-level option.
- The ollama library
- pip install -U ollama. It can build a tool's schema directly from a typed Python function.
#The JSON format for tools
A tool is described with a JSON schema strictly aligned with OpenAI's: an object with type: "function" containing a name, a description, and a parameters object in JSON Schema format. The description matters enormously—it is what the model reads to decide when and how to call the tool. Be explicit.
#First tool call in Python
Let's start with the simplest case: one function, one question, and we observe what the model decides. Type annotations and the docstring are used to generate the schema sent to the model.
At this point, message.content is generally empty: the model returned its request in message.tool_calls. Each tool_call exposes function.name (a string) and function.arguments (already deserialized into a Python dictionary by the library). All that's left is to execute it and return the result.
#The complete execution loop
Here is the reusable skeleton of a tool-calling agent: a dictionary mapping each tool name to its function, execution of the requested calls, adding the results to the history, and then a second call for the final response. Wrap the whole thing in a loop to handle cases where the model chains several tools.
The result message uses the tool role and a tool_name field indicating which call it answers. content must be a string: serialize your objects (json.dumps) before returning them. The model rereads this content as if it were an observation of the world.
#OpenAI parity: the same code with the openai client
Ollama exposes an OpenAI-compatible endpoint at /v1. If your code already uses the openai client, you need to change almost nothing: point base_url to Ollama and use a dummy API key. The tool and tool_calls format is identical—that is the “OpenAI parity” that makes migration trivial.
#Tool-call streaming (0.17 series)
Historically, enabling streaming disabled tool calling: you had to choose. Since the 0.17 series, Ollama has been able to stream tool calls as generation proceeds. In practice, you receive tool_calls in the stream’s chunks, along with any text, allowing you to display a fluid response while triggering tools.
#Structured outputs: enforce a JSON Schema
Tool calling is for taking action; structured outputs guarantee the response format. With the format parameter, you pass a JSON Schema that Ollama applies during decoding: token by token, the model is constrained to produce only output that is valid under the schema. No more fragile parsing of approximate JSON—the structure is guaranteed by construction.
In Python, the most convenient approach is to describe the structure with a Pydantic model, then generate the schema with model_json_schema(). You then get a typed and validated object.
#Error handling and common pitfalls
Tool calling rarely fails loudly; more often, the model silently “goes off the rails.” Here are the common failure modes and how to handle them.
- No tool_call returned
- The model responded with text when a tool was required. Improve the tool description, or switch models: small models often miss the decision to call.
- Missing or incorrect arguments
- function.arguments may omit a required field or type it incorrectly. Validate with Pydantic or a try/except before calling the actual function, and return the error to the model as a tool result.
- Hallucinated tool
- The model invents a nonexistent function name. That's why OUTILS.get(name) returns an error message instead of crashing—the model can then correct itself on the next turn.
- Infinite tool loop
- A model can repeatedly request the same tool forever. Add a maximum iteration counter (e.g., 5) to stop the loop and prevent it from running pointlessly.
- Truncated context
- Ollama sometimes limits the context to 2,048 tokens by default, which crushes tool history during long sessions. Increase num_ctx through the model options.
- Unserialized result
- Returning a raw Python object as content breaks the request. Always serialize it to a string (json.dumps or str) before adding it to the messages.
The key idea: never let a tool error crash the agent. Return the error message to the model as if it were a result. A good model reads “unknown tool” or “missing argument” and adjusts its next call on its own.
#Go further
You now know how to do tool calling Ollama in Python: JSON tool schemas, execution loops, OpenAI parity, streaming, and schema-constrained structured outputs. These guides naturally build on the topic.
- The Ollama REST API
- “Integrate Ollama into a Python application via the REST API” — the basics of endpoint :11434, streaming, and JSON mode, the foundation of this entire guide.
- Agents with LangChain
- “Create a local AI agent in Python with LangChain and Ollama” — orchestrate multiple tools and memory on top of raw tool calling.
- Choose your quantization
- “Choosing your quantization (Q4, Q5, Q8, FP16)”—to balance VRAM and the quality of the model that will control your tools.
Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.