How to use tool calling with Ollama
L'ollama tool call represents a significant advance in turning language models into agents capable of interacting with the outside world. This feature allows an LLM, hosted locally via Ollama, to determine when and how to use external functions (tools) to answer a complex request. In this guide, we will explore in detail the theory behind this capability, practical implementation mechanisms, and how to optimize the use of high-performing open-weights models on your local infrastructure. We will also cover the specifics of the different models available in our catalog to maximize your chances of success with the tool calling.
Understanding Tool Calling
Le tool calling is a sophisticated capability in which an LLM does not merely generate text, but instead produces a specific data structure (usually JSON) indicating that a function should be executed. The process follows an iterative cycle:
- User Query: The user asks a question requiring external information (e.g., “What will the weather be like in Paris tomorrow?”).
- LLM analysis: The model, having been trained with knowledge of your available functions (the schema), determines that a function
get_weather(city)is appropriate. It does not respond directly, but generates a call in the expected format. - External Execution: Your host system intercepts this JSON output and executes the actual function (a weather API call).
- Result Return: The execution result (
{"temperature": "15°C", "condition": "ensoleillé"}) is returned to the LLM as a new context. - Final Generation: The LLM uses this factual result to formulate a natural, complete response to the user.
This architecture bridges the gap between the model's static knowledge and the dynamic data required by real-world applications. For tasks requiring complex reasoning, you might consider using powerful models such as DeepSeek V4 Pro 1.6T spec sheet DeepSeek V4 Pro 1.6T or MiMo V2.5 Pro MiMo V2.5 Pro sheet, whose extended contexts make it easier to manage complex tool schemas, reaching up to 1 000 000 tokens of context.
Choosing the right model for the LLM Agent: Performance vs. Size
The success of theollama tool call inherently depends on the model's ability to follow structured instructions and reason logically. Open-weight models vary considerably in complexity, which directly affects their performance in tool calling.
For tasks requiring high fidelity in tool invocation (where JSON accuracy is critical), favor models with strong instruction alignment. For example, Kimi K2.7 Code Kimi K2.7 Code spec sheet or Qwen3-Coder-Next 80B-A3B Qwen3-Coder-Next 80B-A3B sheet demonstrate a strong aptitude for structured reasoning, which is a prerequisite for tool calling fiable.
On the other hand, if your hardware constraints are severe (for example, running on systems with less than 16GB of VRAM), you’ll need to choose smaller models that still deliver strong performance. Models such as Mistral Medium 3.5 128B Mistral Medium 3.5 128B spec sheet or Llama 4 Scout 109B Llama 4 Scout 109B sheet offer a good compromise between reasoning capacity and memory footprint, enabling effective testing of the tool calling without overloading the GPU.
Technical implementation: The role of the system prompt
How you present tools to the LLM is crucial. You must provide a clear description (the schema) for each available function, including its expected parameters and data type. This documentation must be integrated into the system prompt or provided as context before the user query.
A good example of a tool definition would be:
"description": "Récupère les informations sur un modèle spécifique.", "parameters": {"type": "object", "properties": {"model_name": {"type": "string", "description": "Le nom exact du LLM à rechercher (ex: DeepSeek V4 Pro 1.6T)"}}}
The effectiveness of this description is directly tied to the model’s instruction-tuning level. Models trained on datasets rich in examples of function calling perform better. To compare capabilities, you can consult our LLM comparison guide. We also recommend studying the performance of Inkling Inkling datasheet which has a very large context, ideal for maintaining tool definitions across multiple exchanges.
Managing streams and iterations with Ollama
The implementation of the tool calling requires an interaction loop between your application and the Ollama API. It is essential to manage the following states:
- Initialization: Send the system prompt containing the tool definitions to the LLM via Ollama.
- Output Verification: Analyze the LLM's response. If it contains a tool call (predefined JSON format), proceed to step 3. Otherwise, it is the final response.
- Execution and Injection: Execute the code corresponding to the call, then return the raw result to the LLM in a new request, specifying that it is the result of the tool.
- Finalization: The LLM uses this result to generate the final response for the user.
Models such as GLM 5.2 753B-A40B GLM 5.2 753B-A40B sheet or Mixtral 8x22B Instruct Mixtral 8x22B Instruct sheet are excellent for this iterative reasoning phase because they maintain high consistency between call-and-response steps. For more details on prompt optimization, see our article on prompt engineering.
Concrete use cases for local LLMs
Le tool calling opens the door to autonomous applications based on your open-weights models:
- Specialized Research Agents: An agent can use a web search tool (if you provide one) to find the latest information about a technology, then synthesize that data using a model such as DeepSeek V3.2 DeepSeek V3.2 sheet.
- Contextual Code Assistants: An LLM can call tools to read local documentation or run unit tests on a codebase before proposing a fix using a specialized model such as Kimi K2.5 Kimi K2.5 sheet.
- Entity Management Systems: To extract and classify complex data from documents, the tool can be a classification or schema-validation function. Models like Qwen 3 VL 235B-A22B Qwen 3 VL 235B-A22B spec sheet are particularly well suited to this type of multimodal, structured analysis.
To compare the performance of different models in specific scenarios, we invite you to use our comparison tool.
FAQ on Tool Calling with Ollama
Q: What is the main prerequisite for my LLM supported by Ollama to perform tool calling?
A: The fundamental prerequisite is not just the model’s native capability, but above all the quality of your system prompt. You must describe the available functions (name, description, expected parameters) in a structured format (often JSON Schema) that the LLM must learn to reproduce in its output when it determines that a tool is needed to respond.
Q: Do all open-weight models natively support tool calling?
A: No, it depends on how the model was trained and fine-tuned. Versions optimized or specialized for instruction alignment are much more likely to generate reliable tool calls. Models such as Nemotron 3 Ultra Nemotron 3 Ultra sheet show a strong ability to follow these complex patterns.
Q: What impact does model size have on tool-call success?
A: Generally, the larger the model (in parameters), the better its reasoning ability and its fidelity to the JSON format required for invocation. However, a well-aligned medium-sized model can outperform a very large model that has been poorly fine-tuned for this specific task.
Q: How should errors be handled when running an external tool?
A: When your code runs a tool and encounters an error (e.g., API unavailable), you must send this error message back to the LLM as context. The model must then be able to recognize the error and, if possible, suggest a corrective action or inform the user that the task cannot be completed.
Q: What role does context (the context window) play in the tool calling?
A: A large context allows the LLM to "remember" the complex definitions of all your tools, the complete conversation history, and intermediate results at the same time. Models with an extended context such as DeepSeek V4 Pro 1.6T spec sheet DeepSeek V4 Pro 1.6T are particularly advantageous for multi-step agents.
Conclusion: Toward Autonomous Local LLM Agents
Mastering theollama tool call is the key to evolving a simple text generator into a true agent capable of acting in your local or remote environment. By combining the right choice of a capable model — whether GLM-5 744B-A40B GLM 5 744B-A40B sheet for raw power, or a lighter model such as MiMo V2 Flash MiMo V2 Flash sheet for efficiency—with rigorous prompt engineering, you build robust systems. To explore each LLM's specific capabilities in this area or configure your ideal environment, see our full catalog or our hardware configurator.
External Sources Used for Enrichment (Minimum 3 Required) General guide to Function Calling Ollama API documentation Academic research on LLM agents
The hardware for running an LLM locally
To run these models comfortably locally, a RTX 5070 Ti offers an excellent price/performance ratio. Compare prices:
Affiliate links — BestLLMfor may earn a commission on purchases, at no extra cost to you. As an Amazon Associate, BestLLMfor earns from qualifying purchases.