Function calling and structured JSON outputs with Ollama
Function calling lets an LLM decide on its own to call a function in your code — check the weather, query a database, send an email — by returning the arguments in the correct format. Ollama function calling relies on two building blocks: the format parameter to guarantee valid JSON, and the API’s tools field to declare the available functions. This guide shows both in Python, which local models are actually reliable, and how to harden the whole setup with schema validation and retries.
#Why function calling locally
An LLM generates text, not actions. Function calling bridges that gap: instead of replying in prose, the model returns a structured object that says “call the get_meteo function with the city Paris.” Your code executes the function, retrieves the real result, then passes it back to the model, which writes the final response. This is the basic mechanism behind agents and assistants that interact with the outside world.
Locally, there are two challenges. First, you need to ensure the output is 100% parseable JSON — a chatty model that adds “Here's the JSON:” breaks your entire pipeline. Then you need to ensure the model chooses the right function with the right arguments, which becomes tricky with smaller models. Ollama handles both through its API, but with safeguards you need to know about.
- Guaranteed JSON output
- The format parameter constrains decoding: the model can produce only syntactically valid JSON, or even JSON conforming to a specific schema.
- Function calling
- The tools field declares functions in OpenAI format; the model returns tool_calls with the arguments to pass.
- 100% local
- Everything runs on your machine through the Ollama daemon at http://localhost:11434, with no API key or data leakage.
#Requirements and compatible models
This guide gets you to the model. The kit gets you to the coding copilot in your editor.
- Lifetime online access
- PDF + files
- Lifetime updates
JSON mode (the format parameter) works with any model. Function calling via tools, however, requires a model trained for tool use—otherwise the tool_calls field remains empty. Models aren't all equal: a “compatible” 3B often gets arguments wrong, while a 14B+ holds up on simple schemas.
- Ollama installed
- Daemon started and reachable at http://localhost:11434. Verify with ollama list.
- Python SDK
- pip install ollama pydantic — le client officiel plus Pydantic pour la validation.
- Reliable tool-use model
- qwen3.5:9b, mistral-small (24B), and gpt-oss:20b are excellent starting points in 2026. In Q4: Qwen 3.5 9B ≈ 6.6 GB, a 24B ≈ 14 GB of VRAM.
- Recommended GPU
- A RTX 3060 12GB runs Qwen 3.5 9B comfortably; target a 20–24B (RTX 4070/4080 16 GB) for truly reliable tool use.
#Force valid JSON with the format parameter
The simplest case: you want the model to always respond with JSON, never free-form text. Pass format: 'json' to the chat call. Ollama then constrains decoding token by token to produce a syntactically valid object. Important: keep an explicit instruction in the prompt describing the expected fields; otherwise, the model will invent a structure.
#Schema-based structured JSON (structured outputs)
Since late 2024, Ollama has also accepted a complete JSON schema in format (not just the string 'json'). Decoding is then constrained to follow the schema: types, required fields, and enumerations. This is much more robust than 'json' alone, because the model structurally cannot produce a nonconforming object. With Pydantic, you can generate the schema automatically.
Here, decoding is locked to the Person form, and model_validate_json runs another validation pass on the Python side. Double safety net: the output is guaranteed to be parseable AND conform to the declared types. This is the recommended pattern for any data extraction in local production.
#The tools API step by step in Python
Let's move on to real function calling. Declare the functions in the tools field using the OpenAI format (name, description, parameters in JSON Schema). The model reads these definitions and, if it decides calling a function would be useful, returns one or more tool_calls instead of a text message. You execute the function and return the result.
- 01Describe the functionsFor each function, provide a clear name, a precise description (the model uses it to choose), and a parameters object in JSON Schema listing the arguments and which ones are required.
- 02Send the call with toolsPass the tools list to ollama.chat. The model decides on its own whether to call a function or respond directly.
- 03Read tool_callsInspect resp['message'].get('tool_calls'). If it is present, the model wants to call a function with the provided arguments.
- 04Run and returnCall the actual Python function, then pass its result back to the model in a message with the 'tool' role so it can write the final answer.
#The call → execution → response loop
Function calling is a round trip. First call: the model returns a tool_call. You execute the function. Second call: you return the result, and the model writes the response in natural language. Here is the complete loop, reusable for multiple functions.
#Schema validation and retry patterns
Locally, small models sometimes get things wrong: missing arguments, wrong types, nonexistent functions. Never trust the raw output. Wrap every tool_call in Pydantic validation, and if it fails, retry with the error message in context—often the model corrects itself on the second attempt.
- Validate before running
- A Pydantic model for each function catches missing or incorrectly typed arguments before they reach your code.
- Retry with feedback
- Feeding the error message back into the context guides the model toward correction. 2-3 attempts are almost always enough.
- Function allowlist
- Reject any function name outside dispatch. This is both a security measure and a safeguard against hallucinations.
- Graceful fallback
- After N failures, respond to the user with a clear message instead of crashing—especially with a small model.
#The pitfalls of small models in tool use
Tool use is cognitively demanding: the model must understand the intent, choose the right function, map the arguments, and follow the format. Below 7B, results are fragile. Here's what most often breaks locally and how to fix it.
- empty tool_calls
- The model responds with text instead of calling the function. Often this means a model was not trained for tool use, or that the function description is too vague. Switch to Qwen 3.5 or Mistral Small, and make the descriptions precise.
- Bad arguments
- The model invents or forgets fields. Make them required in the schema, reduce the number of functions exposed at once, and always validate.
- Hallucinated function
- The model calls a function that doesn’t exist. A whitelist is required on the dispatch side.
- Polluted JSON
- Without a format, a small model adds text around the JSON. Always use format='json' or a schema for pure extraction.
- Too many functions
- Beyond 5-6 tools, small models get lost. Segment by subtask or use two-stage routing.
#Go further
Function calling is the foundation of agents and advanced integrations. These site guides build on this one:
- Integrate Ollama via the REST API in Python
- The OpenAI-compatible endpoint on :11434, streaming, and JSON mode in a real FastAPI/Flask app.
- Create a local AI agent with LangChain and Ollama
- Move from raw function calling to a complete agent that chains tools, memory, and reasoning.
- MCP and local LLMs: connect MCP servers to Ollama
- Standardize access to tools (files, web, databases) via the Model Context Protocol instead of defining every function by hand.
Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.