Advanced 13 minAPI

Function calling and structured JSON outputs with Ollama

Function calling lets an LLM decide on its own to call a function in your code — check the weather, query a database, send an email — by returning the arguments in the correct format. Ollama function calling relies on two building blocks: the format parameter to guarantee valid JSON, and the API’s tools field to declare the available functions. This guide shows both in Python, which local models are actually reliable, and how to harden the whole setup with schema validation and retries.

By Mohamed Meguedmi·Update 2026-08-27·Tested on Windows, macOS, and Linux

#Why function calling locally

An LLM generates text, not actions. Function calling bridges that gap: instead of replying in prose, the model returns a structured object that says “call the get_meteo function with the city Paris.” Your code executes the function, retrieves the real result, then passes it back to the model, which writes the final response. This is the basic mechanism behind agents and assistants that interact with the outside world.

Locally, there are two challenges. First, you need to ensure the output is 100% parseable JSON — a chatty model that adds “Here's the JSON:” breaks your entire pipeline. Then you need to ensure the model chooses the right function with the right arguments, which becomes tricky with smaller models. Ollama handles both through its API, but with safeguards you need to know about.

Guaranteed JSON output
The format parameter constrains decoding: the model can produce only syntactically valid JSON, or even JSON conforming to a specific schema.
Function calling
The tools field declares functions in OpenAI format; the model returns tool_calls with the arguments to pass.
100% local
Everything runs on your machine through the Ollama daemon at http://localhost:11434, with no API key or data leakage.

#Requirements and compatible models

The Local Copilot Kit

This guide gets you to the model. The kit gets you to the coding copilot in your editor.

  • Lifetime online access
  • PDF + files
  • Lifetime updates

JSON mode (the format parameter) works with any model. Function calling via tools, however, requires a model trained for tool use—otherwise the tool_calls field remains empty. Models aren't all equal: a “compatible” 3B often gets arguments wrong, while a 14B+ holds up on simple schemas.

Ollama installed
Daemon started and reachable at http://localhost:11434. Verify with ollama list.
Python SDK
pip install ollama pydantic — le client officiel plus Pydantic pour la validation.
Reliable tool-use model
qwen3.5:9b, mistral-small (24B), and gpt-oss:20b are excellent starting points in 2026. In Q4: Qwen 3.5 9B ≈ 6.6 GB, a 24B ≈ 14 GB of VRAM.
Recommended GPU
A RTX 3060 12GB runs Qwen 3.5 9B comfortably; target a 20–24B (RTX 4070/4080 16 GB) for truly reliable tool use.
i
JSON mode ≠ function calling
The format parameter guarantees valid JSON but doesn't trigger a function call: the model fills in an object that YOU interpret. The tools field, on the other hand, triggers genuine function selection by the model. The two are often combined.

#Force valid JSON with the format parameter

The simplest case: you want the model to always respond with JSON, never free-form text. Pass format: 'json' to the chat call. Ollama then constrains decoding token by token to produce a syntactically valid object. Important: keep an explicit instruction in the prompt describing the expected fields; otherwise, the model will invent a structure.

json_mode.py
import ollama
import json

resp = ollama.chat(
    model='qwen3.5:9b',
    messages=[{
        'role': 'user',
        'content': (
            "Extrais le nom, la ville et l'age de ce texte et reponds "
            "UNIQUEMENT en JSON avec les cles nom, ville, age. "
            "Texte : Marie, 34 ans, habite a Lyon."
        ),
    }],
    format='json',  # contraint la sortie a un JSON valide
    options={'temperature': 0},
)

data = json.loads(resp['message']['content'])
print(data)  # {'nom': 'Marie', 'ville': 'Lyon', 'age': 34}
→
Always temperature 0
For structured extraction, set temperature to 0. You want determinism and compliance, not creativity. This significantly reduces field hallucinations.

#Schema-based structured JSON (structured outputs)

Since late 2024, Ollama has also accepted a complete JSON schema in format (not just the string 'json'). Decoding is then constrained to follow the schema: types, required fields, and enumerations. This is much more robust than 'json' alone, because the model structurally cannot produce a nonconforming object. With Pydantic, you can generate the schema automatically.

structured_output.py
import ollama
from pydantic import BaseModel

class Personne(BaseModel):
    nom: str
    ville: str
    age: int

resp = ollama.chat(
    model='mistral-small',
    messages=[{'role': 'user',
               'content': 'Marie, 34 ans, habite a Lyon.'}],
    format=Personne.model_json_schema(),  # schema JSON complet
    options={'temperature': 0},
)

# validation stricte : leve une erreur si non conforme
personne = Personne.model_validate_json(resp['message']['content'])
print(personne)  # nom='Marie' ville='Lyon' age=34

Here, decoding is locked to the Person form, and model_validate_json runs another validation pass on the Python side. Double safety net: the output is guaranteed to be parseable AND conform to the declared types. This is the recommended pattern for any data extraction in local production.

#The tools API step by step in Python

Let's move on to real function calling. Declare the functions in the tools field using the OpenAI format (name, description, parameters in JSON Schema). The model reads these definitions and, if it decides calling a function would be useful, returns one or more tool_calls instead of a text message. You execute the function and return the result.

  1. 01
    Describe the functions
    For each function, provide a clear name, a precise description (the model uses it to choose), and a parameters object in JSON Schema listing the arguments and which ones are required.
  2. 02
    Send the call with tools
    Pass the tools list to ollama.chat. The model decides on its own whether to call a function or respond directly.
  3. 03
    Read tool_calls
    Inspect resp['message'].get('tool_calls'). If it is present, the model wants to call a function with the provided arguments.
  4. 04
    Run and return
    Call the actual Python function, then pass its result back to the model in a message with the 'tool' role so it can write the final answer.
tools_definition.py
def get_meteo(ville: str) -> str:
    # ici un vrai appel API ; on simule
    return f"Il fait 22 C et ensoleille a {ville}."

tools = [{
    'type': 'function',
    'function': {
        'name': 'get_meteo',
        'description': "Renvoie la meteo actuelle d'une ville donnee.",
        'parameters': {
            'type': 'object',
            'properties': {
                'ville': {
                    'type': 'string',
                    'description': 'Nom de la ville, ex: Paris',
                },
            },
            'required': ['ville'],
        },
    },
}]

#The call → execution → response loop

Function calling is a round trip. First call: the model returns a tool_call. You execute the function. Second call: you return the result, and the model writes the response in natural language. Here is the complete loop, reusable for multiple functions.

boucle_tools.py
import ollama

dispatch = {'get_meteo': get_meteo}

messages = [{'role': 'user',
             'content': 'Quel temps fait-il a Marseille ?'}]

resp = ollama.chat(model='mistral-small',
                   messages=messages, tools=tools)
msg = resp['message']
messages.append(msg)

for call in msg.get('tool_calls') or []:
    fn = call['function']['name']
    args = call['function']['arguments']
    resultat = dispatch[fn](**args)  # execution reelle
    messages.append({
        'role': 'tool',
        'name': fn,
        'content': resultat,
    })

# second appel : le modele redige la reponse finale
final = ollama.chat(model='mistral-small', messages=messages)
print(final['message']['content'])
!
Never blindly run arguments
The model controls the function name and its arguments. Use a dispatch dictionary (allowlist) rather than eval or dynamic getattr, and validate every argument before execution. A compromised or hallucinating model must not be able to call anything it wants.

#Schema validation and retry patterns

Locally, small models sometimes get things wrong: missing arguments, wrong types, nonexistent functions. Never trust the raw output. Wrap every tool_call in Pydantic validation, and if it fails, retry with the error message in context—often the model corrects itself on the second attempt.

retry_validation.py
from pydantic import BaseModel, ValidationError

class MeteoArgs(BaseModel):
    ville: str

def valider_appel(call):
    fn = call['function']['name']
    if fn not in dispatch:
        raise ValueError(f"Fonction inconnue: {fn}")
    args = MeteoArgs.model_validate(call['function']['arguments'])
    return fn, args

def appel_avec_retry(messages, max_essais=3):
    for essai in range(max_essais):
        resp = ollama.chat(model='mistral-small',
                           messages=messages, tools=tools)
        try:
            calls = resp['message'].get('tool_calls') or []
            return [valider_appel(c) for c in calls], resp
        except (ValidationError, ValueError) as e:
            messages.append({
                'role': 'user',
                'content': f"Erreur: {e}. Corrige et reessaie.",
            })
    raise RuntimeError('Echec apres retries')
Validate before running
A Pydantic model for each function catches missing or incorrectly typed arguments before they reach your code.
Retry with feedback
Feeding the error message back into the context guides the model toward correction. 2-3 attempts are almost always enough.
Function allowlist
Reject any function name outside dispatch. This is both a security measure and a safeguard against hallucinations.
Graceful fallback
After N failures, respond to the user with a clear message instead of crashing—especially with a small model.

#The pitfalls of small models in tool use

Tool use is cognitively demanding: the model must understand the intent, choose the right function, map the arguments, and follow the format. Below 7B, results are fragile. Here's what most often breaks locally and how to fix it.

empty tool_calls
The model responds with text instead of calling the function. Often this means a model was not trained for tool use, or that the function description is too vague. Switch to Qwen 3.5 or Mistral Small, and make the descriptions precise.
Bad arguments
The model invents or forgets fields. Make them required in the schema, reduce the number of functions exposed at once, and always validate.
Hallucinated function
The model calls a function that doesn’t exist. A whitelist is required on the dispatch side.
Polluted JSON
Without a format, a small model adds text around the JSON. Always use format='json' or a schema for pure extraction.
Too many functions
Beyond 5-6 tools, small models get lost. Segment by subtask or use two-stage routing.
→
The right local compromise
For reliable function calling without a high-end GPU, mistral-small (24B) in Q4 (≈14 GB VRAM) is often the best quality-to-resource ratio on a RTX 4080. Below that, qwen3.5:9b (≈6.6 GB) handles a few well-described functions, and gpt-oss:20b is a very fast alternative. For a strongly agent-oriented workload, glm-4.7-flash (MoE 30B-A3B, ≈19 GB) excels if you have 24 GB of VRAM.

#Go further

Function calling is the foundation of agents and advanced integrations. These site guides build on this one:

Integrate Ollama via the REST API in Python
The OpenAI-compatible endpoint on :11434, streaming, and JSON mode in a real FastAPI/Flask app.
Create a local AI agent with LangChain and Ollama
Move from raw function calling to a complete agent that chains tools, memory, and reasoning.
MCP and local LLMs: connect MCP servers to Ollama
Standardize access to tools (files, web, databases) via the Model Context Protocol instead of defining every function by hand.
Did this guide help you?

Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.