Integrate Ollama into a Python application via the API REST
Ollama exposes two HTTP APIs on port 11434: a native API (/api/generate, /api/chat) and an OpenAI-compatible API (/v1/chat/completions). The latter is the preferred route for Ollama API Python integration: your code uses exactly the same SDK as with GPT-4, but runs on your machine. This guide covers practical patterns—streaming, structured JSON, and function calling—with FastAPI and Flask examples ready to paste into a project.
By Mohamed Meguedmi·Update 2026-08-27·Tested on Windows, macOS, and Linux
#Why use the REST API
The ollama run est CLI is convenient for testing, but it is not designed to be called from an application. The REST API is: it is designed for that purpose, with standard HTTP requests, JSON input and output, and streaming via Server-Sent Events. This is what all interfaces (Open WebUI, Cline, LangChain) use under the hood.
OpenAI compatibility
The /v1/chat/completions endpoint accepts exactly the same payload as api.openai.com/v1/chat/completions. Change the URL and key, and your existing code works.
No reinvention
The official openai Python SDK (or any HTTP client) talks directly to Ollama. There is no specific client to learn.
Runtime decoupling
Your Python app runs in its container, and Ollama runs in its own. When you switch to vLLM or LM Studio, you only change the base_url.
Simultaneous multi-client use
Several Python scripts, a Jupyter notebook, and Open WebUI can all hit the same Ollama instance. The daemon manages the queue on its own.
i
Native API vs. OpenAI-compatible
Ollama supports both. The native API (/api/chat) exposes specific parameters (num_ctx, num_predict, mirostat) but is less portable. The OpenAI-compatible API covers 95% of needs and remains usable with any other provider. By default, start with this one.
#Prerequisites
✓
The Local Copilot Kit
This guide gets you to the model. The kit gets you to the coding copilot in your editor.
The daemon must listen on http://localhost:11434. Check with curl http://localhost:11434—you should see "Ollama is running".
Python 3.10+
Recent SDKs (openai 1.x) require at least Python 3.8, but 3.10+ for modern annotations.
A chat-compatible model
ollama pull qwen3.5:9b ou gemma4:12b. Pour le function calling, choisissez un modèle qui le supporte : Qwen 3.5, Granite 4.2, Mistral Small 24B, Devstral.
Sufficient VRAM
A 9B Q4 model (such as Qwen 3.5 9B) requires about 6–7 GB of VRAM, while a 12B model (Gemma 4 12B) requires about 8 GB. Without a GPU, it also runs, but at 5–10 tok/s.
#1. The two APIs on the Ollama side
Before writing Python, let's look at the endpoints from the terminal to see what's happening. With curl, we talk directly to the daemon, with no abstraction.
The latter returns a payload strictly identical to OpenAI's: choices[0].message.content, id, model, usage fields. That's what makes drop-in use possible.
→
The API key is ignored but required
The openai SDK requires an api_key parameter. Ollama does not verify anything — pass “ollama” or any non-empty string. If you enter your real OpenAI key out of habit, it stays on your machine, but prefer a neutral string to avoid confusion.
#2. The OpenAI SDK pointed at Ollama
The basic pattern for an Ollama Python API integration fits in five lines: install the OpenAI SDK, instantiate it with the local base_url, and call chat.completions.create as usual.
Installation
pip install openai
client.py — basic call
from openai import OpenAI
client = OpenAI(
base_url="http://localhost:11434/v1",
api_key="ollama", # ignoré, mais requis par le SDK
)
reponse = client.chat.completions.create(
model="qwen3.5:9b",
messages=[
{"role": "system", "content": "Tu réponds en français, de façon concise."},
{"role": "user", "content": "Explique en une phrase ce qu'est un LLM."},
],
temperature=0.3,
)
print(reponse.choices[0].message.content)
Run the script. If Ollama is running and the model has been downloaded, you'll get a sentence. If you see a ConnectionRefusedError, check with ollama ps that the daemon is active.
model
The exact name as listed by ollama list (qwen3.5:9b, gemma4:12b, mistral-small, etc.).
messages
The list of conversation turns. Supported roles: system, user, assistant, tool.
temperature
0 for deterministic output, 0.7 for creative output. For data extraction, stay at 0 or 0.1.
max_tokens
Upper limit for the response. Optional — Ollama applies a reasonable num_predict default.
#3. Token-by-token streaming with SSE
For a decent UX (chatbot, long-form generation), you want to display tokens as they arrive instead of waiting for completion. Ollama supports streaming via Server-Sent Events, and the OpenAI SDK makes it a trivial Python loop.
streaming.py
from openai import OpenAI
client = OpenAI(base_url="http://localhost:11434/v1", api_key="ollama")
flux = client.chat.completions.create(
model="qwen3.5:9b",
messages=[{"role": "user", "content": "Raconte une courte histoire de robot."}],
stream=True,
)
for chunk in flux:
delta = chunk.choices[0].delta.content
if delta:
print(delta, end="", flush=True)
print()
Each chunk contains a delta (the added piece of text). The last chunk has delta.content set to None and finish_reason populated—that is the stop signal.
!
Don't forget flush=True
Without flush=True, Python buffers stdout line by line and the streaming effect disappears in the terminal. For an HTTP API, however, the web server (uvicorn, gunicorn) flushes it—you do not need to handle it.
#4. JSON mode for structured outputs
When you need to parse the response (extraction, classification, payload generation), asking for "return JSON" in the prompt is not enough—the model often inserts text around it. JSON mode forces the decoder to output only valid JSON.
json_mode.py
from openai import OpenAI
import json
client = OpenAI(base_url="http://localhost:11434/v1", api_key="ollama")
reponse = client.chat.completions.create(
model="qwen3.5:9b",
messages=[
{"role": "system", "content": (
"Tu extrais des informations structurées. "
"Réponds uniquement avec un objet JSON contenant les clés : "
"nom (string), age (int), ville (string)."
)},
{"role": "user", "content": "Marie a 34 ans, elle habite à Lyon."},
],
response_format={"type": "json_object"},
temperature=0,
)
donnees = json.loads(reponse.choices[0].message.content)
print(donnees)
# {'nom': 'Marie', 'age': 34, 'ville': 'Lyon'}
response_format={"type": "json_object"} enables JSON mode. On the Ollama side, this translates into a constraint at the sampler level: any token that would produce invalid JSON is rejected. It's more reliable than prompting "respond in JSON" and praying.
→
Mention "JSON" in the prompt
As with OpenAI, JSON mode requires at least one mention of the word "JSON" in the conversation (system or user). Without it, some models produce an empty object. Describe the expected schema in the system prompt—that guides the content; JSON mode only guarantees the syntax.
#5. Function calling (tool use)
Function calling lets the model signal that it wants to call a Python function instead of responding directly. Not all models support it—check ollama.com/library to confirm that "tools" appears in the capabilities. Qwen 3.5, Granite 4.2, Mistral Small 24B, and Devstral support it natively.
tools.py
from openai import OpenAI
import json
client = OpenAI(base_url="http://localhost:11434/v1", api_key="ollama")
# 1. Une fonction Python réelle
def meteo(ville: str) -> dict:
# En vrai, vous appelleriez Open-Meteo ou autre
return {"ville": ville, "temperature_c": 18, "conditions": "nuageux"}
# 2. Sa description au format OpenAI
outils = [{
"type": "function",
"function": {
"name": "meteo",
"description": "Donne la météo actuelle d'une ville française.",
"parameters": {
"type": "object",
"properties": {
"ville": {"type": "string", "description": "Nom de la ville"},
},
"required": ["ville"],
},
},
}]
messages = [{"role": "user", "content": "Quel temps fait-il à Bordeaux ?"}]
# 3. Premier appel : le modèle décide d'appeler la fonction
reponse = client.chat.completions.create(
model="qwen3.5:9b",
messages=messages,
tools=outils,
)
appel = reponse.choices[0].message.tool_calls[0]
args = json.loads(appel.function.arguments)
resultat = meteo(**args)
# 4. Second appel : on renvoie le résultat au modèle pour la réponse finale
messages.append(reponse.choices[0].message)
messages.append({
"role": "tool",
"tool_call_id": appel.id,
"content": json.dumps(resultat),
})
finale = client.chat.completions.create(model="qwen3.5:9b", messages=messages)
print(finale.choices[0].message.content)
The loop has two rounds: the first returns a tool_calls (the model says "call meteo with city=Bordeaux"), and the second returns the natural-language response after you've executed the function and injected its result. In production, loop as long as tool_calls is populated.
!
Not all models are created equal
With a model that handles tools poorly (old Llama 2, Mistral 7B v0.1), you'll get malformed calls or hallucinated arguments. If this happens: (1) confirm that the model officially supports tools, (2) lower the temperature to 0, (3) simplify the parameter schema.
#6. Expose Ollama via FastAPI
Typical case: your frontend calls your Python backend, which calls Ollama. FastAPI handles asynchronous execution cleanly, and streaming reaches the browser through a StreamingResponse.
Dependencies
pip install fastapi uvicorn openai
main.py
from fastapi import FastAPI
from fastapi.responses import StreamingResponse
from pydantic import BaseModel
from openai import OpenAI
app = FastAPI()
client = OpenAI(base_url="http://localhost:11434/v1", api_key="ollama")
class Question(BaseModel):
message: str
model: str = "qwen3.5:9b"
@app.post("/chat")
def chat(q: Question):
reponse = client.chat.completions.create(
model=q.model,
messages=[{"role": "user", "content": q.message}],
)
return {"reponse": reponse.choices[0].message.content}
@app.post("/chat/stream")
def chat_stream(q: Question):
def generateur():
flux = client.chat.completions.create(
model=q.model,
messages=[{"role": "user", "content": q.message}],
stream=True,
)
for chunk in flux:
delta = chunk.choices[0].delta.content
if delta:
yield delta
return StreamingResponse(generateur(), media_type="text/plain")
Start the server
uvicorn main:app --reload --port 8000
Testing in the CLI
curl -N -X POST http://localhost:8000/chat/stream \
-H "Content-Type: application/json" \
-d '{"message": "Écris un haïku sur Paris."}'
curl's -N (--no-buffer) option disables client-side buffering so you can see the stream live. On the JS frontend, you read the fetch response's ReadableStream—the same as with the OpenAI API.
#7. Flask chatbot with history
For a complete chatbot, you need to retain the message history between turns. Here's a minimal Flask version that stores the conversation in memory (replace it with a real session/DB in production).
The larger the history grows, the more tokens you consume on every call. For Qwen 3.5, the default window on the Ollama side is 2048 tokens—beyond that, older messages are silently truncated. Increase it through the native API or by overriding it with a Modelfile (num_ctx 8192 or 32768).
#For production
Expose Ollama on the network
By default, the daemon listens only on 127.0.0.1. To allow other machines, run it with OLLAMA_HOST=0.0.0.0 — and put a reverse proxy with authentication in front of it; otherwise, anyone on the network can hammer your models.
Concurrency and queueing
Ollama serializes requests per model. To serve multiple users in parallel, run multiple instances or switch to vLLM, which natively handles dynamic batching.
Client-side timeouts
A request to a model that is not loaded can take 10-30 s (loading into VRAM). Set the OpenAI client timeout: OpenAI(..., timeout=120) rather than the library's 10-minute default, which is often too short on the reverse proxy side.
Keep the model warm
By default, Ollama unloads a model after 5 minutes of inactivity. In the API, pass keep_alive="30m" through the native API /api/chat, or send a periodic ping to avoid a cold start on the first user request.
Observability
Always log model, prompt_tokens, and completion_tokens (available in reponse.usage). These are your inference metrics—useful for spotting a model that is slowing down or a prompt that is ballooning.
→
Migrate from the OpenAI API
If you already have code that talks to api.openai.com, switching to Ollama takes two lines: change base_url="https://api.openai.com/v1" to base_url="http://localhost:11434/v1" and adapt the model name. Everything else—streaming, JSON mode, tools—works the same way. That's the main benefit of the OpenAI-compatible endpoint.
#Go further
You have the basic building blocks. Here are three directions to take things further, depending on your use case:
Build an agent that decides everything on its own
The guide to local AI agents in Python with LangChain takes function calling all the way to the full agent loop, with support for multiple tools and multistep reasoning.
Add RAG to your documents
To have your app respond from an internal corpus (PDFs, notes, code), connect a vector database. The introductory guide to local RAG lays the foundation.
Customize the model's behavior
Rather than repeating the system prompt on every call, create a variant through Modelfile. The Ollama Modelfile customization guide shows how to lock in an FR assistant or coding mode under a reusable model name.
Did this guide help you?
Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.