Advanced 14 minMCP

MCP and local LLMs: connect MCP servers to Ollama

The Model Context Protocol (MCP) standardizes how an LLM calls external tools: file reading, web queries, and database access. Combining MCP with Ollama lets you run an agent that can act on your machine without ever sending your data to a cloud API. This guide shows how to build an MCP-Ollama bridge in Python, which local models actually handle tool use, and where the real limits of small models lie.

By Mohamed Meguedmi·Update 2026-08-27·Tested on Windows, macOS, and Linux

#What is MCP, and why does it change local agents?

MCP (Model Context Protocol) is an open protocol published by Anthropic in late 2024. Its purpose is to provide a single interface between a language model and the tools it can use. Instead of recoding a custom integration for each data source, an “MCP server” exposes tools, resources, and prompts in a standard format. Any compatible client—Claude Desktop, an IDE, or your own bridge—can then connect to it.

In practice, an MCP “filesystem” server exposes tools such as read_file, write_file, or list_directory. A “sqlite” server exposes query or list_tables. The LLM never talks directly to the disk: it issues a tool-call request, the client executes it through the MCP server, then sends the result back to the model. This client/server separation is what makes the protocol reusable.

For local agents, the challenge is twofold: reuse the growing ecosystem of MCP servers (dozens already exist) while keeping inference 100% on your machine through Ollama. You get an agent that reads your files and queries your databases without a single byte leaving your network.

i
MCP ≠ function calling
MCP is not a new LLM API. It is a layer on top of tool use: the model still performs standard function calling; MCP only standardizes tool discovery and execution on the server side.

#Why connect MCP to Ollama

The Local Copilot Kit

This guide gets you to the model. The kit gets you to the coding copilot in your editor.

  • Lifetime online access
  • PDF + files
  • Lifetime updates

Most MCP demos use a cloud model (Claude, GPT). Running mcp ollama locally changes three things: privacy (your files and SQL queries do not leave your system), cost (zero tokens billed regardless of tool-call volume), and control (you choose the model, quantization, and authorized servers).

Privacy
An MCP filesystem server gives the model access to your folders. Locally, this content never passes through a third party.
Zero cost
Agents make many tool round trips. In the cloud, each turn costs tokens; with Ollama, it’s free.
Offline
Once the model is downloaded and the MCP servers are installed, everything works without an internet connection (except web servers, of course).
Sovereignty
You decide which tools are exposed and can audit every call before executing it.
!
MCP provides action capabilities
A filesystem or shell server lets the model write files or run commands. Always limit the scope (root folder, read-only database) and validate sensitive calls before execution.

#Prerequisites

Ollama installed
The daemon listens on http://localhost:11434 by default. Check with "ollama --version".
A tool-use model
Plan on at least one Granite 4.2 8B, ideally a Qwen 3.5 9B for reliability (see the next section).
Python 3.10+
The official MCP SDK and the Ollama client are written in Python.
Node.js (optional)
Many reference MCP servers are launched via npx (@modelcontextprotocol/server-*).
Terminal — prepare the environment
# Vérifier Ollama
ollama --version
curl http://localhost:11434/api/tags

# Tirer un modèle capable de tool-use
ollama pull qwen3.5:9b

# Environnement Python
python -m venv .venv
source .venv/bin/activate
pip install mcp ollama

#Which local models handle tool use correctly

Not all models are equally good at function calling. A model that “knows” the tool format but chooses its arguments poorly will make the agent unusable. In practice, the 2026 generation (Qwen 3.5, Granite 4.2) makes tool use reliable starting at 8–9B, whereas in 2024 you had to target 14B. Here are the benchmarks you can test with Ollama and their Q4_K_M VRAM footprint.

Qwen 3.5 4B / 9B
Excellent tool-use support. The 9B (≈6.6 GB of VRAM in Q4, 256k context, vision) is the best reliability/hardware compromise for a local agent.
Granite 4.2 8B
Solid native tool use with very low token usage (≈5.3 GB in Q4, 128k context). An excellent entry point for 6–8 GB of VRAM.
Mistral Small 24B
Solid function calling and good French-language performance (≈14 GB in Q4). Comfortable with somewhat complex tool schemas.
Qwen 3.6 35B-A3B
Very reliable MoE for call chains (≈23 GB in Q4, only 3B active, so it's fast). Best reserved for 24 GB GPUs such as RTX 4090.
2–3B models
Qwen 3.5 2B or Granite 4.2 3B fit in ≈2 GB, but tool use quickly falls off once there are multiple tools. Avoid them for a real agent.
→
Test tool use before coding
Before connecting MCP, verify that the model actually calls tools with a simple /api/chat test using a dummy tool. If the model returns text instead of a tool_call, switch models rather than debugging the bridge.

#The MCP-Ollama bridge in Python, step by step

Ollama is not a native MCP client. The bridge acts as an intermediary: it starts an MCP server, converts its tools to the format expected by the Ollama API, runs the tool-call loop, and then returns the results to the model. It uses the official MCP SDK (the mcp package) and the Ollama client.

  1. 01
    1. Start an MCP server
    We launch an MCP server as a subprocess via stdio. Here, it’s the official filesystem server, limited to a working directory passed as an argument.
  2. 02
    2. List and convert the tools
    session.list_tools() returns the MCP tools. We transform them into the “tools” format expected by /api/chat from Ollama (name, description, inputSchema → parameters).
  3. 03
    3. Tool-calling loop
    We send the user message and the list of tools. If the model responds with a tool_call, we execute it through MCP, inject the result, and loop until we get a final response.
  4. 04
    4. Return the response
    When the model no longer requests a tool, its last text response is the final result presented to the user.
bridge.py — MCP + Ollama
import asyncio
import ollama
from mcp import ClientSession, StdioServerParameters
from mcp.client.stdio import stdio_client

MODEL = "qwen3.5:9b"

# Serveur MCP filesystem limité au dossier ./workspace
server = StdioServerParameters(
    command="npx",
    args=["-y", "@modelcontextprotocol/server-filesystem", "./workspace"],
)

def to_ollama_tools(mcp_tools):
    return [{
        "type": "function",
        "function": {
            "name": t.name,
            "description": t.description,
            "parameters": t.inputSchema,
        },
    } for t in mcp_tools]

async def run(prompt: str):
    async with stdio_client(server) as (read, write):
        async with ClientSession(read, write) as session:
            await session.initialize()
            tools = (await session.list_tools()).tools
            ollama_tools = to_ollama_tools(tools)
            messages = [{"role": "user", "content": prompt}]

            while True:
                resp = ollama.chat(
                    model=MODEL,
                    messages=messages,
                    tools=ollama_tools,
                )
                msg = resp["message"]
                messages.append(msg)

                if not msg.get("tool_calls"):
                    return msg["content"]

                for call in msg["tool_calls"]:
                    fn = call["function"]
                    result = await session.call_tool(
                        fn["name"], fn.get("arguments", {}),
                    )
                    text = "".join(c.text for c in result.content
                                    if getattr(c, "text", None))
                    messages.append({
                        "role": "tool",
                        "content": text,
                    })

if __name__ == "__main__":
    print(asyncio.run(run("Liste les fichiers du dossier et résume leur contenu.")))
i
Loop safeguard
In production, add an iteration counter (max_iterations) to prevent a model looping on the same tool from running forever. About ten iterations is more than enough for most tasks.

#Examples of useful self-hosted MCP servers

MCP's value comes from its catalog of ready-to-use servers. Here are the ones that add the most value to a local agent, all launchable through npx or pip.

filesystem
Reading/writing files in a mandated root folder. Most useful for an agent working on your documents.
sqlite / postgres
Query a local database in natural language. Switch the connection to read-only mode to prevent any changes.
fetch
Retrieve and convert a web page to text. The only one that assumes an internet connection.
git
Explore a repository: log, diff, status. Useful for a code-review or code-documentation agent.
memory
A persistent key-value store that gives the agent long-term memory across sessions.
Multi-server configuration (excerpt)
{
  "mcpServers": {
    "filesystem": {
      "command": "npx",
      "args": ["-y", "@modelcontextprotocol/server-filesystem", "./workspace"]
    },
    "sqlite": {
      "command": "uvx",
      "args": ["mcp-server-sqlite", "--db-path", "./data/app.db"]
    }
  }
}

#The real limitations of small models

A functional bridge does not guarantee a good agent. The weak link is still the model. With the lightest models (2–4B), several problems consistently appear as soon as the task becomes more complex.

Wrong tool choice
The model calls read_file when it needs list_directory, or invents a tool name. Common below 7B.
Malformed arguments
Incorrect relative paths, invalid JSON in the arguments. A good system prompt and clear tool descriptions mitigate the problem.
Short sequences
Small models struggle beyond 2-3 successive calls and lose track of the objective.
Ignore the result
The model calls a tool and then responds without taking what it received into account. A classic symptom of a model that's too small.
→
The right tier in 2026: Qwen 3.5 9B
For a reliable local MCP agent, target a Qwen 3.5 9B in Q4_K_M (≈6.6 GB of VRAM, manageable on a RTX 3060 12 GB or a 4070). Below 4B, reserve the agent for tightly constrained single-tool tasks.

#Troubleshooting

The model makes no tool_call
Verify that it supports tool use (Qwen 3.5, Granite 4.2) and that the “tools” parameter is correctly passed to /api/chat. An incompatible model ignores tools.
“connection refused” on 11434
The Ollama daemon is not running. Start it (“ollama serve”) and retry “curl http://localhost:11434/api/tags”.
The MCP server won't start
Test the npx / uvx command by itself in a terminal. A missing Node server can be fixed with « npm i -g » for the relevant package.
Infinite tool loop
Set an iteration limit in the loop and log every tool_call to identify the model that keeps repeating the same call.

#Go further

MCP relies on the core building blocks of the local ecosystem. These related guides on the site complement this one:

Create a local AI agent with LangChain and Ollama
The classic agent approach via LangChain, complementary to MCP for orchestrating tools.
Integrate Ollama via the REST API in Python
Understand the OpenAI-compatible endpoint and the function calling underlying the bridge.
Choose your quantization (Q4, Q5, Q8, FP16)
Fit a Qwen 3.5 9B tool-use model on your GPU without sacrificing reliability.
Did this guide help you?

Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.