MCP and local LLMs: connect MCP servers to Ollama
The Model Context Protocol (MCP) standardizes how an LLM calls external tools: file reading, web queries, and database access. Combining MCP with Ollama lets you run an agent that can act on your machine without ever sending your data to a cloud API. This guide shows how to build an MCP-Ollama bridge in Python, which local models actually handle tool use, and where the real limits of small models lie.
#What is MCP, and why does it change local agents?
MCP (Model Context Protocol) is an open protocol published by Anthropic in late 2024. Its purpose is to provide a single interface between a language model and the tools it can use. Instead of recoding a custom integration for each data source, an “MCP server” exposes tools, resources, and prompts in a standard format. Any compatible client—Claude Desktop, an IDE, or your own bridge—can then connect to it.
In practice, an MCP “filesystem” server exposes tools such as read_file, write_file, or list_directory. A “sqlite” server exposes query or list_tables. The LLM never talks directly to the disk: it issues a tool-call request, the client executes it through the MCP server, then sends the result back to the model. This client/server separation is what makes the protocol reusable.
For local agents, the challenge is twofold: reuse the growing ecosystem of MCP servers (dozens already exist) while keeping inference 100% on your machine through Ollama. You get an agent that reads your files and queries your databases without a single byte leaving your network.
#Why connect MCP to Ollama
This guide gets you to the model. The kit gets you to the coding copilot in your editor.
- Lifetime online access
- PDF + files
- Lifetime updates
Most MCP demos use a cloud model (Claude, GPT). Running mcp ollama locally changes three things: privacy (your files and SQL queries do not leave your system), cost (zero tokens billed regardless of tool-call volume), and control (you choose the model, quantization, and authorized servers).
- Privacy
- An MCP filesystem server gives the model access to your folders. Locally, this content never passes through a third party.
- Zero cost
- Agents make many tool round trips. In the cloud, each turn costs tokens; with Ollama, it’s free.
- Offline
- Once the model is downloaded and the MCP servers are installed, everything works without an internet connection (except web servers, of course).
- Sovereignty
- You decide which tools are exposed and can audit every call before executing it.
#Prerequisites
- Ollama installed
- The daemon listens on http://localhost:11434 by default. Check with "ollama --version".
- A tool-use model
- Plan on at least one Granite 4.2 8B, ideally a Qwen 3.5 9B for reliability (see the next section).
- Python 3.10+
- The official MCP SDK and the Ollama client are written in Python.
- Node.js (optional)
- Many reference MCP servers are launched via npx (@modelcontextprotocol/server-*).
#Which local models handle tool use correctly
Not all models are equally good at function calling. A model that “knows” the tool format but chooses its arguments poorly will make the agent unusable. In practice, the 2026 generation (Qwen 3.5, Granite 4.2) makes tool use reliable starting at 8–9B, whereas in 2024 you had to target 14B. Here are the benchmarks you can test with Ollama and their Q4_K_M VRAM footprint.
- Qwen 3.5 4B / 9B
- Excellent tool-use support. The 9B (≈6.6 GB of VRAM in Q4, 256k context, vision) is the best reliability/hardware compromise for a local agent.
- Granite 4.2 8B
- Solid native tool use with very low token usage (≈5.3 GB in Q4, 128k context). An excellent entry point for 6–8 GB of VRAM.
- Mistral Small 24B
- Solid function calling and good French-language performance (≈14 GB in Q4). Comfortable with somewhat complex tool schemas.
- Qwen 3.6 35B-A3B
- Very reliable MoE for call chains (≈23 GB in Q4, only 3B active, so it's fast). Best reserved for 24 GB GPUs such as RTX 4090.
- 2–3B models
- Qwen 3.5 2B or Granite 4.2 3B fit in ≈2 GB, but tool use quickly falls off once there are multiple tools. Avoid them for a real agent.
#The MCP-Ollama bridge in Python, step by step
Ollama is not a native MCP client. The bridge acts as an intermediary: it starts an MCP server, converts its tools to the format expected by the Ollama API, runs the tool-call loop, and then returns the results to the model. It uses the official MCP SDK (the mcp package) and the Ollama client.
- 011. Start an MCP serverWe launch an MCP server as a subprocess via stdio. Here, it’s the official filesystem server, limited to a working directory passed as an argument.
- 022. List and convert the toolssession.list_tools() returns the MCP tools. We transform them into the “tools” format expected by /api/chat from Ollama (name, description, inputSchema → parameters).
- 033. Tool-calling loopWe send the user message and the list of tools. If the model responds with a tool_call, we execute it through MCP, inject the result, and loop until we get a final response.
- 044. Return the responseWhen the model no longer requests a tool, its last text response is the final result presented to the user.
#Examples of useful self-hosted MCP servers
MCP's value comes from its catalog of ready-to-use servers. Here are the ones that add the most value to a local agent, all launchable through npx or pip.
- filesystem
- Reading/writing files in a mandated root folder. Most useful for an agent working on your documents.
- sqlite / postgres
- Query a local database in natural language. Switch the connection to read-only mode to prevent any changes.
- fetch
- Retrieve and convert a web page to text. The only one that assumes an internet connection.
- git
- Explore a repository: log, diff, status. Useful for a code-review or code-documentation agent.
- memory
- A persistent key-value store that gives the agent long-term memory across sessions.
#The real limitations of small models
A functional bridge does not guarantee a good agent. The weak link is still the model. With the lightest models (2–4B), several problems consistently appear as soon as the task becomes more complex.
- Wrong tool choice
- The model calls read_file when it needs list_directory, or invents a tool name. Common below 7B.
- Malformed arguments
- Incorrect relative paths, invalid JSON in the arguments. A good system prompt and clear tool descriptions mitigate the problem.
- Short sequences
- Small models struggle beyond 2-3 successive calls and lose track of the objective.
- Ignore the result
- The model calls a tool and then responds without taking what it received into account. A classic symptom of a model that's too small.
#Troubleshooting
- The model makes no tool_call
- Verify that it supports tool use (Qwen 3.5, Granite 4.2) and that the “tools” parameter is correctly passed to /api/chat. An incompatible model ignores tools.
- “connection refused” on 11434
- The Ollama daemon is not running. Start it (“ollama serve”) and retry “curl http://localhost:11434/api/tags”.
- The MCP server won't start
- Test the npx / uvx command by itself in a terminal. A missing Node server can be fixed with « npm i -g » for the relevant package.
- Infinite tool loop
- Set an iteration limit in the loop and log every tool_call to identify the model that keeps repeating the same call.
#Go further
MCP relies on the core building blocks of the local ecosystem. These related guides on the site complement this one:
- Create a local AI agent with LangChain and Ollama
- The classic agent approach via LangChain, complementary to MCP for orchestrating tools.
- Integrate Ollama via the REST API in Python
- Understand the OpenAI-compatible endpoint and the function calling underlying the bridge.
- Choose your quantization (Q4, Q5, Q8, FP16)
- Fit a Qwen 3.5 9B tool-use model on your GPU without sacrificing reliability.
Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.