MCP: What is it? The Model Context Protocol explained
MCP (Model Context Protocol) is the open standard published by Anthropic in November 2024 and then entrusted to the Agentic AI Foundation (Linux Foundation) in December 2025. It allows an AI to use external tools and data through a host, client, and server shared by all compatible applications.
MCP, short for Model Context Protocol, is an open standard that lets an AI application use external tools and data: read a file, query a database, search the web. Anthropic published it in November 2024; OpenAI, Google, and Microsoft adopted it in 2025. As of September 20, 2026, it is the standard way to connect an assistant to the rest of your system, including when the model runs locally. This page explains the concept, the terminology, and what MCP really costs on local hardware.
#MCP in one paragraph
A language model on its own can only produce text based on what it learned and what the prompt contains. To check your calendar, query the company database, or open a document, it needs a bridge to the outside world. Before MCP, each application built its own bridges: one plugin format per assistant, another per code editor, and another per agent framework. MCP replaces all of that with a single specification. The author of a tool writes an MCP server, and any compatible application can use it. The usual analogy is USB-C: one connector, many devices.
In December 2025, Anthropic entrusted the protocol to the Agentic AI Foundation, hosted by the Linux Foundation. MCP therefore no longer belongs to a vendor, which partly explains the speed of its adoption.
- The official Model Context Protocol specification
- The original announcement by Anthropic (November 2024)
#Host, client, server: who does what
Agents that act on your machine: agentic Cline, MCP, n8n + Ollama, local automations.
- Lifetime online access
- PDF + files
- Lifetime updates
| Item | What it is | Examples |
|---|---|---|
| Host | The AI application you talk to | Claude Desktop, a code editor assistant, Open WebUI, LM Studio |
| Client | The internal connector to the host, one per connected server | Integrated into the host, invisible to the user |
| Server | A small program that exposes capabilities in MCP format | File system, GitHub, PostgreSQL, web search, your internal API |
The messages exchanged are JSON-RPC 2.0. At startup, each client asks its server what it can do. The server responds with a list of capabilities, each with a name, a plain-language description, and a JSON schema describing its parameters. The host passes these descriptions to the model. When the model determines that a tool would be useful, it produces a structured call; the client forwards it to the server, which executes it and returns the result in the conversation.
#What an MCP server can offer
| Primitive | Powered by | Role | Example |
|---|---|---|---|
| Tools (tools) | The model | Live actions and searches | creer_ticket, executer_requete, chercher_web |
| Resources (resources) | L'application | Read-only data to load into context | A file, a database schema, a wiki page |
| Prompts | L'utilisateur | Reusable instruction templates | “Review this pull request,” with parameters |
In practice, most servers provide only tools, and that is what people commonly mean by “MCP.” The protocol also defines functions in the other direction: sampling, where a server asks the host to have its model generate text, and elicitation, where a server asks the user for missing information during execution instead of failing or guessing.
#Adoption and governance, one year later
The shift happened within a few months. In March 2025, Sam Altman wrote that users loved MCP and that OpenAI was adding support to its products; ChatGPT, the Agents SDK, and the Responses API later integrated it. In April 2025, Demis Hassabis confirmed that Gemini and Google DeepMind tools would support MCP, calling the protocol an emerging standard for the agent era. Microsoft followed with integration into Copilot Studio, then announced native secure support in Windows 11.
On December 9, 2025, the Linux Foundation created the Agentic AI Foundation (AAIF): Anthropic, Block, and OpenAI are its founding members, while AWS, Bloomberg, Cloudflare, and Google joined it. MCP, Google's A2A protocol, and a Block project are hosted side by side under the same neutral governance as Kubernetes or PyTorch. This shift addresses a frequent 2025 criticism: a standard that was “open” but whose decisions remained controlled by a single vendor. For readers building a local stack, the practical consequence is twofold: less risk that a unilateral Anthropic decision will break an existing server, and more active contributors to maintain the official Python, TypeScript, Java, C#, and Rust software development kits (SDKs) used by community servers.
#Local or remote: both transports
- stdio
- The host launches the server as a local subprocess and communicates with it through standard input and output. Nothing goes over the network. This is how most desktop installations work, and the natural choice for a private setup.
- Streamable HTTP
- The server runs as a web service, optionally on another machine, with standard authentication. This transport replaced the old HTTP and SSE pair in the 2025 revisions of the specification.
#MCP, function calling, API: the differences
| What it is | Who defines it | |
|---|---|---|
| REST API | How a service exposes its functions to programs | Each service, in its own way |
| Function calling | The capability, acquired during training, to produce a structured call rather than text | Each model vendor, with similar but distinct formats |
| MCP | A standard way to discover, describe, and invoke tools across applications | A single open specification |
MCP doesn’t replace function calling; it builds on it. The model still needs to know how to produce a correct tool call. MCP standardizes everything around that moment: how tools are listed, described, authorized, and executed. An MCP server is very often a thin layer on top of an existing API.
#MCP with a local model: the real cost
MCP is model-agnostic. With a compatible host and a local model trained for tool calling, the entire loop runs on your machine. Two constraints then weigh much more heavily than they do in the cloud.
First constraint: tool descriptions consume context, and context consumes VRAM. Each connected tool injects its name, description, and schema into the prompt on every turn. Plan on 100 to 300 tokens per tool, and 10 to 30 tools per popular server. With a few servers, several thousand tokens are spent before you have typed a word. With a cloud provider, it's a line on the bill. On your GPU, it's KV cache taken from VRAM that does not grow.
| Configuration | Definition tokens | KV cache, Qwen 3 8B | KV cache, Qwen 3 32B |
|---|---|---|---|
| 1 server, 8 tools | ≈ 1 600 | ≈ 0.2 GB | ≈ 0.4 GB |
| 3 servers, 30 tools | ≈ 6 000 | ≈ 0.9 GB | ≈ 1.6 GB |
| 6 servers, 80 tools | ≈ 16 000 | ≈ 2.4 GB | ≈ 4.2 GB |
On a 16 GB card running a 13 GB model, 2.4 GB of tool descriptions can make the difference between a working session and an out-of-memory error. The practical rule for local use: connect only the servers relevant to the task at hand, and prefer those exposing a few well-described tools over those publishing their entire API.
Second constraint: smaller models make poorer choices. Selecting the right tool from thirty and correctly filling out its schema is a capability that depends on model size and training. Below about 8 billion parameters, tool-use errors and invented parameters become common.
- The tutorial: connect MCP servers to Ollama
- MCP and plugins in LM Studio
- The best local LLMs for agents and tool calling
- Reduce context costs: quantize the KV cache
#Security: what tutorials leave out
An MCP server is a program that runs with your permissions and is fed text produced by a model. Three risks are worth understanding before installing anything.
- Prompt injection
- If a tool returns content from outside sources (a web page, an email, a comment), that content may contain instructions intended for the model. A model that reads untrusted text and can take action is a target.
- Tool poisoning
- A malicious server can hide instructions in the description of its own tools, which the model reads and the user generally never sees.
- Overly broad permissions
- A file server pointed at your home directory gives the model your home directory.
The countermeasures are unglamorous: install servers from known sources, restrict each one to the narrowest possible scope, keep confirmation requests for anything that writes or sends, and avoid mixing access to private data with reading untrusted content in the same session.
#Do you need it?
| Your situation | Verdict |
|---|---|
| You chat with a model and paste in what it needs | No. MCP adds moving parts you will not use. |
| You want the assistant to read files, query a database, or search the web | Yes. That’s the standard way to do it in 2026. |
| You’re developing an internal tool that can be used from any AI application | Yes. An MCP server rather than an app integration. |
| You are running a small model on 8 GB of VRAM | With caution: at most one or two lightweight servers. |
Concrete example: the QuelLLM catalog itself is available through an open-source MCP server. An assistant connected to it answers “which model fits on a 16 GB card” using up-to-date data rather than its memory.
#FAQ
What does MCP mean?+
What is an MCP server?+
Is MCP limited to Claude?+
Does MCP work with Ollama or LM Studio?+
Is MCP dangerous?+
Who governs MCP today?+
Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.