Intermediate 10 minIDE

Use Ollama in Claude Code and Cursor (models locaux)

Claude Code and Cursor have become go-to tools for coding with an LLM, but both send your code to Anthropic or OpenAI and cost at least $20/month. Ollama exposes an OpenAI-compatible endpoint at localhost:11434/v1 that lets you connect both IDEs—and any assistant that speaks the OpenAI API—to a local model. This guide shows the exact configuration for Cursor (native) and Claude Code (via proxy), which coding models to prioritize in 2026, and where local models fall short compared with the cloud.

By Mohamed Meguedmi·Update 2026-09-01·Tested on Windows, macOS, and Linux

#Why use Ollama in Claude Code or Cursor?

Three reasons come up again and again. Privacy first: a client project under an NDA, proprietary code, secrets in plain text in the files—none of that should go to a third party. Cost next: Cursor Pro is $20/month, and Claude Code uses Anthropic tokens that quickly add up to $50–100/month with sustained use. With Ollama, it's zero after buying the GPU. Finally, resilience: your assistant doesn't stop when Anthropic's API has an incident or your ADSL connection drops.

The trade-off is straightforward: a local Qwen3-Coder 30B does not match Claude Sonnet 4.6 or GPT-5 for complex agent tasks spanning multiple files. But for 80% of everyday uses—completion, refactoring, writing a test, explaining a block, generating a commit—a 9B to 30B coder in Q4 gets the job done easily. And you can always keep the cloud running in parallel for the heavy lifting.

i
What this guide covers
Connect Ollama (local model) to Cursor and Claude Code through the OpenAI-compatible API. Choosing a code model. The guide does not cover Copilot-style inline completion—for that, see the Continue.dev / Tabby / CodeGeeX guide.

#Ollama’s OpenAI-compatible endpoint

The Local Copilot Kit

This guide gets you to the model. The kit gets you to the coding copilot in your editor.

  • Lifetime online access
  • PDF + files
  • Lifetime updates

Since version 0.1.24, Ollama has exposed an endpoint compatible with the OpenAI ChatCompletions API in addition to its native API. That's what makes everything else possible: any OpenAI client (Python SDK, Node SDK, Cursor, Cline, Aider, Continue, etc.) can call Ollama without modification, simply by changing the base URL.

Verify that it responds
curl http://localhost:11434/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{
    "model": "qwen3-coder:30b",
    "messages": [{"role": "user", "content": "Bonjour"}]
  }'

The response format is identical to OpenAI's: choices[0].message.content, usage with prompt_tokens and completion_tokens, and streaming support via stream: true. The API key is ignored — you can send anything in the Authorization header, and Ollama accepts it. Many clients nevertheless reject an empty field: use ollama or anything to keep them happy.

→
Three endpoints, one daemon
Ollama listens in parallel on /api/* (native API, recommended for real Ollama clients) and /v1/* (OpenAI compatibility). No configuration is needed to enable /v1; it is available immediately after installation. The port remains 11434 in both cases.

#Prerequisites

Ollama 0.5+
ollama --version must return a response. On Windows, the system tray icon must be active. If you are starting from scratch, see the Ollama installation guide for your OS.
GPU with 12 GB of VRAM or 16 GB+ Mac M-series
One Qwen 3.5 9B Q4 = ~6.6 GB, one Devstral 24B Q4 = ~14 GB, one Qwen3-Coder 30B-A3B Q4 = ~19 GB. RTX 3060 12 GB = practical floor (Qwen 3.5 9B); RTX 4070/4080 16 GB or Mac M3/M4 = Devstral 24B and Qwen3-Coder 30B in full.
Cursor 0.40+ or Claude Code CLI
Cursor from cursor.com (built-in OpenAI custom mode). Claude Code via npm install -g @anthropic-ai/claude-code.
Minimal familiarity with environment variables
For Claude Code, we'll work with ANTHROPIC_BASE_URL and ANTHROPIC_AUTH_TOKEN.

#1. Connect Cursor to Ollama

Cursor has an official option for pointing to an OpenAI-compatible endpoint. This works very well for chat (Ctrl+L) and inline editing (Ctrl+K). The Composer agent and Tab autocomplete, however, remain locked to Cursor cloud models—this is a product limitation that Anthysphere has explicitly accepted.

  1. 01
    Download a code model
    With 12 GB of VRAM, Qwen 3.5 9B (qwen3.5:9b) is an excellent starting point: decent French, tool calling, and 256k context. With 16 GB, move up to Devstral 24B (devstral:24b); with 24 GB, use Qwen3-Coder 30B-A3B (qwen3-coder:30b), the reference code MoE.
  2. 02
    Open Cursor settings
    Ctrl+Shift+J (Cmd+, on Mac) → Models tab. You’ll see the list of Cursor models (claude-3.5-sonnet, gpt-5, etc.) with toggles.
  3. 03
    Add a custom model
    Click Add Model at the bottom. Enter the exact model name Ollama: qwen3-coder:30b. Check the box to enable it.
  4. 04
    Configure the base URL
    In the OpenAI API Key section, expand Override OpenAI Base URL. Enter http://localhost:11434/v1 and click Save. For the API Key, enter anything (ollama is enough); Cursor rejects an empty field.
  5. 05
    Check the connection
    Click Verify. Cursor sends a test request to your Ollama. If it returns 200, the button turns green and the model appears in the chat selector.
Model pull on the Ollama side
ollama pull qwen3-coder:30b
!
Privacy mode is not enough
Enabling Privacy Mode in Cursor prevents your code from being used to train models, but it does not change the fact that the code passes through Cursor's servers to call the model. Only switching to a local endpoint (this configuration) guarantees that nothing leaves the machine—verifiable with tcpdump or an outbound firewall.

Once set up, Cursor chat (Ctrl+L) and inline editing (Ctrl+K) work with your local model. The model selector at the top of the chat panel lists qwen3-coder:30b; Cursor cloud models remain available if you want to switch occasionally.

i
What does not work with a custom model
The Composer agent (Ctrl+I in agent mode), Tab autocomplete, and the Cursor Predicts feature remain tied to Cursor’s cloud models — they use internally fine-tuned models and cannot be redirected. For local inline completion, use Continue.dev alongside them.

#2. Connect Claude Code to Ollama

Claude Code (Anthropic’s CLI) natively speaks Anthropic’s Messages API, not OpenAI’s Chat Completions API. The two protocols differ (roles, tool-call format, streaming). To make Claude Code communicate with Ollama, you therefore need a small translator—a proxy that receives Anthropic Messages on one side and emits OpenAI Chat Completions on the other.

The reference project for this is called claude-code-router (musistudio/claude-code-router on GitHub). It installs with one command, runs locally on a port, and accepts per-model routing rules (e.g., send haiku to Ollama, sonnet to the real Claude).

  1. 01
    Install Claude Code
    npm install -g @anthropic-ai/claude-code si ce n'est pas déjà fait. claude --version doit répondre.
  2. 02
    Install claude-code-router
    npm install -g @musistudio/claude-code-router. Le binaire ccr est ajouté au PATH.
  3. 03
    Pull the model in Ollama
    Prefer a model that supports tool calls: qwen3-coder:30b, devstral:24b, or glm-4.7-flash. Without tool support, Claude Code cannot call its internal tools (Read, Edit, Bash, etc.) and loses 90% of its value.
  4. 04
    Configure the router
    Create ~/.claude-code-router/config.json with a Providers entry pointing to Ollama, and a Router rule that maps Claude models to your local model.
  5. 05
    Launch via ccr
    ccr code instead of claude. The wrapper starts the proxy in the background, exports ANTHROPIC_BASE_URL to it, then launches Claude Code. Inside, /model lets you switch between routes.
Pull a model suited to tool calling
ollama pull qwen3-coder:30b
# ou pour du pur agent de code
ollama pull devstral:24b
~/.claude-code-router/config.json
{
  "Providers": [
    {
      "name": "ollama",
      "api_base_url": "http://localhost:11434/v1/chat/completions",
      "api_key": "ollama",
      "models": ["qwen3-coder:30b", "devstral:24b"]
    }
  ],
  "Router": {
    "default": "ollama,qwen3-coder:30b",
    "background": "ollama,qwen3-coder:30b",
    "think": "ollama,devstral:24b",
    "longContext": "ollama,qwen3-coder:30b"
  }
}
Launch
ccr code
# Claude Code démarre, mais les requêtes filent vers Ollama
# Vérifier : /model affiche qwen3-coder:30b
!
Internal tools can drift
Claude Code relies heavily on tool calls for Read, Edit, Bash, Glob, and Grep. Not all Ollama models handle them as reliably as Claude Sonnet—Qwen3-Coder and Devstral do well, while some smaller models forget arguments or hallucinate files. If you see Claude Code looping or asking for the same thing repeatedly, the model is probably missing its tool calls.

Minimalist approach without a router: you can also export ANTHROPIC_BASE_URL and ANTHROPIC_AUTH_TOKEN directly to an Anthropic-compatible proxy (LiteLLM in --anthropic mode or y-router are examples). Lighter, but without the flexibility of task-specific rules.

#3. Which coding model to choose: Qwen3-Coder vs Devstral locally

On Ollama, a few families dominate coding as of summer 2026: Qwen3-Coder (Alibaba), Devstral (Mistral AI), and versatile MoEs such as GLM 4.7 Flash (Z.ai) or gpt-oss (OpenAI). All are open-weight and run in Q4 on consumer hardware.

Qwen3-Coder 30B-A3B (Alibaba, Apache 2.0)
The versatile reference for 2026. Strong tool-calling support, multilingual (good French), good at Python/JS/Go/Rust, with 256k context. It is a MoE with 3B active parameters: as fast as a dense 14B, with 30B-level quality. ≈ 19 GB in Q4. Pull: qwen3-coder:30b.
Devstral 24B (Mistral AI, Apache 2.0)
Designed specifically for SWE agents (multi-file editing, codebase navigation). Excellent in Aider, OpenHands, and—by extension—Claude Code. Dense 24B ≈ 14 GB in Q4, fits on 16 GB. Pull: devstral:24b.
GLM 4.7 Flash (Z.ai, MIT) / gpt-oss 20B (OpenAI)
Two MoEs that excel in agent mode. GLM 4.7 Flash (30B-A3B, ≈ 19 GB) is an excellent tool-call orchestrator; gpt-oss 20B (≈ 14 GB, MXFP4, 131k context) is faster and fits in 16 GB. Pull: glm-4.7-flash or gpt-oss:20b.
Qwen 3.5 9B (8–12 GB GPU fallback)
Only 6.6 GB, a solid generalist for code, with 256k context and vision. The right fallback when VRAM is limited, before moving to 24B+ models. Pull: qwen3.5:9b.
→
Practical recommendation
If you are undecided: qwen3.5:9b on 12 GB of VRAM, devstral:24b or gpt-oss:20b on 16 GB, qwen3-coder:30b or glm-4.7-flash on 24 GB. Devstral and GLM 4.7 Flash shine specifically when the IDE runs in agent mode (Claude Code, Aider); Qwen3-Coder is more versatile for direct conversational chat in Cursor.

VRAM by size in Q4_K_M (the recommended quantization): 7B ≈ 5 GB, 14B ≈ 9 GB, 24B ≈ 14 GB, 32B ≈ 19 GB, 70B ≈ 40 GB. Context requires additional memory: allow +2 to +4 GB for 32k context tokens depending on the model.

#Limits vs. the cloud: where local falls short

Let’s be honest about what local setups don’t do as well as Claude Sonnet 4.6 or GPT-5:

Context window
By default, Ollama limits num_ctx to 4096 tokens; you can raise it to 32k or even 128k depending on the model, but at the cost of VRAM. Cursor with Sonnet in the cloud handles 200k tokens without breaking a sweat. In a large monorepo, the cloud reads everything; the local setup has to choose.
Multi-step reasoning
A 9–14B model loses track when the agent chains 10 tool calls with dependencies between them. Sonnet stays on track. For truly complex agent orchestration, the cloud is still far ahead—and by a wide margin.
Familiarity with recent APIs
Open-weight models have a cutoff (often 2025) and miss APIs released afterward. Cursor in the cloud benefits from continuous updates and web search tools.
First-token latency
Paradoxically, the cloud may be faster to start (no model loading). Locally, the model stays warm between calls — the advantage shifts back after the first prompt.
Electricity cost
A RTX 4090 under load consumes 350 W. 8h/day of sustained coding = ~70 kWh/month = ~€15 in France. Still far below $20/month for Cursor + $50/month for Claude, but it isn't free.
i
The pragmatic hybrid strategy
The approach that works in practice: keep Ollama as the default for the 80% of everyday tasks (completing, explaining, small refactors). Use the cloud (Claude Sonnet or GPT-5) on demand for the 20% that require substantial reasoning or long context. Cursor supports this natively through the model selector; for Claude Code, claude-code-router supports it through /model during a session.

#Tips and troubleshooting

Cursor returns "OpenAI API key invalid"
The API Key field must contain a non-empty value. Enter ollama, sk-anything, or anything plausible. It’s cosmetic: Ollama ignores the header.
Cursor can't see the model
The model name in Cursor (Add Model) must be EXACTLY the one returned by ollama list, including the tag—for example, qwen3-coder:30b. Not qwen3-coder alone, and not Qwen3 Coder.
Claude Code loops or asks for the same thing again
The local model often gets its tool calls wrong. Check that the model supports function calling (ollama show qwen3-coder:30b → look for the tools mention in its capabilities). Switch to a larger model or simplify the task.
Responses truncated after 2-3 sentences
Default num_ctx = 4096. For Claude Code and Cursor with a codebase context, increase it to 16384 or 32768. With Ollama: create a custom Modelfile with PARAMETER num_ctx 32768 and run ollama create coder-32k -f Modelfile. The memory cost is real (+ 2-4 GB for the KV cache).
VRAM saturated, OOM
Check ollama ps during use. If you see >100% loaded in the GPU/CPU split, the model is spilling into RAM and becomes very slow. Solutions: more aggressive quantization (Q3_K_M), a smaller model, or reduce num_ctx.
Latency > 5 s per response
Either the model spills over to the CPU (see the previous point), or Ollama reloads the model for every request. Check OLLAMA_KEEP_ALIVE (5 min by default). Set OLLAMA_KEEP_ALIVE=2h to keep the model warm.
→
Zero cost, but not zero latency
A local coder isn't magically faster than the cloud—it is simply running on your machine. On RTX 4090, qwen3-coder:30b (MoE, 3B active) generates ~60 tokens/s, or ~2–3 seconds for an average response. That's comparable to Claude Sonnet. On a Mac M3 Max or RTX 3090 24 GB, expect closer to 25–30 tokens/s, or 5–7 s—noticeable but usable.

#Go further

You have Ollama serving a code model to Cursor or Claude Code. The natural next directions are:

Complete in the editor (inline FIM)
The free local Copilot guide installs Continue.dev / Tabby / CodeGeeX for Copilot-style inline completion—a natural complement to chat in Cursor.
Choose the right quantization
Q4_K_M, Q5_K_M, Q8_0: the Choosing Your Quantization guide compares the actual quality loss on code models and explains when Q3 remains usable.
Customize a code model
The Customize a model with Ollama Modelfile guide shows how to set num_ctx, the system prompt, and the temperature to get a coder tailored to your stack.
Frequently asked questions
Can you really use Ollama with Claude Code?+
Yes: Claude Code accepts an OpenAI-compatible endpoint, and Ollama exposes one at localhost:11434/v1. Point it there (section 2 of this guide), and your requests go to your own machine instead of the cloud. Capabilities then depend on the local model you choose — Qwen3-Coder 30B-A3B is currently the best speed/quality tradeoff.
Ollama + Claude Code: is it really free?+
Yes: Ollama is free, the open-weight models are too, and there is no longer a $20/month subscription. Your only costs are hardware and electricity. A significant bonus: your code never leaves your machine.
Did this guide help you?

Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.