Use Ollama in Claude Code and Cursor (models locaux)
Claude Code and Cursor have become go-to tools for coding with an LLM, but both send your code to Anthropic or OpenAI and cost at least $20/month. Ollama exposes an OpenAI-compatible endpoint at localhost:11434/v1 that lets you connect both IDEs—and any assistant that speaks the OpenAI API—to a local model. This guide shows the exact configuration for Cursor (native) and Claude Code (via proxy), which coding models to prioritize in 2026, and where local models fall short compared with the cloud.
#Why use Ollama in Claude Code or Cursor?
Three reasons come up again and again. Privacy first: a client project under an NDA, proprietary code, secrets in plain text in the files—none of that should go to a third party. Cost next: Cursor Pro is $20/month, and Claude Code uses Anthropic tokens that quickly add up to $50–100/month with sustained use. With Ollama, it's zero after buying the GPU. Finally, resilience: your assistant doesn't stop when Anthropic's API has an incident or your ADSL connection drops.
The trade-off is straightforward: a local Qwen3-Coder 30B does not match Claude Sonnet 4.6 or GPT-5 for complex agent tasks spanning multiple files. But for 80% of everyday uses—completion, refactoring, writing a test, explaining a block, generating a commit—a 9B to 30B coder in Q4 gets the job done easily. And you can always keep the cloud running in parallel for the heavy lifting.
#Ollama’s OpenAI-compatible endpoint
This guide gets you to the model. The kit gets you to the coding copilot in your editor.
- Lifetime online access
- PDF + files
- Lifetime updates
Since version 0.1.24, Ollama has exposed an endpoint compatible with the OpenAI ChatCompletions API in addition to its native API. That's what makes everything else possible: any OpenAI client (Python SDK, Node SDK, Cursor, Cline, Aider, Continue, etc.) can call Ollama without modification, simply by changing the base URL.
The response format is identical to OpenAI's: choices[0].message.content, usage with prompt_tokens and completion_tokens, and streaming support via stream: true. The API key is ignored — you can send anything in the Authorization header, and Ollama accepts it. Many clients nevertheless reject an empty field: use ollama or anything to keep them happy.
#Prerequisites
- Ollama 0.5+
- ollama --version must return a response. On Windows, the system tray icon must be active. If you are starting from scratch, see the Ollama installation guide for your OS.
- GPU with 12 GB of VRAM or 16 GB+ Mac M-series
- One Qwen 3.5 9B Q4 = ~6.6 GB, one Devstral 24B Q4 = ~14 GB, one Qwen3-Coder 30B-A3B Q4 = ~19 GB. RTX 3060 12 GB = practical floor (Qwen 3.5 9B); RTX 4070/4080 16 GB or Mac M3/M4 = Devstral 24B and Qwen3-Coder 30B in full.
- Cursor 0.40+ or Claude Code CLI
- Cursor from cursor.com (built-in OpenAI custom mode). Claude Code via npm install -g @anthropic-ai/claude-code.
- Minimal familiarity with environment variables
- For Claude Code, we'll work with ANTHROPIC_BASE_URL and ANTHROPIC_AUTH_TOKEN.
#1. Connect Cursor to Ollama
Cursor has an official option for pointing to an OpenAI-compatible endpoint. This works very well for chat (Ctrl+L) and inline editing (Ctrl+K). The Composer agent and Tab autocomplete, however, remain locked to Cursor cloud models—this is a product limitation that Anthysphere has explicitly accepted.
- 01Download a code modelWith 12 GB of VRAM, Qwen 3.5 9B (qwen3.5:9b) is an excellent starting point: decent French, tool calling, and 256k context. With 16 GB, move up to Devstral 24B (devstral:24b); with 24 GB, use Qwen3-Coder 30B-A3B (qwen3-coder:30b), the reference code MoE.
- 02Open Cursor settingsCtrl+Shift+J (Cmd+, on Mac) → Models tab. You’ll see the list of Cursor models (claude-3.5-sonnet, gpt-5, etc.) with toggles.
- 03Add a custom modelClick Add Model at the bottom. Enter the exact model name Ollama: qwen3-coder:30b. Check the box to enable it.
- 04Configure the base URLIn the OpenAI API Key section, expand Override OpenAI Base URL. Enter http://localhost:11434/v1 and click Save. For the API Key, enter anything (ollama is enough); Cursor rejects an empty field.
- 05Check the connectionClick Verify. Cursor sends a test request to your Ollama. If it returns 200, the button turns green and the model appears in the chat selector.
Once set up, Cursor chat (Ctrl+L) and inline editing (Ctrl+K) work with your local model. The model selector at the top of the chat panel lists qwen3-coder:30b; Cursor cloud models remain available if you want to switch occasionally.
#2. Connect Claude Code to Ollama
Claude Code (Anthropic’s CLI) natively speaks Anthropic’s Messages API, not OpenAI’s Chat Completions API. The two protocols differ (roles, tool-call format, streaming). To make Claude Code communicate with Ollama, you therefore need a small translator—a proxy that receives Anthropic Messages on one side and emits OpenAI Chat Completions on the other.
The reference project for this is called claude-code-router (musistudio/claude-code-router on GitHub). It installs with one command, runs locally on a port, and accepts per-model routing rules (e.g., send haiku to Ollama, sonnet to the real Claude).
- 01Install Claude Codenpm install -g @anthropic-ai/claude-code si ce n'est pas déjà fait. claude --version doit répondre.
- 02Install claude-code-routernpm install -g @musistudio/claude-code-router. Le binaire ccr est ajouté au PATH.
- 03Pull the model in OllamaPrefer a model that supports tool calls: qwen3-coder:30b, devstral:24b, or glm-4.7-flash. Without tool support, Claude Code cannot call its internal tools (Read, Edit, Bash, etc.) and loses 90% of its value.
- 04Configure the routerCreate ~/.claude-code-router/config.json with a Providers entry pointing to Ollama, and a Router rule that maps Claude models to your local model.
- 05Launch via ccrccr code instead of claude. The wrapper starts the proxy in the background, exports ANTHROPIC_BASE_URL to it, then launches Claude Code. Inside, /model lets you switch between routes.
Minimalist approach without a router: you can also export ANTHROPIC_BASE_URL and ANTHROPIC_AUTH_TOKEN directly to an Anthropic-compatible proxy (LiteLLM in --anthropic mode or y-router are examples). Lighter, but without the flexibility of task-specific rules.
#3. Which coding model to choose: Qwen3-Coder vs Devstral locally
On Ollama, a few families dominate coding as of summer 2026: Qwen3-Coder (Alibaba), Devstral (Mistral AI), and versatile MoEs such as GLM 4.7 Flash (Z.ai) or gpt-oss (OpenAI). All are open-weight and run in Q4 on consumer hardware.
- Qwen3-Coder 30B-A3B (Alibaba, Apache 2.0)
- The versatile reference for 2026. Strong tool-calling support, multilingual (good French), good at Python/JS/Go/Rust, with 256k context. It is a MoE with 3B active parameters: as fast as a dense 14B, with 30B-level quality. ≈ 19 GB in Q4. Pull: qwen3-coder:30b.
- Devstral 24B (Mistral AI, Apache 2.0)
- Designed specifically for SWE agents (multi-file editing, codebase navigation). Excellent in Aider, OpenHands, and—by extension—Claude Code. Dense 24B ≈ 14 GB in Q4, fits on 16 GB. Pull: devstral:24b.
- GLM 4.7 Flash (Z.ai, MIT) / gpt-oss 20B (OpenAI)
- Two MoEs that excel in agent mode. GLM 4.7 Flash (30B-A3B, ≈ 19 GB) is an excellent tool-call orchestrator; gpt-oss 20B (≈ 14 GB, MXFP4, 131k context) is faster and fits in 16 GB. Pull: glm-4.7-flash or gpt-oss:20b.
- Qwen 3.5 9B (8–12 GB GPU fallback)
- Only 6.6 GB, a solid generalist for code, with 256k context and vision. The right fallback when VRAM is limited, before moving to 24B+ models. Pull: qwen3.5:9b.
VRAM by size in Q4_K_M (the recommended quantization): 7B ≈ 5 GB, 14B ≈ 9 GB, 24B ≈ 14 GB, 32B ≈ 19 GB, 70B ≈ 40 GB. Context requires additional memory: allow +2 to +4 GB for 32k context tokens depending on the model.
#Limits vs. the cloud: where local falls short
Let’s be honest about what local setups don’t do as well as Claude Sonnet 4.6 or GPT-5:
- Context window
- By default, Ollama limits num_ctx to 4096 tokens; you can raise it to 32k or even 128k depending on the model, but at the cost of VRAM. Cursor with Sonnet in the cloud handles 200k tokens without breaking a sweat. In a large monorepo, the cloud reads everything; the local setup has to choose.
- Multi-step reasoning
- A 9–14B model loses track when the agent chains 10 tool calls with dependencies between them. Sonnet stays on track. For truly complex agent orchestration, the cloud is still far ahead—and by a wide margin.
- Familiarity with recent APIs
- Open-weight models have a cutoff (often 2025) and miss APIs released afterward. Cursor in the cloud benefits from continuous updates and web search tools.
- First-token latency
- Paradoxically, the cloud may be faster to start (no model loading). Locally, the model stays warm between calls — the advantage shifts back after the first prompt.
- Electricity cost
- A RTX 4090 under load consumes 350 W. 8h/day of sustained coding = ~70 kWh/month = ~€15 in France. Still far below $20/month for Cursor + $50/month for Claude, but it isn't free.
#Tips and troubleshooting
- Cursor returns "OpenAI API key invalid"
- The API Key field must contain a non-empty value. Enter ollama, sk-anything, or anything plausible. It’s cosmetic: Ollama ignores the header.
- Cursor can't see the model
- The model name in Cursor (Add Model) must be EXACTLY the one returned by ollama list, including the tag—for example, qwen3-coder:30b. Not qwen3-coder alone, and not Qwen3 Coder.
- Claude Code loops or asks for the same thing again
- The local model often gets its tool calls wrong. Check that the model supports function calling (ollama show qwen3-coder:30b → look for the tools mention in its capabilities). Switch to a larger model or simplify the task.
- Responses truncated after 2-3 sentences
- Default num_ctx = 4096. For Claude Code and Cursor with a codebase context, increase it to 16384 or 32768. With Ollama: create a custom Modelfile with PARAMETER num_ctx 32768 and run ollama create coder-32k -f Modelfile. The memory cost is real (+ 2-4 GB for the KV cache).
- VRAM saturated, OOM
- Check ollama ps during use. If you see >100% loaded in the GPU/CPU split, the model is spilling into RAM and becomes very slow. Solutions: more aggressive quantization (Q3_K_M), a smaller model, or reduce num_ctx.
- Latency > 5 s per response
- Either the model spills over to the CPU (see the previous point), or Ollama reloads the model for every request. Check OLLAMA_KEEP_ALIVE (5 min by default). Set OLLAMA_KEEP_ALIVE=2h to keep the model warm.
#Go further
You have Ollama serving a code model to Cursor or Claude Code. The natural next directions are:
- Complete in the editor (inline FIM)
- The free local Copilot guide installs Continue.dev / Tabby / CodeGeeX for Copilot-style inline completion—a natural complement to chat in Cursor.
- Choose the right quantization
- Q4_K_M, Q5_K_M, Q8_0: the Choosing Your Quantization guide compares the actual quality loss on code models and explains when Q3 remains usable.
- Customize a code model
- The Customize a model with Ollama Modelfile guide shows how to set num_ctx, the system prompt, and the temperature to get a coder tailored to your stack.
Can you really use Ollama with Claude Code?+
Ollama + Claude Code: is it really free?+
Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.