IntelliJ and JetBrains + Ollama: the local AI assistant in its IDE
Pairing IntelliJ (or PyCharm, WebStorm, GoLand…) with Ollama gives you a Copilot-style coding assistant that runs on your machine: chat about your project, end-of-line completion, refactoring—all without a single line being sent to a third-party server. This guide sorts through the Ollama-compatible plugins, shows how to connect everything through an “AI proxy,” and explains which coding models to choose based on your GPU.
#Why use a local LLM in your IDE
Cloud assistants (GitHub Copilot, JetBrains AI, Cursor) are convenient, but they send your code's context—sometimes the entire file, sometimes the whole repository—to remote servers. For code under an NDA, proprietary software, or simply on principle, that's a dealbreaker. A local LLM connected to IntelliJ solves the problem at its root: the model runs on your GPU, and the prompt and completion never leave the machine.
The other argument is cost. A Copilot or JetBrains AI subscription is billed monthly, indefinitely. Once Ollama is installed and a code model is downloaded, you can complete code and chat without quotas or per-token charges, including offline. The tradeoff is quality: a small local model does not match a top-tier cloud model, but a Qwen 3.8 27B or Devstral 24B comes close for completion and everyday chat.
- Privacy
- The code, prompts, and responses stay local. Nothing is logged on the cloud side.
- No subscription
- No recurring cost once the model has been downloaded. Unlimited use.
- Hors-ligne
- Works on a train, on an isolated network, or behind a strict corporate proxy.
- Model control
- You choose the size and quantization, and can switch models depending on the task.
#JetBrains plugins that talk to Ollama
This guide gets you to the model. The kit gets you to the coding copilot in your editor.
- Lifetime online access
- PDF + files
- Lifetime updates
The JetBrains ecosystem has no universal native support for Ollama: you use a marketplace plugin. Three options cover nearly all needs, each with a different strength.
- ProxyAI (formerly CodeGPT)
- The most complete local option. Chat, inline completion, selection editing, and an integrated Ollama connector. This is the “AI proxy” mentioned in the title: it bridges the IDE and the Ollama daemon.
- Continue
- Open source and highly configurable through a config file. Chat, code completion, and code actions, with a first-class Ollama provider. Ideal if you want fine-grained control over every model.
- JetBrains AI Assistant
- The official JetBrains assistant. Since 2025, it has been able to point to a local model Ollama or LM Studio for offline mode. Convenient if you want to stay in the in-house tool, but less flexible for local completion.
#Prerequisites
We assume Ollama is already installed and working. All you need beyond that is a recent JetBrains IDE and at least one downloaded coding model.
- A JetBrains IDE 2024.1+
- IntelliJ IDEA, PyCharm, WebStorm, GoLand, Rider, PhpStorm… The plugins above install across the entire product line from the same marketplace.
- Ollama working
- The daemon is installed and reachable at http://localhost:11434. Test it with ollama list before configuring anything.
- A chat/code model
- qwen3.5:9b is a good versatile starting point (256k context, vision). In Q4_K_M, it fits in ~6.6 GB of VRAM.
- A recommended GPU
- Completion must respond in under one second to be useful. A RTX 3060 12 GB runs a 7B comfortably; without a GPU, stick to chat, not autocompletion.
#Configure ProxyAI + Ollama
ProxyAI (formerly CodeGPT) is the shortest path to a complete local assistant in JetBrains. Its Ollama connector handles chat, selection editing, and completion, with no API key or account.
- 01Install the pluginSettings → Plugins → Marketplace, search for « ProxyAI » (or « CodeGPT » depending on the version), and install it. Restart the IDE if prompted.
- 02Choose the Ollama providerSettings → Tools → ProxyAI → Providers. Select Ollama (Local) as the provider instead of the cloud options (OpenAI, Anthropic…).
- 03Check the server URLThe Base URL field must point to http://localhost:11434. If Ollama runs on another machine on the network, enter its IP address instead of localhost.
- 04Select the modelIn the model list, choose the one you downloaded (e.g., qwen3.5:9b). ProxyAI queries Ollama to list the available models.
- 05Test the chatOpen the ProxyAI panel (the icon in the sidebar) and ask a question about an open file. The answer should arrive locally, with no external connection warning.
#Continue and the native AI Assistant
If you prefer to tune every detail, Continue exposes its configuration in a file rather than in menus. There, you explicitly declare the Ollama provider, the chat model, and—separately—the completion model. It is more verbose but much more precise, especially when assigning a small, fast model to autocomplete and a larger one to chat.
For JetBrains AI Assistant, the process is more structured: in the assistant's settings, enable the use of local models and point it to Ollama. This is useful if you want to keep a single tool, but the native assistant is still primarily designed for the JetBrains cloud; its local completion is less mature than ProxyAI or Continue. For serious 100% local use, we recommend the latter two instead.
#Which coding models fit your VRAM
The rule is simple: the larger the model, the better its responses, but the more VRAM it consumes and the slower it responds. The Qwen 3.5 / Qwen 3.8 families cover the full range and are the benchmark in 2026 for local coding; Devstral 24B (a coding-agent specialist) and Granite 4.2 are good alternatives. Here are the reference points for Q4_K_M quantization (the best size/quality tradeoff).
- Lightweight 3B — ~2 GB VRAM
- granite4.2:3b. Quick troubleshooting, simple questions, very lightweight. Runs even on a small GPU or on the CPU for non-interactive use.
- Versatile 9B — ~6.6 GB VRAM
- qwen3.5:9b. The sweet spot: solid chat, editing, and refactoring, 256k context, and vision. Target for a RTX 3060 12 GB / 4070 12 GB.
- Comfortable 12B — ~7.6 GB VRAM
- gemma4:12b. Significantly better at reasoning and multi-file refactoring, multimodal, and licensed under Apache 2.0. Comfortable on a RTX 4080 16 GB.
- High-end 27B — ~18 GB VRAM
- qwen3.8:27b. The closest thing to the cloud for local coding (262k context, vision). Requires a RTX 4090 24 GB or a Mac Apple Silicon with 32 GB+ of unified memory. Tip: set its reasoning to “low”; it tends to overthink by default.
#Local completion: watch out for FIM
Copilot-style completion relies on “fill-in-the-middle” (FIM): the model must complete code in the middle while knowing what comes before AND after the cursor. Not all models support it. Instruct variants are trained for chat, not FIM—for autocompletion, use base variants specifically designed for it.
- Base model, not instruct
- For completion, choose qwen2.5-coder:7b-base, which remains the FIM benchmark in 2026. An instruct model will produce verbose or incorrectly formatted completions.
- Speed above all
- For autocomplete, speed matters more than nuance. A dedicated base model that responds in under half a second is more useful than a large chat model that takes 2 seconds.
- GPU nearly mandatory
- Without hardware acceleration, completion arrives too late to keep up with typing. In that case, reserve local AI for chat.
#What local AI still cannot do in the IDE
Local models have improved, but gaps remain compared with high-end cloud assistants. It is better to know what they are so you can set your expectations instead of being disappointed.
- Multi-file reasoning
- Large repositories exceed the context window of local models. The model sees the files you give it, not your entire architecture. Copilot Workspace or Cursor index the entire project; locally, it’s still a hands-on process.
- Advanced agent modes
- Having commands executed, tests run, and iterations performed in a loop, as with the Cline/Cursor Agent, requires a model that handles tool use reliably. Small local models often fail at this; target a dedicated coding model (Devstral 24B, Qwen3-Coder 30B, or GLM 4.7 Flash) and accept some failures.
- Raw quality on complex code
- For advanced algorithms or newer frameworks that are underrepresented in training, a local 8B to 12B model still lags behind a cutting-edge cloud model.
- Fine-grained product integration
- Automatic language detection, numerous contextual actions, PR resolution… local tooling is catching up but remains a step behind the polished experience of commercial assistants.
In practice, local models excel at completion, chatting about a file, code explanation, and localized refactoring. For heavy agentic workloads and analyzing an entire repository, the cloud retains the advantage — hence the value of keeping both and routing based on code sensitivity.
#Troubleshooting
- The plugin lists no models
- It can’t reach Ollama. Check that the daemon is running (ollama list) and that the URL is http://localhost:11434. If Ollama is on another machine, launch it with OLLAMA_HOST=0.0.0.0 and point to its IP.
- “Connection refused”
- Ollama is not running, or a firewall is blocking port 11434. Test with curl http://localhost:11434/api/tags from the same machine as the IDE.
- The model does not appear in the list
- The tag doesn't match. Copy the exact name returned by ollama list, including the tag (qwen3.5:9b, not qwen3.5).
- Very slow responses
- The model spills over to the CPU. Switch to a smaller size or Q4_K_M, and use nvidia-smi to verify that the GPU is actually being used.
- Empty or nonsensical completion
- You're using an instruct model for FIM. Switch to a -base variant (qwen2.5-coder:7b-base).
- The IDE lags during generation
- Two loaded models max out VRAM. Reduce one model's size, or activate only one plugin at a time.
#Go further
The local IDE assistant is only as good as the daemon and GPU powering it. These guides complete the setup.
- Install Ollama
- The basics: install and start the daemon that serves your code models on port 11434.
- Free local copilot: Cline, Tabby & CodeGeeX in VS Code
- The VS Code equivalent, for comparing an editor’s approaches and plugins with those of another.
- Use Ollama in Claude Code and Cursor
- To connect the same local models to other assistants and choose between local and cloud processing depending on the task.
Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.