Obsidian + local LLM: Copilot and Smart Connections with Ollama
Obsidian stores your notes in plain-text Markdown on your disk. It is ideal ground for an AI assistant—as long as it does not send your entire vault to the cloud. This guide connects a local LLM through the obsidian + ollama pairing: configuring the Copilot plugin, chatting with all your notes, semantic search with Smart Connections, and embeddings that stay on your machine.
#Why connect a local LLM to Obsidian
An Obsidian vault often contains your most personal information: journals, meeting notes, research, and drafts. Sending this corpus to a cloud service to “chat with your notes” means entrusting your entire second memory to a third party. A local LLM solves the problem at its root: the model, embeddings, and index stay on your disk.
In practical terms, obsidian + ollama enables three use cases: generating and rewriting text directly in the editor (autocomplete, summarization, translation), querying the entire vault in natural language (“what did I note about project X?”), and automatically surfacing notes semantically similar to the one you're writing. Two plugins cover these needs: Copilot for chat and generation, Smart Connections for similarity search.
- Privacy
- No note or embedding leaves the machine. Ideal for medical records or customer data.
- Offline
- Works on a train or plane once the models have been downloaded.
- Zero cost
- No subscription or per-request billing, regardless of the volume of notes.
- Model control
- You choose the size, quantization, and context based on your hardware.
#Prerequisites
Your private ChatGPT, free, on your own machine in an hour — LM Studio, Ollama, Open WebUI, your documents, no cloud.
- Lifetime online access
- PDF + files
- Lifetime updates
The setup consists of three components: Obsidian, Ollama as a local daemon, and a community plugin. Nothing else is required—no API key, no account.
- Obsidian 1.5+
- Desktop version (Windows, macOS, or Linux). Community plugins do not work the same way on mobile; this guide targets desktop.
- Ollama installed
- The daemon listens on http://localhost:11434 by default. If you haven't done so yet, see the Ollama installation guide referenced at the end of the article.
- A chat model
- For example, qwen3.5:9b with Q4_K_M quantization (≈6.6 GB of VRAM), the reference 8 GB choice in 2026 (256k context, vision).
- An embeddings model
- nomic-embed-text or mxbai-embed-large, essential for chatting over a vault and Smart Connections.
- RAM/VRAM
- 8 GB of VRAM is enough for a modern 9B model (Qwen 3.5 9B, ≈6.6 GB). On pure CPU, plan instead on a 2–3B model (Qwen 3.5 2B or Granite 4.2 3B) and a little patience.
#Prepare Ollama for Obsidian
Before touching Obsidian, download the models and verify that the daemon responds. Ollama exposes an OpenAI-compatible API on port 11434, which the plugins know how to consume.
One Obsidian-specific point: plugins run in a browser-like context (Electron), and requests to Ollama may be blocked by the CORS policy. You therefore need to allow Obsidian to call the API. Set the OLLAMA_ORIGINS environment variable before starting the daemon.
#Configure the Copilot plugin on Ollama
Copilot (by logancyang) is the most complete chat and generation plugin for connecting a local model. It handles everything from editor-based rewriting to conversational chat and “Vault QA” mode. Here's the step-by-step installation.
- 01Install the pluginSettings → Community plugins → Browse, search for “Copilot” (logancyang), install it, then enable it. First accept the use of community plugins if Obsidian asks.
- 02Open Copilot settingsIn Settings → Copilot, under “Model.” Copilot offers predefined providers; select Ollama as the provider for chat.
- 03Enter the URL and modelBase URL: http://localhost:11434 . Chat model name: qwen3.5:9b (exactly the tag shown by ollama list). Leave the API key blank; it is unnecessary locally.
- 04Configure embeddingsIn the “Embedding Model” section, select Ollama again as the provider and enter nomic-embed-text. This is what enables chat across the entire vault.
- 05TestOpen the Copilot panel (the icon in the sidebar or the “Copilot: Open Chat” command) and ask a simple question. A local response confirms that the connection is working.
Copilot also adds context-aware commands: select a passage, open the palette (Ctrl/Cmd+P), and run “Copilot: Summarize,” “Simplify,” or “Translate.” The selected text is sent to the local model, and the response is inserted or displayed depending on the command.
#Chat with your entire vault locally
Copilot's most powerful mode is “Vault QA” (also called QA mode): instead of chatting only with the model, you query your notes. Behind the scenes, Copilot builds a vector index of your vault with the embedding model, then retrieves the relevant passages and injects them into the prompt. This is RAG applied to your personal notes.
- 01Enable Vault QA modeIn the Copilot chat panel, switch the mode selector from “Chat” to “Vault QA” (or “QA”).
- 02Build the indexRun the command « Copilot: Index (refresh) vault for QA ». The plugin scans your notes and computes embeddings via Ollama. The duration depends on the vault size and the model.
- 03Ask cross-cutting questionsExamples: “Summarize everything I noted about hexagonal architecture” or “Which meetings discussed the Q3 budget?” The answers cite the source notes.
- 04Reindex after major changesThe index isn't magic: after adding a lot of notes, rerun indexing so they are included.
#Smart Connections for finding related notes
Smart Connections (by Brian Petro) addresses a different need: instead of asking questions, it continuously displays notes semantically similar to the one you are editing. It is a living sidebar that surfaces links you would never have created manually—an automatic equivalent of “related notes,” calculated by vector similarity.
- 01Install Smart ConnectionsSettings → Third-party add-ons → Browse, search for “Smart Connections,” install, and enable.
- 02Point embeddings to OllamaIn the plugin settings, under embeddings, choose Ollama as the adapter, URL http://localhost:11434, and the nomic-embed-text model (or mxbai-embed-large for greater precision).
- 03Let indexing runOn first launch, Smart Connections calculates embeddings for all your notes. The “Smart Connections” panel then fills in as you browse.
- 04Optional: Smart ChatSmart Connections also includes a chat mode (Smart Chat) that uses these same embeddings and can use your Ollama chat model to answer based on nearby notes.
Smart Connections' advantage over Copilot Vault QA is that it is passive and continuous. You do not have to ask for anything; as you write a note about a topic, related older notes surface automatically, encouraging serendipity in a large vault. The two plugins complement each other and can share the same Ollama embedding model.
#Recommended models by vault size
The right model depends first on your hardware, then on the volume of notes. For text generation, a modern 8–9B model in Q4_K_M (Qwen 3.5 9B) offers the best quality/VRAM compromise for most setups. For embeddings, nomic-embed-text is lightweight and sufficient; mxbai-embed-large is more accurate on large vaults, at the cost of a heavier index.
- Small vault (< 500 notes)
- Chat: qwen3.5:9b (Q4, ≈6.6 GB VRAM). Embeddings: nomic-embed-text. Fits on a RTX 3060 12GB.
- Medium vault (500–3000 notes)
- Chat: qwen3.5:9b. Embeddings: mxbai-embed-large for better link relevance. RTX 4070/4080 are comfortable.
- Large vault (> 3000 notes)
- Chat: Qwen 3.8 27B (qwen3.8:27b, ≈18 GB, 262k context) for more refined summaries. Embeddings: mxbai-embed-large. RTX 4090 24GB or Mac M4 Pro with unified memory.
- Without a GPU (CPU only)
- Chat: a 2-3B model (qwen3.5:2b, ≈1.9 GB, or very lightweight granite4.2:3b). Embeddings: nomic-embed-text, which remains fast on CPU. Slower but usable responses.
- Mac Apple Silicon
- Unified memory is an advantage: an M4 Pro 24–48 GB runs a Qwen 3.8 27B and embeddings indexing without a hitch.
#Troubleshooting
Most problems come down to three causes: CORS, an incorrect model name, or a stopped Ollama daemon. Here is the quick triage.
- “Connection failed” / network error
- OLLAMA_ORIGINS does not include Obsidian. Add app://obsidian.md* and restart the daemon (not just Obsidian).
- “model not found”
- The tag entered in the plugin does not match. Check with ollama list and copy the exact name, including the tag (qwen3.5:9b, not qwen3.5).
- Empty or truncated responses
- Context too short. Increase num_ctx via a Modelfile, or reduce the size of the notes being sent.
- Very slow indexing
- The embedding model is running on the CPU or the vault is huge. Check ollama ps; consider the lighter nomic-embed-text.
- Empty Smart Connections
- The index hasn't finished building, or the adapter points to a chat model instead of an embeddings model.
#Go further
Once Obsidian is connected to your local model, these guides on the site can help you fine-tune the stack and understand the components under the hood:
- Install Ollama: Windows, macOS, and Linux
- The foundation of the whole setup: clean daemon installation and model management.
- Choose your quantization (Q4, Q5, Q8, FP16)
- For balancing generation quality against the VRAM available on your card.
- AnythingLLM: production-ready RAG locally
- If you want to take RAG beyond Obsidian, with workspaces and an API.
Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.