Intermediate 11 minObsidian

Obsidian + local LLM: Copilot and Smart Connections with Ollama

Obsidian stores your notes in plain-text Markdown on your disk. It is ideal ground for an AI assistant—as long as it does not send your entire vault to the cloud. This guide connects a local LLM through the obsidian + ollama pairing: configuring the Copilot plugin, chatting with all your notes, semantic search with Smart Connections, and embeddings that stay on your machine.

By Léa B.·Update 2026-08-27·Tested on Windows, macOS, and Linux

#Why connect a local LLM to Obsidian

An Obsidian vault often contains your most personal information: journals, meeting notes, research, and drafts. Sending this corpus to a cloud service to “chat with your notes” means entrusting your entire second memory to a third party. A local LLM solves the problem at its root: the model, embeddings, and index stay on your disk.

In practical terms, obsidian + ollama enables three use cases: generating and rewriting text directly in the editor (autocomplete, summarization, translation), querying the entire vault in natural language (“what did I note about project X?”), and automatically surfacing notes semantically similar to the one you're writing. Two plugins cover these needs: Copilot for chat and generation, Smart Connections for similarity search.

Privacy
No note or embedding leaves the machine. Ideal for medical records or customer data.
Offline
Works on a train or plane once the models have been downloaded.
Zero cost
No subscription or per-request billing, regardless of the volume of notes.
Model control
You choose the size, quantization, and context based on your hardware.

#Prerequisites

The Local AI Kit

Your private ChatGPT, free, on your own machine in an hour — LM Studio, Ollama, Open WebUI, your documents, no cloud.

  • Lifetime online access
  • PDF + files
  • Lifetime updates

The setup consists of three components: Obsidian, Ollama as a local daemon, and a community plugin. Nothing else is required—no API key, no account.

Obsidian 1.5+
Desktop version (Windows, macOS, or Linux). Community plugins do not work the same way on mobile; this guide targets desktop.
Ollama installed
The daemon listens on http://localhost:11434 by default. If you haven't done so yet, see the Ollama installation guide referenced at the end of the article.
A chat model
For example, qwen3.5:9b with Q4_K_M quantization (≈6.6 GB of VRAM), the reference 8 GB choice in 2026 (256k context, vision).
An embeddings model
nomic-embed-text or mxbai-embed-large, essential for chatting over a vault and Smart Connections.
RAM/VRAM
8 GB of VRAM is enough for a modern 9B model (Qwen 3.5 9B, ≈6.6 GB). On pure CPU, plan instead on a 2–3B model (Qwen 3.5 2B or Granite 4.2 3B) and a little patience.
i
Chat ≠ embeddings
Two distinct models are involved. The chat model generates responses; the embeddings model turns your notes into vectors for search. Don't configure a chat model where an embeddings model is expected, or vice versa — it's the most common mistake.

#Prepare Ollama for Obsidian

Before touching Obsidian, download the models and verify that the daemon responds. Ollama exposes an OpenAI-compatible API on port 11434, which the plugins know how to consume.

Terminal
# Modèle de chat (généraliste, ~6,6 Go en Q4)
ollama pull qwen3.5:9b

# Modèle d'embeddings (obligatoire pour le RAG)
ollama pull nomic-embed-text

# Vérifier que le daemon tourne
ollama ps

One Obsidian-specific point: plugins run in a browser-like context (Electron), and requests to Ollama may be blocked by the CORS policy. You therefore need to allow Obsidian to call the API. Set the OLLAMA_ORIGINS environment variable before starting the daemon.

Linux / macOS
# Autoriser les origines Obsidian (app://) et relancer Ollama
export OLLAMA_ORIGINS="app://obsidian.md*"
ollama serve
Windows (PowerShell)
# Définir la variable puis relancer Ollama
setx OLLAMA_ORIGINS "app://obsidian.md*"
# Quitter Ollama depuis la barre système, puis le relancer
!
CORS: the No. 1 cause of “connection failed” errors
If Copilot or Smart Connections displays a network error while ollama ps works, OLLAMA_ORIGINS is almost always missing. On macOS with the Ollama app, use launchctl setenv OLLAMA_ORIGINS "app://obsidian.md*" and then restart the application.

#Configure the Copilot plugin on Ollama

Copilot (by logancyang) is the most complete chat and generation plugin for connecting a local model. It handles everything from editor-based rewriting to conversational chat and “Vault QA” mode. Here's the step-by-step installation.

  1. 01
    Install the plugin
    Settings → Community plugins → Browse, search for “Copilot” (logancyang), install it, then enable it. First accept the use of community plugins if Obsidian asks.
  2. 02
    Open Copilot settings
    In Settings → Copilot, under “Model.” Copilot offers predefined providers; select Ollama as the provider for chat.
  3. 03
    Enter the URL and model
    Base URL: http://localhost:11434 . Chat model name: qwen3.5:9b (exactly the tag shown by ollama list). Leave the API key blank; it is unnecessary locally.
  4. 04
    Configure embeddings
    In the “Embedding Model” section, select Ollama again as the provider and enter nomic-embed-text. This is what enables chat across the entire vault.
  5. 05
    Test
    Open the Copilot panel (the icon in the sidebar or the “Copilot: Open Chat” command) and ask a simple question. A local response confirms that the connection is working.

Copilot also adds context-aware commands: select a passage, open the palette (Ctrl/Cmd+P), and run “Copilot: Summarize,” “Simplify,” or “Translate.” The selected text is sent to the local model, and the response is inserted or displayed depending on the command.

→
Context window
To summarize long notes, increase the model's context. Create a Modelfile with PARAMETER num_ctx 8192 (or more) and recreate the model through ollama create. A context that is too short silently truncates your notes before the model even reads them.

#Chat with your entire vault locally

Copilot's most powerful mode is “Vault QA” (also called QA mode): instead of chatting only with the model, you query your notes. Behind the scenes, Copilot builds a vector index of your vault with the embedding model, then retrieves the relevant passages and injects them into the prompt. This is RAG applied to your personal notes.

  1. 01
    Enable Vault QA mode
    In the Copilot chat panel, switch the mode selector from “Chat” to “Vault QA” (or “QA”).
  2. 02
    Build the index
    Run the command « Copilot: Index (refresh) vault for QA ». The plugin scans your notes and computes embeddings via Ollama. The duration depends on the vault size and the model.
  3. 03
    Ask cross-cutting questions
    Examples: “Summarize everything I noted about hexagonal architecture” or “Which meetings discussed the Q3 budget?” The answers cite the source notes.
  4. 04
    Reindex after major changes
    The index isn't magic: after adding a lot of notes, rerun indexing so they are included.
i
The index stays local
The vector index is stored in the plugin’s configuration folder inside your vault. Nothing is sent outside: embeddings are computed by Ollama on your machine and written to your disk.

#Smart Connections for finding related notes

Smart Connections (by Brian Petro) addresses a different need: instead of asking questions, it continuously displays notes semantically similar to the one you are editing. It is a living sidebar that surfaces links you would never have created manually—an automatic equivalent of “related notes,” calculated by vector similarity.

  1. 01
    Install Smart Connections
    Settings → Third-party add-ons → Browse, search for “Smart Connections,” install, and enable.
  2. 02
    Point embeddings to Ollama
    In the plugin settings, under embeddings, choose Ollama as the adapter, URL http://localhost:11434, and the nomic-embed-text model (or mxbai-embed-large for greater precision).
  3. 03
    Let indexing run
    On first launch, Smart Connections calculates embeddings for all your notes. The “Smart Connections” panel then fills in as you browse.
  4. 04
    Optional: Smart Chat
    Smart Connections also includes a chat mode (Smart Chat) that uses these same embeddings and can use your Ollama chat model to answer based on nearby notes.

Smart Connections' advantage over Copilot Vault QA is that it is passive and continuous. You do not have to ask for anything; as you write a note about a topic, related older notes surface automatically, encouraging serendipity in a large vault. The two plugins complement each other and can share the same Ollama embedding model.


#Recommended models by vault size

The right model depends first on your hardware, then on the volume of notes. For text generation, a modern 8–9B model in Q4_K_M (Qwen 3.5 9B) offers the best quality/VRAM compromise for most setups. For embeddings, nomic-embed-text is lightweight and sufficient; mxbai-embed-large is more accurate on large vaults, at the cost of a heavier index.

Small vault (< 500 notes)
Chat: qwen3.5:9b (Q4, ≈6.6 GB VRAM). Embeddings: nomic-embed-text. Fits on a RTX 3060 12GB.
Medium vault (500–3000 notes)
Chat: qwen3.5:9b. Embeddings: mxbai-embed-large for better link relevance. RTX 4070/4080 are comfortable.
Large vault (> 3000 notes)
Chat: Qwen 3.8 27B (qwen3.8:27b, ≈18 GB, 262k context) for more refined summaries. Embeddings: mxbai-embed-large. RTX 4090 24GB or Mac M4 Pro with unified memory.
Without a GPU (CPU only)
Chat: a 2-3B model (qwen3.5:2b, ≈1.9 GB, or very lightweight granite4.2:3b). Embeddings: nomic-embed-text, which remains fast on CPU. Slower but usable responses.
Mac Apple Silicon
Unified memory is an advantage: an M4 Pro 24–48 GB runs a Qwen 3.8 27B and embeddings indexing without a hitch.
→
Embedding consistency
Don’t switch embedding models casually: vectors from one model aren’t comparable to those from another. If you switch from nomic-embed-text to mxbai-embed-large, you must fully reindex the vault; otherwise, similarities become inconsistent.

#Troubleshooting

Most problems come down to three causes: CORS, an incorrect model name, or a stopped Ollama daemon. Here is the quick triage.

“Connection failed” / network error
OLLAMA_ORIGINS does not include Obsidian. Add app://obsidian.md* and restart the daemon (not just Obsidian).
“model not found”
The tag entered in the plugin does not match. Check with ollama list and copy the exact name, including the tag (qwen3.5:9b, not qwen3.5).
Empty or truncated responses
Context too short. Increase num_ctx via a Modelfile, or reduce the size of the notes being sent.
Very slow indexing
The embedding model is running on the CPU or the vault is huge. Check ollama ps; consider the lighter nomic-embed-text.
Empty Smart Connections
The index hasn't finished building, or the adapter points to a chat model instead of an embeddings model.
Terminal
# Vérifier les modèles disponibles et leurs tags exacts
ollama list

# Confirmer que le daemon répond et voir ce qui est chargé en mémoire
ollama ps

# Test brut de l'API embeddings (doit renvoyer un vecteur)
curl http://localhost:11434/api/embeddings -d '{
  "model": "nomic-embed-text",
  "prompt": "note de test"
}'

#Go further

Once Obsidian is connected to your local model, these guides on the site can help you fine-tune the stack and understand the components under the hood:

Install Ollama: Windows, macOS, and Linux
The foundation of the whole setup: clean daemon installation and model management.
Choose your quantization (Q4, Q5, Q8, FP16)
For balancing generation quality against the VRAM available on your card.
AnythingLLM: production-ready RAG locally
If you want to take RAG beyond Obsidian, with workspaces and an API.
Did this guide help you?

Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.