Nextcloud Assistant + Ollama: AI in your cloud self-hosted
Nextcloud now has its own AI assistant: email summaries, rewriting, text generation, and document chat, directly in the interface. By default, it pushes you toward cloud services, but you don't have to use them. This guide shows how to connect Nextcloud Assistant to a self-hosted Ollama to keep 100% of your data at home: step-by-step AppAPI and ExApp configuration, OpenAI-compatible integration, recommended models, and server sizing for family or small-business use.
#Why connect Nextcloud to Ollama
Nextcloud Assistant is the AI interface integrated into Nextcloud (the “magic wand” icon in the top bar): it brings together tasks such as summarizing text, rephrasing it, generating a draft, extracting key points, and translating. Since Nextcloud Hub, a “contextual chat” assistant also lets you interact with a model and, with the right context, query your own files.
The problem: the backends offered by default or in the demo often point to cloud APIs (OpenAI or Nextcloud's hosted services). For a self-hosted cloud, this reintroduces exactly what we wanted to avoid—sending private content to a third party. Connecting the assistant to Ollama, which runs on your own machine and listens on http://localhost:11434, keeps your document text strictly on your infrastructure.
- Data sovereignty
- No file, email, or note content leaves your server. Useful for GDPR compliance and simple peace of mind.
- Zero usage cost
- No per-token billing: once the hardware has paid for itself, summaries and writing are free and unlimited.
- Models to choose from
- You choose the model (Qwen, Gemma, Granite) and its quantization based on your VRAM, instead of being stuck with a prescribed model.
- Works offline
- The assistant remains available even without internet access, as long as the Nextcloud server and Ollama are running.
#How it works: AppAPI and ExApp
Deploy local AI at work: privacy, compliance, multi-user architecture, costs, the one-page memo for leadership.
- Lifetime online access
- PDF + files
- Lifetime updates
Nextcloud does not talk directly to Ollama. An external-app mechanism sits between them. Understanding the three components prevents a lot of configuration guesswork.
- Assistant
- The Nextcloud-integrated front-end app that exposes tasks (summarization, rewriting, chat) in the user interface.
- AppAPI
- The Nextcloud app that manages the life cycle of external apps (ExApps): installation, deployment, and communication. It is the foundation to install first.
- AI back-end ExApp
- An external application (container) that implements task providers and knows how to communicate with an inference engine. The most commonly used for local deployments are: «integration_openai» (compatible with any OpenAI endpoint, including Ollama) and «LocalAI» through the dedicated ExApp.
- Deploy Daemon
- The service (often Docker) that AppAPI controls to start and stop ExApp containers. Without a working Deploy Daemon, no ExApp can be deployed.
The path of a request is therefore: you click “Summarize” in the Assistant → Nextcloud delegates the task to the provider → the ExApp (for example, integration_openai) formats an OpenAI-format call → that call is sent to the Ollama endpoint → the response follows the same path back. The key is to point the integration ExApp to the Ollama URL, in OpenAI-compatible mode.
#Prerequisites
- Recent Nextcloud Hub
- A version supporting AppAPI and Assistant (Hub 6/7 or later). Ideally a Docker/AIO installation or a server you administer, so you can install AppAPI and a Deploy Daemon.
- Ollama installed and accessible
- On the same machine or a machine on the network. The daemon listens on http://localhost:11434 by default. Check with “ollama --version” and “ollama ps”.
- Docker for the Deploy Daemon
- AppAPI deploys ExApps through a Docker daemon. On Nextcloud AIO, this is integrated; on a manual installation, you need Docker access (socket or API).
- A little VRAM or RAM
- Q4 rule of thumb: a 7B fits in ~5 GB, and a 14B in ~9 GB. Without a GPU, RAM takes over, more slowly—acceptable for occasional summaries.
- Nextcloud admin access
- Assistant, AppAPI, and integration settings are configured under “Administration settings.”
#1. Prepare Ollama and choose a model
Start with a versatile, French-speaking instruct model. For family use or a small team, Qwen 3.5 9B in Q4_K_M offers the best quality/memory tradeoff: it summarizes and reformulates very well in French, handles a very large context (256k tokens), and fits in ~6.6 GB of VRAM. Even leaner, Granite 4.2 8B (IBM, Apache 2.0) drops to ~5.3 GB for a more modest server.
Critical point: Ollama must be reachable from Nextcloud. If Nextcloud runs in Docker, “localhost” refers to the container, not your machine. Have Ollama listen on all interfaces and point the integration to the host's IP (or host.docker.internal, depending on the platform).
#2. Install AppAPI and a Deploy Daemon
In Nextcloud, AI runs through AppAPI. This step lays the foundation for deploying AI task ExApps.
- 01Install AppAPIAdministration Settings → Applications → “Tools” category (or “Outils”) → search for “AppAPI” and enable it. It then appears as a dedicated entry in the admin settings.
- 02Configure a Deploy DaemonIn Administration Settings → AppAPI, add a Docker-type “Deploy Daemon.” On Nextcloud AIO, the daemon is pre-provisioned; otherwise, enter access to the Docker socket (e.g. /var/run/docker.sock) and run the built-in connection test.
- 03Check statusAppAPI displays the daemon status (green = ready). Until it turns green, the ExApps cannot deploy—there is no point going further before fixing this.
- 04Install the Assistant appStill in Applications, enable “Assistant” if it is not already there. This is what displays the magic wand in the interface.
#3. Connect the OpenAI/LocalAI integration
The “OpenAI and LocalAI integration” application can communicate with any endpoint that follows the OpenAI API. Since Ollama exposes /v1, you only need to provide the correct URL and choose the model.
- 01Install the integrationApplications → search for “OpenAI and LocalAI integration” and enable it. It registers task providers (text generation, summarization, etc.) that the Assistant can use.
- 02Enter Ollama's URLAdministration settings → “Connected accounts” (or the integration section) → “Service URL” field: enter the OpenAI endpoint for Ollama, for example http://IP_DE_L_HOTE:11434/v1. Leave the API key blank or enter a dummy value: Ollama does not verify it.
- 03Choose the default modelOn the same page, select the model to use for text generation (e.g., qwen3.5:9b). If it does not appear in the list, enter its exact name as returned by “ollama list”.
- 04Save and testValidate. The integration queries the endpoint to list the models; an error here almost always indicates a URL or network problem (see Troubleshooting).
#4. Summarize, write, and chat about your files
Once the provider is connected, the Assistant’s magic wand comes to life. Here are the practical uses, from simplest to richest.
- Summarize a text
- Open Assistant (magic wand), then paste or select text and choose the “Summary” task. Ideal for a long email in Nextcloud Mail or an endless note.
- Rephrase / correct
- Rewrite a paragraph in a more formal tone, fix the style, shorten it. Useful in Nextcloud Text and Deck cards.
- Generate a draft
- Given an instruction (“write an invitation to Friday's meeting”), the assistant produces a first draft to revise.
- Extract the key points
- Turn meeting notes into a list of actions or decisions.
- Contextual chat
- The Assistant's chat tab enables free-form dialogue with the Ollama model, useful for brainstorming without leaving Nextcloud.
For “chatting with your files” in the strict sense—asking questions whose answers are in your documents—you need to provide context to the model. Two approaches: paste the file contents into the chat for a one-off document, or deploy a “Context Chat” ExApp that indexes your files and feeds the model through semantic search—the equivalent of RAG built into Nextcloud. The first is immediate; the second requires an additional ExApp and more resources.
#Size the server
The right hardware depends on the number of simultaneous users and your ambitions (occasional summaries vs. permanent RAG). Here are realistic guidelines.
- Family use (1–5 people)
- 7–8B model in Q4_K_M. An entry-level GPU such as RTX 3060 12 GB, or a Apple Apple Silicon Mac with unified memory, is more than enough for summaries and drafting. Without a GPU, it runs on the CPU in a few seconds to tens of seconds per request.
- Small business (5–20 people)
- 7–14B model. A RTX 4070/4080 16 GB handles spikes; a 14B in Q4 fits in ~9 GB. Allow plenty of system RAM for the rest of the Nextcloud stack.
- Sustained use / Context Chat
- Add the embeddings and vector store overhead. Target a RTX 4090 24 GB or an M4 Pro Mac with 24–48 GB of unified memory, and monitor “ollama ps” to verify that the model stays on the GPU.
- CPU only
- Suitable for a low-demand household: favor a 3B–7B and accept the latency. Beyond that, a GPU quickly becomes essential for a smooth experience.
#Troubleshooting
- “Connection refused” / no models listed
- Nextcloud cannot attach Ollama. Most often, the URL points to localhost while Nextcloud is running in Docker. Test the URL with curl from the Nextcloud container and use host.docker.internal or the host’s IP address.
- The Deploy Daemon stays red
- AppAPI cannot access Docker. Check the /var/run/docker.sock mount and permissions; on AIO, restart the master container. Without a green daemon, no ExApp can be deployed.
- The Assistant displays no task
- No provider is registered. Check that “OpenAI and LocalAI integration” (or the ExApp back end) is enabled and configured correctly, then reload the page.
- The requested model cannot be found
- The name doesn't match “ollama list.” Enter the exact tag (e.g., qwen3.5:9b) and verify that it has been pulled (“ollama pull”).
- Very slow responses
- The model is running on the CPU due to insufficient VRAM. Check with “ollama ps” (GPU vs. CPU) and drop down one model size or quantization level.
- English responses
- Some models respond in English by default. Choose a strong model in French (Qwen, Mistral) or add the instruction "respond in French" to the integration's system prompt.
#Go further
This setup reuses components already covered on the site. These guides build on this one:
- Install Ollama: Windows, macOS, and Linux
- The basic installation guide if you're starting from scratch before connecting Nextcloud to it.
- Choose your quantization (Q4, Q5, Q8, FP16)
- To balance model size against the VRAM available on your Nextcloud server.
- Deploy an AI chatbot for your team on the intranet
- To move toward a broader Ollama multi-user stack around your cloud.
Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.