Intermediate 12 minNextcloud

Nextcloud Assistant + Ollama: AI in your cloud self-hosted

Nextcloud now has its own AI assistant: email summaries, rewriting, text generation, and document chat, directly in the interface. By default, it pushes you toward cloud services, but you don't have to use them. This guide shows how to connect Nextcloud Assistant to a self-hosted Ollama to keep 100% of your data at home: step-by-step AppAPI and ExApp configuration, OpenAI-compatible integration, recommended models, and server sizing for family or small-business use.

By Marie L.·Update 2026-08-27·Tested on Windows, macOS, and Linux

#Why connect Nextcloud to Ollama

Nextcloud Assistant is the AI interface integrated into Nextcloud (the “magic wand” icon in the top bar): it brings together tasks such as summarizing text, rephrasing it, generating a draft, extracting key points, and translating. Since Nextcloud Hub, a “contextual chat” assistant also lets you interact with a model and, with the right context, query your own files.

The problem: the backends offered by default or in the demo often point to cloud APIs (OpenAI or Nextcloud's hosted services). For a self-hosted cloud, this reintroduces exactly what we wanted to avoid—sending private content to a third party. Connecting the assistant to Ollama, which runs on your own machine and listens on http://localhost:11434, keeps your document text strictly on your infrastructure.

Data sovereignty
No file, email, or note content leaves your server. Useful for GDPR compliance and simple peace of mind.
Zero usage cost
No per-token billing: once the hardware has paid for itself, summaries and writing are free and unlimited.
Models to choose from
You choose the model (Qwen, Gemma, Granite) and its quantization based on your VRAM, instead of being stuck with a prescribed model.
Works offline
The assistant remains available even without internet access, as long as the Nextcloud server and Ollama are running.
i
Nextcloud Assistant vs Nextcloud AI
“Assistant” is the interface (the magic wand and the chat). It does nothing by itself: it delegates to AI task “providers” supplied by back-end applications. This is the back end we’ll point to Ollama—the interface itself does not change.

#How it works: AppAPI and ExApp

The AI at Work Kit

Deploy local AI at work: privacy, compliance, multi-user architecture, costs, the one-page memo for leadership.

  • Lifetime online access
  • PDF + files
  • Lifetime updates

Nextcloud does not talk directly to Ollama. An external-app mechanism sits between them. Understanding the three components prevents a lot of configuration guesswork.

Assistant
The Nextcloud-integrated front-end app that exposes tasks (summarization, rewriting, chat) in the user interface.
AppAPI
The Nextcloud app that manages the life cycle of external apps (ExApps): installation, deployment, and communication. It is the foundation to install first.
AI back-end ExApp
An external application (container) that implements task providers and knows how to communicate with an inference engine. The most commonly used for local deployments are: «integration_openai» (compatible with any OpenAI endpoint, including Ollama) and «LocalAI» through the dedicated ExApp.
Deploy Daemon
The service (often Docker) that AppAPI controls to start and stop ExApp containers. Without a working Deploy Daemon, no ExApp can be deployed.

The path of a request is therefore: you click “Summarize” in the Assistant → Nextcloud delegates the task to the provider → the ExApp (for example, integration_openai) formats an OpenAI-format call → that call is sent to the Ollama endpoint → the response follows the same path back. The key is to point the integration ExApp to the Ollama URL, in OpenAI-compatible mode.

→
The integration_openai shortcut
Ollama exposes an OpenAI-compatible API at /v1 (http://IP:11434/v1). Nextcloud’s “OpenAI and LocalAI integration” app, configured with this URL, is sufficient in many cases and avoids managing a complete containerized ExApp. We detail both approaches below.

#Prerequisites

Recent Nextcloud Hub
A version supporting AppAPI and Assistant (Hub 6/7 or later). Ideally a Docker/AIO installation or a server you administer, so you can install AppAPI and a Deploy Daemon.
Ollama installed and accessible
On the same machine or a machine on the network. The daemon listens on http://localhost:11434 by default. Check with “ollama --version” and “ollama ps”.
Docker for the Deploy Daemon
AppAPI deploys ExApps through a Docker daemon. On Nextcloud AIO, this is integrated; on a manual installation, you need Docker access (socket or API).
A little VRAM or RAM
Q4 rule of thumb: a 7B fits in ~5 GB, and a 14B in ~9 GB. Without a GPU, RAM takes over, more slowly—acceptable for occasional summaries.
Nextcloud admin access
Assistant, AppAPI, and integration settings are configured under “Administration settings.”

#1. Prepare Ollama and choose a model

Start with a versatile, French-speaking instruct model. For family use or a small team, Qwen 3.5 9B in Q4_K_M offers the best quality/memory tradeoff: it summarizes and reformulates very well in French, handles a very large context (256k tokens), and fits in ~6.6 GB of VRAM. Even leaner, Granite 4.2 8B (IBM, Apache 2.0) drops to ~5.3 GB for a more modest server.

Terminal — retrieve a model
# Modèle polyvalent, bon en français, léger
ollama pull qwen3.5:9b

# Alternative très sobre (IBM, Apache 2.0)
ollama pull granite4.2:8b

# Vérifier qu'il répond
ollama run qwen3.5:9b "Résume en une phrase : Nextcloud est un cloud auto-hébergé."

Critical point: Ollama must be reachable from Nextcloud. If Nextcloud runs in Docker, “localhost” refers to the container, not your machine. Have Ollama listen on all interfaces and point the integration to the host's IP (or host.docker.internal, depending on the platform).

Terminal — expose Ollama on the network (systemd)
# Éditer le service
sudo systemctl edit ollama

# Ajouter dans la section [Service] :
#   Environment="OLLAMA_HOST=0.0.0.0:11434"

# Recharger et redémarrer
sudo systemctl daemon-reload
sudo systemctl restart ollama

# Depuis le serveur Nextcloud, tester l'accès
curl http://IP_DE_L_HOTE:11434/api/tags
!
Do not expose Ollama to the internet
OLLAMA_HOST=0.0.0.0 opens port 11434 across the entire network—convenient on a LAN, dangerous if the machine is exposed. The Ollama API has no authentication. Stay on the local network, or put an authenticated firewall / reverse proxy in front of it. Never forward this port from your router.

#2. Install AppAPI and a Deploy Daemon

In Nextcloud, AI runs through AppAPI. This step lays the foundation for deploying AI task ExApps.

  1. 01
    Install AppAPI
    Administration Settings → Applications → “Tools” category (or “Outils”) → search for “AppAPI” and enable it. It then appears as a dedicated entry in the admin settings.
  2. 02
    Configure a Deploy Daemon
    In Administration Settings → AppAPI, add a Docker-type “Deploy Daemon.” On Nextcloud AIO, the daemon is pre-provisioned; otherwise, enter access to the Docker socket (e.g. /var/run/docker.sock) and run the built-in connection test.
  3. 03
    Check status
    AppAPI displays the daemon status (green = ready). Until it turns green, the ExApps cannot deploy—there is no point going further before fixing this.
  4. 04
    Install the Assistant app
    Still in Applications, enable “Assistant” if it is not already there. This is what displays the magic wand in the interface.
i
Two possible approaches
You can either deploy a containerized back-end ExApp via AppAPI (e.g., the ExApp LocalAI/Ollama), or keep it simpler with the “OpenAI and LocalAI integration” app, which only needs a URL. To get started with Ollama, the latter approach is the most direct—it's the one detailed in the next step.

#3. Connect the OpenAI/LocalAI integration

The “OpenAI and LocalAI integration” application can communicate with any endpoint that follows the OpenAI API. Since Ollama exposes /v1, you only need to provide the correct URL and choose the model.

  1. 01
    Install the integration
    Applications → search for “OpenAI and LocalAI integration” and enable it. It registers task providers (text generation, summarization, etc.) that the Assistant can use.
  2. 02
    Enter Ollama's URL
    Administration settings → “Connected accounts” (or the integration section) → “Service URL” field: enter the OpenAI endpoint for Ollama, for example http://IP_DE_L_HOTE:11434/v1. Leave the API key blank or enter a dummy value: Ollama does not verify it.
  3. 03
    Choose the default model
    On the same page, select the model to use for text generation (e.g., qwen3.5:9b). If it does not appear in the list, enter its exact name as returned by “ollama list”.
  4. 04
    Save and test
    Validate. The integration queries the endpoint to list the models; an error here almost always indicates a URL or network problem (see Troubleshooting).
Terminal — validate Ollama's OpenAI endpoint
# Depuis le serveur Nextcloud, l'endpoint /v1 doit répondre
curl http://IP_DE_L_HOTE:11434/v1/models

# Et une complétion de chat au format OpenAI
curl http://IP_DE_L_HOTE:11434/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{
    "model": "qwen3.5:9b",
    "messages": [{"role": "user", "content": "Dis bonjour en une phrase."}]
  }'
→
URL depending on the location of Ollama
Ollama on the same machine outside Docker: http://localhost:11434/v1. Nextcloud in Docker and Ollama on the host: http://host.docker.internal:11434/v1 (or the host's LAN IP). Ollama on another machine: its network IP. Always test the URL with curl from the Nextcloud environment, not from your workstation.

#4. Summarize, write, and chat about your files

Once the provider is connected, the Assistant’s magic wand comes to life. Here are the practical uses, from simplest to richest.

Summarize a text
Open Assistant (magic wand), then paste or select text and choose the “Summary” task. Ideal for a long email in Nextcloud Mail or an endless note.
Rephrase / correct
Rewrite a paragraph in a more formal tone, fix the style, shorten it. Useful in Nextcloud Text and Deck cards.
Generate a draft
Given an instruction (“write an invitation to Friday's meeting”), the assistant produces a first draft to revise.
Extract the key points
Turn meeting notes into a list of actions or decisions.
Contextual chat
The Assistant's chat tab enables free-form dialogue with the Ollama model, useful for brainstorming without leaving Nextcloud.

For “chatting with your files” in the strict sense—asking questions whose answers are in your documents—you need to provide context to the model. Two approaches: paste the file contents into the chat for a one-off document, or deploy a “Context Chat” ExApp that indexes your files and feeds the model through semantic search—the equivalent of RAG built into Nextcloud. The first is immediate; the second requires an additional ExApp and more resources.

i
Context Chat = homegrown RAG
The “Context Chat” ExApp (via AppAPI) builds a vector index of your files and combines it with the Ollama model to answer with citations from your documents. It’s powerful but resource-intensive (embeddings + vector store): reserve it for a properly sized server, and start with simple chat to validate the pipeline.

#Size the server

The right hardware depends on the number of simultaneous users and your ambitions (occasional summaries vs. permanent RAG). Here are realistic guidelines.

Family use (1–5 people)
7–8B model in Q4_K_M. An entry-level GPU such as RTX 3060 12 GB, or a Apple Apple Silicon Mac with unified memory, is more than enough for summaries and drafting. Without a GPU, it runs on the CPU in a few seconds to tens of seconds per request.
Small business (5–20 people)
7–14B model. A RTX 4070/4080 16 GB handles spikes; a 14B in Q4 fits in ~9 GB. Allow plenty of system RAM for the rest of the Nextcloud stack.
Sustained use / Context Chat
Add the embeddings and vector store overhead. Target a RTX 4090 24 GB or an M4 Pro Mac with 24–48 GB of unified memory, and monitor “ollama ps” to verify that the model stays on the GPU.
CPU only
Suitable for a low-demand household: favor a 3B–7B and accept the latency. Beyond that, a GPU quickly becomes essential for a smooth experience.
→
Only one model loaded at a time
Ollama unloads and reloads models on demand. If Nextcloud and another service share the same Ollama, ideally keep only one model “hot” to avoid costly reloads. “OLLAMA_KEEP_ALIVE” keeps it in memory between requests.

#Troubleshooting

“Connection refused” / no models listed
Nextcloud cannot attach Ollama. Most often, the URL points to localhost while Nextcloud is running in Docker. Test the URL with curl from the Nextcloud container and use host.docker.internal or the host’s IP address.
The Deploy Daemon stays red
AppAPI cannot access Docker. Check the /var/run/docker.sock mount and permissions; on AIO, restart the master container. Without a green daemon, no ExApp can be deployed.
The Assistant displays no task
No provider is registered. Check that “OpenAI and LocalAI integration” (or the ExApp back end) is enabled and configured correctly, then reload the page.
The requested model cannot be found
The name doesn't match “ollama list.” Enter the exact tag (e.g., qwen3.5:9b) and verify that it has been pulled (“ollama pull”).
Very slow responses
The model is running on the CPU due to insufficient VRAM. Check with “ollama ps” (GPU vs. CPU) and drop down one model size or quantization level.
English responses
Some models respond in English by default. Choose a strong model in French (Qwen, Mistral) or add the instruction "respond in French" to the integration's system prompt.

#Go further

This setup reuses components already covered on the site. These guides build on this one:

Install Ollama: Windows, macOS, and Linux
The basic installation guide if you're starting from scratch before connecting Nextcloud to it.
Choose your quantization (Q4, Q5, Q8, FP16)
To balance model size against the VRAM available on your Nextcloud server.
Deploy an AI chatbot for your team on the intranet
To move toward a broader Ollama multi-user stack around your cloud.
Did this guide help you?

Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.