Intermediate 13 minNo-code

Self-hosted Dify: create AI apps on your LLM local

Dify is an open-source platform for building AI applications—chatbots, workflows, agents, RAG assistants—without writing a single line of code, through a visual interface. Self-hosted and connected to Ollama, it gives you a complete no-code layer over your local models: your prompts, documents, and data never leave your machine. This guide covers deployment with Docker Compose, connecting to a local LLM, and building a functional first RAG assistant.

By Marie L.·Update 2026-08-27·Tested on Windows, macOS, and Linux

#Why Dify instead of a chat interface

Open WebUI or LM Studio is enough to chat with a model. Dify aims a level higher: building reusable applications. You define a system prompt, connect a knowledge base to it, expose everything through an API or embeddable chat widget, and version your iterations. It is the local, open-source equivalent of OpenAI Assistants or Coze.

The Dify + Ollama combination has two main benefits. First, privacy: unlike Dify connected to GPT-4 or Claude, indexed documents and conversations stay on your infrastructure. Second, cost: no tokens are billed, so you can iterate over hundreds of prompts without watching an API bill.

Chatbot
A conversational assistant with a system prompt, input variables, and conversation memory.
Agent
An assistant capable of calling tools (web search, calculations, APIs) in a loop until it solves the task.
Workflow
A visual sequence of nodes (LLM, condition, extraction, HTTP) for deterministic processing.
Knowledge / RAG
Indexing your documents so apps can answer using them, with source citations.
i
no-code, not zero-config
Dify avoids writing application code, but it remains a technical tool: you work with prompts, variables, and document-splitting parameters. Allow a good hour to get comfortable with the tool beyond the basic installation.

#Prerequisites

The Local Copilot Kit

This guide gets you to the model. The kit gets you to the coding copilot in your editor.

  • Lifetime online access
  • PDF + files
  • Lifetime updates

Dify is a multi-container application (API, worker, frontend, PostgreSQL database, Redis, vector database). It is deployed via Docker Compose. On the model side, we assume that Ollama is already running and listening on its default port.

Docker + Docker Compose
Recent Docker Engine with the Compose plugin (docker compose command). On Windows/macOS, Docker Desktop is sufficient.
Ollama working
The Ollama daemon is installed and reachable at http://localhost:11434. Test it with ollama list before you begin.
A chat model
For example, qwen3.5:9b (256k context, multimodal, Apache 2.0), the reference 8 GB choice in 2026. In Q4, it fits in ~6.6 GB of VRAM (RTX 3060 12 GB is more than sufficient).
An embeddings model
Essential for RAG: nomic-embed-text is the default choice, lightweight and effective.
Resources
Plan for ~8 GB of RAM for the Dify stack itself, in addition to the VRAM/RAM consumed by your Ollama models.
Terminal — prepare the models
# Vérifier qu'Ollama répond
ollama list

# Modèle de chat (choisissez selon votre VRAM)
ollama pull qwen3.5:9b

# Modèle d'embeddings pour le RAG
ollama pull nomic-embed-text

#Install Dify with Docker Compose

Dify provides an official repository with a ready-to-use Docker folder. Clone it, copy the example environment file, and launch the stack.

  1. 01
    Clone the repository
    Get the project's latest stable version from GitHub, then go to the docker folder containing the docker-compose.yaml file.
  2. 02
    Create the .env file
    Copy .env.example to .env. The default values are sufficient for local use; this is where you can adjust ports or secrets later.
  3. 03
    Start the stack
    Run docker compose up -d. The first startup downloads the images and initializes the PostgreSQL database; allow a few minutes.
  4. 04
    Create the admin account
    Open http://localhost/install in the browser and enter the first administrator's email address and password. This account will manage the workspace.
Terminal — deployment
git clone https://github.com/langgenius/dify.git
cd dify/docker

cp .env.example .env

docker compose up -d

# Vérifier que les conteneurs tournent
docker compose ps
→
The interface is on port 80
By default, the Dify frontend is exposed on http://localhost (port 80), not on an unusual application port. If port 80 is already occupied, modify EXPOSE_NGINX_PORT in the .env file before rerunning docker compose up -d.

#Connect Ollama as a model provider

Dify only knows about your local models if you declare Ollama as a provider. This is done in the model settings, not in a config file. The main thing to watch is the URL: from a Docker container, localhost refers to the container itself, not the host machine where Ollama is running.

  1. 01
    Open provider settings
    Click your avatar in the upper-right corner → Settings → Model Providers. Find Ollama in the list and select it.
  2. 02
    Enter the server URL
    In the Base URL field, enter the address reachable from the container (see the callout below). The model name must exactly match what ollama list returns, for example qwen3.5:9b.
  3. 03
    Add the chat model
    Model type: LLM. Enter the context size (for example, 8192 or more depending on the model) and confirm. Dify tests the connection to the entry.
  4. 04
    Add the embedding model
    Repeat the process with nomic-embed-text, choosing the Text Embedding type. Without it, you won't be able to build a knowledge base.
  5. 05
    Set the default models
    Still in Settings, set qwen3.5:9b as the default system model and nomic-embed-text as the default embedding model.
!
localhost doesn’t work from inside the container
Ollama runs on the host, but Dify runs in Docker. Use http://host.docker.internal:11434 with Docker Desktop (Windows/macOS). On Linux, use the Docker gateway IP (often http://172.17.0.1:11434) or run Ollama with OLLAMA_HOST=0.0.0.0, then point it to the host machine's IP address on the local network.
Terminal — expose Ollama on the network (Linux)
# Rendre Ollama joignable au-delà de localhost
# (à ajouter dans le service systemd ou l'environnement)
OLLAMA_HOST=0.0.0.0 ollama serve

# Vérifier depuis l'hôte que l'API répond
curl http://localhost:11434/api/tags

#Build a first RAG assistant without coding

RAG (Retrieval-Augmented Generation) lets the model answer from your documents rather than relying solely on its training knowledge. In Dify, this uses a knowledge base (Knowledge) that you then attach to an application.

  1. 01
    Create a knowledge base
    Knowledge tab → Create. Import your files (PDF, Markdown, TXT, DOCX). Dify automatically splits them into chunks.
  2. 02
    Adjust chunking and indexing
    Choose the “high quality” indexing mode, which uses your nomic-embed-text embedding model. Adjust the chunk size if your documents are highly structured (tables, code).
  3. 03
    Create the chat application
    Studio tab → Create an app → Chatbot. Give it a name and description.
  4. 04
    Write the system prompt
    In the editor, describe the assistant’s role: “You answer only from the provided documents. If the information isn’t there, say so.” This is what limits hallucinations.
  5. 05
    Attach the knowledge base
    In the app's Context panel, add the database created in step 1. Enable source citations so responses display the excerpts used.
  6. 06
    Test, then publish
    Use the debugging panel on the right to ask questions. When you're satisfied with the behavior, click Publish to get a chat URL and an API key.
→
The embedding model matters as much as the chat model
RAG quality depends first on the relevance of the retrieved passages. A good embedding model (nomic-embed-text, or mxbai-embed-large if you have the headroom) significantly improves responses, even with a modest chat model. Don't invest everything in a large LLM while neglecting embeddings.

#Workflows and agents: how far can no-code go?

Beyond a simple chatbot, Dify offers two more advanced modes. Workflow mode provides a visual canvas where you connect nodes: user input, an LLM call, a condition (if/else), parameter extraction, an HTTP request, and iteration over a list. This lets you build deterministic pipelines—for example: receive an email, classify it, extract the entities, then draft a standard reply.

The Agent mode, by contrast, lets the model decide which tools to call and in what order, in a loop, until it reaches a result. It is more powerful but more fragile: quality depends heavily on the model's ability to reason and follow the tool-call format. Very small local models (2–4B) struggle; target a model designed for tool use, such as GLM 4.7 Flash (30B-A3B MoE, MIT, ~19 GB) or Qwen 3.8 27B (~18 GB) if your VRAM allows it.

What no-code does well
Rapidly prototype a RAG chatbot, chain together a few LLM steps, expose an API without writing a backend, and iterate on prompts as a team.
Where local deployment hits its limits
Demanding multi-tool agents require a model capable of reliable tool use, which means VRAM. A Qwen 3.5 9B is enough for RAG, but rarely for a complex agent.
When to switch to code
Fine-grained business logic, custom integrations, and full control over the retrieval pipeline: a library like LangChain takes over where the visual canvas reaches its limits.

#Troubleshooting

“Connection refused” when adding the model
Dify cannot reach Ollama. The problem is almost always the URL: replace localhost with host.docker.internal (Docker Desktop) or the Docker gateway IP (Linux), and check that Ollama is listening on 0.0.0.0.
The model does not appear
The specified name does not match ollama list. Copy the exact name, including the tag (qwen3.5:9b, not qwen3.5).
Error indexing documents
The embedding model is not configured or has not been downloaded. Run ollama pull nomic-embed-text and select it as the default embedding model.
Very slow responses
The chat model is probably spilling onto the CPU. Switch to a lighter quantization (Q4_K_M) or a smaller model, and verify that the GPU is actually being used on the Ollama side.
Port 80 is already in use
Another service is using the port. Change EXPOSE_NGINX_PORT in the .env (e.g., 8080), then run docker compose up -d again.
Off-topic responses despite RAG
Make sure the knowledge base is properly attached to the app and that the system prompt constrains the model to rely on the context. Also adjust the chunk size.

#Go further

Dify is only as good as the model and infrastructure supporting it. These guides strengthen the building blocks it relies on.

Install Ollama
The daemon that serves your chat model and embeddings on port 11434—the foundation of the entire local Dify stack.
Local RAG with ChromaDB and Ollama
To understand what Dify automates under the hood and take back control in Python when no-code reaches its limits.
n8n + Ollama: automate locally
Another no-code approach focused on task automation, complementary to Dify workflows.
Did this guide help you?

Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.