Self-hosted Dify: create AI apps on your LLM local
Dify is an open-source platform for building AI applications—chatbots, workflows, agents, RAG assistants—without writing a single line of code, through a visual interface. Self-hosted and connected to Ollama, it gives you a complete no-code layer over your local models: your prompts, documents, and data never leave your machine. This guide covers deployment with Docker Compose, connecting to a local LLM, and building a functional first RAG assistant.
#Why Dify instead of a chat interface
Open WebUI or LM Studio is enough to chat with a model. Dify aims a level higher: building reusable applications. You define a system prompt, connect a knowledge base to it, expose everything through an API or embeddable chat widget, and version your iterations. It is the local, open-source equivalent of OpenAI Assistants or Coze.
The Dify + Ollama combination has two main benefits. First, privacy: unlike Dify connected to GPT-4 or Claude, indexed documents and conversations stay on your infrastructure. Second, cost: no tokens are billed, so you can iterate over hundreds of prompts without watching an API bill.
- Chatbot
- A conversational assistant with a system prompt, input variables, and conversation memory.
- Agent
- An assistant capable of calling tools (web search, calculations, APIs) in a loop until it solves the task.
- Workflow
- A visual sequence of nodes (LLM, condition, extraction, HTTP) for deterministic processing.
- Knowledge / RAG
- Indexing your documents so apps can answer using them, with source citations.
#Prerequisites
This guide gets you to the model. The kit gets you to the coding copilot in your editor.
- Lifetime online access
- PDF + files
- Lifetime updates
Dify is a multi-container application (API, worker, frontend, PostgreSQL database, Redis, vector database). It is deployed via Docker Compose. On the model side, we assume that Ollama is already running and listening on its default port.
- Docker + Docker Compose
- Recent Docker Engine with the Compose plugin (docker compose command). On Windows/macOS, Docker Desktop is sufficient.
- Ollama working
- The Ollama daemon is installed and reachable at http://localhost:11434. Test it with ollama list before you begin.
- A chat model
- For example, qwen3.5:9b (256k context, multimodal, Apache 2.0), the reference 8 GB choice in 2026. In Q4, it fits in ~6.6 GB of VRAM (RTX 3060 12 GB is more than sufficient).
- An embeddings model
- Essential for RAG: nomic-embed-text is the default choice, lightweight and effective.
- Resources
- Plan for ~8 GB of RAM for the Dify stack itself, in addition to the VRAM/RAM consumed by your Ollama models.
#Install Dify with Docker Compose
Dify provides an official repository with a ready-to-use Docker folder. Clone it, copy the example environment file, and launch the stack.
- 01Clone the repositoryGet the project's latest stable version from GitHub, then go to the docker folder containing the docker-compose.yaml file.
- 02Create the .env fileCopy .env.example to .env. The default values are sufficient for local use; this is where you can adjust ports or secrets later.
- 03Start the stackRun docker compose up -d. The first startup downloads the images and initializes the PostgreSQL database; allow a few minutes.
- 04Create the admin accountOpen http://localhost/install in the browser and enter the first administrator's email address and password. This account will manage the workspace.
#Connect Ollama as a model provider
Dify only knows about your local models if you declare Ollama as a provider. This is done in the model settings, not in a config file. The main thing to watch is the URL: from a Docker container, localhost refers to the container itself, not the host machine where Ollama is running.
- 01Open provider settingsClick your avatar in the upper-right corner → Settings → Model Providers. Find Ollama in the list and select it.
- 02Enter the server URLIn the Base URL field, enter the address reachable from the container (see the callout below). The model name must exactly match what ollama list returns, for example qwen3.5:9b.
- 03Add the chat modelModel type: LLM. Enter the context size (for example, 8192 or more depending on the model) and confirm. Dify tests the connection to the entry.
- 04Add the embedding modelRepeat the process with nomic-embed-text, choosing the Text Embedding type. Without it, you won't be able to build a knowledge base.
- 05Set the default modelsStill in Settings, set qwen3.5:9b as the default system model and nomic-embed-text as the default embedding model.
#Build a first RAG assistant without coding
RAG (Retrieval-Augmented Generation) lets the model answer from your documents rather than relying solely on its training knowledge. In Dify, this uses a knowledge base (Knowledge) that you then attach to an application.
- 01Create a knowledge baseKnowledge tab → Create. Import your files (PDF, Markdown, TXT, DOCX). Dify automatically splits them into chunks.
- 02Adjust chunking and indexingChoose the “high quality” indexing mode, which uses your nomic-embed-text embedding model. Adjust the chunk size if your documents are highly structured (tables, code).
- 03Create the chat applicationStudio tab → Create an app → Chatbot. Give it a name and description.
- 04Write the system promptIn the editor, describe the assistant’s role: “You answer only from the provided documents. If the information isn’t there, say so.” This is what limits hallucinations.
- 05Attach the knowledge baseIn the app's Context panel, add the database created in step 1. Enable source citations so responses display the excerpts used.
- 06Test, then publishUse the debugging panel on the right to ask questions. When you're satisfied with the behavior, click Publish to get a chat URL and an API key.
#Workflows and agents: how far can no-code go?
Beyond a simple chatbot, Dify offers two more advanced modes. Workflow mode provides a visual canvas where you connect nodes: user input, an LLM call, a condition (if/else), parameter extraction, an HTTP request, and iteration over a list. This lets you build deterministic pipelines—for example: receive an email, classify it, extract the entities, then draft a standard reply.
The Agent mode, by contrast, lets the model decide which tools to call and in what order, in a loop, until it reaches a result. It is more powerful but more fragile: quality depends heavily on the model's ability to reason and follow the tool-call format. Very small local models (2–4B) struggle; target a model designed for tool use, such as GLM 4.7 Flash (30B-A3B MoE, MIT, ~19 GB) or Qwen 3.8 27B (~18 GB) if your VRAM allows it.
- What no-code does well
- Rapidly prototype a RAG chatbot, chain together a few LLM steps, expose an API without writing a backend, and iterate on prompts as a team.
- Where local deployment hits its limits
- Demanding multi-tool agents require a model capable of reliable tool use, which means VRAM. A Qwen 3.5 9B is enough for RAG, but rarely for a complex agent.
- When to switch to code
- Fine-grained business logic, custom integrations, and full control over the retrieval pipeline: a library like LangChain takes over where the visual canvas reaches its limits.
#Troubleshooting
- “Connection refused” when adding the model
- Dify cannot reach Ollama. The problem is almost always the URL: replace localhost with host.docker.internal (Docker Desktop) or the Docker gateway IP (Linux), and check that Ollama is listening on 0.0.0.0.
- The model does not appear
- The specified name does not match ollama list. Copy the exact name, including the tag (qwen3.5:9b, not qwen3.5).
- Error indexing documents
- The embedding model is not configured or has not been downloaded. Run ollama pull nomic-embed-text and select it as the default embedding model.
- Very slow responses
- The chat model is probably spilling onto the CPU. Switch to a lighter quantization (Q4_K_M) or a smaller model, and verify that the GPU is actually being used on the Ollama side.
- Port 80 is already in use
- Another service is using the port. Change EXPOSE_NGINX_PORT in the .env (e.g., 8080), then run docker compose up -d again.
- Off-topic responses despite RAG
- Make sure the knowledge base is properly attached to the app and that the system prompt constrains the model to rely on the context. Also adjust the chunk size.
#Go further
Dify is only as good as the model and infrastructure supporting it. These guides strengthen the building blocks it relies on.
- Install Ollama
- The daemon that serves your chat model and embeddings on port 11434—the foundation of the entire local Dify stack.
- Local RAG with ChromaDB and Ollama
- To understand what Dify automates under the hood and take back control in Python when no-code reaches its limits.
- n8n + Ollama: automate locally
- Another no-code approach focused on task automation, complementary to Dify workflows.
Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.