Flowise: build AI agents with drag-and-drop on Ollama
Flowise turns LLM agent building into block assembly on a canvas. Instead of writing LangChain code, you drag nodes, draw connections, and test live in an integrated chat. Connected to Ollama, the flowise ollama combination runs chatflows, tool-using agents, and RAG pipelines entirely locally, without a single request leaving your machine. This guide starts with installation, builds a first chatflow, adds document RAG, then publishes an embeddable chatbot on your site.
#Why Flowise
Flowise is an open-source visual builder (Apache 2.0 license) built on LangChain and LangGraph. It exposes the same building blocks—models, memory, retrievers, and tools—as connectable nodes, turning an agent prototype from several hundred lines of Python into a graph you can understand at a glance.
Compared with Dify, the site’s other no-code platform, Flowise focuses more on fine-grained orchestration: where Dify offers a highly structured product experience (apps, prompt sets, team management), Flowise exposes the LangChain plumbing and is better suited when you want to understand and adjust every step of an agent. Both connect to Ollama in the same way.
- 100% local
- Paired with Ollama, no token or document is sent to a cloud API. Ideal for sensitive data or offline use.
- Fast iteration
- The test chat is integrated into the canvas: edit a node and see the effect immediately, without redeploying anything.
- API and widget exposure
- Each chatflow automatically becomes a REST endpoint and an embeddable web widget, without writing a backend.
- Extensible
- Template marketplace, custom JavaScript nodes, and tool integration (web search, calculator, HTTP calls).
#Prerequisites
This guide gets you to the model. The kit gets you to the coding copilot in your editor.
- Lifetime online access
- PDF + files
- Lifetime updates
Flowise is lightweight: Ollama and the loaded model are what consume memory. Plan for a machine capable of running the target LLM comfortably.
- Ollama installed
- The daemon must be running and listening on http://localhost:11434 (the default value). Check with “ollama list”.
- A chat model
- A model capable of tool calling for agents: Qwen 3.5 8B, Granite 4.2 8B, or Mistral Small 24B depending on your VRAM (7B ≈ 5 GB, 14B ≈ 9 GB, 24B ≈ 16 GB in Q4_K_M).
- An embeddings model
- For RAG: “nomic-embed-text” or “mxbai-embed-large,” both lightweight and available through Ollama.
- Node.js 18.15+ or Docker
- Flowise can be installed via npm or in a container. Docker is recommended for a clean, persistent deployment.
#Install Flowise
Two methods. The fastest way to test is npx; for sustained use with persistent data, prefer Docker.
Then open http://localhost:3000 in your browser: the canvas interface appears. The mounted volume (~/.flowise) preserves your chatflows and keys across container restarts.
#Connect Ollama to Flowise
The “ChatOllama” node is the bridge between Flowise and your local daemon. The key point is the base URL: it depends on how Flowise is launched.
- 01Add the ChatOllama nodeIn the canvas, open the node panel, select the “Chat Models” category, and drag “ChatOllama” onto the workspace.
- 02Enter the base URLIf Flowise runs natively (npx/npm), use http://localhost:11434. If it runs in Docker, use http://host.docker.internal:11434—the container cannot see the host’s “localhost.”
- 03Choose the modelIn the “Model Name” field, enter the exact name of the model loaded in Ollama, for example “qwen3.5:8b” or “mistral-small”.
- 04Adjust the settingsAdjust Temperature (0.7 for chat, 0.1–0.3 for factual RAG) and, if needed, the context size via num_ctx in the advanced options.
#The essential nodes of a chatflow
A chatflow reads from left to right: provider nodes (model, memory, tools) feed a “chain” or “agent” node that produces the response. Here are the handful of blocks that appear in almost every flow.
- Chat Model (ChatOllama)
- The brain: the LLM that generates the responses. Always present.
- Memory (Buffer Memory)
- Keeps the conversation history so the agent can maintain context from one turn to the next.
- Prompt Template
- Defines the system instructions and the question's formatting. This is where you give the model a role and a framework.
- Chain / Agent
- The terminal node. “Conversation Chain” for a simple chat; “Tool Agent” when the model needs to decide whether to call tools.
- Tools
- External capabilities: calculator, web search, HTTP requests, file reading. Connected to an agent, they expand what it can do.
- Document Loaders & Vector Store
- The RAG component: loads documents, splits them, indexes them, and makes them searchable (see below).
#Build a first chatflow
Let’s start with the simplest option: a conversational assistant with memory, connected to Ollama. It will serve as the foundation for everything else.
- 01Create a new ChatflowFrom the home screen, click “Add New” in the Chatflows tab. A blank canvas opens.
- 02Set up the three core nodesDrag “ChatOllama,” “Buffer Memory,” and “Conversation Chain” onto the canvas.
- 03Wire up the connectionsConnect the ChatOllama output to the Conversation Chain's “Chat Model” input, and the Buffer Memory output to the “Memory” input of the same chain.
- 04Customize the system promptIn the Conversation Chain, open the System Message field and assign a role: “You are a concise technical assistant that responds in French.”
- 05Test it in the embedded chatClick the chat icon in the upper-right corner, ask a question, then ask a second question that depends on the first to verify that memory works.
#A complete drag-and-drop RAG
RAG (Retrieval-Augmented Generation) lets the model answer from your own documents. In Flowise, everything happens on the canvas: load the files, vectorize them, and connect the retriever to a question-answering chain.
- 01Load the documentsAdd a “Document Loader” node suited to your source: PDF File, Text File, or Folder. It reads the raw content.
- 02Split into chunksConnect a “Recursive Character Text Splitter” to the loader. Set the chunk size to around 1000 characters with an overlap of 100 to preserve context between chunks.
- 03Generate embeddingsAdd an “Ollama Embeddings” node with the “nomic-embed-text” model and the same base URL as ChatOllama.
- 04Index in a vector storeAdd an “In-Memory Vector Store” (simple, to get started) or “Chroma” for persistence. Connect the splitter and embeddings to it.
- 05Connect a Retrieval QA ChainConnect the vector store (as a retriever) and ChatOllama to a “Retrieval QA Chain” or “Conversational Retrieval QA Chain” node to preserve conversation memory.
- 06Query your documentsSave your files, open the chat, and ask a question whose answer is in them. The model now cites your content.
#Publish and integrate the chatbot
Once the chatflow is satisfactory, Flowise exposes it in two ways without a line of backend code: a REST API and a web widget you can paste into your site.
Click “API Endpoint” or the </> icon at the top right of the canvas. Flowise generates the prediction URL and ready-to-copy examples. The endpoint always follows the same pattern, with the chatflow's unique identifier.
For web integration, the “Embed” tab provides a script snippet to place before the closing </body> tag of your pages. The widget displays a fully configurable floating chat bubble (colors, welcome message, avatar).
#Troubleshooting common pitfalls
- “fetch failed” in ChatOllama
- Incorrect base URL. In Docker, use host.docker.internal:11434, not localhost. Also check that Ollama is actually running (ollama list).
- The model never calls tools
- The LLM does not support tool calling, or supports it poorly. Switch to a recent model trained for it (Qwen 3.5, Mistral Small) and verify that you are using a “Tool Agent,” not a simple chain.
- Slow or truncated responses
- Context too large for VRAM: the model spills over to the CPU. Reduce num_ctx, the RAG chunk size, or choose a smaller model.
- RAG finds nothing relevant
- Poorly sized chunks or the wrong embedding model. Adjust the chunk size, increase the number of documents returned (top-k), and verify that indexing completed correctly.
- Data disappears on restart
- In Docker, without a mounted volume, everything is lost. Make sure you have « -v ~/.flowise:/root/.flowise » and use persistent vector storage rather than in-memory storage.
#Go further
Flowise is only a visual layer: the quality of your agents depends mainly on the model and the stack underneath. Three guides on the site build on this one — properly install Ollama before connecting Flowise, compare Dify’s no-code approach with Flowise’s, and understand how RAG works in Python so you can fine-tune what the canvas automates.
Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.