Intermediate 12 minNo-code

Flowise: build AI agents with drag-and-drop on Ollama

Flowise turns LLM agent building into block assembly on a canvas. Instead of writing LangChain code, you drag nodes, draw connections, and test live in an integrated chat. Connected to Ollama, the flowise ollama combination runs chatflows, tool-using agents, and RAG pipelines entirely locally, without a single request leaving your machine. This guide starts with installation, builds a first chatflow, adds document RAG, then publishes an embeddable chatbot on your site.

By Marie L.·Update 2026-09-17·Tested on Windows, macOS, and Linux

#Why Flowise

Flowise is an open-source visual builder (Apache 2.0 license) built on LangChain and LangGraph. It exposes the same building blocks—models, memory, retrievers, and tools—as connectable nodes, turning an agent prototype from several hundred lines of Python into a graph you can understand at a glance.

Compared with Dify, the site’s other no-code platform, Flowise focuses more on fine-grained orchestration: where Dify offers a highly structured product experience (apps, prompt sets, team management), Flowise exposes the LangChain plumbing and is better suited when you want to understand and adjust every step of an agent. Both connect to Ollama in the same way.

100% local
Paired with Ollama, no token or document is sent to a cloud API. Ideal for sensitive data or offline use.
Fast iteration
The test chat is integrated into the canvas: edit a node and see the effect immediately, without redeploying anything.
API and widget exposure
Each chatflow automatically becomes a REST endpoint and an embeddable web widget, without writing a backend.
Extensible
Template marketplace, custom JavaScript nodes, and tool integration (web search, calculator, HTTP calls).

#Prerequisites

The Local Copilot Kit

This guide gets you to the model. The kit gets you to the coding copilot in your editor.

  • Lifetime online access
  • PDF + files
  • Lifetime updates

Flowise is lightweight: Ollama and the loaded model are what consume memory. Plan for a machine capable of running the target LLM comfortably.

Ollama installed
The daemon must be running and listening on http://localhost:11434 (the default value). Check with “ollama list”.
A chat model
A model capable of tool calling for agents: Qwen 3.5 8B, Granite 4.2 8B, or Mistral Small 24B depending on your VRAM (7B ≈ 5 GB, 14B ≈ 9 GB, 24B ≈ 16 GB in Q4_K_M).
An embeddings model
For RAG: “nomic-embed-text” or “mxbai-embed-large,” both lightweight and available through Ollama.
Node.js 18.15+ or Docker
Flowise can be installed via npm or in a container. Docker is recommended for a clean, persistent deployment.
i
Which model for agents
Flowise agents rely on tool calling: the model must decide when to call a tool. 3B models are often unable to do this reliably. Aim for at least a recent 7–8B model, ideally one explicitly trained for function calling.

#Install Flowise

Two methods. The fastest way to test is npx; for sustained use with persistent data, prefer Docker.

Quick installation via npm
# Lancement direct, sans installation permanente
npx flowise start

# Ou installation globale puis démarrage
npm install -g flowise
flowise start
Installation via Docker (recommended)
docker run -d \
  --name flowise \
  -p 3000:3000 \
  --add-host=host.docker.internal:host-gateway \
  -v ~/.flowise:/root/.flowise \
  flowiseai/flowise

Then open http://localhost:3000 in your browser: the canvas interface appears. The mounted volume (~/.flowise) preserves your chatflows and keys across container restarts.

→
Secure access from the start
By default, the interface is open. Add « -e FLOWISE_USERNAME=admin -e FLOWISE_PASSWORD=motdepasse » to the Docker command to protect access with a username and password, especially if the machine is accessible on the network.

#Connect Ollama to Flowise

The “ChatOllama” node is the bridge between Flowise and your local daemon. The key point is the base URL: it depends on how Flowise is launched.

  1. 01
    Add the ChatOllama node
    In the canvas, open the node panel, select the “Chat Models” category, and drag “ChatOllama” onto the workspace.
  2. 02
    Enter the base URL
    If Flowise runs natively (npx/npm), use http://localhost:11434. If it runs in Docker, use http://host.docker.internal:11434—the container cannot see the host’s “localhost.”
  3. 03
    Choose the model
    In the “Model Name” field, enter the exact name of the model loaded in Ollama, for example “qwen3.5:8b” or “mistral-small”.
  4. 04
    Adjust the settings
    Adjust Temperature (0.7 for chat, 0.1–0.3 for factual RAG) and, if needed, the context size via num_ctx in the advanced options.
!
The “fetch failed” error in Docker
Nine times out of ten, a chatflow that returns “fetch failed” or “ECONNREFUSED” is caused by a localhost URL inside a container. Replace it with host.docker.internal:11434 and make sure the Docker command includes --add-host=host.docker.internal:host-gateway on Linux.

#The essential nodes of a chatflow

A chatflow reads from left to right: provider nodes (model, memory, tools) feed a “chain” or “agent” node that produces the response. Here are the handful of blocks that appear in almost every flow.

Chat Model (ChatOllama)
The brain: the LLM that generates the responses. Always present.
Memory (Buffer Memory)
Keeps the conversation history so the agent can maintain context from one turn to the next.
Prompt Template
Defines the system instructions and the question's formatting. This is where you give the model a role and a framework.
Chain / Agent
The terminal node. “Conversation Chain” for a simple chat; “Tool Agent” when the model needs to decide whether to call tools.
Tools
External capabilities: calculator, web search, HTTP requests, file reading. Connected to an agent, they expand what it can do.
Document Loaders & Vector Store
The RAG component: loads documents, splits them, indexes them, and makes them searchable (see below).

#Build a first chatflow

Let’s start with the simplest option: a conversational assistant with memory, connected to Ollama. It will serve as the foundation for everything else.

  1. 01
    Create a new Chatflow
    From the home screen, click “Add New” in the Chatflows tab. A blank canvas opens.
  2. 02
    Set up the three core nodes
    Drag “ChatOllama,” “Buffer Memory,” and “Conversation Chain” onto the canvas.
  3. 03
    Wire up the connections
    Connect the ChatOllama output to the Conversation Chain's “Chat Model” input, and the Buffer Memory output to the “Memory” input of the same chain.
  4. 04
    Customize the system prompt
    In the Conversation Chain, open the System Message field and assign a role: “You are a concise technical assistant that responds in French.”
  5. 05
    Test it in the embedded chat
    Click the chat icon in the upper-right corner, ask a question, then ask a second question that depends on the first to verify that memory works.
→
Back up before testing
Flowise only runs a chatflow if it has been saved. Give it a name and click the save icon (floppy disk) before opening the chat; otherwise, your latest changes won’t be applied.

#A complete drag-and-drop RAG

RAG (Retrieval-Augmented Generation) lets the model answer from your own documents. In Flowise, everything happens on the canvas: load the files, vectorize them, and connect the retriever to a question-answering chain.

  1. 01
    Load the documents
    Add a “Document Loader” node suited to your source: PDF File, Text File, or Folder. It reads the raw content.
  2. 02
    Split into chunks
    Connect a “Recursive Character Text Splitter” to the loader. Set the chunk size to around 1000 characters with an overlap of 100 to preserve context between chunks.
  3. 03
    Generate embeddings
    Add an “Ollama Embeddings” node with the “nomic-embed-text” model and the same base URL as ChatOllama.
  4. 04
    Index in a vector store
    Add an “In-Memory Vector Store” (simple, to get started) or “Chroma” for persistence. Connect the splitter and embeddings to it.
  5. 05
    Connect a Retrieval QA Chain
    Connect the vector store (as a retriever) and ChatOllama to a “Retrieval QA Chain” or “Conversational Retrieval QA Chain” node to preserve conversation memory.
  6. 06
    Query your documents
    Save your files, open the chat, and ask a question whose answer is in them. The model now cites your content.
Retrieve the models required for RAG
ollama pull qwen3.5:8b
ollama pull nomic-embed-text
i
In-Memory vs Chroma
An in-memory vector store is perfect for prototyping: quick to wire up, but it is emptied on every restart and re-indexes the documents each time. As soon as the corpus grows or needs to persist, switch to Chroma or Qdrant so you only index it once.

#Publish and integrate the chatbot

Once the chatflow is satisfactory, Flowise exposes it in two ways without a line of backend code: a REST API and a web widget you can paste into your site.

Click “API Endpoint” or the </> icon at the top right of the canvas. Flowise generates the prediction URL and ready-to-copy examples. The endpoint always follows the same pattern, with the chatflow's unique identifier.

Call the chatflow via cURL
curl http://localhost:3000/api/v1/prediction/<CHATFLOW_ID> \
  -H "Content-Type: application/json" \
  -d '{"question": "Résume la politique de retour en une phrase."}'

For web integration, the “Embed” tab provides a script snippet to place before the closing </body> tag of your pages. The widget displays a fully configurable floating chat bubble (colors, welcome message, avatar).

Embeddable widget to paste into your page
<script type="module">
  import Chatbot from 'https://cdn.jsdelivr.net/npm/flowise-embed/dist/web.js'
  Chatbot.init({
    chatflowid: '<CHATFLOW_ID>',
    apiHost: 'http://localhost:3000',
  })
</script>
!
Do not expose localhost in production
The apiHost in http://localhost:3000 only works on your machine. For a public site, Flowise must run on an accessible server (behind an HTTPS reverse proxy), and apiHost must point to that domain. Also enable an API key on the chatflow so anyone can't call your model.

#Troubleshooting common pitfalls

“fetch failed” in ChatOllama
Incorrect base URL. In Docker, use host.docker.internal:11434, not localhost. Also check that Ollama is actually running (ollama list).
The model never calls tools
The LLM does not support tool calling, or supports it poorly. Switch to a recent model trained for it (Qwen 3.5, Mistral Small) and verify that you are using a “Tool Agent,” not a simple chain.
Slow or truncated responses
Context too large for VRAM: the model spills over to the CPU. Reduce num_ctx, the RAG chunk size, or choose a smaller model.
RAG finds nothing relevant
Poorly sized chunks or the wrong embedding model. Adjust the chunk size, increase the number of documents returned (top-k), and verify that indexing completed correctly.
Data disappears on restart
In Docker, without a mounted volume, everything is lost. Make sure you have « -v ~/.flowise:/root/.flowise » and use persistent vector storage rather than in-memory storage.

#Go further

Flowise is only a visual layer: the quality of your agents depends mainly on the model and the stack underneath. Three guides on the site build on this one — properly install Ollama before connecting Flowise, compare Dify’s no-code approach with Flowise’s, and understand how RAG works in Python so you can fine-tune what the canvas automates.


Did this guide help you?

Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.