Flowise: Visual LLM Workflows Over a Local Model
Flowise is a drag-and-drop builder for chains and agents, self-hosted and model-agnostic. Wired to Ollama it keeps everything local. What the nodes actually do, where visual building helps, and where it starts fighting you.
Key takeaways
- Flowise is a visual builder for LLM applications: you connect nodes on a canvas — a model, a retriever, a memory, a tool — and get a working chat endpoint.
- It is self-hosted and model-agnostic, so pointing it at Ollama or any OpenAI-compatible server keeps the whole pipeline on your machine.
- Every flow can be published as an API or an embeddable widget, which is what makes it more than a prototyping toy.
- Its document store handles ingestion and retrieval, and like every such feature it needs a separate embedding model — the most common reason a new flow returns nothing useful.
- The canvas is an advantage while a flow is simple and a liability once it has branches, retries and conditions. That crossover point arrives sooner than people expect.
What it actually is
Replace GitHub Copilot and Cursor with a code assistant that runs 100% on your machine — the reference guide, configs included.
- Lifetime online access
- PDF + files
- 30-day refund
Every LLM application is a small graph: take the question, maybe retrieve context, maybe call a tool, send a prompt, return the answer. Writing that graph in code is unremarkable work — and it is work you repeat, in a language not everyone on the team reads.
Flowise turns that graph into a canvas. Nodes are the components you would otherwise import; edges are the data flowing between them. Two families of flow exist: chatflows, a fixed path through the graph, and agentflows, where a model decides which tools to call and in what order. The first is predictable and cheap; the second is flexible and costs an unknown number of generations.
Wiring it to a local model
- Self-host it. A container or a Node process; either way it runs on your infrastructure with its own database for flows and credentials.
- Add the model node pointing at your local server. From inside a container,
localhostmeans the container — use the host address Docker exposes and make sure Ollama listens beyond loopback. - Add an embedding model node if the flow retrieves documents. A chat model cannot embed; leaving this wrong produces a flow that runs and answers from nothing.
- Attach a vector store. The in-memory one is fine to prototype and forgets everything on restart. For anything real, point it at a persistent store — pgvector if you already have PostgreSQL.
- Test in the built-in chat, then publish as an API endpoint or a widget.
Agentflows multiply cost invisibly. On a hosted API you would notice in the bill. Locally you notice as a fan: an agent that decides to call three tools and re-read its history each time is a dozen generations for one question. Cap the iterations in the node settings before the first real test.
Model size decides which flows work
| Flow type | Minimum realistic model | Why |
|---|---|---|
| Simple chat with a system prompt | 7B–8B | No structure to respect beyond prose |
| Retrieval then answer | 7B–8B | The model summarises supplied text; retrieval does the hard part |
| Conditional routing | 12B–14B | It must choose a branch and say so in a fixed format |
| Tool-calling agent | 24B–32B | Malformed calls break the flow rather than degrading it |
This is the same ceiling every local agent framework hits: prose is easy, strict formats are not. The tool-use ranking is the shortlist worth starting from.
Flowise, Dify, or code
| Situation | Best fit |
|---|---|
| One or two flows, quick iteration, small team | Flowise |
| Many apps, user accounts, versioned prompts, logging | Dify |
| A single application you own and will maintain | Code — the Ollama API in Python is a short read |
| Agent teams with defined roles | CrewAI |
| A chat interface and nothing more | A front end such as Open WebUI |
Where the canvas stops helping
- Version control. Flows export as JSON, which technically diffs and practically does not review well. Code wins the moment several people change the same logic.
- Complex branching. Past a handful of conditions, the graph becomes harder to read than the equivalent function.
- Testing. There is no natural unit test for a canvas; you test by clicking, which does not scale.
- Debugging retrieval. When answers are poor, you need to see the retrieved chunks. Build that visibility in deliberately or you will be guessing.
- Upgrades. Node definitions evolve; a flow built a year ago may need adjusting. Export your flows and keep the exports.
Verdict
Flowise is at its best in the early middle of a project: past the notebook, before the codebase. It gets a retrieval chatbot running over a local model in an afternoon and publishes it as an API someone else can call, which is exactly what most internal tools need. Keep flows small, configure the embedding model, cap agent iterations, and be willing to rewrite the successful ones in code once they matter.
Frequently asked questions
Is Flowise free?
The self-hosted version is open source and free to run; a paid cloud offering exists alongside it. Check the current licence terms before embedding it in a commercial product.
Can Flowise run entirely offline?
Yes, with a local model node and a local embedding model, provided no node in the flow calls an external service.
Why does my flow return irrelevant answers?
Most often the embedding model: either not configured, or unsuited to your documents' language. After that, check the vector store actually persisted your documents and that the retriever returns enough chunks.
Flowise or Dify?
Flowise is lighter and faster to iterate in. Dify is the fuller platform — accounts, knowledge base management, logging, several apps. Choose by how many applications and people are involved.
Can I call a Flowise flow from my own application?
Yes, each flow can be published as an API endpoint with a key, or embedded as a chat widget. That is the normal way to use it in production.
What model do I need for an agent flow?
A 24B–32B model with genuine tool-calling ability. Smaller models write plausible tool calls that fail to parse, which breaks the flow rather than degrading the answer.
A current option for local AI: GMKtec EVO-X2 64GB / 1TB (Ryzen AI Max+ 395). Match memory to your model and software. A mini PC is a complete PC alternative; Mac/MLX and CUDA instructions require compatible hardware.
Amazon Check GMKtec EVO-X2 64GB / 1TB (Ryzen AI Max+ 395) price →As an Amazon Associate, BestLLMfor earns from qualifying purchases, at no extra cost to you. It does not influence our independent rankings.
Found an error or have feedback? Let us know — it helps everyone who reads this guide.