Intermediate 18 minn8n

n8n + Ollama: automate with 100% AI locale

n8n is the most credible no-code automation tool compared with Zapier or Make, with one decisive advantage: it is self-hosted. Combined with Ollama, it gives you a complete pipeline where data never leaves your machine—translated emails, summarized articles, sorted leads, all without an external API call. This n8n ollama automation guide covers installation, the native Ollama node, and three concrete workflows ready to copy.

By Mohamed Meguedmi·Update 2026-08-27·Tested on Windows, macOS, and Linux

#Why pair n8n with Ollama?

n8n is a visual orchestrator: you connect nodes (trigger, HTTP, database, LLM…) without writing code, and the workflow runs in the background on a cron job, webhook, or mailbox. All mainstream competitors (Zapier, Make, Power Automate) charge by volume and require your data to pass through their servers.

Ollama exposes a local LLM on http://localhost:11434 with a compatible HTTP API. Since late 2023, n8n has included a native Ollama node (and an OpenAI-compatible node that works too). Result: a workflow that summarizes client PDFs or sorts HR emails never sends a single byte outside your infrastructure.

i
The combo in one sentence
n8n = the brain that orchestrates. Ollama = the hand that writes, translates, or classifies. They both run on your machine and communicate over local HTTP.

#Prerequisites

The Local Agents Kit

Agents that act on your machine: agentic Cline, MCP, n8n + Ollama, local automations.

  • Lifetime online access
  • PDF + files
  • Lifetime updates
Ollama installed
Daemon active on port 11434. Check with curl http://localhost:11434. If you’re starting from scratch, first follow the site’s Ollama installation guide.
A downloaded model
ollama pull qwen3.5:9b for versatility, or qwen3.5:4b if you want something lighter. For classification workflows, granite4.2:8b is solid and very token-efficient.
Docker + Docker Compose
n8n deploys in a few seconds with Docker. Docker Desktop on Windows/macOS, Docker Engine on Linux.
8 GB of free RAM
n8n uses ~500 MB, Ollama loads the model into VRAM (~6.6 GB for a 9B Q4). Allow plenty of headroom if you want several workflows running simultaneously.
A little patience for no-code
n8n is visual but not magic: understanding JSON and expressions like {{ $json.field }} remains useful.
→
Recommended model to get started
qwen3.5:9b in Q4_K_M. Good in French, 256k context, supports tools (function calling) and vision, and fits in 6.6 GB of VRAM. If you have 12 GB or more, move up to qwen3.5:9b-q8_0 (11 GB) for classification—the accuracy clearly jumps.

#1. Install n8n self-hosted via Docker

The official and simplest method is a Docker image maintained by the n8n team. Create a dedicated directory and a docker-compose.yml file:

docker-compose.yml
services:
  n8n:
    image: docker.n8n.io/n8nio/n8n
    container_name: n8n
    restart: unless-stopped
    ports:
      - "5678:5678"
    environment:
      - N8N_HOST=localhost
      - N8N_PORT=5678
      - N8N_PROTOCOL=http
      - GENERIC_TIMEZONE=Europe/Paris
      - N8N_SECURE_COOKIE=false
    volumes:
      - ./n8n_data:/home/node/.n8n
    extra_hosts:
      - "host.docker.internal:host-gateway"

The extra_hosts line is the key: from the n8n container, you'll call Ollama via http://host.docker.internal:11434 rather than localhost (which would point to the container itself). On Linux only, this trick requires the host-gateway directive.

Start n8n
docker compose up -d
docker compose logs -f n8n

Open http://localhost:5678. On the first connection, n8n asks you to create an owner account—it remains local, with no mandatory telemetry. Once you're logged in, you arrive at the blank canvas.

!
Expose n8n to the Internet?
For personal use, keep n8n behind your LAN. If you need to expose it (public webhooks), you must use a reverse proxy with HTTPS (Caddy, Traefik) and enable N8N_BASIC_AUTH_ACTIVE=true. An n8n instance exposed without authentication means full access to your workflows and credentials.

#2. Configure the Ollama node in n8n

n8n offers two ways to call Ollama: the dedicated “Ollama Chat Model” node (built into n8n’s LangChain branch) and a generic HTTP call. The dedicated node is cleaner for agents and chains; the HTTP call gives you more control. Both work.

  1. 01
    Add a Ollama node
    On the canvas, type "/" and search for "Ollama". Choose "Ollama Chat Model" (under AI > Language Models).
  2. 02
    Create a credential
    In the Credential field, click "Create new." Base URL: http://host.docker.internal:11434 (from the Docker container) or http://localhost:11434 if n8n runs natively. No API key — it’s local.
  3. 03
    Test the connection
    n8n tests automatically and lists the available models. If the list is empty, check with docker exec n8n curl http://host.docker.internal:11434/api/tags from the host.
  4. 04
    Choose the model
    In the Model dropdown, select qwen3.5:9b (or the one you pulled). Leave Temperature at the default 0.7; lower it to 0.2 for deterministic tasks (classification, extraction).
i
The HTTP Request node is still useful
For more unusual calls (streaming, forced JSON format, advanced num_ctx or stop parameters), an HTTP Request POST node on http://host.docker.internal:11434/api/generate gives you 100% control. The body is the same JSON as in the Ollama documentation.

#3. First workflow: summarize an RSS feed

Typical case: you follow 15 tech blogs and want a daily summary in French without reading everything. n8n fetches new articles, Ollama summarizes them, and the result is sent by email or written to a Markdown file.

  1. 01
    Trigger: Schedule Trigger
    Set it to run every day at 7 a.m. Cron expression mode: 0 7 * * *.
  2. 02
    Retrieve the feed: RSS Feed Trigger or RSS Read
    Feed URL (e.g., https://blog.lemondeinformatique.fr/feed/). Enable "Return only new items" to avoid re-summarizing older ones.
  3. 03
    Split if needed: Split In Batches
    If the pipeline returns 10 articles, limit it to 5 to avoid saturating Ollama. Batch size: 1, to process articles one at a time.
  4. 04
    Call Ollama
    Ollama Chat Model node with this system prompt and the article as a user message.
Contents of node Ollama
{
  "system": "Tu es un assistant éditorial. Résume l'article en 3 puces de 15 mots maximum, en français, ton neutre. Pas d'intro, juste les puces.",
  "prompt": "Titre : {{ $json.title }}\n\nContenu : {{ $json.content }}"
}
  1. 01
    Aggregate summaries: Merge or Code
    Combine the article outputs into a single email body. A Code node (JavaScript) with items.map(i => `**${i.json.title}**\n${i.json.response}`).join('\n\n') is enough.
  2. 02
    Send: Send Email or write file
    Email (SMTP) node to your address, or Write Binary File node to /home/node/.n8n/daily-digest.md (persisted in the Docker volume).
→
Tip: shorter context = speed
On a 9B model in Q4, an 800-word article can be summarized in 4–6 seconds. If you want to speed things up, explicitly ask, “Read only the first 500 words,” and truncate it on the n8n side with an expression {{ $json.content.slice(0, 2000) }}.

#4. Workflow: translating incoming emails

Typical case: you receive supplier emails in English, German, and Spanish. You want to read them in French without Google Translate (which siphons off the content). n8n monitors your IMAP inbox, Ollama translates, and the result lands in a "Translated" label or as a drafted reply.

  1. 01
    Trigger: IMAP Email
    Configure your IMAP credentials (Gmail requires an app password). Filter: all unread emails in the INBOX. Polling every 10 minutes.
  2. 02
    Detect language: Ollama #1
    First Ollama call with system prompt: "Respond only with the ISO 639-1 code of the language of this text (fr, en, de, es...). No sentence, just two letters." Prompt: {{ $json.subject }} {{ $json.textPlain }}.
  3. 03
    Connect: IF node
    Condition: {{ $json.response }} ≠ "fr". If true → translate. If false → ignore (already in French).
  4. 04
    Translate: Ollama #2
    System prompt: "Translate the following text into natural French. Preserve the formatting (paragraphs, lists). Do not comment; translate directly." Prompt: the body of the email.
  5. 05
    Final action
    Two options: a Gmail “Add Label” node followed by writing the translation at the top of the message via “Reply Draft,” or simply sending a Slack/Telegram notification with the translated version.
!
Classic pitfall: HTML
Emails are often HTML (signatures, images, tracking tags). Use the textPlain field instead of html, or add an HTML Extract node to clean it before sending it to Ollama. Otherwise, your LLM will try to translate <table>...

#5. Workflow: classify inbound leads

Typical case: a contact form populates an Airtable / Notion / Postgres database. You want to automatically classify each lead: Hot / Warm / Cold, industry, urgency. n8n listens to the database, Ollama reads the free-form message and writes the structured tags.

This workflow uses Ollama's JSON mode, which forces the output to follow a schema. Essential whenever structured data is involved.

HTTP Request node body to Ollama
{
  "model": "qwen3.5:9b",
  "format": "json",
  "stream": false,
  "options": { "temperature": 0.1 },
  "system": "Tu es un assistant CRM. Réponds uniquement en JSON valide avec les clés : temperature (Hot/Warm/Cold), secteur (string court), urgence (1-5).",
  "prompt": "Message du lead : {{ $json.message }}\n\nEntreprise : {{ $json.company }}"
}
  1. 01
    Trigger: Airtable/Postgres Trigger
    On “New row in table Leads.” Polling every 5 min, or a webhook if your source supports it (more responsive).
  2. 02
    Ollama call in JSON mode
    HTTP Request POST node on http://host.docker.internal:11434/api/generate with the body above. The response arrives in $json.response (JSON string).
  3. 03
    Parser: Code or Set node
    JSON.parse($input.first().json.response) extracts temperature, sector, and urgency. If parsing fails (because the model hallucinates), retry with a temperature of 0.
  4. 04
    Write to the CRM
    Airtable Update / Postgres Update node with the three fields. Optional: if temperature = Hot and urgency ≥ 4, trigger a Slack node that pings commercial@boite.fr.
→
Why temperature 0.1 here
For classification, you want consistency: two identical leads should receive the same tag. A low temperature (0 to 0.2) removes sampling randomness. Reserve 0.7+ for creative tasks (writing, brainstorming).

#Tips & troubleshooting

n8n does not connect to Ollama
ECONNREFUSED on localhost:11434? You are in Docker—use host.docker.internal. On Linux, verify that extra_hosts is present in the compose file and that Ollama is listening on 0.0.0.0 (OLLAMA_HOST=0.0.0.0:11434).
Slow workflow
Ollama loads the model cold (5–15 sec the first time). To keep it warm, run a cron job in parallel that pings /api/generate every 4 minutes, or increase OLLAMA_KEEP_ALIVE=30m.
The model hallucinates during classification
Always use format: "json" plus a system prompt that lists the allowed values. If the 9B in Q4 remains unclear, moving up to qwen3.5:9b-q8_0 significantly improves JSON robustness.
Lost Docker volumes
The workflow and credentials live in ./n8n_data. Never delete this folder without a backup. Export your workflows as JSON via the menu (Settings > Workflows > Download) so you can version them in Git.
Too many calls = VRAM saturation
If a workflow triggers 20 Ollama calls in parallel, the GPU becomes saturated. Use Split In Batches with batch size 1 or add OLLAMA_NUM_PARALLEL=1 to serialize them.
n8n versions
n8n releases a major version every ~10 days. docker compose pull && docker compose up -d updates it. The Ollama node has been stable since 1.18+.
i
Version your workflows
An exported n8n workflow is clean JSON. Commit it to a private Git repo: you keep the history, can restore it if you break everything, and can easily share it with a colleague.

#Go further

You have a working n8n + Ollama pipeline. Some natural next steps:

Connect a RAG
The site’s ChromaDB RAG guide shows how to index your PDFs and notes. You can call this RAG pipeline from n8n through an HTTP endpoint—your workflows become aware of the context in your knowledge base.
Optimize the model
If your workflows process significant volume, read the Q4/Q5/Q8 quantization guide: dropping to Q4_K_M can halve VRAM usage with no visible loss for classification.
Choose a dedicated GPU
If n8n + Ollama need to run 24/7 on one machine, the site's GPU buying guide gives the 12/16/24 GB thresholds for your target workflows.
Did this guide help you?

Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.