Intermediate 12 minAgents

Hermes Agent with Ollama: memory, tools, and limites

Direct response

Yes: Nous Research's Hermes Agent connects to a local Ollama by declaring a custom provider with the address http://127.0.0.1:11434 and an empty API key. But local inference does not make the agent offline by default: web search, text-to-speech, and cloud browsing remain third-party services until you disable them. Choose a model advertised for tool calling, or it may describe an action instead of executing it.

Hermes Agent is the autonomous agent framework published by Nous Research, with a terminal, persistent memory, and messaging connectors. This guide covers connecting it to a local Ollama, choosing a model capable of calling tools, what its memory actually retains, and the security and reliability limitations to know before using it daily.

By Mohamed Meguedmi·Update 2026-09-28·Tested on Windows, macOS, and Linux

#What Hermes Agent is

Hermes Agent (NousResearch/hermes-agent repository) is an autonomous command-line agent with a TUI, a messaging gateway (Telegram, Discord, Slack, WhatsApp, Signal), and a skills system that creates and improves skills through use. It does not replace Hermes 4, which is a language model: Hermes Agent is the software that orchestrates tool calls, memory, and the terminal, regardless of which model is connected behind it.

The repository is moving at a brisk pace: the tagged v2026.9.24 release (Hermes Agent v0.21.5), published on September 24, 2026, includes several hundred fixes since the previous release. Nous Research publishes the project under the MIT license, authorizing professional use without royalties, including for a modified deployment.

The agent runs in a single gateway process that can serve multiple channels at once: local terminal, Telegram, Discord, Slack, WhatsApp, and Signal share the same memory and conversation history. This lets you start a task from your workstation and then follow it from a phone, without duplicating the model or memory configuration between channels.

i
Difference from Hermes 4
Hermes Agent is an agent framework that accepts any model as a back end. Hermes 4 is a family of models that you can specifically run behind it via Ollama. Both carry the name Hermes, but they are not interchangeable.

#Connect Ollama locally

The Local Agents Kit

Agents that act on your machine: agentic Cline, MCP, n8n + Ollama, local automations.

  • Lifetime online access
  • PDF + files
  • Lifetime updates

Hermes Agent routes self-hosted OpenAI-compatible providers (Ollama, vLLM, llama.cpp) under their own provider name. The official documentation is explicit about the expected format: a base_url pointing to the server and an empty API key used as a simple placeholder.

~/.hermes/config.yaml
auxiliary:
  compression:
    provider: ollama
    model: qwen3.6:27b-q4_k_m
    base_url: http://127.0.0.1:11434

A base_url reduced to a simple host:port (without /v1) automatically receives the suffix expected by the OpenAI-compatible API: there is no need to add it manually. The same mechanism works for vLLM and llama.cpp by changing only the provider name and port.

This same configuration block can target different roles: the primary conversation model, the context compression model, or the model used to generate a session title. Each can point to a separate provider—for example, a local Ollama model for conversation and a faster cloud model reserved for auxiliary tasks—so you can keep the core of your exchanges private while preventing a small local model from slowing down secondary tasks.

!
Local does not mean offline
Routing the main model to Ollama does not disable auxiliary services. Web search, image generation, voice synthesis, and the cloud browser use external APIs (Nous Portal or your own keys) unless you disable these tools individually in the configuration.

#Which model for tool calling

Hermes Agent exposes more than 40 tools (terminal, files, browser, memory, cron) that the model must invoke through structured calls, not a simple text description. A Ollama model that wasn't trained for tool calling — or an overly aggressive quantization of the same model — often responds with an explanation instead of triggering the call.

Primary conversation model
Prefer a Ollama variant explicitly listed as tool-compatible (tools tag on ollama.com/library) over a generic chat model.
Auxiliary model (titles, compression)
It can remain lightweight: Hermes Agent calls it after the main turn’s response on a single-slot local provider, not in parallel, to prevent a single-slot server from mixing the two requests.
Identifier format
A named provider may prefix the model, for example ollama-local/qwen3.6:27b-q4_k_m; this prefix also applies when only the raw request (without a prefix) reaches the server.

One fact worth watching closely: with a custom local provider, the title-model call is sent after the turn's response arrives, not at the same time. On a local server with a single processing slot, this prevents a structured-JSON title request from interrupting the response generation itself—a scheduling detail rarely documented elsewhere.

#Memory: what really persists

Hermes Agent's memory is based on simple text files that can be read and edited directly: MEMORY.md and USER.md in the profile directory. The model decides, through a dedicated tool, what is worth writing there—preferences, persistent facts, and learned procedures.

Terminal
cat ~/.hermes/memories/MEMORY.md
cat ~/.hermes/memories/USER.md

The weakness documented by the project itself: a small local model may say “it’s saved” without actually calling the writing tool. The official documentation recommends that, for a local model with fewer than approximately 30 billion parameters or unreliable tool calling, you explicitly request use of the memory tool and then verify the file—instead of multiplying instructions in the prompt.

→
Verify rather than believe
After requesting that something be remembered, open MEMORY.md or USER.md to confirm that the entry is actually there. If writing is enabled with human approval (write_approval), it remains pending until it is explicitly approved.

#Security: eight layers, not a wall

Nous Research documents an eight-layer security model: user authorization, human approval for dangerous commands, safeguards on file writing, container isolation (Docker, Singularity, Modal), filtering of credentials sent to MCP servers, prompt-injection detection in context files, isolation between sessions, and validation of terminal-tool working paths.

Approval of dangerous commands is configured through approvals.mode in the configuration file, with three available modes: automatic, manual, or smart (the recommended default mode, which approves low-risk actions and blocks the others for validation).

!
Container isolation isn't automatic
These eight layers exist in the software, but Docker/Singularity/Modal isolation must be selected as the backend terminal during configuration. An agent launched with a local backend executes commands directly on your machine, without the container sandbox.

#Limitations of a small local model for this agent

An agent as tool-rich as Hermes Agent places heavy demands on the model’s instruction-following capacity: memory, planning, chained tool calls, and sometimes parallel sub-agents. With a 7B to 14B model in Q4_K_M quantization, expect missed tool calls, confirmations without actual execution, and more frequent context compression during long sessions.

What degrades first with a small local model
FunctionTypical symptom
Memory (dedicated tool)Verbal confirmation without actually writing the file
Tool chainingThe model describes the next step instead of executing it
Parallel subagentsPoorly transmitted instructions, results requiring manual cross-checking
Long sessionsMore frequent context compression, with loss of older details

None of this is specific to Hermes Agent: it is a known limitation of small quantized models when handling structured output formats. The project’s documented answer is not to add more instructions, but to use a stronger model at least for the initial configuration phase. Once the inputs have been placed in memory, a lighter model can read them back without problems because they arrive directly in the system prompt.

Subagents and Python scripts that call tools via RPC—two capabilities highlighted by the project for parallelizing tasks—rely on the same instruction-following ability as a simple tool call. A model that already struggles with an isolated call will have an even harder time orchestrating multiple subagents without mixing up their respective contexts: it is better to validate the basic behavior before enabling these advanced features.

#MCP servers: catalog and permissions

Hermes Agent includes a catalog of MCP servers reviewed and integrated by the Nous Research team (including Linear, GitHub, Figma, and Asana), but each server remains disabled by default: you must explicitly install it before it becomes accessible to the agent. Once the credentials are entered, Hermes Agent queries the server to list the tools it exposes and displays a checklist, avoiding the need to enable all the capabilities of an unfamiliar third-party server at once.

A detail few tutorials mention: for stdio-mode MCP servers, Hermes Agent does not pass the workstation's full shell environment to the subprocess. Only variables explicitly declared in the server configuration, plus a minimal baseline (including PATH, HOME, USER, LANG, TERM, and SHELL), are passed through; other environment variables, including API keys and tokens, are filtered out. This mechanism reduces the risk that a poorly audited third-party MCP server will siphon off a secret present in the shell.

#Subagents: what delegation enables

The delegate_task tool lets Hermes Agent assign a task to a child agent: each sub-agent starts its own conversation, with its own terminal session, and inherits only the tools already enabled for the parent agent—it cannot grant itself an additional capability that the parent does not have. Only the sub-agent’s final summary is returned to the main agent’s context; intermediate tool calls never enter it, which limits context consumption on long tasks and prevents execution details from cluttering the main conversation.

By default, up to ten sub-agents can run in parallel—a configurable limit with no hard cap imposed by the software—but the default delegation depth remains set to 1: a sub-agent cannot create other sub-agents in turn until this behavior is explicitly enabled in the configuration. On a local 7 to 14B model, it is best to keep this depth at 1: each additional delegation level multiplies the number of tool calls that must be chained correctly, exactly the weakness identified above for small quantized models.

#Troubleshooting: symptoms, cause, fix

Most issues with a local Ollama endpoint fall into four symptom categories, each with an identifiable cause in the official documentation rather than a guess-and-check hypothesis.

Common symptoms with a local Ollama endpoint
SymptomLikely causeCorrection
Stream interruptions or timeout with a long contextThe default read timeout does not cover the latency of a large local modelIncrease HERMES_STREAM_READ_TIMEOUT (e.g., 1800 seconds) in the configuration
The model describes the action instead of executing itModel not advertised as tool-calling compatible, or overly aggressive quantization of the same modelChoose a variant labeled “tools” on the model's ollama.com/library page
Confirmation that it was memorized without writing the fileLocal model under about 30B with unreliable tool callingCheck MEMORY.md/USER.md after every request; if needed, temporarily switch back to a more capable model
A third-party MCP server cannot see an expected variableEnvironment filtering: only declared variables and the minimal base environment are passed to the subprocessExplicitly declare the variable in this MCP server's configuration

Hermes Agent automatically detects local endpoints and adjusts its read timeouts accordingly, but a large context on a slow Ollama model can still exceed the default timeout; the official documentation recommends adjusting the HERMES_STREAM_READ_TIMEOUT variable rather than arbitrarily reducing the context sent to the model.

#Practical installation

  1. 01
    Install Hermes Agent
    On Linux, macOS, or WSL2: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash, then reload the shell and run hermes.
  2. 02
    Prepare the Ollama model
    Verify that Ollama is running on port 11434 and that the selected model is listed as tool-calling compatible before declaring it as a provider.
  3. 03
    Declare the local provider
    Use hermes config set to specify provider: ollama, the model name, and the base_url http://127.0.0.1:11434, with an empty API key.
  4. 04
    Test a simple tool call
    Request a verifiable action (list a directory, read a file) and confirm in the terminal that the tool actually ran, rather than merely describing it.
  5. 05
    Check memory
    Explicitly request a memory to be saved, then open MEMORY.md to verify that the entry appears there before trusting the agent on this point.
Frequently asked questions
Does Hermes Agent work without an internet connection once connected to Ollama?+
Not by default. The primary conversational model runs well locally once the Ollama provider is declared, but web search, speech synthesis, image generation, and the cloud browser remain external services until you explicitly disable them in the tool configuration. For truly offline use, you must turn off these tools one by one instead of relying solely on the primary model's routing.
Which Ollama model should you choose for Hermes Agent?+
Choose a model explicitly advertised as tool-calling compatible on its ollama.com/library page, such as qwen3.6:27b-q4_k_m mentioned in the official documentation. A generic chat model, even a large one, often responds in natural language instead of triggering the structured tool call expected by the framework, which blocks entire features such as memory, the terminal, or installed MCP servers.
Is Hermes Agent the same thing as the Hermes 4 model?+
No. Hermes Agent is the orchestration software: terminal, memory, tool calling, MCP servers, and delegation to subagents. Hermes 4 is a language model that you can run behind it through Ollama or another provider, just like any other tool-calling-compatible model. They share the same name but serve different roles.
How can you verify that Hermes Agent's memory is actually working?+
Explicitly request that something be remembered, then open ~/.hermes/memories/MEMORY.md or USER.md to confirm that the entry actually appears on disk. With a small local model of less than about 30 billion parameters, a verbal confirmation without the file actually being written is a scenario the project documentation itself recognizes as common, which is why it is worth checking instead of trusting the agent's response.
Is container isolation enabled automatically?+
No. Container sandboxing (Docker, Singularity, Modal) is a terminal backend choice made explicitly during configuration, among seven available backends, including local execution. An agent launched with the default local backend runs commands directly on the host machine without this isolation; you must select a containerized backend to actually use it.
How do you enable an MCP server in Hermes Agent?+
MCP servers from the Nous Research catalog are disabled by default: you have to install them one by one in the configuration before they become accessible to the agent. Once you enter the credentials, Hermes Agent queries the server to list the tools it exposes and displays a checklist, so you don't enable all the capabilities of an unfamiliar third-party server at once.
How many sub-agents can Hermes Agent launch in parallel?+
By default, up to ten sub-agents can run simultaneously through the delegate_task tool, with a configurable limit and no software-imposed cap. Delegation depth nevertheless remains set to 1 by default: a sub-agent cannot create others until this behavior is explicitly enabled, preventing an uncontrolled explosion in the number of agents and token consumption.

Did this guide help you?

Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.