Hermes Agent with Ollama: memory, tools, and limites
Yes: Nous Research's Hermes Agent connects to a local Ollama by declaring a custom provider with the address http://127.0.0.1:11434 and an empty API key. But local inference does not make the agent offline by default: web search, text-to-speech, and cloud browsing remain third-party services until you disable them. Choose a model advertised for tool calling, or it may describe an action instead of executing it.
Hermes Agent is the autonomous agent framework published by Nous Research, with a terminal, persistent memory, and messaging connectors. This guide covers connecting it to a local Ollama, choosing a model capable of calling tools, what its memory actually retains, and the security and reliability limitations to know before using it daily.
#What Hermes Agent is
Hermes Agent (NousResearch/hermes-agent repository) is an autonomous command-line agent with a TUI, a messaging gateway (Telegram, Discord, Slack, WhatsApp, Signal), and a skills system that creates and improves skills through use. It does not replace Hermes 4, which is a language model: Hermes Agent is the software that orchestrates tool calls, memory, and the terminal, regardless of which model is connected behind it.
The repository is moving at a brisk pace: the tagged v2026.9.24 release (Hermes Agent v0.21.5), published on September 24, 2026, includes several hundred fixes since the previous release. Nous Research publishes the project under the MIT license, authorizing professional use without royalties, including for a modified deployment.
The agent runs in a single gateway process that can serve multiple channels at once: local terminal, Telegram, Discord, Slack, WhatsApp, and Signal share the same memory and conversation history. This lets you start a task from your workstation and then follow it from a phone, without duplicating the model or memory configuration between channels.
#Connect Ollama locally
Agents that act on your machine: agentic Cline, MCP, n8n + Ollama, local automations.
- Lifetime online access
- PDF + files
- Lifetime updates
Hermes Agent routes self-hosted OpenAI-compatible providers (Ollama, vLLM, llama.cpp) under their own provider name. The official documentation is explicit about the expected format: a base_url pointing to the server and an empty API key used as a simple placeholder.
A base_url reduced to a simple host:port (without /v1) automatically receives the suffix expected by the OpenAI-compatible API: there is no need to add it manually. The same mechanism works for vLLM and llama.cpp by changing only the provider name and port.
This same configuration block can target different roles: the primary conversation model, the context compression model, or the model used to generate a session title. Each can point to a separate provider—for example, a local Ollama model for conversation and a faster cloud model reserved for auxiliary tasks—so you can keep the core of your exchanges private while preventing a small local model from slowing down secondary tasks.
#Which model for tool calling
Hermes Agent exposes more than 40 tools (terminal, files, browser, memory, cron) that the model must invoke through structured calls, not a simple text description. A Ollama model that wasn't trained for tool calling — or an overly aggressive quantization of the same model — often responds with an explanation instead of triggering the call.
- Primary conversation model
- Prefer a Ollama variant explicitly listed as tool-compatible (tools tag on ollama.com/library) over a generic chat model.
- Auxiliary model (titles, compression)
- It can remain lightweight: Hermes Agent calls it after the main turn’s response on a single-slot local provider, not in parallel, to prevent a single-slot server from mixing the two requests.
- Identifier format
- A named provider may prefix the model, for example ollama-local/qwen3.6:27b-q4_k_m; this prefix also applies when only the raw request (without a prefix) reaches the server.
One fact worth watching closely: with a custom local provider, the title-model call is sent after the turn's response arrives, not at the same time. On a local server with a single processing slot, this prevents a structured-JSON title request from interrupting the response generation itself—a scheduling detail rarely documented elsewhere.
#Memory: what really persists
Hermes Agent's memory is based on simple text files that can be read and edited directly: MEMORY.md and USER.md in the profile directory. The model decides, through a dedicated tool, what is worth writing there—preferences, persistent facts, and learned procedures.
The weakness documented by the project itself: a small local model may say “it’s saved” without actually calling the writing tool. The official documentation recommends that, for a local model with fewer than approximately 30 billion parameters or unreliable tool calling, you explicitly request use of the memory tool and then verify the file—instead of multiplying instructions in the prompt.
#Security: eight layers, not a wall
Nous Research documents an eight-layer security model: user authorization, human approval for dangerous commands, safeguards on file writing, container isolation (Docker, Singularity, Modal), filtering of credentials sent to MCP servers, prompt-injection detection in context files, isolation between sessions, and validation of terminal-tool working paths.
Approval of dangerous commands is configured through approvals.mode in the configuration file, with three available modes: automatic, manual, or smart (the recommended default mode, which approves low-risk actions and blocks the others for validation).
#Limitations of a small local model for this agent
An agent as tool-rich as Hermes Agent places heavy demands on the model’s instruction-following capacity: memory, planning, chained tool calls, and sometimes parallel sub-agents. With a 7B to 14B model in Q4_K_M quantization, expect missed tool calls, confirmations without actual execution, and more frequent context compression during long sessions.
| Function | Typical symptom |
|---|---|
| Memory (dedicated tool) | Verbal confirmation without actually writing the file |
| Tool chaining | The model describes the next step instead of executing it |
| Parallel subagents | Poorly transmitted instructions, results requiring manual cross-checking |
| Long sessions | More frequent context compression, with loss of older details |
None of this is specific to Hermes Agent: it is a known limitation of small quantized models when handling structured output formats. The project’s documented answer is not to add more instructions, but to use a stronger model at least for the initial configuration phase. Once the inputs have been placed in memory, a lighter model can read them back without problems because they arrive directly in the system prompt.
Subagents and Python scripts that call tools via RPC—two capabilities highlighted by the project for parallelizing tasks—rely on the same instruction-following ability as a simple tool call. A model that already struggles with an isolated call will have an even harder time orchestrating multiple subagents without mixing up their respective contexts: it is better to validate the basic behavior before enabling these advanced features.
#MCP servers: catalog and permissions
Hermes Agent includes a catalog of MCP servers reviewed and integrated by the Nous Research team (including Linear, GitHub, Figma, and Asana), but each server remains disabled by default: you must explicitly install it before it becomes accessible to the agent. Once the credentials are entered, Hermes Agent queries the server to list the tools it exposes and displays a checklist, avoiding the need to enable all the capabilities of an unfamiliar third-party server at once.
A detail few tutorials mention: for stdio-mode MCP servers, Hermes Agent does not pass the workstation's full shell environment to the subprocess. Only variables explicitly declared in the server configuration, plus a minimal baseline (including PATH, HOME, USER, LANG, TERM, and SHELL), are passed through; other environment variables, including API keys and tokens, are filtered out. This mechanism reduces the risk that a poorly audited third-party MCP server will siphon off a secret present in the shell.
#Subagents: what delegation enables
The delegate_task tool lets Hermes Agent assign a task to a child agent: each sub-agent starts its own conversation, with its own terminal session, and inherits only the tools already enabled for the parent agent—it cannot grant itself an additional capability that the parent does not have. Only the sub-agent’s final summary is returned to the main agent’s context; intermediate tool calls never enter it, which limits context consumption on long tasks and prevents execution details from cluttering the main conversation.
By default, up to ten sub-agents can run in parallel—a configurable limit with no hard cap imposed by the software—but the default delegation depth remains set to 1: a sub-agent cannot create other sub-agents in turn until this behavior is explicitly enabled in the configuration. On a local 7 to 14B model, it is best to keep this depth at 1: each additional delegation level multiplies the number of tool calls that must be chained correctly, exactly the weakness identified above for small quantized models.
#Troubleshooting: symptoms, cause, fix
Most issues with a local Ollama endpoint fall into four symptom categories, each with an identifiable cause in the official documentation rather than a guess-and-check hypothesis.
| Symptom | Likely cause | Correction |
|---|---|---|
| Stream interruptions or timeout with a long context | The default read timeout does not cover the latency of a large local model | Increase HERMES_STREAM_READ_TIMEOUT (e.g., 1800 seconds) in the configuration |
| The model describes the action instead of executing it | Model not advertised as tool-calling compatible, or overly aggressive quantization of the same model | Choose a variant labeled “tools” on the model's ollama.com/library page |
| Confirmation that it was memorized without writing the file | Local model under about 30B with unreliable tool calling | Check MEMORY.md/USER.md after every request; if needed, temporarily switch back to a more capable model |
| A third-party MCP server cannot see an expected variable | Environment filtering: only declared variables and the minimal base environment are passed to the subprocess | Explicitly declare the variable in this MCP server's configuration |
Hermes Agent automatically detects local endpoints and adjusts its read timeouts accordingly, but a large context on a slow Ollama model can still exceed the default timeout; the official documentation recommends adjusting the HERMES_STREAM_READ_TIMEOUT variable rather than arbitrarily reducing the context sent to the model.
#Practical installation
- 01Install Hermes AgentOn Linux, macOS, or WSL2: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash, then reload the shell and run hermes.
- 02Prepare the Ollama modelVerify that Ollama is running on port 11434 and that the selected model is listed as tool-calling compatible before declaring it as a provider.
- 03Declare the local providerUse hermes config set to specify provider: ollama, the model name, and the base_url http://127.0.0.1:11434, with an empty API key.
- 04Test a simple tool callRequest a verifiable action (list a directory, read a file) and confirm in the terminal that the tool actually ran, rather than merely describing it.
- 05Check memoryExplicitly request a memory to be saved, then open MEMORY.md to verify that the entry appears there before trusting the agent on this point.
- Run the Hermes 4 model locally
- Connect MCP servers to Ollama
- AI agent: general definitions and limitations
- Source: official Hermes Agent GitHub repository
- Source: provider configuration documentation
- Source: security model documentation
Does Hermes Agent work without an internet connection once connected to Ollama?+
Which Ollama model should you choose for Hermes Agent?+
Is Hermes Agent the same thing as the Hermes 4 model?+
How can you verify that Hermes Agent's memory is actually working?+
Is container isolation enabled automatically?+
How do you enable an MCP server in Hermes Agent?+
How many sub-agents can Hermes Agent launch in parallel?+
Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.