BestLLMfor Your hardware. Your LLM. Your call.
◆ The kits◆ Kits APIOpen data Find my LLM
Guide · 2026-09-28

Hermes Agent With a Local Ollama Endpoint

◆ Local Agents — Agents that act on your machine, without the cloud · $24 · or all kits $49 →

Nous Research's Hermes Agent can run its conversation model through a local Ollama server. Here is the exact configuration, which models actually call tools, what its memory keeps, and where local inference stops being private by default.

By Mohamed Meguedmi·Last updated 2026-09-28·10 min read·Tested on Windows, macOS, Linux

Key takeaways

  • Hermes Agent routes to a local Ollama server by declaring a custom provider with base_url: http://127.0.0.1:11434 and an empty API key — no code change required.
  • Local inference is not offline by default. Web search, image generation, text-to-speech and the cloud browser still call external services unless each tool is disabled explicitly.
  • Model choice matters more than size. Pick an Ollama model advertised for tool calling — a generic chat model often narrates an action instead of executing it.
  • Hermes Agent's own documentation states that small local models (under roughly 30B) can claim a memory save that never happened — the tool call silently didn't fire.
  • Security runs in eight layers, from command approval to container isolation, but container sandboxing is a terminal-backend choice you have to make, not a default.

What Hermes Agent actually is

The Local Agents Kit

Agents that act on your machine: agentic Cline, MCP, n8n + Ollama, local automations.

  • Lifetime online access
  • PDF + files
  • 30-day refund

Hermes Agent (repository NousResearch/hermes-agent) is Nous Research's self-improving agent framework: a terminal UI, a messaging gateway (Telegram, Discord, Slack, WhatsApp, Signal), persistent memory, and a skills system that creates and refines its own procedures from experience. It is distinct from the Hermes model family — Hermes Agent is the orchestration software, and any model, local or hosted, can sit behind it.

The project ships fast: tag v2026.9.24 (Hermes Agent v0.21.5), published September 24, 2026, rolls up several hundred merged pull requests since the prior tagged release. The repository is released under the MIT license by Nous Research.

The agent runs as a single gateway process that can serve several channels from the same memory and conversation state — terminal, Telegram, Discord, Slack, WhatsApp, Signal. You can start a task at your desk and keep following it from a phone without reconfiguring the model or losing context.

Connecting a local Ollama endpoint

Hermes Agent routes self-hosted OpenAI-compatible servers — Ollama, vLLM, llama.cpp — under their own provider names. The official configuration documentation is explicit about the expected shape: a base_url pointing at the server, and an empty API key acting as a placeholder.

auxiliary:
  compression:
    provider: ollama
    model: qwen3.6:27b-q4_k_m
    base_url: http://127.0.0.1:11434

A bare host:port base_url gets the OpenAI-compatible /v1 suffix appended automatically — no need to type it. The same mechanism applies to vLLM and llama.cpp by swapping the provider name and port.

Different roles can point at different providers in the same config: the main conversation model, the context-compression model, and the session-title model. A common pattern is a local Ollama model for the conversation itself, with a faster cloud model reserved for auxiliary tasks — keeping the bulk of the conversation private without letting a small local model slow down background jobs.

Local is not offline. Routing the main model to Ollama does not cut the rest of the stack from the network. Web search, image generation, text-to-speech and the cloud browser still call external APIs (Nous Portal or your own keys) until you disable each tool individually.

Which model actually calls tools

Hermes Agent exposes more than 40 tools — terminal, files, browser, memory, cron — that the model must invoke through structured calls, not describe in prose. An Ollama model that was never fine-tuned for tool calling, or the same model pushed to an aggressive quantization, frequently answers with an explanation instead of triggering the call.

  • Main conversation model: prefer an Ollama variant explicitly listed as tool-capable (the tools tag on ollama.com/library) over a generic chat build.
  • Auxiliary model (titles, compression): can stay light. On a custom local provider, Hermes Agent sends this call after the turn's reply has arrived, not concurrently — a single-slot local server otherwise risks answering the reply with the title request's JSON instead.
  • Model id format: a named provider can prefix the model, e.g. ollama-local/qwen3.6:27b-q4_k_m; that prefix also applies when only the bare model id reaches the server.

That sequencing detail — title generation deferred until after the reply — is easy to miss and rarely documented elsewhere, but it is exactly what keeps a single-slot local Ollama instance from mixing up two concurrent requests.

Memory: what actually persists

Hermes Agent's memory is plain, editable text: MEMORY.md and USER.md under the profile directory. The model decides, through a dedicated tool, what is worth writing there — preferences, durable facts, learned procedures.

cat ~/.hermes/memories/MEMORY.md
cat ~/.hermes/memories/USER.md

The documented weak point: a small local model can say "saved" without the write tool ever firing. For a local model under roughly 30B parameters, or one with weak tool-calling, the official guidance is to ask explicitly for the memory tool and confirm the entry landed in the file — not to pile on more instructions in the prompt. Once the entries exist, a smaller model reads them fine, since they arrive directly in the system prompt.

Verify, don't trust. After asking the agent to remember something, open MEMORY.md or USER.md to confirm the entry is actually there. With write_approval enabled, writes outside the interactive CLI are held pending review until explicitly approved.

Security: eight layers, not a wall

Nous Research documents an eight-layer security model: user authorization, human-in-the-loop approval for dangerous commands, file-write safeguards, container isolation (Docker, Singularity, Modal), MCP credential filtering, prompt-injection scanning in context files, cross-session isolation, and input sanitization on terminal working directories.

Dangerous command approval is set via approvals.mode in the config file, with three modes: fully automatic, manual, or the recommended default, smart, which auto-approves low-risk actions and flags the rest.

Container isolation is not automatic. The eight layers exist in the software, but Docker/Singularity/Modal sandboxing is a terminal-backend choice made at setup time. An agent started on the local backend runs commands directly on your machine, without that container sandbox.

What a small local model doesn't do well here

An agent this rich in tools stresses a model's instruction-following hard: memory, planning, chained tool calls, and sometimes parallel subagents. On a 7B–14B model at Q4_K_M, expect missed tool calls, confirmations with no actual execution, and more frequent context compression on long sessions.

FeatureTypical failure with a small local model
Memory toolVerbal confirmation with no actual file write
Chained tool useThe model narrates the next step instead of running it
Parallel subagentsInstructions get lost between agents, results need manual cross-checking
Long sessionsMore frequent context compression, older details drop out

None of this is specific to Hermes Agent — it is the well-known ceiling of small quantized models against structured output formats. The project's own fix is not more prompting, but a stronger model at least for initial setup; once memory entries exist, a lighter model can read them without issue. Subagents and RPC-driven Python scripts — both used to parallelize work — rely on that same instruction-following capacity, so validate plain single-tool behavior before turning them on.

Setting it up in practice

  1. Install Hermes Agent. On Linux, macOS or WSL2: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash, then reload the shell and run hermes.
  2. Prepare the Ollama model. Confirm Ollama is listening on port 11434 and that the chosen model is advertised as tool-capable before wiring it up as a provider.
  3. Declare the local provider. Use hermes config set to set provider: ollama, the model name, and base_url: http://127.0.0.1:11434 with an empty API key.
  4. Test one tool call. Ask for a verifiable action (list a folder, read a file) and confirm in the terminal that the tool actually ran, not just that the agent described it.
  5. Check memory for real. Ask the agent to remember something, then open MEMORY.md to confirm the entry is there before trusting it on this point.

For giving a local model tools outside Hermes Agent's own tool system, see Ollama + MCP: Give Your Local Model Tools and What Is MCP? Model Context Protocol Explained Simply. If tool calling itself is unreliable regardless of the agent, Claude Desktop + Local LLM via MCP covers a comparable setup with a different client.

Primary sources used in this guide: the Hermes Agent GitHub repository, the provider configuration docs, the security model docs, and the memory docs.

Frequently asked questions

Does Hermes Agent work offline once it's wired to Ollama?

Not by default. The conversation model runs locally, but web search, text-to-speech, image generation and the cloud browser remain external services until you disable those tools explicitly in the configuration.

Which Ollama model should I use with Hermes Agent?

One explicitly listed as tool-capable on its Ollama model page. A generic chat model, even a large one, often replies in prose instead of triggering the structured tool call the framework expects.

Is Hermes Agent the same thing as the Hermes model?

No. Hermes Agent is the orchestration software — terminal, memory, tools. The Hermes model family is a set of language models you can run behind it, through Ollama or any other provider, like any other model.

How do I confirm Hermes Agent's memory actually works?

Ask it to remember something, then open ~/.hermes/memories/MEMORY.md or USER.md to check the entry landed. A verbal confirmation with no real write is a documented failure mode on small local models.

Is container isolation turned on automatically?

No. Docker, Singularity or Modal sandboxing is a terminal-backend choice made during setup. An agent started on the local backend runs commands directly on the host machine, without that isolation.

Recommended hardware

A current option for local AI: GMKtec EVO-X2 64GB / 1TB (Ryzen AI Max+ 395). Match memory to your model and software. A mini PC is a complete PC alternative; Mac/MLX and CUDA instructions require compatible hardware.

Amazon Check GMKtec EVO-X2 64GB / 1TB (Ryzen AI Max+ 395) price →

As an Amazon Associate, BestLLMfor earns from qualifying purchases, at no extra cost to you. It does not influence our independent rankings.

Did this guide help?

Found an error or have feedback? Let us know — it helps everyone who reads this guide.