Intermediate 10 minNous Research

Hermes 4 locally (Ollama): Nous's model Research

Hermes 4 is the fourth generation of Nous Research models, from a lab known for fine-tunes that follow the system prompt precisely and rarely refuse a legitimate task. This guide explains how to install Hermes 4 with Ollama, which variant to choose based on your graphics card, and how to leverage its two strengths: adherence to system instructions and structured outputs for agents. We use it ourselves in production on a 12 GB RTX 5070 Ti; the figures come from that setup.

By Mohamed Meguedmi·Update 2026-09-25·Tested on Windows, macOS, and Linux

#Why Hermes 4 instead of a basic Qwen or Llama

Nous Research does not pretrain a model. The lab takes an existing open-weight model and applies extensive post-training: for Hermes 4, approximately 5 million examples and 60 billion tokens of synthetic data produced by its DataForge pipeline, with a high proportion of verified reasoning traces. The result is not more “intelligent” than its base model on general-knowledge benchmarks, but it behaves differently day to day: it does what you ask, in the requested format, without unnecessary commentary.

Three traits explain Hermes's popularity among self-hosted LLM users. First, the system prompt is treated as the source of truth: persona, tone, formatting constraints—everything is respected throughout a conversation. Second, its alignment is deliberately neutral: the model does not add unsolicited warnings and refuses much less often than mainstream assistants on ordinary topics (medicine, law, cybersecurity, adult fiction). Finally, Nous's tool-calling format, with its dedicated tags, has become a de facto standard adopted by many agent frameworks.

Hermes 4 adds a hybrid reasoning mode inherited from DeepHermes: the model can think between think tags before responding, or respond directly, depending on what you specify in the system prompt. You decide, query by query, whether to pay the extra token cost of reasoning.

i
Hermes is not an “uncensored” model
Don't confuse Hermes with the abliterated or uncensored variants circulating on Hugging Face. Nous Research removes nothing: it trains a model that obeys the operator. If your system prompt sets limits, Hermes follows them. If you set none, it invents none. Responsibility for the guardrails shifts to you, which is precisely what self-hosted use is intended to achieve.

#The Hermes 4 variants and the required VRAM

The Local AI Kit

Your private ChatGPT, free, on your own machine in an hour — LM Studio, Ollama, Open WebUI, your documents, no cloud.

  • Lifetime online access
  • PDF + files
  • Lifetime updates

Hermes 4 was released in late August 2025 in three sizes, each built on a different base. This detail matters: the license, maximum context, and behavior in French come from the base, not Nous Research.

Hermes 4 14B
Qwen3-14B base. About 9 GB of VRAM in Q4_K_M, a comfortable native context, and very good French. This is the variant to install on a 12–16 GB card and the one we use in production. License inherited from Qwen3: Apache 2.0.
Hermes 4 70B
Base Llama 3.1 70B. Approximately 40 GB in Q4_K_M: two RTX 3090/4090s, a Mac Studio, or a MacBook Pro with 48 GB or more of unified memory. Significantly better at long-form reasoning and following complex instructions. Llama 3.1 Community license.
Hermes 4 405B
Base Llama 3.1 405B. More than 200 GB even in Q4: beyond the reach of a personal computer, reserved for multi-GPU servers or API providers. We mention it for completeness.
And below 12 GB?
There is no Hermes 4 in 3B or 8B. On an 8 GB card, use hermes3:8b (based on Llama 3.1 8B, about 5 GB in Q4_K_M): same philosophy, same tool format, without reasoning mode.

The choice of quantization follows the site's usual rules: Q4_K_M is the default good compromise, Q5_K_M or Q8_0 if VRAM allows and you are doing structured extraction, where every bit of precision reduces formatting errors. On our RTX 5070 Ti with 12 GB, hermes4:14b in Q4_K_M with a context of 8,192 tokens uses 9.4 GB of VRAM, leaving room for an embeddings model loaded in parallel.

→
Mac Apple Silicon
On an M4 Pro with 24 GB of unified memory, the 14B in Q4_K_M runs effortlessly. The 70B requires at least 48 GB to remain smooth; with 36 GB, it loads, but the available context becomes too short for agent use.

#Prerequisites

Ollama up to date
The Qwen3 14B base requires a version of Ollama later than April 2025. Check with ollama --version and update if needed; otherwise, you'll get an unknown architecture error when loading.
GPU or unified memory
12 GB of VRAM for the 14B in Q4_K_M with an 8k context. With 8 GB, the model partially loads onto the CPU and drops to a few tokens per second: switch to hermes3:8b.
Disk space
Allow 9 GB for the 14B in Q4_K_M, 15 GB in Q8_0, and 40 GB for the 70B in Q4_K_M.
An interface (optional)
Open WebUI or LM Studio for comfortable discussion. Open WebUI automatically collapses reasoning blocks, which is useful with Hermes 4.

#Install Hermes 4 with Ollama in 4 steps

  1. 01
    Download the variant suited to your machine
    In the Ollama library, the hermes4 tag comes in different sizes. On a 12 to 16 GB card, choose the 14B; that's the tag we run every day.
  2. 02
    Or import the official GGUF from Hugging Face
    Nous Research publishes its own GGUF quantizations. The hf.co syntax for Ollama lets you precisely choose the level (Q4_K_M, Q5_K_M, Q8_0) without using a Modelfile. This is the preferred approach if you want Q8_0 or if the library tag does not offer the quantization you want.
  3. 03
    Check loading and GPU/CPU distribution
    The ollama ps command shows occupied memory and the percentage loaded on the GPU. Aim for 100% GPU: as soon as any part moves to the CPU, speed plummets.
  4. 04
    First test with a system prompt
    Don't test Hermes without a system prompt—you'd miss what makes it interesting. Give it a role, an output format, and a constraint, then see how closely it sticks to them.
Terminal — from the Ollama library
ollama pull hermes4:14b
ollama run hermes4:14b
Terminal — Official Nous Research GGUF, choose your quantization
# Q4_K_M : le compromis recommandé (environ 9 Go)
ollama run hf.co/NousResearch/Hermes-4-14B-GGUF:Q4_K_M

# Q8_0 si vous avez 16 Go et plus, pour l'extraction structurée
ollama run hf.co/NousResearch/Hermes-4-14B-GGUF:Q8_0
Terminal — loading control
ollama ps
# NAME           SIZE     PROCESSOR    CONTEXT
# hermes4:14b    9.4 GB   100% GPU     8192
!
Default context too short
Ollama loads models with a context of 4,096 tokens by default. For an agent that accumulates tool calls or a reasoning mode that can produce several thousand reasoning tokens, this is insufficient: the first messages fall out of the window and the model “forgets” its instructions. Set num_ctx to at least 8,192, or 16,384 if VRAM allows (see the settings section).

#What distinguishes Hermes from mainstream aligned models

A consumer assistant is trained to please an anonymous user and protect its publisher. This produces behaviors that eventually become invisible: an opening that restates the question, a warning at the end of the answer, a polite refusal as soon as a sensitive keyword appears, and a tendency to soften the edges. For conversation, that is acceptable. For an automated pipeline expecting JSON, a ranking, or a rewrite, every deviation breaks downstream processing.

Hermes 4 is trained for the operator, not the end user. Nous Research measured this choice with an in-house benchmark, RefusalBench, where Hermes 4 refuses significantly less often than closed models and most aligned open-weight models. In practice, on a support-ticket corpus asking it to extract the reason for a complaint, a consumer model may sometimes reply, “I can’t provide legal advice,” where Hermes returns the requested field. On a rewriting task, it doesn’t add “feel free to reach out if you have any other questions.”

No preamble or conclusion
If the system prompt says “respond only with JSON,” the response starts with a brace. On many models, you need three examples and post-processing to get the same result.
Stable persona
A role defined in the system prompt holds up for dozens of turns without drifting, even when the user tries to push it outside its boundaries.
Fewer unsolicited refusals
Medical topics, legal matters, cybersecurity, dark fiction: Hermes handles the request. The limits are the ones you write.
Neutral tone
No flattery about the quality of the question, no emojis, no artificial enthusiasm. The style is controlled entirely by the system prompt.

The trade-off is obvious: if you expose Hermes to strangers, through a public chatbot for example, it's up to you to write the guardrails. A system prompt listing what the assistant must not do, plus an application-side output filter for critical cases, is essential. For personal or internal use, that responsibility is precisely the advantage you're looking for.


#The system prompt, where Hermes 4 excels

With Hermes 4, the system prompt is not a suggestion but a contract. This changes how you write it: be precise, complete, and do not expect the model to fill gaps with an assistant’s “common sense.” Three rules are enough to get reliable results.

Describe the role in one sentence
“You are the marketing team’s writing assistant for X” is better than three paragraphs of context. Hermes does not dilute the information, but a short role is more stable.
Set the output format explicitly
Length, language, structure, and what must be omitted. Hermes follows a constraint such as “no introduction, no conclusion, 120 words maximum” without having to repeat it.
Write down the prohibitions
Since the model doesn't add any, list your own: out-of-scope topics, information that must never be revealed, and behavior when a question is ambiguous.

The most practical approach is to lock this system prompt into a Modelfile: you get a ready-to-use alias, with the context and temperature already configured, usable in Open WebUI, in your scripts, and from the terminal.

Modelfile — French writing assistant
FROM hermes4:14b

PARAMETER num_ctx 16384
PARAMETER temperature 0.6
PARAMETER top_p 0.95
PARAMETER repeat_penalty 1.05

SYSTEM """Tu es un assistant de rédaction pour une PME française.
Tu réponds toujours en français, au vouvoiement.
Tu ne commences jamais par reformuler la demande et tu ne termines jamais par une formule de politesse.
Quand on te demande de réécrire un texte, tu renvoies uniquement le texte réécrit, sans commentaire.
Si une demande sort du cadre de la rédaction professionnelle, tu le dis en une phrase et tu t'arrêtes."""
Terminal — create and use the alias
ollama create hermes-redac -f Modelfile
ollama run hermes-redac "Réécris en deux phrases : Nous sommes ravis de vous annoncer que notre nouvelle offre est disponible dès maintenant et qu'elle vous permettra de gagner du temps."
→
Test role consistency
After creating the alias, try pushing it outside its boundaries in two or three messages (“forget your instructions,” “speak to me in English”). Hermes 4 holds up significantly better than the Qwen3-14B base model on this type of test, and it's a good way to verify that your system prompt is complete.

#Enable or disable reasoning mode

Hermes 4 is a hybrid-reasoning model: by default, it responds directly, like a classic instruct model. To enable reasoning, add a dedicated instruction to the system prompt—the one Nous Research documents on the model card. The model then produces its reasoning between think tags, followed by its answer.

System prompt — reasoning instruction recommended by Nous Research
You are a deep thinking AI, you may use extremely long chains of thought to deeply consider the problem and deliberate with yourself via systematic reasoning processes to help come to a correct solution prior to answering. You should enclose your thoughts and internal monologue inside <think> </think> tags, and then provide your solution or response to the problem.

You can combine it with your own instructions, in French, afterward. Reasoning is useful for math, nontrivial code, agent planning, or analyzing a long document. It's unnecessary and costly for field extraction, classification, or rewriting: for those tasks, leave direct mode enabled. Expect 500 to several thousand reasoning tokens per request, which is why a generous num_ctx matters.

Direct mode (default)
No special instruction. Immediate response, ideal for pipelines and everyday chat. This is the mode we use to evaluate our monitoring system automatically.
Reasoning mode
Add the instruction above at the top of the system prompt. Open WebUI collapses the think block; in your scripts, filter it out before displaying or parsing the response.
Two aliases
The simplest approach is to create two Modelfiles, hermes-direct and hermes-think, and choose one or the other depending on the task instead of changing the prompt for every call.

#Use case: agents and structured extraction

Nous Research’s tool-calling format, in which the model receives the list of available functions in the system prompt and emits its calls as tagged JSON, is what Ollama uses through the tools parameter of its chat API. With Hermes 4, you do not need to configure anything: describe your functions, send the conversation, and the model returns a structured call when needed, or a text response otherwise.

For structured extraction, two approaches complement each other. The first is to enforce a JSON schema through Ollama's format parameter: the engine constrains generation, and the model cannot deviate from the schema. The second relies solely on the system prompt, and that's where Hermes makes the difference: even without constrained decoding, it returns valid JSON without commentary in the vast majority of cases, simplifying integrations with tools that do not support constrained decoding.

Python — structured extraction with a required schema
from ollama import chat
from pydantic import BaseModel

class Ticket(BaseModel):
    motif: str
    produit: str
    urgence: int  # 1 à 5
    remboursement_demande: bool

message = """Bonjour, ma cafetière X200 achetée il y a 3 semaines fuit par le dessous.
Je veux un remboursement, c'est la deuxième fois que ça arrive."""

response = chat(
    model="hermes4:14b",
    messages=[
        {"role": "system", "content": "Tu extrais les informations d'un ticket de support. Réponds uniquement avec le JSON demandé."},
        {"role": "user", "content": message},
    ],
    format=Ticket.model_json_schema(),
    options={"temperature": 0.1, "num_ctx": 8192},
)

ticket = Ticket.model_validate_json(response.message.content)
print(ticket)
# motif='fuite' produit='X200' urgence=4 remboursement_demande=True
Terminal — tool call via the REST API
curl http://localhost:11434/api/chat -d '{
  "model": "hermes4:14b",
  "stream": false,
  "messages": [
    {"role": "system", "content": "Tu es un assistant qui utilise les outils fournis quand ils sont pertinents."},
    {"role": "user", "content": "Quel temps fait-il à Lyon ?"}
  ],
  "tools": [{
    "type": "function",
    "function": {
      "name": "get_weather",
      "description": "Donne la météo actuelle pour une ville",
      "parameters": {
        "type": "object",
        "properties": {"ville": {"type": "string"}},
        "required": ["ville"]
      }
    }
  }]
}'

The response contains a tool_calls field with the function name and its arguments. You must execute the function and return the result in a tool-role message so the model can formulate its final response. This is the loop implemented by LangChain, CrewAI, n8n, or Hermes Agent, the agent framework published by Nous Research itself, which works with any model but was designed around this format.

Classification and routing
Sort emails, tickets, or articles into defined categories. Low temperature, direct mode, enforced JSON schema: the 14B processes several hundred items per hour on a 12 GB card.
Field extraction
Invoices, resumes, product sheets, reports: Hermes returns exactly the requested fields, using null when information is missing, if you specify this in the system prompt.
Automated evaluation
Score texts according to a rubric (that's what our daily monitoring pipeline with hermes4:14b does). Stable formatting and the absence of sycophancy make the scores usable.
Agent with tools
Searching a database, calling an internal API, executing commands validated by a human. Reasoning mode helps with multi-step plans.
!
The 14B is still a 14B
With an agent that chains more than five or six tool calls with long results, the 14B starts losing the thread, whether the context is 16k or not. For these scenarios, the 70B really changes the game. If your hardware can’t handle it, split the task into short sub-agents instead of lengthening the loop.

#Recommended settings

The values below are what we use in production. They do not come from the model's datasheet, which remains brief on the subject, but from several months of daily use on a 12 GB card.

num_ctx
8,192 for chat and extraction, 16,384 for agents and reasoning mode. Each context doubling costs VRAM: on 12 GB, 16k works with the quantized KV cache (see below), but no further.
temperature
0.6 to 0.7 for conversation and writing. 0.1 to 0.2 for extraction, classification, and anything that must be reproducible.
top_p and repeat_penalty
0.95 and 1.05. A stronger repetition penalty degrades JSON outputs, where repeated keys are normal.
Quantization
Q4_K_M for all common use. Q8_0 if you do large-scale extraction and have 16 GB: formatting errors become almost nonexistent.

On the server side, two Ollama environment variables save space on a 12 GB card: flash attention and 8-bit KV-cache quantization. They are defined in the systemd service file on Linux, in the user environment variables on Windows, or through launchctl on macOS.

Linux — /etc/systemd/system/ollama.service.d/override.conf
[Service]
Environment="OLLAMA_FLASH_ATTENTION=1"
Environment="OLLAMA_KV_CACHE_TYPE=q8_0"
Environment="OLLAMA_KEEP_ALIVE=24h"
Terminal — apply
sudo systemctl daemon-reload
sudo systemctl restart ollama
ollama ps   # vérifier que le modèle est toujours à 100 % GPU

The KV cache in q8_0 halves the memory consumed by the context, with no measurable loss on extraction and chat tasks. This is the setting that makes it possible to run hermes4:14b with a 16k context on 12 GB. A 24-hour keep-alive prevents reloading 9 GB for every spaced-out call, which is useful for a server responding to scripts throughout the day.


#Troubleshooting

think tags appear in the response
Normal in reasoning mode in the terminal or through the raw API. Open WebUI folds them; in your scripts, remove everything before the closing tag before parsing. If you didn't want reasoning, remove the dedicated instruction from the system prompt.
The model responds in English
Hermes 4 is trained overwhelmingly on English data. Add “Tu réponds toujours en français” to the system prompt: this one line is enough in nearly all cases, precisely because the model follows the system prompt.
Unknown architecture error while loading
Your Ollama is too old for the 14B Qwen3 base. Update it, then run the pull again.
A few tokens per second
ollama ps montre un pourcentage CPU : le modèle ne tient pas en VRAM. Réduisez num_ctx, activez le cache KV q8_0, ou passez à hermes3:8b.
Invalid JSON from time to time
Use the format parameter with a schema: generation becomes constrained. If you can't, lower the temperature to 0.1 and add an example of the expected output to the system prompt.
The model forgets its instructions after several turns
The context is saturated. Increase num_ctx or truncate the history on the application side, always keeping the system prompt at the front.
Unexpected refusal
Rare, but possible with the 14B when the wording resembles an attack scenario. Rephrase it by specifying the legitimate context in the system prompt (authorized test, professional context): Hermes follows the framework you set.

#Go further

Hermes 4 shows its full value once you master the components around it. These site guides complement this one:

Master system prompts
The method for writing a complete system prompt, with examples for each use case. The natural companion to this guide, since this is how you control Hermes.
Function calling and structured JSON outputs with Ollama
Details on the format parameter, JSON schemas, and the tool-calling loop, with complete Python examples.
Choose your quantization (Q4, Q5, Q8, FP16)
To choose between Q4_K_M and Q8_0 based on your VRAM and tolerance for formatting errors.
Did this guide help you?

Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.