Hermes 4 locally (Ollama): Nous's model Research
Hermes 4 is the fourth generation of Nous Research models, from a lab known for fine-tunes that follow the system prompt precisely and rarely refuse a legitimate task. This guide explains how to install Hermes 4 with Ollama, which variant to choose based on your graphics card, and how to leverage its two strengths: adherence to system instructions and structured outputs for agents. We use it ourselves in production on a 12 GB RTX 5070 Ti; the figures come from that setup.
#Why Hermes 4 instead of a basic Qwen or Llama
Nous Research does not pretrain a model. The lab takes an existing open-weight model and applies extensive post-training: for Hermes 4, approximately 5 million examples and 60 billion tokens of synthetic data produced by its DataForge pipeline, with a high proportion of verified reasoning traces. The result is not more “intelligent” than its base model on general-knowledge benchmarks, but it behaves differently day to day: it does what you ask, in the requested format, without unnecessary commentary.
Three traits explain Hermes's popularity among self-hosted LLM users. First, the system prompt is treated as the source of truth: persona, tone, formatting constraints—everything is respected throughout a conversation. Second, its alignment is deliberately neutral: the model does not add unsolicited warnings and refuses much less often than mainstream assistants on ordinary topics (medicine, law, cybersecurity, adult fiction). Finally, Nous's tool-calling format, with its dedicated tags, has become a de facto standard adopted by many agent frameworks.
Hermes 4 adds a hybrid reasoning mode inherited from DeepHermes: the model can think between think tags before responding, or respond directly, depending on what you specify in the system prompt. You decide, query by query, whether to pay the extra token cost of reasoning.
#The Hermes 4 variants and the required VRAM
Your private ChatGPT, free, on your own machine in an hour — LM Studio, Ollama, Open WebUI, your documents, no cloud.
- Lifetime online access
- PDF + files
- Lifetime updates
Hermes 4 was released in late August 2025 in three sizes, each built on a different base. This detail matters: the license, maximum context, and behavior in French come from the base, not Nous Research.
- Hermes 4 14B
- Qwen3-14B base. About 9 GB of VRAM in Q4_K_M, a comfortable native context, and very good French. This is the variant to install on a 12–16 GB card and the one we use in production. License inherited from Qwen3: Apache 2.0.
- Hermes 4 70B
- Base Llama 3.1 70B. Approximately 40 GB in Q4_K_M: two RTX 3090/4090s, a Mac Studio, or a MacBook Pro with 48 GB or more of unified memory. Significantly better at long-form reasoning and following complex instructions. Llama 3.1 Community license.
- Hermes 4 405B
- Base Llama 3.1 405B. More than 200 GB even in Q4: beyond the reach of a personal computer, reserved for multi-GPU servers or API providers. We mention it for completeness.
- And below 12 GB?
- There is no Hermes 4 in 3B or 8B. On an 8 GB card, use hermes3:8b (based on Llama 3.1 8B, about 5 GB in Q4_K_M): same philosophy, same tool format, without reasoning mode.
The choice of quantization follows the site's usual rules: Q4_K_M is the default good compromise, Q5_K_M or Q8_0 if VRAM allows and you are doing structured extraction, where every bit of precision reduces formatting errors. On our RTX 5070 Ti with 12 GB, hermes4:14b in Q4_K_M with a context of 8,192 tokens uses 9.4 GB of VRAM, leaving room for an embeddings model loaded in parallel.
#Prerequisites
- Ollama up to date
- The Qwen3 14B base requires a version of Ollama later than April 2025. Check with ollama --version and update if needed; otherwise, you'll get an unknown architecture error when loading.
- GPU or unified memory
- 12 GB of VRAM for the 14B in Q4_K_M with an 8k context. With 8 GB, the model partially loads onto the CPU and drops to a few tokens per second: switch to hermes3:8b.
- Disk space
- Allow 9 GB for the 14B in Q4_K_M, 15 GB in Q8_0, and 40 GB for the 70B in Q4_K_M.
- An interface (optional)
- Open WebUI or LM Studio for comfortable discussion. Open WebUI automatically collapses reasoning blocks, which is useful with Hermes 4.
#Install Hermes 4 with Ollama in 4 steps
- 01Download the variant suited to your machineIn the Ollama library, the hermes4 tag comes in different sizes. On a 12 to 16 GB card, choose the 14B; that's the tag we run every day.
- 02Or import the official GGUF from Hugging FaceNous Research publishes its own GGUF quantizations. The hf.co syntax for Ollama lets you precisely choose the level (Q4_K_M, Q5_K_M, Q8_0) without using a Modelfile. This is the preferred approach if you want Q8_0 or if the library tag does not offer the quantization you want.
- 03Check loading and GPU/CPU distributionThe ollama ps command shows occupied memory and the percentage loaded on the GPU. Aim for 100% GPU: as soon as any part moves to the CPU, speed plummets.
- 04First test with a system promptDon't test Hermes without a system prompt—you'd miss what makes it interesting. Give it a role, an output format, and a constraint, then see how closely it sticks to them.
#What distinguishes Hermes from mainstream aligned models
A consumer assistant is trained to please an anonymous user and protect its publisher. This produces behaviors that eventually become invisible: an opening that restates the question, a warning at the end of the answer, a polite refusal as soon as a sensitive keyword appears, and a tendency to soften the edges. For conversation, that is acceptable. For an automated pipeline expecting JSON, a ranking, or a rewrite, every deviation breaks downstream processing.
Hermes 4 is trained for the operator, not the end user. Nous Research measured this choice with an in-house benchmark, RefusalBench, where Hermes 4 refuses significantly less often than closed models and most aligned open-weight models. In practice, on a support-ticket corpus asking it to extract the reason for a complaint, a consumer model may sometimes reply, “I can’t provide legal advice,” where Hermes returns the requested field. On a rewriting task, it doesn’t add “feel free to reach out if you have any other questions.”
- No preamble or conclusion
- If the system prompt says “respond only with JSON,” the response starts with a brace. On many models, you need three examples and post-processing to get the same result.
- Stable persona
- A role defined in the system prompt holds up for dozens of turns without drifting, even when the user tries to push it outside its boundaries.
- Fewer unsolicited refusals
- Medical topics, legal matters, cybersecurity, dark fiction: Hermes handles the request. The limits are the ones you write.
- Neutral tone
- No flattery about the quality of the question, no emojis, no artificial enthusiasm. The style is controlled entirely by the system prompt.
The trade-off is obvious: if you expose Hermes to strangers, through a public chatbot for example, it's up to you to write the guardrails. A system prompt listing what the assistant must not do, plus an application-side output filter for critical cases, is essential. For personal or internal use, that responsibility is precisely the advantage you're looking for.
#The system prompt, where Hermes 4 excels
With Hermes 4, the system prompt is not a suggestion but a contract. This changes how you write it: be precise, complete, and do not expect the model to fill gaps with an assistant’s “common sense.” Three rules are enough to get reliable results.
- Describe the role in one sentence
- “You are the marketing team’s writing assistant for X” is better than three paragraphs of context. Hermes does not dilute the information, but a short role is more stable.
- Set the output format explicitly
- Length, language, structure, and what must be omitted. Hermes follows a constraint such as “no introduction, no conclusion, 120 words maximum” without having to repeat it.
- Write down the prohibitions
- Since the model doesn't add any, list your own: out-of-scope topics, information that must never be revealed, and behavior when a question is ambiguous.
The most practical approach is to lock this system prompt into a Modelfile: you get a ready-to-use alias, with the context and temperature already configured, usable in Open WebUI, in your scripts, and from the terminal.
#Enable or disable reasoning mode
Hermes 4 is a hybrid-reasoning model: by default, it responds directly, like a classic instruct model. To enable reasoning, add a dedicated instruction to the system prompt—the one Nous Research documents on the model card. The model then produces its reasoning between think tags, followed by its answer.
You can combine it with your own instructions, in French, afterward. Reasoning is useful for math, nontrivial code, agent planning, or analyzing a long document. It's unnecessary and costly for field extraction, classification, or rewriting: for those tasks, leave direct mode enabled. Expect 500 to several thousand reasoning tokens per request, which is why a generous num_ctx matters.
- Direct mode (default)
- No special instruction. Immediate response, ideal for pipelines and everyday chat. This is the mode we use to evaluate our monitoring system automatically.
- Reasoning mode
- Add the instruction above at the top of the system prompt. Open WebUI collapses the think block; in your scripts, filter it out before displaying or parsing the response.
- Two aliases
- The simplest approach is to create two Modelfiles, hermes-direct and hermes-think, and choose one or the other depending on the task instead of changing the prompt for every call.
#Use case: agents and structured extraction
Nous Research’s tool-calling format, in which the model receives the list of available functions in the system prompt and emits its calls as tagged JSON, is what Ollama uses through the tools parameter of its chat API. With Hermes 4, you do not need to configure anything: describe your functions, send the conversation, and the model returns a structured call when needed, or a text response otherwise.
For structured extraction, two approaches complement each other. The first is to enforce a JSON schema through Ollama's format parameter: the engine constrains generation, and the model cannot deviate from the schema. The second relies solely on the system prompt, and that's where Hermes makes the difference: even without constrained decoding, it returns valid JSON without commentary in the vast majority of cases, simplifying integrations with tools that do not support constrained decoding.
The response contains a tool_calls field with the function name and its arguments. You must execute the function and return the result in a tool-role message so the model can formulate its final response. This is the loop implemented by LangChain, CrewAI, n8n, or Hermes Agent, the agent framework published by Nous Research itself, which works with any model but was designed around this format.
- Classification and routing
- Sort emails, tickets, or articles into defined categories. Low temperature, direct mode, enforced JSON schema: the 14B processes several hundred items per hour on a 12 GB card.
- Field extraction
- Invoices, resumes, product sheets, reports: Hermes returns exactly the requested fields, using null when information is missing, if you specify this in the system prompt.
- Automated evaluation
- Score texts according to a rubric (that's what our daily monitoring pipeline with hermes4:14b does). Stable formatting and the absence of sycophancy make the scores usable.
- Agent with tools
- Searching a database, calling an internal API, executing commands validated by a human. Reasoning mode helps with multi-step plans.
#Recommended settings
The values below are what we use in production. They do not come from the model's datasheet, which remains brief on the subject, but from several months of daily use on a 12 GB card.
- num_ctx
- 8,192 for chat and extraction, 16,384 for agents and reasoning mode. Each context doubling costs VRAM: on 12 GB, 16k works with the quantized KV cache (see below), but no further.
- temperature
- 0.6 to 0.7 for conversation and writing. 0.1 to 0.2 for extraction, classification, and anything that must be reproducible.
- top_p and repeat_penalty
- 0.95 and 1.05. A stronger repetition penalty degrades JSON outputs, where repeated keys are normal.
- Quantization
- Q4_K_M for all common use. Q8_0 if you do large-scale extraction and have 16 GB: formatting errors become almost nonexistent.
On the server side, two Ollama environment variables save space on a 12 GB card: flash attention and 8-bit KV-cache quantization. They are defined in the systemd service file on Linux, in the user environment variables on Windows, or through launchctl on macOS.
The KV cache in q8_0 halves the memory consumed by the context, with no measurable loss on extraction and chat tasks. This is the setting that makes it possible to run hermes4:14b with a 16k context on 12 GB. A 24-hour keep-alive prevents reloading 9 GB for every spaced-out call, which is useful for a server responding to scripts throughout the day.
#Troubleshooting
- think tags appear in the response
- Normal in reasoning mode in the terminal or through the raw API. Open WebUI folds them; in your scripts, remove everything before the closing tag before parsing. If you didn't want reasoning, remove the dedicated instruction from the system prompt.
- The model responds in English
- Hermes 4 is trained overwhelmingly on English data. Add “Tu réponds toujours en français” to the system prompt: this one line is enough in nearly all cases, precisely because the model follows the system prompt.
- Unknown architecture error while loading
- Your Ollama is too old for the 14B Qwen3 base. Update it, then run the pull again.
- A few tokens per second
- ollama ps montre un pourcentage CPU : le modèle ne tient pas en VRAM. Réduisez num_ctx, activez le cache KV q8_0, ou passez à hermes3:8b.
- Invalid JSON from time to time
- Use the format parameter with a schema: generation becomes constrained. If you can't, lower the temperature to 0.1 and add an example of the expected output to the system prompt.
- The model forgets its instructions after several turns
- The context is saturated. Increase num_ctx or truncate the history on the application side, always keeping the system prompt at the front.
- Unexpected refusal
- Rare, but possible with the 14B when the wording resembles an attack scenario. Rephrase it by specifying the legitimate context in the system prompt (authorized test, professional context): Hermes follows the framework you set.
#Go further
Hermes 4 shows its full value once you master the components around it. These site guides complement this one:
- Master system prompts
- The method for writing a complete system prompt, with examples for each use case. The natural companion to this guide, since this is how you control Hermes.
- Function calling and structured JSON outputs with Ollama
- Details on the format parameter, JSON schemas, and the tool-calling loop, with complete Python examples.
- Choose your quantization (Q4, Q5, Q8, FP16)
- To choose between Q4_K_M and Q8_0 based on your VRAM and tolerance for formatting errors.
Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.