Intermediate 10 minCommunity

Qwythos 9B: the small reasoning model trained on Claude, tested in local

Qwythos-9B is a community model published by Empero AI in June 2026: a Qwen3.5-9B base under the Apache 2.0 license, with a 1M context, fine-tuned on reasoning traces generated by Claude—hence the portmanteau “Qwythos” (Qwen + Mythos). There is no polished marketing page behind it: just GGUF weights on Hugging Face and community-driven Ollama tags. This guide sorts out its origins, shows how to install it properly, and puts it through a real-world test against its Qwen3.5-9B base on a RTX 5070 Ti.

By Mohamed Meguedmi·Update 2026-08-21·Tested on Windows, macOS, and Linux

#What exactly is Qwythos?

Qwythos-9B isn't a “from scratch” model. It's a fine-tune: someone took a strong open base—Alibaba's Qwen3.5-9B—and retrained it on a specific corpus of reasoning traces. These traces (detailed chains of thought, question → step-by-step reasoning → answer) reportedly came from outputs by Claude, Anthropic's model family, on a set of logic, math, and coding problems. The name “Qwythos” condenses the idea: the initial letter of Qwen, and “Mythos,” the informal name given to Claude's reasoning style by the community that assembled the dataset.

The project’s appeal comes down to one sentence: distill a verbose, structured reasoning style into a model small enough to run on a consumer GPU. A 9B model fits in 6 GB of VRAM in Q4, while large reasoning models require tens of GB. If the transfer of a “way of thinking” works, we get a pocket-sized reasoner. That’s exactly what we’re going to verify.

i
Community model, not an official product
Qwythos is neither an Alibaba model nor an Anthropic model. It is a community fine-tune distributed by Empero AI. Neither original company maintains or endorses it. Treat it like an open-source project: the weights are available, the quality needs to be verified, and support depends on the goodwill of its authors.

#Qwythos origins: the two sources you need to know

The Local AI Kit

Your private ChatGPT, free, on your own machine in an hour — LM Studio, Ollama, Open WebUI, your documents, no cloud.

  • Lifetime online access
  • PDF + files
  • Lifetime updates

Before installing anything, you need to understand where the weights actually come from. For a community model, provenance is not a detail: it determines how much you can trust the file you download. Qwythos is available in two places, and they are not equivalent.

Source 1 — Hugging Face (original repository)
Empero AI publishes the weights on Hugging Face, both in transformers format (safetensors) and quantized GGUF. This is the source of truth: model card, displayed Apache 2.0 license, mention of the Qwen3.5-9B base, and the reasoning-trace dataset. Always start there to verify the hash and read the card.
Source 2 — Community tags Ollama
In the Ollama registry, Qwythos appears under user namespaces (such as username/qwythos), not in the official library. These tags are convenient—a single command to install—but they are repackaged by third parties from Hugging Face GGUFs. Check who publishes the tag before trusting it.
!
A Ollama tag is unsigned
Anyone can push a model to the Ollama registry under their own name. A “qwythos” tag does not guarantee that the GGUF behind it matches Empero AI's official weights. For serious use, prefer downloading the GGUF from Hugging Face and importing it yourself via a Modelfile (see below), rather than blindly pulling a community tag.
Base
Qwen3.5-9B (Alibaba), Apache 2.0 license
Type
Reasoning fine-tune, explicit chain of thought
Editor
Empero AI (community), published June 2026
License
Apache 2.0 — commercial use permitted, inherited from the base
Context
1M tokens claimed (inherited from Qwen3.5), to be tempered according to actual VRAM
Formats
safetensors + GGUF (Q4_K_M, Q5_K_M, Q8_0) on Hugging Face

#Hardware requirements

A 9-billion-parameter model is comfortable on recent hardware. In Q4_K_M quantization—the recommended quality/memory compromise—Qwythos-9B uses about 6 GB of VRAM once loaded, plus the context cache (KV cache), which grows with prompt length. Our test bench runs on a RTX 5070 Ti (16 GB), well above the minimum.

Q4_K_M (recommended)
≈ 6 GB of VRAM. Fits on a RTX 3060 12 GB, a 4070, or a 5070 Ti—and even partly on 8 GB with a modest context.
Q5_K_M
≈ 7 GB. Slight quality improvement; preferred if you have 12 GB or more.
Q8_0
≈ 10 GB. Nearly lossless compared with FP16, relevant for 16 GB and above.
Long context
The announced 1M is theoretical. A KV cache for 128k tokens already adds several GB; target 16 GB+ of VRAM to comfortably exceed 32k.
→
A reasoning model generates a lot
Reasoning models “think” out loud before answering: they sometimes produce several hundred reasoning tokens for a one-line answer. Leave headroom in the output context (num_predict), and expect longer response times than with a conventional model of the same size.

#Install Qwythos with Ollama

Two paths. The fast one: pull a community tag. The safe one: download the official GGUF from Hugging Face and import it through a Modelfile. I’ll cover both; for a serious workstation, the second is preferable because you know exactly which file you’re running.

  1. 01
    Verify that Ollama is running
    Ollama exposes its daemon on http://localhost:11434. A simple call confirms that it responds before you go further. If you haven't installed it yet, see the Ollama installation guide.
  2. 02
    Fast track — community tag
    Find the tag on ollama.com (user namespace), then run ollama run. Note the publisher’s exact name: it is your only guarantee of provenance.
  3. 03
    Safe route — GGUF from Hugging Face
    Download the .gguf file in Q4_K_M from the Empero AI repository, verify its SHA256 checksum against the one displayed on the model card, then create a Modelfile that points to it.
  4. 04
    Create the local model
    A ollama create builds a named model from your Modelfile. You get a local tag that you control end to end.
  5. 05
    Launch and chat
    ollama run d starts the interactive session. The first load puts the model in VRAM; subsequent loads are instant as long as it remains resident.
Terminal — check Ollama
curl http://localhost:11434/api/version
# {"version":"0.x.x"}
Terminal — fast path (community tag)
# Remplacez <publieur> par le nom réel du dépôt Ollama
ollama run <publieur>/qwythos:9b-q4_k_m
Terminal—safe path (Hugging Face GGUF)
# 1. Télécharger le GGUF officiel depuis le dépôt Empero AI
huggingface-cli download EmperoAI/Qwythos-9B-GGUF \
  qwythos-9b-Q4_K_M.gguf --local-dir ./qwythos

# 2. Vérifier l'empreinte contre celle de la carte du modèle
sha256sum ./qwythos/qwythos-9b-Q4_K_M.gguf
Modelfile — qwythos.Modelfile
FROM ./qwythos/qwythos-9b-Q4_K_M.gguf

PARAMETER temperature 0.6
PARAMETER top_p 0.95
PARAMETER num_ctx 32768

# Encourage la chaîne de pensée explicite
SYSTEM """Tu es un assistant de raisonnement. Décompose chaque problème étape par étape avant de conclure."""
Terminal — import and run
ollama create qwythos-9b -f qwythos.Modelfile
ollama run qwythos-9b
i
Lower temperature for reasoning
For reasoning models, a temperature around 0.6 produces more stable chains of thought than the usual 0.8. Too high, and the model goes off on tangents; too low (0.1-0.2), and it repeats itself. The 0.6 / top_p 0.95 combination is a good starting point; adjust it to your tasks.

#First reasoning test

We start with a classic trick problem, the kind that trips up small models not trained for reasoning. The goal is not just to get the right answer, but to see whether the model follows a coherent process instead of guessing.

Test prompt
ollama run qwythos-9b "Un batte et une balle coûtent 1,10 € au total. La batte coûte 1 € de plus que la balle. Combien coûte la balle ? Raisonne étape par étape."

Expected answer: €0.05. The intuitive trap leads to €0.10. In our benchmark, Qwythos-9B sets up the equation (ball = x, bat = x + 1, total 2x + 1 = 1.10), isolates x = 0.05, and verifies that 0.05 + 1.05 = 1.10 before concluding. The process is verbose—about ten lines of reasoning for a one-digit answer—but it is correct and traceable. That is precisely the behavior we expect from a reasoning fine-tune.

→
Test on YOUR problems
A model may have seen this kind of classic puzzle during training. To evaluate it honestly, create two or three problems from your actual domain—a logistics constraint, a piece of code to debug, or a business calculation—and compare the reasoning process, not just the final result.

#Qwythos vs. base Qwen3.5-9B: what fine-tuning changes

The real question: does the fine-tune bring anything compared with the base model it came from? To measure this, we install the base Qwen3.5-9B in parallel and give both models the same series of logic and math problems at the same temperature. Here is what emerges from our RTX 5070 Ti benchmark.

Terminal — install the baseline for comparison
ollama pull qwen3.5:9b
ollama run qwen3.5:9b "Un batte et une balle coûtent 1,10 €..."
Reasoning length
Qwythos consistently lays out an explicit chain of thought; the Qwen3.5 base model often responds more briefly, sometimes directly, without showing the intermediate steps.
Trick questions
On counterintuitive puzzles (bat-and-ball, logic sequences), Qwythos makes fewer mistakes because it checks its answer. The base model falls into the intuitive trap more easily.
Speed
An advantage from the outset: it generates fewer tokens for an equivalent response. Qwythos is slower in practice because it “thinks” before responding—the price of reasoning.
Raw knowledge
A draw. Reasoning fine-tuning does not add factual knowledge; on general-knowledge questions, the two models are equal (and share the same gaps).
French
Equivalent. The quality of the French comes from the Qwen3.5 base; Qwythos neither improves nor degrades fluency—it changes the response structure, not the language.

Comparison verdict: Qwythos clearly wins on tasks where a process matters (math, logic, debugging, planning), and loses on speed and short answers. If you want a fast assistant for rewriting or summarization, the Qwen3.5-9B base model is sufficient. If you want a small reasoner that shows its work, Qwythos justifies its existence.

!
Reasoning isn’t free
Chain of thought = more generated tokens = more latency and more context VRAM consumed. For simple tasks, forcing a model to reason slows it down without adding anything. Reserve Qwythos for problems that genuinely benefit from decomposition, and keep a more direct model for everything else.

#The bigger sibling: Qwythos 27B

Empero AI also publishes a 27B variant, built on a larger Qwen3.5 base according to the same principle: fine-tuning on reasoning traces. It targets those with enough VRAM to go further in reasoning quality, at the cost of a significantly larger memory budget.

VRAM (Q4_K_M)
≈ 19 GB. You need a RTX 4090 24 GB, a professional card, or a Mac with unified memory (M4 Pro/Max) to run it comfortably.
What you gain
Longer, more robust reasoning chains on multistep problems; fewer cascading arithmetic errors. The gap widens mainly on genuinely difficult tasks.
What you pay for
Three times more VRAM and slower generation. On 12–16 GB hardware, the 27B simply won't fit in Q4 without CPU offload, which severely hurts speed.
When to choose it
If the 9B regularly fails on your real-world problems AND you have 24 GB of VRAM. Otherwise, the 9B remains the best quality-to-resource ratio.
i
Start with the 9B
Don't jump to 27B by default. Test 9B on your real tasks first: in many cases, it's enough. Move up to 27B only if you identify a concrete quality ceiling on problems you actually need to solve—not on principle.

#Troubleshooting

The Ollama tag does not exist / has disappeared
Community tags come and go. Switch to the safe route: download the GGUF from Hugging Face and import it through Modelfile. You are no longer dependent on a third party's availability.
The model doesn't reason; it gives terse answers
Add an explicit system instruction (“break it down step by step”) and check that num_ctx is large enough to leave room for reasoning. A temperature that is too low also suppresses the chain of thought.
Truncated responses
Reasoning consumes the output budget. Increase num_predict (or your client's equivalent parameter) to let the model finish its reasoning AND provide its answer.
Very slow output
Make sure the model fits comfortably in VRAM (ollama ps). If it spills into CPU RAM, speed collapses: step down one quantization level or reduce num_ctx.
File integrity in doubt
Compare the SHA256 of the downloaded GGUF with the one published on the Hugging Face model card. A different hash means the file was repackaged or corrupted—do not run it.

#Go further

To install the components around Qwythos and understand the tradeoffs discussed here:

Install Ollama (Windows, macOS, Linux)
The prerequisite for this guide: set up the Ollama daemon on port 11434 before pulling any model.
Choose your quantization (Q4, Q5, Q8, FP16)
To choose between Q4_K_M, Q5_K_M, and Q8_0 based on your VRAM and quality requirements for Qwythos and its siblings.
Install DeepSeek R1 with Ollama
Another open-source chain-of-thought reasoning model, useful for comparing Qwythos’s approach with a well-established reference.
Did this guide help you?

Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.

Prices in euros (€) are French market prices including VAT, checked by QuelLLM. US prices differ: the Amazon buttons show the current US price.