Qwythos 9B: the small reasoning model trained on Claude, tested in local
Qwythos-9B is a community model published by Empero AI in June 2026: a Qwen3.5-9B base under the Apache 2.0 license, with a 1M context, fine-tuned on reasoning traces generated by Claude—hence the portmanteau “Qwythos” (Qwen + Mythos). There is no polished marketing page behind it: just GGUF weights on Hugging Face and community-driven Ollama tags. This guide sorts out its origins, shows how to install it properly, and puts it through a real-world test against its Qwen3.5-9B base on a RTX 5070 Ti.
#What exactly is Qwythos?
Qwythos-9B isn't a “from scratch” model. It's a fine-tune: someone took a strong open base—Alibaba's Qwen3.5-9B—and retrained it on a specific corpus of reasoning traces. These traces (detailed chains of thought, question → step-by-step reasoning → answer) reportedly came from outputs by Claude, Anthropic's model family, on a set of logic, math, and coding problems. The name “Qwythos” condenses the idea: the initial letter of Qwen, and “Mythos,” the informal name given to Claude's reasoning style by the community that assembled the dataset.
The project’s appeal comes down to one sentence: distill a verbose, structured reasoning style into a model small enough to run on a consumer GPU. A 9B model fits in 6 GB of VRAM in Q4, while large reasoning models require tens of GB. If the transfer of a “way of thinking” works, we get a pocket-sized reasoner. That’s exactly what we’re going to verify.
#Qwythos origins: the two sources you need to know
Your private ChatGPT, free, on your own machine in an hour — LM Studio, Ollama, Open WebUI, your documents, no cloud.
- Lifetime online access
- PDF + files
- Lifetime updates
Before installing anything, you need to understand where the weights actually come from. For a community model, provenance is not a detail: it determines how much you can trust the file you download. Qwythos is available in two places, and they are not equivalent.
- Source 1 — Hugging Face (original repository)
- Empero AI publishes the weights on Hugging Face, both in transformers format (safetensors) and quantized GGUF. This is the source of truth: model card, displayed Apache 2.0 license, mention of the Qwen3.5-9B base, and the reasoning-trace dataset. Always start there to verify the hash and read the card.
- Source 2 — Community tags Ollama
- In the Ollama registry, Qwythos appears under user namespaces (such as username/qwythos), not in the official library. These tags are convenient—a single command to install—but they are repackaged by third parties from Hugging Face GGUFs. Check who publishes the tag before trusting it.
- Base
- Qwen3.5-9B (Alibaba), Apache 2.0 license
- Type
- Reasoning fine-tune, explicit chain of thought
- Editor
- Empero AI (community), published June 2026
- License
- Apache 2.0 — commercial use permitted, inherited from the base
- Context
- 1M tokens claimed (inherited from Qwen3.5), to be tempered according to actual VRAM
- Formats
- safetensors + GGUF (Q4_K_M, Q5_K_M, Q8_0) on Hugging Face
#Hardware requirements
A 9-billion-parameter model is comfortable on recent hardware. In Q4_K_M quantization—the recommended quality/memory compromise—Qwythos-9B uses about 6 GB of VRAM once loaded, plus the context cache (KV cache), which grows with prompt length. Our test bench runs on a RTX 5070 Ti (16 GB), well above the minimum.
- Q4_K_M (recommended)
- ≈ 6 GB of VRAM. Fits on a RTX 3060 12 GB, a 4070, or a 5070 Ti—and even partly on 8 GB with a modest context.
- Q5_K_M
- ≈ 7 GB. Slight quality improvement; preferred if you have 12 GB or more.
- Q8_0
- ≈ 10 GB. Nearly lossless compared with FP16, relevant for 16 GB and above.
- Long context
- The announced 1M is theoretical. A KV cache for 128k tokens already adds several GB; target 16 GB+ of VRAM to comfortably exceed 32k.
#Install Qwythos with Ollama
Two paths. The fast one: pull a community tag. The safe one: download the official GGUF from Hugging Face and import it through a Modelfile. I’ll cover both; for a serious workstation, the second is preferable because you know exactly which file you’re running.
- 01Verify that Ollama is runningOllama exposes its daemon on http://localhost:11434. A simple call confirms that it responds before you go further. If you haven't installed it yet, see the Ollama installation guide.
- 02Fast track — community tagFind the tag on ollama.com (user namespace), then run ollama run. Note the publisher’s exact name: it is your only guarantee of provenance.
- 03Safe route — GGUF from Hugging FaceDownload the .gguf file in Q4_K_M from the Empero AI repository, verify its SHA256 checksum against the one displayed on the model card, then create a Modelfile that points to it.
- 04Create the local modelA ollama create builds a named model from your Modelfile. You get a local tag that you control end to end.
- 05Launch and chatollama run d starts the interactive session. The first load puts the model in VRAM; subsequent loads are instant as long as it remains resident.
#First reasoning test
We start with a classic trick problem, the kind that trips up small models not trained for reasoning. The goal is not just to get the right answer, but to see whether the model follows a coherent process instead of guessing.
Expected answer: €0.05. The intuitive trap leads to €0.10. In our benchmark, Qwythos-9B sets up the equation (ball = x, bat = x + 1, total 2x + 1 = 1.10), isolates x = 0.05, and verifies that 0.05 + 1.05 = 1.10 before concluding. The process is verbose—about ten lines of reasoning for a one-digit answer—but it is correct and traceable. That is precisely the behavior we expect from a reasoning fine-tune.
#Qwythos vs. base Qwen3.5-9B: what fine-tuning changes
The real question: does the fine-tune bring anything compared with the base model it came from? To measure this, we install the base Qwen3.5-9B in parallel and give both models the same series of logic and math problems at the same temperature. Here is what emerges from our RTX 5070 Ti benchmark.
- Reasoning length
- Qwythos consistently lays out an explicit chain of thought; the Qwen3.5 base model often responds more briefly, sometimes directly, without showing the intermediate steps.
- Trick questions
- On counterintuitive puzzles (bat-and-ball, logic sequences), Qwythos makes fewer mistakes because it checks its answer. The base model falls into the intuitive trap more easily.
- Speed
- An advantage from the outset: it generates fewer tokens for an equivalent response. Qwythos is slower in practice because it “thinks” before responding—the price of reasoning.
- Raw knowledge
- A draw. Reasoning fine-tuning does not add factual knowledge; on general-knowledge questions, the two models are equal (and share the same gaps).
- French
- Equivalent. The quality of the French comes from the Qwen3.5 base; Qwythos neither improves nor degrades fluency—it changes the response structure, not the language.
Comparison verdict: Qwythos clearly wins on tasks where a process matters (math, logic, debugging, planning), and loses on speed and short answers. If you want a fast assistant for rewriting or summarization, the Qwen3.5-9B base model is sufficient. If you want a small reasoner that shows its work, Qwythos justifies its existence.
#The bigger sibling: Qwythos 27B
Empero AI also publishes a 27B variant, built on a larger Qwen3.5 base according to the same principle: fine-tuning on reasoning traces. It targets those with enough VRAM to go further in reasoning quality, at the cost of a significantly larger memory budget.
- VRAM (Q4_K_M)
- ≈ 19 GB. You need a RTX 4090 24 GB, a professional card, or a Mac with unified memory (M4 Pro/Max) to run it comfortably.
- What you gain
- Longer, more robust reasoning chains on multistep problems; fewer cascading arithmetic errors. The gap widens mainly on genuinely difficult tasks.
- What you pay for
- Three times more VRAM and slower generation. On 12–16 GB hardware, the 27B simply won't fit in Q4 without CPU offload, which severely hurts speed.
- When to choose it
- If the 9B regularly fails on your real-world problems AND you have 24 GB of VRAM. Otherwise, the 9B remains the best quality-to-resource ratio.
#Troubleshooting
- The Ollama tag does not exist / has disappeared
- Community tags come and go. Switch to the safe route: download the GGUF from Hugging Face and import it through Modelfile. You are no longer dependent on a third party's availability.
- The model doesn't reason; it gives terse answers
- Add an explicit system instruction (“break it down step by step”) and check that num_ctx is large enough to leave room for reasoning. A temperature that is too low also suppresses the chain of thought.
- Truncated responses
- Reasoning consumes the output budget. Increase num_predict (or your client's equivalent parameter) to let the model finish its reasoning AND provide its answer.
- Very slow output
- Make sure the model fits comfortably in VRAM (ollama ps). If it spills into CPU RAM, speed collapses: step down one quantization level or reduce num_ctx.
- File integrity in doubt
- Compare the SHA256 of the downloaded GGUF with the one published on the Hugging Face model card. A different hash means the file was repackaged or corrupted—do not run it.
#Go further
To install the components around Qwythos and understand the tradeoffs discussed here:
- Install Ollama (Windows, macOS, Linux)
- The prerequisite for this guide: set up the Ollama daemon on port 11434 before pulling any model.
- Choose your quantization (Q4, Q5, Q8, FP16)
- To choose between Q4_K_M, Q5_K_M, and Q8_0 based on your VRAM and quality requirements for Qwythos and its siblings.
- Install DeepSeek R1 with Ollama
- Another open-source chain-of-thought reasoning model, useful for comparing Qwythos’s approach with a well-established reference.
Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.