Intermediate 9 minOllama

Import a GGUF model from Hugging Face into Ollama

The official Ollama library covers only a fraction of the available models. On Hugging Face, tens of thousands of GGUF files are waiting—community fine-tunes, recent models, and versions not yet packaged. This guide shows how to import any Hugging Face GGUF into Ollama: the direct ollama run hf.co command, the Modelfile FROM method for a local file, how to choose quantization based on your VRAM, and how to fix a broken chat template that makes responses inconsistent.

By Marie L.·Update 2026-08-27·Tested on Windows, macOS, and Linux

#Why import a GGUF from Hugging Face

Ollama maintains a library of practical models (ollama.com/library), but it is intentionally limited: the maintainers publish the most requested models there, using quantizations selected for them. As soon as you are looking for a specialized fine-tune, a newly released version, a French-language model, or a specific quantization, you have to find it on Hugging Face—the largest platform for sharing open-weight models.

The GGUF format (successor to GGML) is the format that Ollama understands natively: a single file containing the quantized weights, tokenizer, and model metadata. Contributors such as TheBloke, bartowski, and unsloth publish thousands of ready-to-use GGUFs, often in around ten quantizations per model. Knowing how to import them unlocks this entire ecosystem in your Ollama installation.

Recent models
A model published yesterday on Hugging Face can be used before it even appears in Ollama's official library.
Niche fine-tunes
Specialized models (code, medicine, role-playing, French) that no one has taken the trouble to package officially.
Precise quantization
Choose exactly the level (Q4_K_M, Q5_K_M, Q8_0…) that fits in your VRAM, rather than settling for the default variant.
Private models
Your own fine-tunes or downloaded GGUF files, imported locally through a Modelfile.
i
GGUF, GGML, safetensors?
Ollama reads GGUF, not safetensors (the PyTorch training format). If a repository contains only .safetensors files, you first need to convert it to GGUF with llama.cpp—or find a “GGUF” version already converted by the community (search Hugging Face for the model name + “GGUF”).

#Prerequisites

The Local AI Kit

Your private ChatGPT, free, on your own machine in an hour — LM Studio, Ollama, Open WebUI, your documents, no cloud.

  • Lifetime online access
  • PDF + files
  • Lifetime updates
Ollama installed
Recent version (0.5+) for native hf.co support. The daemon listens on http://localhost:11434 by default. Check with “ollama --version”.
An internet connection
For the direct method that downloads from Hugging Face. After that, the model runs 100% locally.
Enough VRAM or RAM
Q4 rule of thumb: a 7B fits in ~5 GB, a 14B in ~9 GB, a 32B in ~19 GB, and a 70B in ~40 GB. Without a GPU, RAM is what matters, and it will be slower.
The name of a GGUF repository
For example, bartowski/Qwen3.5-9B-Instruct-GGUF. Find it in the URL of the model's Hugging Face page.

To find a GGUF repository, Hugging Face search supports filtering by format. Search for the model name and add “GGUF,” or filter for the “GGUF” library in the sidebar. Open the “Files and versions” tab: you'll see the list of .gguf files, one per quantization, with their size in GB — valuable information for what comes next.

#Direct method: ollama run hf.co/...

This is by far the simplest way to import a GGUF model from Hugging Face into Ollama. Since version 0.5, Ollama can pull a GGUF directly from a Hugging Face repository with a single command, without downloading the file manually or writing a Modelfile. The syntax uses the repository path prefixed with hf.co/.

Terminal — run a GGUF from Hugging Face
# Format : ollama run hf.co/{utilisateur}/{depot}
ollama run hf.co/bartowski/Qwen3.5-9B-Instruct-GGUF

Without a specification, Ollama chooses a default quantization (usually Q4_K_M if it exists in the repository). To target a specific quantization, add it after a colon, exactly like a standard model tag. The tag name matches the .gguf filename suffix and is case-insensitive.

Terminal — target a quantization
# Choisir explicitement Q5_K_M
ollama run hf.co/bartowski/Qwen3.5-9B-Instruct-GGUF:Q5_K_M

# Ou une version plus légère pour une petite carte
ollama run hf.co/bartowski/Qwen3.5-9B-Instruct-GGUF:Q4_K_M

Ollama downloads the file, saves it to its local storage, and starts the conversation. The model then appears in « ollama list » under its full hf.co/... name and can be restarted instantly. You can give it a shorter alias with « ollama cp » if the name feels too long to type.

Terminal — shorten the name
# Copier vers un alias court
ollama cp hf.co/bartowski/Qwen3.5-9B-Instruct-GGUF:Q4_K_M qwen-fr

# Désormais utilisable simplement
ollama run qwen-fr
→
Private or gated repositories
For a private or gated repository, authenticate first. Add your Hugging Face key to your SSH keys on the site, or export an access token. Most public community GGUFs do not require authentication.

#Choosing the right quantization for your VRAM

The same model is released in several quantizations: this is the central trade-off between quality and memory. The more aggressive the quantization (fewer bits per weight), the smaller the file and the more easily it fits on a modest card—at the cost of a slight loss in precision. The right approach is to choose the highest quantization that fits comfortably in your VRAM.

Q4_K_M — recommended
The best trade-off for the vast majority of use cases. Almost imperceptible quality loss, with a smaller memory footprint. Choose it by default if you're unsure.
Q5_K_M — one step up
Slightly heavier, slightly more accurate. Interesting if you have VRAM to spare and want maximum quality without moving to 8-bit.
Q8_0 — nearly lossless
Very close to the unquantized model, but about twice as large as Q4. Reserved for cases where even the slightest degradation matters and VRAM is plentiful.
FP16 — full precision
The unquantized model, and the heaviest one. Rarely necessary for local inference: Q8_0 is almost always sufficient and cuts memory usage in half.

To estimate whether a quantization will fit, use the .gguf file size shown on Hugging Face, plus a margin of about 1 to 2 GB for the context and system. Here are the Q4_K_M VRAM guidelines by model size, along with the typical GPUs that can run them.

3B ≈ 2 GB
Runs everywhere, even on an entry-level card or on the CPU. Ideal RTX 3060 12 GB with context to spare.
7B ≈ 5 GB
Comfortable on RTX 3060 12 GB, RTX 4070 12 GB. The most versatile format for everyday use.
14B ≈ 9 GB
RTX 4070 12 GB (barely), RTX 4080 16 GB comfortably. A good quality tier for reasoning and coding.
32B ≈ 19 GB
RTX 4090 24 GB, or a Mac M4 Pro with unified memory. The accessible high end for a workstation.
70B ≈ 40 GB
Requires 48 GB+: a Mac Studio with large unified memory or a multi-GPU setup. Use Q4 or an even more aggressive quantization.
!
Don't go below Q4 without a reason
Q3, Q2, or IQ2 quantizations let large models fit into limited VRAM, but the degradation becomes clearly noticeable (less coherent responses, reasoning errors). A 7B in Q4_K_M is better than a 14B in Q2 on the same card. The dedicated quantization guide goes into these trade-offs in detail.

#Modelfile method: FROM filename.gguf

The direct method assumes that the GGUF is on Hugging Face and accessible online. But if you have already downloaded a .gguf file manually, built your own with llama.cpp, or want to customize the model (system prompt, parameters), you need to use a Modelfile. It is a small text file, similar to a Dockerfile, that describes how to build a Ollama model from a local GGUF.

The core directive is FROM, which points to the .gguf file path. Create a file named “Modelfile” (with no extension) next to your GGUF, containing at least this line.

Modelfile — minimal
FROM ./mon-modele.Q4_K_M.gguf

Then build the model with « ollama create », giving it any name you like. Ollama reads the GGUF, saves it to its storage, and makes it available like any other model.

Terminal — create and run
# Construire le modèle depuis le Modelfile du dossier courant
ollama create mon-modele -f ./Modelfile

# Le lancer
ollama run mon-modele

A complete Modelfile can go much further: set a system prompt, tune sampling parameters, and above all define the TEMPLATE—the chat format expected by the model. This is where most quality issues are fixed, as we'll see in the next section.

Modelfile — complete
FROM ./mon-modele.Q4_K_M.gguf

# System prompt par défaut
SYSTEM """Tu es un assistant francophone concis et précis."""

# Paramètres d'inférence
PARAMETER temperature 0.7
PARAMETER top_p 0.9
PARAMETER num_ctx 8192
PARAMETER stop "<|im_end|>"

# Template de chat (exemple format ChatML)
TEMPLATE """{{ if .System }}<|im_start|>system
{{ .System }}<|im_end|>
{{ end }}{{ if .Prompt }}<|im_start|>user
{{ .Prompt }}<|im_end|>
{{ end }}<|im_start|>assistant
{{ .Response }}<|im_end|>
"""
i
One GGUF, multiple variants
The Modelfile is also the clean way to derive several assistants from the same GGUF: a “translator,” a “coder,” and a “French assistant,” each with its own system prompt and parameters, without duplicating the weight file. The site's Modelfile guide details this workflow.

#Fix a broken chat template

This is the number-one trap when importing GGUF files. An imported model may respond unpredictably: sentences that never end, strange tags in the output (<|im_end|>, [INST], <end_of_turn>), answers that ignore the question or go into a loop. Nine times out of ten, the model is not bad—the chat template does not match the one used during training.

Each model family expects a specific conversation format: ChatML (<|im_start|>) for Qwen and many fine-tunes, [INST]...[/INST] for Mistral and Llama 2, <start_of_turn> for Gemma, and a specific format for Llama 3. If the GGUF embeds the wrong template in its metadata, or if Ollama guesses an incorrect one, responses degrade. The most common symptom is end-of-turn tags appearing as plain text in the response instead of stopping generation.

Symptom: visible tags
The model displays <|im_end|> or <|eot_id|> in its response. A corresponding PARAMETER stop is missing, or the template does not emit the correct end token.
Symptom: endless generation
The model never stops and keeps going through turns on its own. The expected stop token is not declared.
Symptom: inconsistent responses
The model ignores the system prompt or gives an off-target response. The role format (system/user/assistant) does not match the one used during training.

The fix is to provide the correct TEMPLATE and the correct PARAMETER stop settings in a Modelfile. The source of truth is the original model's Hugging Face “model card”: look for the “prompt format” or “chat template” section, which specifies the exact format. For a ChatML model (Qwen and derivatives), the template and stops look like this.

Modelfile — fixing a ChatML template
FROM ./mon-modele.Q4_K_M.gguf

TEMPLATE """{{ if .System }}<|im_start|>system
{{ .System }}<|im_end|>
{{ end }}{{ if .Prompt }}<|im_start|>user
{{ .Prompt }}<|im_end|>
{{ end }}<|im_start|>assistant
{{ .Response }}<|im_end|>
"""

PARAMETER stop "<|im_start|>"
PARAMETER stop "<|im_end|>"

Then rebuild with “ollama create” and test it. An effective way to retrieve the right template without rewriting it: start from an official model from the same family that is already present in Ollama and inspect its generated Modelfile, then reuse its TEMPLATE block.

Terminal — retrieve an existing template
# Voir le Modelfile complet d'un modèle officiel de la même famille
ollama show --modelfile qwen3.5:9b

# Copiez-en le bloc TEMPLATE et les PARAMETER stop
# dans votre propre Modelfile, puis reconstruisez
ollama create mon-modele -f ./Modelfile
→
Check the inherited template first
Before rewriting everything, run “ollama show --modelfile hf.co/...” on your imported model: Ollama displays the template it inferred from the GGUF. If it’s correct, there’s no need to redo it; if it’s missing or wrong, you know what to fix. Always compare it with the original model card.

#Troubleshooting

“Error: pull model manifest”
The hf.co path is misspelled, the repository is private/gated, or your version of Ollama is too old. Check the repository's exact URL and update Ollama.
The quantization tag does not exist
Ollama says the tag cannot be found: open “Files and versions” on Hugging Face and copy the exact suffix of the .gguf file (e.g., Q4_K_M, IQ4_XS). Case is ignored, but the name must match.
Very slow model / choppy responses
The model spills over into RAM/CPU because it lacks enough VRAM. Check with “ollama ps” whether it’s running on the GPU or CPU, and drop down one quantization level or model size.
The repository contains only safetensors
No .gguf available: look for a community-converted “GGUF” version, or convert the model yourself using the llama.cpp scripts.
Outputs polluted by tags
Incorrect chat template: see the previous section, and provide the correct TEMPLATE and PARAMETER stop values in a Modelfile.

#Go further

Importing a GGUF relies on two core skills from the Ollama ecosystem. These site guides build on this one:

Choose your quantization (Q4, Q5, Q8, FP16)
Understand the quality/memory trade-off in detail so you can choose the right GGUF variant for your GPU.
Customize a model with the Ollama Modelfile
Go further with the Modelfile: system prompts, parameters, templates, and multiple variants of the same model.
Install Ollama: Windows, macOS, and Linux
The up-to-date basic installation guide, if you're starting from scratch before importing your first GGUFs.
Did this guide help you?

Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.